Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “science data analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Machine learning in materials research: Developments over the last decade and challenges for the future

The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials science methods, and ML methods) used within papers that apply ML to materials science. The analysis demonstrates that despite the growth of deep learning techniques, the use of classical machine learning is still dominant as a whole. It also demonstrates how new research can effectively build upon past research, particular in the domain of ML models trained on density functional theory calculation data. Next, I present the progression of best scores as a function of time on the matbench materials science benchmark for formation enthalpy prediction. In particular, a dramatic improvement of 7 times reduction in error is obtained when progressing from feature-based methods that use conventional ML (random forest, support vector regression, etc.) to the use of graph neural network techniques. Finally, I provide views on future challenges and opportunities, focusing on data size and complexity, extrapolation, interpretation, access, and relevance.

36 MATERIALS SCIENCE↗

Using AI to Reproduce Neutrino Cross Section Analysis - Prototyping the Neutrino Discovery Platform

The Neutrino Discovery Platform (NDP) aims to accelerate DUNE-era science by making the neutrino program's existing datasets analyzable through fast, reproducible, and auditable workflows. We report a working version of two of its layers, data curation and agentic orchestration, built and tested end to end on MINERvA open data. The guiding lesson throughout is that a cross section is a measurement, and not just a plotted shape, only if it carries a defensible systematic-uncertainty budget, a trustworthy unfolding, and a reproducible record. Using a single medium-energy playlist pair from the MINERvA open-data release (about $2.05\times10^{17}$ protons on target of data), we first reproduced the shapes of two published charged-current inclusive $\nu_\mu$ measurements through a complete extraction ladder: selection, background subtraction, D'Agostini unfolding, efficiency correction, and flux normalization. These shape-level reproductions ran and tracked the published results, but they lacked the systematic-uncertainty machinery that defines a MINERvA cross section. To supply it, we vendored and built the MINERvA Analysis Toolkit and developed a many-universe systematic-uncertainty tool that produces a portable covariance artifact, a parallel event-loop runner, and a per-run auditability harness. Validated against a published covariance release, the toolchain reproduces the released statistical, flux, and muon-energy-scale terms and shows that they account for roughly 63\% of the total variance, with the remainder unreleased. Using this same infrastructure, we then performed a measurement of our own design, the hadronic recoil-energy distribution of low-energy ($E_\nu<2.5$~GeV) charged-current inclusive events, and found data/simulation shape agreement of $\chi^2/\mathrm{ndf}=1.26$. Together these results show that the platform supports original physics and not only reproductions.

Breaux, Auto [Tulane U. (main)]↗

Erratum to: Measurements of higher-order cumulants of multiplicity and net-electric charge distributions in inelastic proton-proton interactions by NA61/SHINE

This Erratum replaces, due to a discovery of coding mistakes, the following quantities: κ 3 /κ 1 of the h + − h − distribution presented in Fig. 6 and Table 6, κ 4 listed in Table 4, and $\hat{C}$ 4 presented in Fig. 7 and Table 5. All mentioned figures and tables were updated.

High-Energy Particle Collision Data Analysis↗

Elucidating Photoinduced Processes of Photosystem I Via Multidimensional Electronic and Vibrational Spectroscopies

This project was motivated by an overarching goal to elucidate the mechanism of energy and electron transfer that governs the efficient charge separation in photosystem I (PSI) complexes. PSI is a natural light harvesting complex that drives oxygenic photosynthesis in plants, algae, and cyanobacteria. It uses ~300 tightly packed chlorophylls (Chls) to absorb photons, transfer the excitation energy to the reaction center (RC), and generate a charge separated state with near unity quantum efficiency (QE). A better understanding of the mechanism of energy transfer and charge separation in PSI is required for understanding the high QE of natural light harvesting complexes, and it could lead to the further development of artificial photosynthetic systems for solar energy conversion and modification of light harvesting complexes to improve crop yields. We applied two-dimensional optical spectroscopies to different cyanobacterial photosystem I complexes, including PSI complexes that contain Chl f molecules, to map energy transfer pathways and gain insight into the efficient light harvesting of PSI. We used two-dimensional electronic spectroscopies (2DES) to map energy transfer in Chl a and Chl f containing PSI complexes. To investigate the Chl f PSI complexes, we modified our spectrometer to probe the lower energy states associated with Chl f molecules. We interpreted the 2DES spectra through global analysis procedures to generate maps of energy transfer. We also constructed a two-dimensional electronic vibrational (2DEV) spectrometer that will be used to investigate charge transfer transitions and dynamics within PSI complexes. Measurements were performed on model systems to establish general data analysis procedures for interpreting 2D spectra and gain insight into protein cofactor interactions.

14 SOLAR ENERGY↗

Human Host Cellular Response to HCoV-229E Infection Transcriptomics (ACS-DP1)

The purpose of this experiment was to evaluate the human host cellular response to wild-type Human coronavirus strain 229E (HCoV-229E) infection. Sample data was obtained for mock and infected immortalized human lung epithelial cells (A549) (MOI 5), immortalized human lung fibroblasts cells (MRC5) (MOI5), and primary human airway epithelial (HAE) (MOI 3) cells from lung tissue. Sample data was acquired using an Illumina HiSeq 2000 sequencer system and processed for RNA sequencing (RNA-Seq) expression analysis.

59 BASIC BIOLOGICAL SCIENCES↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

Protocols and methodologies for acquiring and analyzing critical-current versus longitudinal-strain data in Bi 2 Sr 2 CaCu 2 O 8+x wires

Abstract In the literature on Bi 2 Sr 2 CaCu 2 O 8+ x (Bi-2212) superconducting wires, it is evident that measurement protocols for transport critical-current I c versus longitudinal strain ϵ and definitions of the so-called ‘strain limit’ are generally dissimilar. Yet, values obtained for the ‘strain limit’ are frequently assimilated to being those of the irreversible strain limit ϵ irr , regardless of the I c degradation-criterion used to define it. In effect, ϵ irr should correspond specifically to the I c ( ϵ ) irreversibility onset , where crack formation in Bi-2212 filaments presumably starts. Because I c ( ϵ ) degradation remains progressive over a fairly wide strain range beyond ϵ irr , the different I c degradation-criteria in use do not yield to the same result and, thus, are not equivalent from metrology perspective. Indeed, in studying densified samples of a modern Bi-2212 round wire, we found ϵ irr ≈ 0.4% and ϵ 5% ≈ 0.6% ( ϵ 5% being the strain where I c degrades by 5%). In this paper, we outline and suggest I c ( ϵ )-measurement protocols and data-analysis methodologies in the hope to converge the various approaches taken for studying Bi-2212 strain properties and, thus, remove related result discrepancies. A unified approach would enable more objective data comparisons among laboratories and among different Bi-2212 conductors. It would pave the way for more rigorous studies of effects potentially associated with wire design, powder, heat treatments, and other such parameters on the conductor’s strain properties.

protocols↗

Review of Particle Physics - Scalar Mesons below 1 GeV

The summarizes much of particle physics and cosmology. Using data from previous editions, plus 2,717 new measurements from 869 papers, we list, evaluate, and average measured properties of gauge bosons and the recently discovered Higgs boson, leptons, quarks, mesons, and baryons. We summarize searches for hypothetical particles such as supersymmetric particles, heavy bosons, axions, dark photons, etc. Particle properties and search limits are listed in Summary Tables. We give numerous tables, figures, formulae, and reviews of topics such as Higgs Boson Physics, Supersymmetry, Grand Unified Theories, Neutrino Mixing, Dark Energy, Dark Matter, Cosmology, Particle Detectors, Colliders, Probability and Statistics. Most of the 120 reviews are updated, including many that are heavily revised. The is divided into two volumes. Volume 1 includes the Summary Tables and 97 review articles. Volume 2 consists of the Particle Listings and contains also 23 reviews that address specific aspects of the data presented in the Listings. The complete (both volumes) is published online on the website of the Particle Data Group () and in a journal. Volume 1 is available in print as the . A with the Summary Tables and essential tables, figures, and equations from selected review articles is available in print, as a web version optimized for use on phones, and as an Android app. The 2024 edition of the Review of Particle Physics should be cited as: S. Navas et al. (Particle Data Group), Phys. Rev. D 110, 030001 (2024)© 20242024

Navas, S. [Universidad de Granada]↗

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

A universal language for finding mass spectrometry data patterns

Despite being information rich, the vast majority of untargeted mass spectrometry data are underutilized; most analytes are not used for downstream interpretation or reanalysis after publication. The inability to dive into these rich raw mass spectrometry datasets is due to the limited flexibility and scalability of existing software tools. Here, in this study, we introduce a new language, the Mass Spectrometry Query Language (MassQL), and an accompanying software ecosystem that addresses these issues by enabling the community to directly query mass spectrometry data with an expressive set of user-defined mass spectrometry patterns. Illustrated by real-world examples, MassQL provides a data-driven definition of chemical diversity by enabling the reanalysis of all public untargeted metabolomics data, empowering scientists across many disciplines to make new discoveries. MassQL has been widely implemented in multiple open-source and commercial mass spectrometry analysis tools, which enhances the ability, interoperability and reproducibility of mining of mass spectrometry data for the research community.

Damiani, Tito [Czech Academy of Sciences (CAS), Pr↗

Comparative Analysis of DNA LLM Classification Techniques Using Intra-Layer Feature Extraction with Autoencoder Stacks [Poster]

This project conducts a comparative analysis of DNA LLM classification techniques using Evo2, Grover, and UTRML, focusing on intra-layer feature extraction in Evo2. By extracting features from multiple layers of Evo2 and integrating them into an autoencoder stack with a binary classification head, we evaluate its effectiveness in classifying genomic sequences compared to smaller DNA language models. My findings demonstrate that Evo2 outperforms Grover and UTRML in classification accuracy on a dataset provided by department 08625, CAO2021, while UTRML offers competitive performance with lower computational costs. This study highlights the potential of advanced embedding techniques in enhancing genomic data analysis and informs future research in bioinformatics.

59 BASIC BIOLOGICAL SCIENCES↗

Computer Vision Pipeline for Image Analysis for Freeze‐Fracture Electron Microscopy: Rosette Cellulose Synthase Complexes Case

In materials science, plant biology, agriculture, and environmental research, the automated analysis of high-magnification, complex microscopy images, such as those generated by freeze-fracture electron microscopy (FF-TEM), remains a critical challenge that limits the scalability of data interpretation. We present a deep learning computer vision pipeline for high-throughput detection and morphological characterization analysis of cellulose synthase complexes (CSCs, or rosettes) in FF-TEM images. The pipeline integrates preprocessing, detection, human-in-the-loop verification, and semantic segmentation to quantify features such as rosette diameter and inter-lobe spacing. The approach was trained and tested on a curated dataset of high-resolution FF-TEM micrographs of Physcomitrium patens, expanded via strategic tiling and augmentation to over 650 images. We compare YOLOv8 and YOLOv9 architectures and demonstrate that YOLOv9 achieves superior performance in both localization accuracy (mAP50-95 = 0.854) and inference speed. The resulting distributions revealed biological variability consistent with prior manual studies, validating the approach for high-throughput applications. Our results show that the pipeline achieves human-expert level accuracy while dramatically reducing analysis time, enabling scalable, reproducible structural characterization of intramembrane protein complexes. The pipeline is broadly applicable to other domains requiring precise interpretation of complex microscopy data and establishes a foundation for future artificial intelligence (AI)-assisted workflows in biological imaging.

59 BASIC BIOLOGICAL SCIENCES↗

Partial wave analysis of 𝑒 + ⁢𝑒 − → 𝜋 + ⁢𝜋 − ⁢𝐽/𝜓 and cross section measurement of 𝑒 + ⁢𝑒 − → 𝜋 ± ⁢𝑍 𝑐 ⁢(3900) ∓ from 4.1271 to 4.3583 GeV

Based on 12.0 fb −1 of 𝑒 + ⁢𝑒 − collision data samples collected by the BESIII detector at center-of-mass energies from 4.1271 to 4.3583 GeV, a partial wave analysis is performed for the process 𝑒 + ⁢𝑒 − → 𝜋 + ⁢𝜋 − ⁢𝐽/𝜓. The cross sections for the subprocesses 𝑒 + ⁢𝑒 − → 𝜋 + ⁢𝑍 𝑐 ⁢(3900) − + c.c. → 𝜋 + ⁢𝜋 − ⁢𝐽/𝜓, 𝑓 0 ⁡(980)⁢(→ 𝜋 + ⁢𝜋 − )⁢𝐽/𝜓, and (𝜋 + ⁢𝜋 − ) S−wave⁢ 𝐽/𝜓 are measured for the first time. The mass and width of the 𝑍 𝑐 ⁢(3900) ± are determined to be 3884.6 ± 0.7 ± 3.3 MeV/𝑐 2 and 37.2 ± 1.3 ± 6.6 MeV, respectively. The first errors are statistical and the second systematic. The final state (𝜋 + ⁢𝜋 − ) S−wave ⁢𝐽/𝜓 dominates the process 𝑒 + ⁢𝑒 − → 𝜋 + ⁢𝜋 − ⁢𝐽/𝜓. By analyzing the cross sections of 𝜋 ±⁢ 𝑍 𝑐 ⁢(3900) ∓ and 𝑓 0 ⁡(980)⁢𝐽/𝜓, 𝑌⁡(4220) has been observed. Its mass and width are determined to be 4225.7 ± 4.1 ± 3.4 MeV/𝑐 2 and 57.5 ± 9.4 ± 12.1 MeV, respectively.

lepton colliders↗

Challenges of open data in aquatic sciences: issues faced by data users and data providers

Free use and redistribution of data (i.e., Open Data) increases the reproducibility, transparency, and pace of aquatic sciences research. However, barriers to both data users and data providers may limit the adoption of Open Data practices. Here, we describe common Open Data challenges faced by data users and data providers within the aquatic sciences community (i.e., oceanography, limnology, hydrology, and others). These challenges were synthesized from literature, authors’ experiences, and a broad survey of 174 data users and data providers across academia, government agencies, industry, and other sectors. Through this work, we identified seven main challenges: 1) metadata shortcomings, 2) variable data quality and reusability, 3) open data inaccessibility, 4) lack of standardization, 5) authorship and acknowledgement issues 6) lack of funding, and 7) unequal barriers around the globe. Our key recommendation is to improve resources to advance Open Data practices. This includes dedicated funds for capacity building, hiring and maintaining of skilled personnel, and robust digital infrastructures for preparation, storage, and long-term maintenance of Open Data. Further, to incentivize data sharing we reinforce the need for standardized best practices to handle data acknowledgement and citations for both data users and data providers. We also highlight and discuss regional disparities in resources and research practices within a global perspective.

54 ENVIRONMENTAL SCIENCES↗

Prediction of the Cu oxidation state from EELS and XAS spectra using supervised machine learning

Abstract Electron energy loss spectroscopy (EELS) and X-ray absorption spectroscopy (XAS) provide detailed information about bonding, distributions and locations of atoms, and their coordination numbers and oxidation states. However, analysis of XAS/EELS data often relies on matching an unknown experimental sample to a series of simulated or experimental standard samples. This limits analysis throughput and the ability to extract quantitative information from a sample. In this work, we have trained a random forest model capable of predicting the oxidation state of copper based on its L-edge spectrum. Our model attains an R 2 score of 0.85 and a root mean square error of 0.24 on simulated data. It has also successfully predicted experimental L-edge EELS spectra taken in this work and XAS spectra extracted from the literature. We further demonstrate the utility of this model by predicting simulated and experimental spectra of mixed valence samples generated by this work. This model can be integrated into a real-time EELS/XAS analysis pipeline on mixtures of copper-containing materials of unknown composition and oxidation state. By expanding the training data, this methodology can be extended to data-driven spectral analysis of a broad range of materials.

36 MATERIALS SCIENCE↗

New constraint on the Np 237 ( n , γ ) Np 238 integral cross section using the Godiva-IV critical assembly

Accurate knowledge of the 237 Np(n, γ) 238 Np cross section at fast neutron energies is important for applied nuclear science. The presently available experimental data has large disagreements in the fast neutron region. Perform a model-independent measurement of the 237 Np(n, γ) 238 Np integral cross section using a well characterized fast neutron source and compare the result with previous measurements and current nuclear data evaluations. Provide an integral measurement that can be used as a benchmark for current evaluations. Multiple samples of 237 Np were irradiated in the Godiva-IV critical assembly. Following the irradiation, the samples placed in a γ-ray counting setup and the γ-rays emitted from the decay of 238 Np were measured over a time period of approximately 7 days. Multiple γ-ray decay branches of 238 Np were observed. The observed activity of 238 Np was used to calculate the amount of 238 Np produced during the irradiation via the 237 Np(n, γ) 238 Np reaction and an integral cross section of 342(11) mb was measured for the Godiva-IV neutron spectrum. Further, the 238 Np half-life has been measured with a result of 50.31(5) hours. The 237 Np(n, γ) 238 Np integral cross section measured in this work is in agreement with overlapping 1σ error bands to ENDF/B-VIII.0. However, the measured value is 3σ away from the calculated integral cross section using JENDL-5. This measurement offers a reliable benchmark for future 237 Np(n, γ) 238 Np cross section evaluations.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗