Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scarce data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Precision Measurement of the Neutron Magnetic Form Factor via the Ratio Method at Jefferson Lab Hall A

Protons and neutrons, collectively known as nucleons, are composed of quarks and gluons. The Sachs electromagnetic form factors encode information about the spatial distributions of charge and magnetization in the nucleon, particularly at low momentum transfer. In particular, the neutron magnetic form factor (GMn) provides crucial information about the distribution of magnetization inside the neutron and helps constrain theoretical models of nucleon structure. Quasi-elastic electron scattering from deuterium was measured up to Q^2=13.5 GeV^2 using the Super BigBite Spectrometer in Hall A at Jefferson Lab. In this work, the neutron magnetic form factor GMn was extracted at Q^2 = 3.0 GeV^2 and Q^2=4.5 GeV^2 using the Ratio Method. These results represent a subset of the full dataset collected in this experiment, which extended to significantly higher Q^2. The extracted GMn values agree with the existing global fit within approximately two standard deviations at Q^2=3.0 and show excellent agreement at Q^2=4.5. The measurements achieved systematic uncertainties of about 2% and statistical uncertainties below 0.5%, among the most precise determinations of GMn at these kinematics. These results demonstrate the robustness of the experimental technique and provide an important validation point for future extractions at higher Q^2, where data remain scarce. In addition, the GRINCH heavy gas Cherenkov detector—a key component of the experimental apparatus—was commissioned and achieved an electron detection efficiency of approximately 97%, supporting reliable particle identification. Together, the analysis presented here advances both our understanding of nucleon structure and the validation of the experimental methods and instrumentation used to access it.

Satnik, Maria [College of William and Mary, Willia↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Predicting Kyasanur forest disease in resource-limited settings using event-based surveillance and transfer learning

In recent years, the reports of Kyasanur forest disease (KFD) breaking endemic barriers by spreading to new regions and crossing state boundaries is alarming. Effective disease surveillance and reporting systems are lacking for this emerging zoonosis, hence hindering control and prevention efforts. We compared time-series models using weather data with and without Event-Based Surveillance (EBS) information, i.e., news media reports and internet search trends, to predict monthly KFD cases in humans. We fitted Extreme Gradient Boosting (XGB) and Long Short-Term Memory models at the national and regional levels. We utilized the rich epidemiological data from endemic regions by applying Transfer Learning (TL) techniques to predict KFD cases in new outbreak regions where disease surveillance information was scarce. Overall, the inclusion of EBS data, in addition to the weather data, substantially increased the prediction performance across all models. The XGB method produced the best predictions at the national and regional levels. The TL techniques outperformed baseline models in predicting KFD in new outbreak regions. Novel sources of data and advanced machine-learning approaches, e.g., EBS and TL, show great potential towards increasing disease prediction capabilities in data-scarce scenarios and/or resource-limited settings, for better-informed decisions in the face of emerging zoonotic threats.

60 APPLIED LIFE SCIENCES↗

Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles

The sciences struggle with poor integration across disciplines, the absence of coordination within and across data generation and modeling activities, scarce or disconnected open data, and weaknesses of networks to engage diverse stakeholders within and beyond the scientific community. The American Geophysical Union (AGU) is divided into 25 sections intended to encompass the breadth of the geosciences. Here, we introduce a special collection of commentary articles spanning 19 AGU sections on the challenges and opportunities associated with the use of ICON science principles. These principles focus on research intentionally designed to be Integrated, Coordinated, Open, and Networked (ICON) with the goal of maximizing mutual benefit (among stakeholders) and cross-system transferability of science outcomes. This article summarizes the ICON principles; discusses the crowdsourced approach to creating the collection; and explores insights from across the articles. There were multiple common themes among the commentary articles, including the broad agreement that the benefits of using ICON principles outweigh the costs, but that using ICON principles has important risks that need to be understood and mitigated. It was also clear that the ICON principles are not monolithic or static, but should instead be considered a heuristic tool that can and should be modified to meet changing needs. As a whole, the collection is intended as a resource for scientists pursuing ICON science and represents an important inflection point in which the geosciences community has come together around ICON principles as a unified approach for improving how science is done across the geosciences and beyond.

58 GEOSCIENCES↗

Leveraging explainable AI to characterize floating-point exceptions in linear solvers

Linear solver packages are central to many scientific, engineering, and machine learning applications. When floating-point exceptions occur in these solvers, e.g., division by zero or overflow, numerical results are compromised and become unreliable. Existing static and dynamic analysis tools can detect such exceptions, but they do not explain why the exceptions occur in terms of the solver inputs. Here, we present a study to characterize the inputs that cause numerical exceptions in linear solver packages. Our approach uses explainable AI (XAI) to find the most relevant characteristics of input matrices that explain the occurrence of exceptions in the solvers. Since training data in this domain is scarce, we perform extensive data gathering and data augmentation to obtain exception-inducing inputs. Our approach uses a repair strategy on the features blamed by XAI to validate that such features indeed explain the exceptions. We compare the LIME and SHAP XAI techniques using a dozen matrix features with three classifiers. We evaluate the approach on three widely used linear solver packages and find that some input characteristics can explain the occurrence of exceptions 100% of the time, in specific solvers and preconditioners.

Explainable AI↗

Image processing workflow yielding high contrast synchrotron nanoscale computed tomography data from Ni-YSZ electrodes

The operating lifetime of Ni-YSZ fuel electrodes used in solid oxide electrolysis cells and fuel cells (SOECs and SOFCs) is limited by Ni redistribution, one of the primary degradation mechanisms that must be overcome to extend the longevity and maximize the performance of SOECs and SOFCs. To achieve this, 3D microstructural data is needed to relate both initial performance and performance loss over time to microstructural properties and their evolution throughout operation under various conditions. However, 3D microstructure data remains relatively scarce within the literature due to multiple challenges in acquiring and analyzing such data reliably. This work presents a workflow for acquiring and processing synchrotron X-ray nanoscale computed tomography (nano-CT) data from Ni-YSZ electrodes. Parameters for each step in the nano-CT workflow are described up to the final result (a 3D reconstruction), with particular emphasis on image alignment using freely available software. Following the results of a parametric sweep of the image alignment step, high contrast, low signal-to-noise 3D nano-CT data is obtained with relatively short compute times. While the exact methods best suited to samples with different microstructural qualities, or similar Ni-YSZ nano-CT data obtained from other sources may deviate from the solution found herein, this work also generalizes the decision points and evaluation of each step to provide a starting point to adapt this workflow to other datasets.

08 HYDROGEN↗

Pluminate: Quantifying aerosol injection behavior from simulation, experimentation and observations

Marine aerosol injections are a key component in further understanding of both the potentials of deliberate injection for marine cloud brightening (MCB), a potential climate intervention (CI) strategy, and key aerosol-cloud interaction behaviors that currently form the largest uncertainty in global climate model (GCM) predictions of our climate. Since the rate of spread of aerosols in a marine environment directly translates to the effectiveness and ability of aerosol injections in impacting cloud radiative forcing, it is crucial to understand the spatial and temporal extent of injected-aerosol effects following direct injection into marine environments. The ubiquity of ship-injected aerosol tracks from satellite imagery renders observational validation of new parameterizations possible in 2D, however, 3D compatible data is more scarce, and necessary for the development of subgrid scale parameterizations of aerosol-cloud interactions in GCMs. This report introduces two novel parameterizations of atmospheric aerosol injection behavior suitable for both 3D (GCM-compatible) and 2D (observation-related) modeling. Their applicability is highlighted using a wealth of different observational data: small and larger scale salt-aerosol injection experiments conducted at SNL, 3D large eddy simulations of ship-injected aerosol tracks and 2D satellite images of ship tracks. The power of experimental data in enhancing knowledge of aerosol-cloud interactions is in particular emphasized by studying key aerosol microphysical and optical properties as observed through their mixing in cloud-like environments.

54 ENVIRONMENTAL SCIENCES↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗

Comprehensive assessment of metrology techniques for heliostat efficiency and performance evaluation

Concentrating solar power plants, specifically central receiver type systems and their heliostat field, are struggling with negative reputation in the USA, due to perceived underperformance and reliability issues. This is in part due to a lack of standards for performance assessment as well as overly simplified techno-economical models. A better understanding of influences and losses along the solar radiation path from the sun, across the solar collector to the receiver, increases the fidelity of heliostat efficiency assessment as well as solar field performance predictions. Such data are currently scarce and require a complete set of metrology capabilities to evaluate direct solar irradiance, sun shape, atmospheric attenuation, reflectance, collector shape, slope errors and total beam dispersion. In preparation for establishing a 3rd party metrology platform in collaboration with Sandia National Labs, NLR conducted a scoping study on available metrology. We present an extensive overview of techniques and commercial systems for each category. Our work includes an analysis to increase understanding of strengths and limitations of the many techniques used for surface shape and slope measurement. This applies to a controlled, indoor or outdoor laboratory environment assessing a single heliostat.

14 SOLAR ENERGY↗

Multi-Scale Modeling of the Evolution of Structure and Properties in Materials for Nuclear Energy Applications [Slides]

Nuclear energy is an important component of an overall strategy to address climate change. Idaho National Laboratory (INL) is the U.S. Department of Energy’s primary facility for research and development in nuclear science and technology for energy generation, supporting the improvement and life extension of the existing reactor fleet and the development and licensing of new reactor designs. Computational modeling is an important component of these activities, particularly in the area of materials for nuclear applications, where experimental data can be very challenging and expensive to acquire, and where data is especially scarce for new reactor designs. INL has used multi-scale modeling – linking atomistic, mesoscale, and engineering scales – to improve the ability to predict the performance of materials for nuclear energy applications. In this talk, I will give an overview of the approach and tools used, and several examples of application, including performance of nuclear fuels, understanding radiation-driven formation of nanoscale void and gas bubble superlattices, and powder densification through electric field assisted sintering.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Multi-scale modeling of the evolution of structure and properties in materials for nuclear energy applications

Nuclear energy is an important component of an overall strategy to address climate change. Idaho National Laboratory (INL) is the U.S. Department of Energy’s primary facility for research and development in nuclear science and technology for energy generation, supporting the improvement and life extension of the existing reactor fleet and the development and licensing of new reactor designs. Computational modeling is an important component of these activities, particularly in the area of materials for nuclear applications, where experimental data can be very challenging and expensive to acquire, and where data is especially scarce for new reactor designs. INL has used multi-scale modeling – linking atomistic, mesoscale, and engineering scales – to improve the ability to predict the performance of materials for nuclear energy applications. These modeling efforts make extensive of MOOSE (Multiphysics Object-Oriented Simulation Environment), a general-purpose open source finite element framework developed at INL. In this talk, I will give an overview of the approach and tools used, and several examples of application, including performance of nuclear fuels, understanding radiation-driven formation of nanoscale void and gas bubble superlattices, and powder densification through electric field assisted sintering.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Multi-scale modeling of the evolution of structure and properties in materials for nuclear energy applications

Nuclear energy is an important component of an overall strategy to address climate change. Idaho National Laboratory (INL) is the U.S. Department of Energy’s primary facility for research and development in nuclear science and technology for energy generation, supporting the improvement and life extension of the existing reactor fleet and the development and licensing of new reactor designs. Computational modeling is an important component of these activities, particularly in the area of materials for nuclear applications, where experimental data can be very challenging and expensive to acquire, and where data is especially scarce for new reactor designs. INL has used multi-scale modeling – linking atomistic, mesoscale, and engineering scales – to improve the ability to predict the performance of materials for nuclear energy applications. These modeling efforts make extensive of MOOSE (Multiphysics Object-Oriented Simulation Environment), a general-purpose open source finite element framework developed at INL. In this talk, I will give an overview of the approach and tools used, and several examples of application, including performance of nuclear fuels, understanding radiation-driven formation of nanoscale void and gas bubble superlattices, and powder densification through electric field assisted sintering.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Rubidium and potassium isotopic variations in chondrites and Mars: Accretion signatures and planetary overprints

As moderately volatile elements, isotopes of Rb and K can trace volatilization processes in planetary bodies. Rubidium isotopic data are however very scarce, especially for non-carbonaceous meteorites. Here, in this study, we report combined Rb and K isotopic data (δ 87/85 Rb and δ 41/39 Κ) for 7 ordinary, 6 enstatite, and 4 Martian meteorite falls to understand the causes for the variations in volatile abundances and isotopic compositions. Bulk Rb and K isotopic compositions of planetary bodies are estimated to be (Table 1): Mars +0.10 ± 0.03 ‰ for Rb and -0.26 ± 0.05 ‰ for K, bulk OCs $-0.12^{+0.15}_{-0.24}$ ‰ for Rb and $-0.72^{+0.28}_{-0.41}$ ‰ for K, bulk ECs $-0.02^{+0.29}_{-0.26}$ ‰ for Rb and $-0.33^{+0.37}_{-0.23}$ ‰ for K. The bulk K isotopic compositions of subgroup OCs are estimated to be $-0.72^{+0.26}_{-0.55}$ ‰ for H chondrites, $-0.71^{+0.23}_{-0.39}$ ‰ for L chondrites, and $-0.77^{+0.63}_{-0.30}$ ‰ for LL chondrites. A broad correlation between the Rb and K isotopic compositions of planetary bodies is observed. The correlation follows a slope that is consistent with kinetic evaporation and condensation processes, suggesting volatility-controlled mass-dependent isotope fractionation (as opposed to nucleosynthetic anomalies). Individual ordinary and enstatite chondrites show large Rb and K isotopic variations (-1.02 to +0.29 ‰ for Rb and -0.91 to -0.15 ‰ for K). Samples of lower metamorphic grades display correlated elemental and isotopic fractionations between Rb and K, while samples of higher metamorphic grades show great scatter, suggesting that chondrite parent-body processes have decoupled the two elements and their isotopes at the sample scale. Several processes could have contributed to the observed isotopic variations of Rb and K, including (i) chondrule “nugget effect”, (ii) volatilization during parent-body thermal metamorphism (heat-induced vaporization and gas transport within parent bodies), (iii) thermal diffusion during parent-body metamorphism, and (iv) impact/shock heating. Quantitative modeling of the first two processes suggests that neither of them could produce isotopic variations large enough to explain the observed isotopic variations. Volatilization during parent-body thermal metamorphism [the scenario (ii)], which has been commonly invoked to explain the isotopic variations of volatile elements, is gas transport-limited and its effect on isotopic fractionations of moderately volatile elements should be negligible. Modeling of diffusion processes suggests that (iii) could produce K isotopic variation comparable to the observed variation. The large isotopic variations in non-carbonaceous meteorites are thus most likely due to diffusive redistribution of K and Rb during metamorphism and/or shock-induced heating and vaporization.

58 GEOSCIENCES↗

Enhancing Electron Microscopy Image Classification Using Data Augmentation

Manual labeling for machine learning tasks such as image classification is tedious and labor-intensive; as a result, scientific datasets suitable for deep learning applications are scarce and limited. While data augmentation techniques have shown promise for extending image datasets, very little work has been done to understand the impact of combining multiple augmentation methods sequentially or the limits of their effectiveness when combined. Our work addresses this gap by examining how standard and combinatorial data augmentation affects the performance of machine learning models when trained on small datasets for label classification tasks. For our analysis, we generate single, double and quadruple-augmented datasets for a microscopy image classification task using six standard augmentation methods, and compare the resultant improvements observed in binary classification accuracy with three standard image classification models (DenseNet169, MobileNetV2, ResNet101V2). Our experiments show a non-monotonic relationship between the number of simultaneous augmentation methods and classification accuracy, indicating that there is a trade-off between the degree of augmentation and the model performance. These findings suggest that the optimal number of augmentation methods will vary by domain and use case. We also find that the order in which augmentation methods are applied to a limited dataset matters when combining augmentation schemes, with our use case showing performance differences up to 2.6% when the augmentation order is reversed for double-augmented datasets. Our work offers insights to the limits of data augmentation when working on image classification tasks with limited datasets.

Welsman, Jordan A↗

Measurements of Gamow-Teller transitions from 59 Co via the 59 Co ⁢(𝑡, 3 He +𝛾) charge-exchange reaction and its application to the stellar electron-capture rates

Electron-capture reactions on iron-group nuclei play a crucial role in the late stages of massive star evolution. Since stellar evolution simulations depend on accurate electron-capture rates—which are highly sensitive to the detailed Gamow-Teller (GT) strength distributions—reliable theoretical models are essential. However, experimental data on GT strength distributions are scarce. High-resolution measurements are therefore vital for benchmarking and improving these theoretical calculations. To provide high-resolution data on Gamow-Teller strength distributions of iron-group nuclei and to compare these results with theoretical calculations within this mass region. Differential cross sections for the 59 Co ⁢(𝑡, 3 He)⁢ 59 Fe charge-exchange reaction at 115 MeV/u were measured using the S800 spectrometer. Furthermore, to resolve individual levels that are not distinguishable in the S800 particle singles data, coincident 𝛾 rays from the 59 Fe residual nucleus were detected by using the Gamma-Ray Energy Tracking In-beam Nuclear Array 𝛾-ray tracking array. Here, the Gamow-Teller transition strength distribution from the ground state of 59 Co to 59 Fe was extracted up to an excitation energy of 10 MeV. Additionally, transition strengths for several low-lying states were determined from coincident 𝛾-ray measurements. Electron-capture rates calculated using the present data indicate that these low-lying states contribute significantly to the overall rates in relevant stellar environments. The experimental results show reasonable agreement with theoretical predictions based on both shell-model and projected shell-model calculations. High-resolution data on Gamow-Teller strength distributions—particularly for individual low-lying states—are essential for accurately determining electron-capture rates in iron-group nuclei. Coincident 𝛾-ray measurements provide a powerful tool for obtaining such detailed information. While the present work demonstrates that shell-model calculations successfully reproduce the experimental results, such comparisons are scarce and more experimental data are desirable.

59 ≤ A ≤ 89↗

Comparison of venous and pooled capillary hemoglobin levels for the detection of anemia among adolescent girls

Introduction: Blood source is a known preanalytical factor affecting hemoglobin (Hb) concentrations, and there is evidence that capillary and venous blood may yield disparate Hb levels and anemia prevalence. However, data from adolescents are scarce. Objective: To compare Hb and anemia prevalence measured by venous and individual pooled capillary blood among a sample of girls aged 10–19 years from 232 schools in four regions of Ghana in 2022. Methods: Among girls who had venous blood draws, a random subsample was selected for capillary blood. Hb was measured using HemoCue® Hb-301. We used Lin’s concordance correlation coefficient (CCC) to quantify the strength of the bivariate relationship between venous and capillary Hb and a paired t-test for difference in means. We used McNemar’s test for discordance in anemia cases by blood source and weighted Kappa to quantify agreement by anemia severity. A multivariate generalized estimating equation was used to quantify adjusted population anemia prevalence and assess the association between blood source and predicted anemia risk. Results: We found strong concordance between Hb measures (CCC = 0.86). The difference between mean venous Hb (12.8 g/dL, ± 1.1) and capillary Hb (12.9 g/dL, ± 1.2) was not significant (p = 0.26). Crude anemia prevalence by venous and capillary blood was 20.6% and 19.5%, respectively. Adjusted population anemia prevalence was 23.5% for venous blood and 22.5% for capillary ( = 0.45). Blood source was not associated with predicted anemia risk (risk ratio: 0.99, 95% CI: 0.96, 1.02). Discordance in anemia cases by blood source was not significant (McNemar p = 0.46). Weighted Kappa demonstrated moderate agreement by severity (k =0.67). Among those with anemia by either blood source (n = 111), 59% were identified by both sources. Conclusion: In Ghanaian adolescent girls, there was no difference in mean Hb, anemia prevalence, or predicted anemia risk by blood source. However, only 59% of girls with anemia by either blood source were identified as having anemia by both sources. These findings suggest that pooled capillary blood may be useful for estimating Hb and anemia at the population level, but that caution is needed when interpreting individual-level data.

59 BASIC BIOLOGICAL SCIENCES↗

Structural constraint integration in a generative model for the discovery of quantum materials

Billions of organic molecules have been computationally generated, yet functional inorganic materials remain scarce due to limited data and structural complexity. Here, in this work, we introduce Structural Constraint Integration in a GENerative model (SCIGEN), a framework that enforces geometric constraints, such as honeycomb and kagome lattices, within diffusion-based generative models to discover stable quantum materials candidates. SCIGEN enables conditional sampling from the original distribution, preserving output validity while guiding structural motifs. This approach generates ten million inorganic compounds with Archimedean and Lieb lattices, over 10% of which pass multistage stability screening. High-throughput density functional theory calculations on 26,000 candidates shows over 95% convergence and 53% structural stability. A graph neural network classifier detects magnetic ordering in 41% of relaxed structures. Furthermore, we synthesize and characterize two predicted materials, TiPd 0.22 Bi 0.88 and Ti 0.5 Pd 1.5 Sb, which display paramagnetic and diamagnetic behaviour, respectively. Our results indicate that SCIGEN provides a scalable path for generating quantum materials guided by lattice geometry.

36 MATERIALS SCIENCE↗