Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data scarce”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Coaxial laser absorption and optical emission spectroscopy of high-pressure aluminum monoxide

This work advances laser absorption spectroscopy with measurements of aluminum monoxide (AlO) temperature and column density in extreme pressure ( P > 60 bar) and temperature ( T > 4000 K) environments. Measurements of the AlO A 2 Π i – X 2 Σ + transition are made using a microelectromechanical system, tunable vertical cavity surface emitting laser (MEMS-VCSEL). Simultaneous emission measurements of the AlO B 2 Σ + – X 2 Σ + transition are made along a line of sight that is coaxial with the laser absorption. Absorption temperature fits agree with emission spectra for a T = 3200 K, P = 9 bar case. In cases with T > 4000 K, P > 60 bar, absorption fits match the ambient temperature while emission fits over-estimate it, owing to high optical depths. These data juxtapose passive and active spectroscopic methods and demonstrate the versatility of AlO laser absorption in high-pressure and high-temperature environments where experimental data remain scarce, and engineering models will benefit from refined measurements.

Daniel, K. A.↗

Phosphorus sorption and its environmental predictors across pantropical forest soils sampled over the past decade

Tropical forest productivity is frequently constrained by soil phosphorus (P) availability, yet global Land Surface Model (LSM), which are used to simulate ecosystem processes, still represent P cycling in tropical regions only in a limited way, largely because of scarce observational data. Phosphorus adsorption and desorption of dissolved inorganic P to and from soil minerals (hereafter termed sorption), is an important process for predicting how much P is available to plants. This dataset was created to improve predictions of soil P sorption in tropical soils by identifying the isotherm equation that best describes pantropical soils. It includes raw measurements of environmental variables, such as soil properties and climate, together with P sorption data collected from 40 forest soil pits from 9 Forest Global Earth Observatory (ForestGEO) sites across 7 tropical countries during 2018-2022. The data are organized by site, country, and continent. Each site may include several soil pits. For each pit, P sorption was measured across a range of soil P concentrations to build sorption isotherm curves, typically with about 6 to 8 measurements per curve. While sorbed P varies across these concentration levels, the other environmental variables remain constant at the plot level.

Aluminum oxide↗

Development of a Annual Air Handling Unit Fault Dataset for FDD Tools: Lessons Learned and Considerations for FDD Developers

As energy management and information systems (e.g., automated fault detection and diagnostics [AFDD] tools) become more prevalent in the commercial building stock, it is important to determine the effectiveness of these technologies by benchmarking their performance. The authors have been working to develop the largest publicly available dataset of HVAC fault data for performance benchmarking applications, covering the most common HVAC systems and designs including chiller plants, rooftop packaged units, dual duct air handling units and single duct air handling units. This study covers the development, modeling, and validation of a synthetic fault dataset for a single duct air handling unit (AHU), one of the most common HVAC configurations found in the commercial building stock. Despite this being a common system, real-world time series data are scarce and usually do not span a wide range of weather conditions. Due to this limitation, a detailed AHU model was employed to carry out annual simulations of numerous common sensor and mechanical faults, which were then validated by comparing their effects on system performance to expected symptoms. We summarize the nature of each fault and their impacts under different weather and operation conditions. Finally, we highlight considerations for FDD developers that may want to use this dataset to assess their algorithms’ performance and their improvement over time.

Casillas, Armando↗

Bridging Power System Protection Gaps with Data-driven Approaches

Protection is a critical function in power systems to avoid equipment damage, maintain personnel safety, and support system reliability. However, current protective relay technology cannot adequately protect equipment and personnel from effects of some events; these deficiencies are termed protection gaps. In this research, a data-driven approach is proposed to complement traditional protection technology and distinguish fault conditions from transients caused by normal operations. A convolutional neural network (CNN) based fault detection approach is implemented to achieve data translation invariance of the time-series input data. As a result, the data-driven method can accurately detect system faults despite variation and noise in the input data. In addition, using the CNN–based method avoids the complicated manual feature extraction procedure required by many traditional data-driven methods. The effectiveness of the proposed approach is tested on four kinds of protection gaps: high impedance faults, transformer/generator inter-turn faults, distribution system PV circuit faults, and the mis-operation situations of Zone 3 line protection relays operating under system stress. Finally, a transfer learning method is also proposed to address the common issue of data-driven methods for which real-world training data are scarce. Extensive study results demonstrate that the proposed approach can accurately bridge power system protection gaps.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Load Forecasting for the Moroccan Electricity Sector

The Moroccan electricity sector is undergoing rapid transformation as it seeks to increase its utilization of renewable energy from its abundant domestic supply. Key to implementing variable renewable energy is understanding current electricity demand and forecasting this demand on the long, medium, and short timescales. This report leverages existing Moroccan electricity sector data to build basic load forecasts on these timescales. Taking these forecasts, the report recommends next steps in terms of additional algorithms, mathematical models, data collection, and scenarios (such as vehicle electrification or high levels of distributed generation) that should be examined for advanced load forecasts.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

In-situ thermodynamics measurements at metal oxides-solution interfaces using Flow Adsorption Microcalorimetry.

Mineral-fluid interfaces are the principal sites of geochemical processes near Earth’s surface, hosting chemical reactions that play a fundamental role in (bio)geochemical cycles and in the fate and transport of anthropogenic contaminants, and that effectively control the compositions of soil and water environments. These complex interfaces are critical for our energy and environmental future. The mineral-fluid interface has been studied in an unprecedented level of detail with both experimental and computational approaches, separately and in combination. However, conspicuously missing from studies of the mineral-fluid interface are direct measurements of the energies of ion sorption and exchange. The literature on energetics and enthalpies of exchange, adsorption, dissolution, precipitation, and surface protonation reactions, especially those directly supported by experimental data, remains scarce despite their fundamental nature. The overarching goal of this project is to complete a systematic study of the thermodynamics properties of interfacial reactions at four MO surfaces (Rutile (α-TiO2), Quartz (SiO2), boehmite (γ-AlOOH) and goethite (α-FeOOH)) through the application and construction of novel flow adsorption microcalorimetry techniques and instrumentations. These unique and specialized microcalorimeters will operate at various temperatures and solution chemical compositions allowing for in-situ measurements across metal oxides and ligands of various characteristics. The overall research goal will be accomplished by completing the following three specific objectives (O): O1) Determine the energetics of surface protonation and deprotonation, ion exchange and ligand sorption reactions; O2) Investigate the surface charge thermodynamic properties under a range of temperature and solution chemical compositions; and O3) Develop predictive trends of the interplay between MO structure, surface coverage and surface reactivity. In addition to key thermodynamics parameters, calorimetric measurements provide a wealth of mechanistic information about reactions energetics and kinetics, surface charge characteristics, and structure-reactivity or selectivity relationships, all obtained in-situ and in real-time. This report includes science highlights from various projects completed over the performance period of the project. Also listed are dissemination opportunities, people supported on the grant and the impact on available physical resources and the discipline as a whole.

58 GEOSCIENCES↗

Predicted thermophysical properties of UN, PuN and (U,Pu)N

Molecular dynamics and density functional theory simulations are used to predict the lattice and electronic contributions of thermophysical properties for UN, PuN, and mixed (U,Pu)N systems. The properties predicted include the lattice parameter, linear thermal expansion, enthalpy, and specific heat capacity, as a function of temperature. The simulation predictions for high temperature specific heat capacity are compared against experimental measurements to understand the behavior, and why differences in the experimental measurements are observed. The influence of adding U vacancies, N interstitials, and Pu to UN is also examined. For this, a new PuN potential parameter set is developed and used with the Kocevski UN potential, enabling the dynamics of mixed (U,Pu)N systems to be studied. How defects impact the thermophysical properties is important for understanding fuel behavior under different reactor conditions, and these mechanistic predictions can be used to support fuel performance codes where data is scarce.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An Ab Initio Molecular Dynamics Study of Key Thermodynamic Input Parameters for Computer Simulation of U-6Nb Solidification

The key to metallic fuel development is the fabrication of uranium metal and alloys into fuel forms. U-Nb alloys are one of the best candidates for a metallic fuel alloy with high-temperature strength sufficient to support the core, acceptable nuclear properties, good fabricability, and compatibility with usable coolant media. Melt processing has been a key component of the metallic fuel cycle, and process models require thermophysical parameters at elevated temperatures, particularly above the melting temperatures, regarding which experimental data are scarce, for accurate simulations and process development. By means of ab initio density-functional theory (DFT) quantum molecular dynamics (QMD), we have calculated the main thermophysical parameters—the density, thermal expansion coefficient, specific heat, thermal conductivity, melting temperature, latent heat of fusion, and viscosity—used in the modeling of the U-6 wt.% Nb alloy casting. The melting temperature of the U-6 wt.% Nb alloy at ambient pressure is obtained by means of QMD simulations using the Z-method. The ambient volume change and latent heat of melting of U-6 wt.% Nb are also derived from QMD simulations in conjunction with analytical fitting for the energy and pressure. The thermal conductivity for the solid U-Nb alloy is calculated from the semi-classical Boltzmann transport equation combined with an estimate of the electron relaxation time obtained from DFT simulations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Missing Wedge Completion via Unsupervised Learning with Coordinate Networks

Cryogenic electron tomography (cryoET) is a powerful tool in structural biology, enabling detailed 3D imaging of biological specimens at a resolution of nanometers. Despite its potential, cryoET faces challenges such as the missing wedge problem, which limits reconstruction quality due to incomplete data collection angles. Recently, supervised deep learning methods leveraging convolutional neural networks (CNNs) have considerably addressed this issue; however, their pretraining requirements render them susceptible to inaccuracies and artifacts, particularly when representative training data is scarce. To overcome these limitations, we introduce a proof-of-concept unsupervised learning approach using coordinate networks (CNs) that optimizes network weights directly against input projections. This eliminates the need for pretraining, reducing reconstruction runtime by 3–20× compared to supervised methods. Our in silico results show improved shape completion and reduction of missing wedge artifacts, assessed through several voxel-based image quality metrics in real space and a novel directional Fourier Shell Correlation (FSC) metric. Our study illuminates benefits and considerations of both supervised and unsupervised approaches, guiding the development of improved reconstruction strategies.

42 ENGINEERING↗

Runoff evaluation in an Earth System Land Model for permafrost regions in Alaska

Modeling of hydrological runoff is essential for accurately capturing spatiotemporal feedbacks within the land–atmosphere system, particularly in sensitive regions such as permafrost landscapes. However, substantial uncertainties persist in the terrestrial runoff parameterization schemes used in Earth system and land surface models. This is particularly true in permafrost regions, where landscape heterogeneity is high and reliable observational data are scarce. In this study, we evaluate the performance of runoff parameterization schemes in the Energy Exascale Earth System Model (E3SM) land model (ELM). Our proposed framework leverages simulation results from the Advanced Terrestrial Simulator (ATS), which is a physics-based integrated surface/subsurface hydrologic model that has been successfully evaluated previously in Arctic tundra regions. We used ATS to simulate runoff from 22 representative hillslopes in the Sagavanirktok River basin, located on the North Slope of Alaska, then compared the output with ELM's parameterized representation of total runoff. Results show that (1) ELM's total runoff was the same order of magnitude as the ATS simulations, and both models were similarly variable over time; (2) minor adjustments to coefficients in ELM's runoff parameterization improved the match between the ATS simulation and ELM's parameterized representation of annual and seasonal total runoff; (3) overall, runoff responses in ATS and ELM are more similar in flat hillslope environments compared to steep hillslopes; and (4) shallower active layer thicknesses and higher precipitation simulations resulted in lower correlations between the two models due to greater total runoff. By incorporating the optimized runoff coefficients from the Sagavanirktok River basin into ELM, the simulated total runoff better matched the streamflow observations at a small watershed located on the Seward Peninsula of Alaska. Our findings revealed important insights into the effectiveness of runoff parameterizations in land surface models and pathways for improving runoff coefficients in typical Arctic regions.

54 ENVIRONMENTAL SCIENCES↗

Precision Measurement of the Neutron Magnetic Form Factor via the Ratio Method at Jefferson Lab Hall A

Protons and neutrons, collectively known as nucleons, are composed of quarks and gluons. The Sachs electromagnetic form factors encode information about the spatial distributions of charge and magnetization in the nucleon, particularly at low momentum transfer. In particular, the neutron magnetic form factor (GMn) provides crucial information about the distribution of magnetization inside the neutron and helps constrain theoretical models of nucleon structure. Quasi-elastic electron scattering from deuterium was measured up to Q^2=13.5 GeV^2 using the Super BigBite Spectrometer in Hall A at Jefferson Lab. In this work, the neutron magnetic form factor GMn was extracted at Q^2 = 3.0 GeV^2 and Q^2=4.5 GeV^2 using the Ratio Method. These results represent a subset of the full dataset collected in this experiment, which extended to significantly higher Q^2. The extracted GMn values agree with the existing global fit within approximately two standard deviations at Q^2=3.0 and show excellent agreement at Q^2=4.5. The measurements achieved systematic uncertainties of about 2% and statistical uncertainties below 0.5%, among the most precise determinations of GMn at these kinematics. These results demonstrate the robustness of the experimental technique and provide an important validation point for future extractions at higher Q^2, where data remain scarce. In addition, the GRINCH heavy gas Cherenkov detector—a key component of the experimental apparatus—was commissioned and achieved an electron detection efficiency of approximately 97%, supporting reliable particle identification. Together, the analysis presented here advances both our understanding of nucleon structure and the validation of the experimental methods and instrumentation used to access it.

Satnik, Maria [College of William and Mary, Willia↗

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Predicting Kyasanur forest disease in resource-limited settings using event-based surveillance and transfer learning

In recent years, the reports of Kyasanur forest disease (KFD) breaking endemic barriers by spreading to new regions and crossing state boundaries is alarming. Effective disease surveillance and reporting systems are lacking for this emerging zoonosis, hence hindering control and prevention efforts. We compared time-series models using weather data with and without Event-Based Surveillance (EBS) information, i.e., news media reports and internet search trends, to predict monthly KFD cases in humans. We fitted Extreme Gradient Boosting (XGB) and Long Short-Term Memory models at the national and regional levels. We utilized the rich epidemiological data from endemic regions by applying Transfer Learning (TL) techniques to predict KFD cases in new outbreak regions where disease surveillance information was scarce. Overall, the inclusion of EBS data, in addition to the weather data, substantially increased the prediction performance across all models. The XGB method produced the best predictions at the national and regional levels. The TL techniques outperformed baseline models in predicting KFD in new outbreak regions. Novel sources of data and advanced machine-learning approaches, e.g., EBS and TL, show great potential towards increasing disease prediction capabilities in data-scarce scenarios and/or resource-limited settings, for better-informed decisions in the face of emerging zoonotic threats.

60 APPLIED LIFE SCIENCES↗

Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles

The sciences struggle with poor integration across disciplines, the absence of coordination within and across data generation and modeling activities, scarce or disconnected open data, and weaknesses of networks to engage diverse stakeholders within and beyond the scientific community. The American Geophysical Union (AGU) is divided into 25 sections intended to encompass the breadth of the geosciences. Here, we introduce a special collection of commentary articles spanning 19 AGU sections on the challenges and opportunities associated with the use of ICON science principles. These principles focus on research intentionally designed to be Integrated, Coordinated, Open, and Networked (ICON) with the goal of maximizing mutual benefit (among stakeholders) and cross-system transferability of science outcomes. This article summarizes the ICON principles; discusses the crowdsourced approach to creating the collection; and explores insights from across the articles. There were multiple common themes among the commentary articles, including the broad agreement that the benefits of using ICON principles outweigh the costs, but that using ICON principles has important risks that need to be understood and mitigated. It was also clear that the ICON principles are not monolithic or static, but should instead be considered a heuristic tool that can and should be modified to meet changing needs. As a whole, the collection is intended as a resource for scientists pursuing ICON science and represents an important inflection point in which the geosciences community has come together around ICON principles as a unified approach for improving how science is done across the geosciences and beyond.

58 GEOSCIENCES↗

Leveraging explainable AI to characterize floating-point exceptions in linear solvers

Linear solver packages are central to many scientific, engineering, and machine learning applications. When floating-point exceptions occur in these solvers, e.g., division by zero or overflow, numerical results are compromised and become unreliable. Existing static and dynamic analysis tools can detect such exceptions, but they do not explain why the exceptions occur in terms of the solver inputs. Here, we present a study to characterize the inputs that cause numerical exceptions in linear solver packages. Our approach uses explainable AI (XAI) to find the most relevant characteristics of input matrices that explain the occurrence of exceptions in the solvers. Since training data in this domain is scarce, we perform extensive data gathering and data augmentation to obtain exception-inducing inputs. Our approach uses a repair strategy on the features blamed by XAI to validate that such features indeed explain the exceptions. We compare the LIME and SHAP XAI techniques using a dozen matrix features with three classifiers. We evaluate the approach on three widely used linear solver packages and find that some input characteristics can explain the occurrence of exceptions 100% of the time, in specific solvers and preconditioners.

Explainable AI↗

Image processing workflow yielding high contrast synchrotron nanoscale computed tomography data from Ni-YSZ electrodes

The operating lifetime of Ni-YSZ fuel electrodes used in solid oxide electrolysis cells and fuel cells (SOECs and SOFCs) is limited by Ni redistribution, one of the primary degradation mechanisms that must be overcome to extend the longevity and maximize the performance of SOECs and SOFCs. To achieve this, 3D microstructural data is needed to relate both initial performance and performance loss over time to microstructural properties and their evolution throughout operation under various conditions. However, 3D microstructure data remains relatively scarce within the literature due to multiple challenges in acquiring and analyzing such data reliably. This work presents a workflow for acquiring and processing synchrotron X-ray nanoscale computed tomography (nano-CT) data from Ni-YSZ electrodes. Parameters for each step in the nano-CT workflow are described up to the final result (a 3D reconstruction), with particular emphasis on image alignment using freely available software. Following the results of a parametric sweep of the image alignment step, high contrast, low signal-to-noise 3D nano-CT data is obtained with relatively short compute times. While the exact methods best suited to samples with different microstructural qualities, or similar Ni-YSZ nano-CT data obtained from other sources may deviate from the solution found herein, this work also generalizes the decision points and evaluation of each step to provide a starting point to adapt this workflow to other datasets.

08 HYDROGEN↗

Pluminate: Quantifying aerosol injection behavior from simulation, experimentation and observations

Marine aerosol injections are a key component in further understanding of both the potentials of deliberate injection for marine cloud brightening (MCB), a potential climate intervention (CI) strategy, and key aerosol-cloud interaction behaviors that currently form the largest uncertainty in global climate model (GCM) predictions of our climate. Since the rate of spread of aerosols in a marine environment directly translates to the effectiveness and ability of aerosol injections in impacting cloud radiative forcing, it is crucial to understand the spatial and temporal extent of injected-aerosol effects following direct injection into marine environments. The ubiquity of ship-injected aerosol tracks from satellite imagery renders observational validation of new parameterizations possible in 2D, however, 3D compatible data is more scarce, and necessary for the development of subgrid scale parameterizations of aerosol-cloud interactions in GCMs. This report introduces two novel parameterizations of atmospheric aerosol injection behavior suitable for both 3D (GCM-compatible) and 2D (observation-related) modeling. Their applicability is highlighted using a wealth of different observational data: small and larger scale salt-aerosol injection experiments conducted at SNL, 3D large eddy simulations of ship-injected aerosol tracks and 2D satellite images of ship tracks. The power of experimental data in enhancing knowledge of aerosol-cloud interactions is in particular emphasized by studying key aerosol microphysical and optical properties as observed through their mixing in cloud-like environments.

54 ENVIRONMENTAL SCIENCES↗

Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery

Machine learning (ML)-accelerated discovery requires large amounts of high-fidelity data to reveal predictive structure–property relationships. For many properties of interest in materials discovery, the challenging nature and high cost of data generation has resulted in a data landscape that is both scarcely populated and of dubious quality. Data-driven techniques starting to overcome these limitations include the use of consensus across functionals in density functional theory, the development of new functionals or accelerated electronic structure theories, and the detection of where computationally demanding methods are most necessary. When properties cannot be reliably simulated, large experimental data sets can be used to train ML models. In the absence of manual curation, increasingly sophisticated natural language processing and automated image analysis are making it possible to learn structure–property relationships from the literature. Finally, models trained on these data sets will improve as they incorporate community feedback.

36 MATERIALS SCIENCE↗