Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

GenAI4UQ: A software for forward and inverse uncertainty quantification using conditional generative AI

We introduce GenAI4UQ, a software package for forward and inverse uncertainty quantification in model calibration, parameter estimation, and ensemble forecasting. GenAI4UQ leverages a generative AI-based conditional modeling framework to address limitations of traditional inverse modeling techniques, such as Markov Chain Monte Carlo (MCMC) methods. By replacing computationally intensive iterative processes with a direct, learned mapping, GenAI4UQ enables efficient calibration of input parameters and generation of predictions directly from observations. The software supports rapid ensemble forecasting with robust uncertainty quantification while maintaining computational and storage efficiency. Built-in auto-tuning of hyperparameters simplifies model training, ensuring accessibility for users with varying expertise. Its versatile conditional generative framework is applicable across diverse scientific domains. While GenAI4UQ offers significant advantages in flexibility and efficiency, users should interpret its uncertainty estimates with caution in data-sparse scenarios, as the model may overestimate uncertainty—an effect common to all surrogate-based approaches including MCMC with surrogate models. Despite this, GenAI4UQ transforms inverse modeling by providing a fast, reliable, and user-friendly solution. It empowers researchers and practitioners to quickly estimate parameter distributions and generate model predictions for new observations, facilitating efficient decision-making and advancing the state of uncertainty quantification in computational modeling.

97 MATHEMATICS AND COMPUTING↗

Selecting Appropriate Model Complexity: An Example of Tracer Inversion for Thermal Prediction in Enhanced Geothermal Systems

Abstract A major challenge in the inversion of subsurface parameters is the ill‐posedness issue caused by the inherent subsurface complexities and the generally spatially sparse data. Appropriate simplifications of inversion models are thus necessary to make the inversion process tractable and meanwhile preserve the predictive ability of the inversion results. In this study, we investigate the effect of model complexity on fracture aperture inversion and thermal performance prediction in a field‐scale EGS model. Principal component analysis was used to map the aperture field to a low‐dimensional latent space. The complexity of the inversion model was quantitatively represented by the percentage of total variance in the original aperture fields preserved by the latent space. Tracer, pressure and flow rate data were used to invert for fracture aperture through an ensemble‐based inversion method, and the inferred aperture field was used to predict thermal performance. With an over‐simplified aperture model, ensemble collapse occurred. The inverted aperture models failed to resolve necessary flow and transport features, leading to a biased thermal performance prediction. A complex aperture model involved excessive features and was prone to overinterpreting the inversion data. Both the tracer/pressure/flow rate data reproduction and thermal prediction showed significant uncertainties, making it difficult to properly estimate long‐term thermal performance. Fortunately, our results indicate that there exists an appropriate model complexity which can simultaneously match inversion data and predict thermal performance with an acceptable uncertainty. The quality of the fit of tracer data appears to be a useful indicator of such an appropriate model complexity.

15 GEOTHERMAL ENERGY↗

MetaFlux: Meta-learning global carbon fluxes from sparse spatiotemporal observations

We provide a global, long-term carbon flux dataset of gross primary production and ecosystem respiration generated using meta-learning, called MetaFlux. The idea behind meta-learning stems from the need to learn efficiently given sparse data by learning how to learn broad features across tasks to better infer other poorly sampled ones. Using meta-trained ensemble of deep models, we generate global carbon products on daily and monthly timescales at a 0.25-degree spatial resolution from 2001 to 2021, through a combination of reanalysis and remote-sensing products. Site-level validation finds that MetaFlux ensembles have lower validation error by 5–7% compared to their non-meta-trained counterparts. In addition, they are more robust to extreme observations, with 4–24% lower errors. We also checked for seasonality, interannual variability, and correlation to solar-induced fluorescence of the upscaled product and found that MetaFlux outperformed other machine-learning based carbon product, especially in the tropics and semi-arids by 10–40%. Overall, MetaFlux can be used to study a wide range of biogeochemical processes.

54 ENVIRONMENTAL SCIENCES↗

Deep learning with mixup augmentation for improved pore detection during additive manufacturing

In additive manufacturing (AM), process defects such as keyhole pores are difficult to anticipate, affecting the quality and integrity of the AM-produced materials. Hence, considerable efforts have aimed to predict these process defects by training machine learning (ML) models using passive measurements such as acoustic emissions. This work considered a dataset in which keyhole pores of a laser powder bed fusion (LPBF) experiment were identified using X-ray radiography and then registered both in space and time to acoustic measurements recorded during the LPBF experiment. Due to AM’s intrinsic process controls, where a pore-forming event is relatively rare, the acoustic datasets collected during monitoring include more non-pores than pores. In other words, the dataset for ML model development is imbalanced. Moreover, this imbalanced and sparse data phenomenon remains ubiquitous across many AM monitoring schemes since training data is nontrivial to collect. Hence, we propose a machine learning approach to improve this dataset imbalance and enhance the prediction accuracy of pore-labeled data. Specifically, we investigate how data augmentation helps predict pores and non-pores better. This imbalance is improved using recent advances in data augmentation called Mixup, a weak-supervised learning method. Convolutional neural networks (CNNs) are trained on original and augmented datasets, and an appreciable increase in performance is reported when testing on five different experimental trials. When ML models are trained on original and augmented datasets, they achieve an accuracy of 95% and 99% on test datasets, respectively. We also provide information on how dataset size affects model performance. Lastly, we investigate the optimal Mixup parameters for augmentation in the context of CNN performance.

36 MATERIALS SCIENCE↗

Uncertainty Quantification for Multiphase Computational Fluid Dynamics Closure Relations with a Physics-Informed Bayesian Approach

Multiphase Computational Fluid Dynamics (MCFD) based on the two-fluid model is considered a promising tool to model complex two-phase flow systems. MCFD simulation can predict local flow features without resolving interfacial information. As a result, the MCFD solver relies on closure relations to describe the interaction between the two phases. Those empirical or semi-mechanistic closure relations constitute a major source of uncertainty for MCFD predictions. In this paper, we leverage a physics-informed uncertainty quantification (UQ) approach to inversely quantify the closure relations’ model form uncertainty in a physically consistent manner. This proposed approach considers the model form uncertainty terms as stochastic fields that are additive to the closure relation outputs. Combining dimensionality reduction and Gaussian processes, the posterior distribution of the stochastic fields can be effectively quantified within the Bayesian framework with the support of experimental measurements. As this UQ approach is fully integrated into the MCFD solving process, the physical constraints of the system can be naturally preserved in the UQ results. Here, in a case study of adiabatic bubbly flow, we demonstrate that this UQ approach can quantify the model form uncertainty of the MCFD interfacial force closure relations, thus effectively improving the simulation results with relatively sparse data support.

42 ENGINEERING↗

APSO-enhanced algebraic derivative estimation approach for real-time traffic flow prediction on critical road sections during wildfire evacuation

In rapid-onset disaster scenarios such as wildfires, evacuation traffic often significantly deviates from historical patterns, rendering conventional data-driven forecasting methods less effective. To address this challenge, we propose an improved algebraic derivative estimation (ADE) incorporating particle swarm optimization (PSO) for real-time traffic flow prediction. Our approach dynamically adjusts the ADE prediction time window at each step by minimizing a cost function based on the mean and variance of accumulated forecasting errors within the window, thereby balancing bias and variability. We evaluate the method using traffic data from the January 2025 California wildfires, focusing on key road segments critical for large-scale evacuations. The results demonstrate that our approach surpasses established machine learning and deep learning models—XGBoost, LSTM, and GRU—in predictive accuracy and maintains high computational efficiency. Notably, the proposed method eliminates the need for offline model training. Moreover, rapid PSO-based tuning enables real-time deployment, which provides a crucial advantage in scenarios where evacuation timings and road closures change dynamically. In conclusion, these findings highlight the benefits of the PSO-enhanced ADE framework for emergency traffic management, where rapid, data-sparse forecasts are essential for effective evacuation planning.

Algebraic derivative estimation↗

Increasing freshwater and dissolved organic carbon flows to Northwest Alaska’s Elson lagoon

Manifestations of climate change in the Arctic are numerous and include hydrological cycle intensification and permafrost thaw, both expected as a result of atmospheric and surface warming. Across the terrestrial Arctic dissolved organic carbon (DOC) entrained in arctic rivers may be providing a carbon subsidy to coastal food webs. Yet, data from field sampling is too often of limited duration to confidently ascertain impacts of climate change on freshwater and DOC flows to coastal waters. This study applies numerical modeling to investigate trends in freshwater and DOC exports from land to Elson Lagoon in Northwest Alaska over the period 1981-2020. While the modeling approach has limitations, the results point to significant increases in freshwater and DOC exports to the lagoon over the past four decades. The model simulation reveals significant increases in surface, subsurface (suprapermafrost), and total freshwater exports. Significant increases are also noted in surface and subsurface DOC production and export, influenced by warming soils and associated active-layer thickening. The largest changes in subsurface components are noted in September, which has experienced a ~50% increase in DOC export emanating from suprapermafrost flow. Direct coastal suprapermafrost freshwater and DOC exports in late summer more than doubled between the first and last five years of the simulation period, with a large anomaly in September 2019 representing a more than fourfold increase over September direct coastal export during the early 1980s. These trends highlight the need for dedicated measurement programs that will enable improved understanding of climate change impacts on coastal zone processes in this data sparse region of Northwest Alaska.

54 ENVIRONMENTAL SCIENCES↗

Deep learning based event reconstruction for cyclotron radiation emission spectroscopy

The objective of the cyclotron radiation emission spectroscopy (CRES) technology is to build precise particle energy spectra. This is achieved by identifying the start frequencies of charged particle trajectories which, when exposed to an external magnetic field, leave semi-linear profiles (called tracks) in the time–frequency plane. Due to the need for excellent instrumental energy resolution in application, highly efficient and accurate track reconstruction methods are desired. Deep learning convolutional neural networks (CNNs) - particularly suited to deal with information-sparse data and which offer precise foreground localization—may be utilized to extract track properties from measured CRES signals (called events) with relative computational ease. In this work, we develop a novel machine learning based model which operates a CNN and a support vector machine in tandem to perform this reconstruction. A primary application of our method is shown on simulated CRES signals which mimic those of the Project 8 experiment—a novel effort to extract the unknown absolute neutrino mass value from a precise measurement of tritium β - -decay energy spectrum. When compared to a point-clustering based technique used as a baseline, we show a relative gain of 24.1% in event reconstruction efficiency and comparable performance in accuracy of track parameter reconstruction.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Transformational Regional-Scale Earthquake Simulations with the DOE EarthQuake SIMulation Exascale Framework

Earthquakes present worldwide risk to economic and human safety. The 2023 earthquakes in Turkiye provided a reminder of the potential for catastrophic consequences with 50,700 deaths and 15.7 million people affected. The ability to predict ground motions and infrastructure damage for earthquakes continues to be a challenging problem for scientists and engineers. Until now, estimates of ground motions have been performed empirically by looking at sparse data from past earthquakes. This approach can provide statistical information on intensity amplitudes but cannot inform site-specific ground motions essential to developing the most effective resilience. Interest has grown in large-scale computational models to simulate earthquakes at regional scale. The U.S. Department of Energy EarthQuake SIMulation (EQSIM) framework was developed for regional-scale earthquake simulations at unprecedented fidelity, taking advantage of emerging GPU-accelerated systems. This article describes the EQSIM workflow and demonstrates regional-scale simulations with the new computational capability available to scientists in their quest to mitigate future disasters.

58 GEOSCIENCES↗

GraphSAGE-Sparse v.0.2.0

SAND2022-12899 O GraphSAGE-Sparse is an implementation of the GraphSAGE Graph Neural Network that adds support for sparse data structures, as well as improved functionality through the Tensorflow 2 functional API. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Matrix and Array (MATAR) Library

MATAR is a C++ software library to allow developers to easily create and use dense and sparse data representations that are also portable across disparate architectures using Kokkos.

Morgan, Nathaniel↗

Uncertainty Propagation from Flux to α and from Time to α in Reaction History

This report documents measurement uncertainty propagation from flux to α and from time to α in gamma reaction history. Analytical formulas are derived, and on the basis of them, their discrete versions are provided to serve for current reaction history software verification, and for development of the future versions of the software. The discrete version is provided for the general case: sparse data with varying time intervals between the data points.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Deep learning with mixup augmentation for improved pore detection during additive manufacturing

In additive manufacturing (AM), process defects such as keyhole pores are difficult to anticipate, affecting the quality and integrity of the AM-produced materials. Hence, considerable efforts have aimed to predict these process defects by training machine learning (ML) models using passive measurements such as acoustic emissions. This work considered a dataset in which keyhole pores of a laser powder bed fusion (LPBF) experiment were identified using X-ray radiography and then registered both in space and time to acoustic measurements recorded during the LPBF experiment. Due to AM’s intrinsic process controls, where a pore-forming event is relatively rare, the acoustic datasets collected during monitoring include more non-pores than pores. In other words, the dataset for ML model development is imbalanced. Moreover, this imbalanced and sparse data phenomenon remains ubiquitous across many AM monitoring schemes since training data is nontrivial to collect. Hence, we propose a machine learning approach to improve this dataset imbalance and enhance the prediction accuracy of pore-labeled data. Specifically, we investigate how data augmentation helps predict pores and non-pores better. This imbalance is improved using recent advances in data augmentation called Mixup, a weak-supervised learning method. Convolutional neural networks (CNNs) are trained on original and augmented datasets, and an appreciable increase in performance is reported when testing on five different experimental trials. When ML models are trained on original and augmented datasets, they achieve an accuracy of 95% and 99% on test datasets, respectively. We also provide information on how dataset size affects model performance. Lastly, we investigate the optimal Mixup parameters for augmentation in the context of CNN performance.

42 ENGINEERING↗

Uncertainty in Atlantic Multidecadal Oscillation derived from different observed datasets and their possible causes

As a leading mode of sea surface temperature (SST) variability over the North Atlantic in both observations and model simulations, the Atlantic Multidecadal Oscillation (AMO) can have a substantial influence on regional and global climate. By using Low-Frequency Component Analysis, we explore the uncertainties of the resulting AMO indices and the corresponding spatial patterns derived from three observational SST datasets. We found that the known coherent spatial pattern of the AMO at the basin scale over the North Atlantic appears in two out of the three datasets. Further analysis indicates that both the warming trend and the different techniques used to construct these observed gridded SSTs contribute to the AMO’s spatial coherence over the North Atlantic, especially during periods of sparse data sampling. The SST in the Extended Reconstructed SST dataset version 5 (ERSSTv5), changes from being systematically below the other datasets during the dense sampling periods on either side of the Second World War (WWII), to systematically above the other datasets during WWII, thereby introducing an artificial 10–20-year variability that affects the AMO’s spatial coherence. This coherence in the AMO’s spatial pattern is also affected by bias adjustment in ERSSTv5 at relative cool (i.e., non-summer) seasons, and by the heterogeneous North Atlantic warming pattern. The different AMO patterns can induce the different effects of wind, surface heat fluxes, and then drive ocean circulation and its heat transport convergence, especially for some seasons. For AMO indices, both the different detrending methods and different observational data result in uncertainty for the period 1935–1950. Such SST uncertainty is important to detect the relative role of the atmosphere and ocean in shaping the AMO.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Prior Concept Learning Improves Generalization From Few Examples in Computational Models of Human Object Recognition

Humans quickly and accurately learn new visual concepts from sparse data, sometimes just a single example. The impressive performance of artificial neural networks which hierarchically pool afferents across scales and positions suggests that the hierarchical organization of the human visual system is critical to its accuracy. These approaches, however, require magnitudes of order more examples than human learners. We used a benchmark deep learning model to show that the hierarchy can also be leveraged to vastly improve the speed of learning. We specifically show how previously learned but broadly tuned conceptual representations can be used to learn visual concepts from as few as two positive examples; reusing visual representations from earlier in the visual hierarchy, as in prior approaches, requires significantly more examples to perform comparably. These results suggest techniques for learning even more efficiently and provide a biologically plausible way to learn new visual concepts from few examples.

Rule, Joshua S.↗

Large spatiotemporal variability in aerosol properties over central Argentina during the CACTI field campaign

Abstract. Few field campaigns with extensive aerosol measurements have been conducted over continental areas in the Southern Hemisphere. To address this data gap and better understand the interactions of convective clouds and the surrounding environment, extensive in situ and remote sensing measurements were collected during the Cloud, Aerosol, and Complex Terrain Interactions (CACTI) field campaign conducted between October 2018 and April 2019 over the Sierras de Córdoba range of central Argentina. This study describes measurements of aerosol number, size, composition, mixing state, and cloud condensation nuclei (CCN) collected on the ground and from a research aircraft during 7 weeks of the campaign. Large spatial and multiday variations in aerosol number, size, composition, and CCN were observed due to transport from upwind sources controlled by mesoscale to synoptic-scale meteorological conditions. Large vertical wind shears, back trajectories, single-particle measurements, and chemical transport model predictions indicate that different types of emissions and source regions, including biogenic emissions and biomass burning from the Amazon and anthropogenic emissions from Chile and eastern Argentina, contribute to aerosols observed during CACTI. Repeated aircraft measurements near the boundary layer top reveal strong spatial and temporal variations in CCN and demonstrate that understanding the complex co-variability of aerosol properties and clouds is critical to quantify the impact of aerosol–cloud interactions. In addition to quantifying aerosol properties in this data-sparse region, these measurements will be valuable to evaluate predictions over the midlatitudes of South America and improve parameterized aerosol processes in local, regional, and global models.

54 ENVIRONMENTAL SCIENCES↗

Discrete-Direct Model Calibration and Uncertainty Propagation Method Confirmed on Multi-Parameter Plasticity Model Calibrated to Sparse Random Field Data

A discrete direct (DD) model calibration and uncertainty propagation approach is explained and demonstrated on a 4-parameter Johnson-Cook (J-C) strain-rate dependent material strength model for an aluminum alloy. The methodology's performance is characterized in many trials involving four random realizations of strain-rate dependent material-test data curves per trial, drawn from a large synthetic population. The J-C model is calibrated to particular combinations of the data curves to obtain calibration parameter sets which are then propagated to “Can Crush” structural model predictions to produce samples of predicted response variability. These are processed with appropriate sparse-sample uncertainty quantification (UQ) methods to estimate various statistics of response with an appropriate level of conservatism. This is tested on 16 output quantities (von Mises stresses and equivalent plastic strains) and it is shown that important statistics of the true variabilities of the 16 quantities are bounded with a high success rate that is reasonably predictable and controllable. The DD approach has several advantages over other calibration-UQ approaches like Bayesian inference for capturing and utilizing the information obtained from typically small numbers of replicate experiments in model calibration situations—especially when sparse replicate functional data are involved like force–displacement curves from material tests. The DD methodology is straightforward and efficient for calibration and propagation problems involving aleatory and epistemic uncertainties in calibration experiments, models, and procedures.

42 ENGINEERING↗