Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Three-dimensional reconstruction of x-ray emission volumes in magnetized liner inertial fusion from sparse projection data using a learned basis

The ability to visualize x-ray and neutron emission from fusion plasmas in 3D is critical to understand the origin of the complex shapes of the plasmas in experiments. Unfortunately, this remains challenging in experiments that study a fusion concept known as Magnetized Liner Inertial Fusion (MagLIF) due to a small number of available diagnostic views. Here, we present a basis function-expansion approach to reconstruct MagLIF stagnation plasmas from a sparse set of x-ray emission images. A set of natural basis functions is “learned” from training volumes containing quasi-helical structures whose projections are qualitatively similar to those observed in experimental images. Tests on several known volumes demonstrate that the learned basis outperforms both a cylindrical harmonic basis and a simple voxel basis with additional regularization, according to several metrics. Two-view reconstructions with the learned basis can estimate emission volumes to within 11% and those with three views recover morphology to a high degree of accuracy. The technique is applied to experimental data, producing the first 3D reconstruction of a MagLIF stagnation column from multiple views, providing additional indications of liner instabilities imprinting onto the emitting plasma.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Processing Aleatory and Epistemic Uncertainties in Experimental Data From Sparse Replicate Tests of Stochastic Systems for Real-Space Model Validation

This paper presents a practical methodology for propagating and processing uncertainties associated with random measurement and estimation errors (that vary from test-to-test) and systematic measurement and estimation errors (uncertain but similar from test-to-test) in inputs and outputs of replicate tests to characterize response variability of stochastically varying test units. Also treated are test condition control variability from test-to-test and sampling uncertainty due to limited numbers of replicate tests. These aleatory variabilities and epistemic uncertainties result in uncertainty on computed statistics of output response quantities. The methodology was developed in the context of processing experimental data for “real-space” (RS) model validation comparisons against model-predicted statistics and uncertainty thereof. The methodology is flexible and sufficient for many types of experimental and data uncertainty, offering the most extensive data uncertainty quantification (UQ) treatment of any model validation method the authors are aware of. It handles both interval and probabilistic uncertainty descriptions and can be performed with relatively little computational cost through use of simple and effective dimension- and order-adaptive polynomial response surfaces in a Monte Carlo (MC) uncertainty propagation approach. A key feature of the progressively upgraded response surfaces is that they enable estimation of propagation error contributed by the surrogate model. Sensitivity analysis of the relative contributions of the various uncertainty sources to the total uncertainty of statistical estimates is also presented. Finally, the methodologies are demonstrated on real experimental validation data involving all the mentioned sources and types of error and uncertainty in five replicate tests of pressure vessels heated and pressurized to failure. Simple spreadsheet procedures are used for all processing operations.

97 MATHEMATICS AND COMPUTING↗

Development of a Framework and Methodology for an Advanced Reactor Materials Environmental Effects Design Guide

Advanced non-light-water reactor components may operate at elevated temperature while experiencing cyclic loading, significant neutron irradiation, and exposure to reactor coolant. ASME Boiler and Pressure Vessel Code, Section III, Division 5, provides design rules for elevated-temperature service but does not include specific procedures to account for environmental effects on material properties. This report develops an initial framework and methodology for an Environmental Effects Design Guide (EEDG) focused on neutron irradiation; coolant-environment effects are reserved for future work. The proposed approach treats irradiation as a property-based overlay on the existing Division 5 design process, with two routes: a sparse-data route applying two reduction factors — FCR on creep-rupture strength and FF on fatigue life — for the creep-fatigue evaluations that typically control the design of advanced high-temperature reactor components, and a fuller framework developing the property-to-rule chain across the four Division 5 checks (primary load, strain limits and ratcheting, creep-fatigue, and buckling), together with swelling and weldments as scope items. Both routes are scoped by an in-pile qualification that restricts the use of post-irradiation-examination-derived properties in regimes where an in-pile mechanism could control the design outcome. Illustrative outputs derived on a compiled annealed Type 316 database — FCR ≈ 0.78–0.86 and FF ≈ 0.4 — demonstrate the calculation method within that specific dataset. The framework is an initial, testable design-rule concept; it identifies a practical path for preliminary design evaluations under sparse data and the material data and testing needed to develop the framework further.

Barua, Bipul (ORCID:0000000247184113)↗

National population mapping from sparse survey data: A hierarchical Bayesian modeling framework to account for uncertainty

Population estimates are critical for government services, development projects, and public health campaigns. Such data are typically obtained through a national population and housing census. However, population estimates can quickly become inaccurate in localized areas, particularly where migration or displacement has occurred. Some conflict-affected and resource-poor countries have not conducted a census in over 10 y. We developed a hierarchical Bayesian model to estimate population numbers in small areas based on enumeration data from sample areas and nationwide information about administrative boundaries, building locations, settlement types, and other factors related to population density. We demonstrated this model by estimating population sizes in every 10- m grid cell in Nigeria with national coverage. These gridded population estimates and areal population totals derived from them are accompanied by estimates of uncertainty based on Bayesian posterior probabilities. The model had an overall error rate of 67 people per hectare (mean of absolute residuals) or 43% (using scaled residuals) for predictions in out-of-sample survey areas (approximately 3 ha each), with increased precision expected for aggregated population totals in larger areas. This statistical approach represents a significant step toward estimating populations at high resolution with national coverage in the absence of a complete and recent census, while also providing reliable estimates of uncertainty to support informed decision making.

99 GENERAL AND MISCELLANEOUS↗

Rapid Analyses of Sparse Seismoacoustic Data Reveals the Timing and Size of the Accurate Energetic Systems Explosion

On 10 October 2025 an explosion occurred at a facility operated by Accurate Energetic Systems in Humphreys County, Tennessee. The incident resulted in 16 fatalities and created a debris field over several square kilometers. To address remaining questions about explosion timing and size, we collected about 20 seismic and 19 acoustic records of the blast from sensors up to hundreds of kilometers away. We then deployed 10 distinct physics-based, reduced order models (ROMs) that used validated geological structure and atmospheric conditions from the time of the event, along with observations of body- and surface-wave energy, as well as acoustic overpressure and phase duration. Each ROM predicted either timing, yield estimates, or both. We binned these estimates and their uncertainties according to each ROMs’ assumptions about confinement (aboveground, buried fully coupled, and buried partially coupled) and combined these estimates with other forensic data to conclude that the event occurred as a single, aboveground explosion on 10 December 2025 12:47:50.8 ±0.1 s with a yield equivalent to 11.8 [2.3,16.5] tons of Trinitrotoluene. Our estimates align with the Bureau of Alcohol, Tobacco, Firearms and Explosives inventory reports of 11–13 tons. This multimethod approach demonstrates the use of remotely observed geophysical data to rapidly aid conventional forensic investigations of accidental explosions.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Zero-Truncated Poisson Tensor Decomposition for Sparse Count Data

We propose a novel statistical inference paradigm for zero-inflated multiway count data that dispenses with the need to distinguish between true and false zero counts. Our approach ignores all zero entries and applies zero-truncated Poisson regression on the positive counts. Inference is accomplished via tensor completion that imposes low-rank structure on the Poisson parameter space. Our main result shows that an $\textit{N}$-way rank-R parametric tensor 𝓜 ϵ (0, ∞) $I$Χ∙∙∙Χ$I$ generating Poisson observations can be accurately estimated from approximately $IR^2 \text{log}^2_2(I)$ non-zero counts for a nonnegative canonical polyadic decomposition. Several numerical experiments are presented demonstrating that our zero-truncated paradigm is comparable to the ideal scenario where the locations of false zero counts are known $\textit{a priori}$.

97 MATHEMATICS AND COMPUTING↗

Bayesian Physics Informed Spatio-Temporal Network for Streamflow Data Imputation

Reliable reconstruction of incomplete streamflow records is critical for improving hydrological forecasting, flood preparedness, and water resource management. However, large observational gaps and uncertainties in governing physical parameters limit the accuracy of traditional statistical and machinelearning imputation frameworks. To address these challenges, we develop a Bayesian Physics-Informed Spatio-Temporal Network (BPI-STNet) that jointly captures spatial and temporal dependencies while enforcing hydrologic consistency through embedded physical constraints. The framework integrates a GraphSAGE-LSTM architecture to model spatial connectivity across gauges and temporal flow dynamics, coupled with a Bayesian update mechanism to estimate uncertain parameters in a simplified water-balance framework. Unlike conventional physics-informed networks that rely on sampling-based posterior estimation, BPI-STNet derives an analytic solution to the inverse problem, allowing closed-form Bayesian updates of uncertain parameters Λ={α,β,k} using Gaussian priors and likelihoods. Applied to daily observations from the Susquehanna River Basin (1980-2022), BPI-STNet achieves substantial improvements over a purely data-driven RGNN baseline, which reduced RMSE by 23 % and MAE by 9 %, and achieving an average NSE values up to 0.96. The results demonstrate that coupling Bayesian inference with physics-informed learning yields physically consistent, uncertainty-aware reconstructions that preserve the temporal persistence and statistical distribution of observed flows. The proposed framework establishes a generalizable paradigm for data-sparse hydrologic systems where both data fidelity and physical interpretability are essential.

Krishnan Kutty Ambika, Anukesh [ORNL] (ORCID:00000↗

Multiscale Modeling Meets Machine Learning: What Can We Learn?

Machine learning is increasingly recognized as a promising technology in the biological, biomedical, and behavioral sciences. There can be no argument that this technique is incredibly successful in image recognition with immediate applications in diagnostics including electrophysiology, radiology, or pathology, where we have access to massive amounts of annotated data. However, machine learning often performs poorly in prognosis, especially when dealing with sparse data. This is a field where classical physics-based simulation seems to remain irreplaceable. In this review, we identify areas in the biomedical sciences where machine learning and multiscale modeling can mutually benefit from one another: Machine learning can integrate physics-based knowledge in the form of governing equations, boundary conditions, or constraints to manage ill-posted problems and robustly handle sparse and noisy data; multiscale modeling can integrate machine learn- ing to create surrogate models, identify system dynamics and parameters, analyze sensitivities, and quantify uncertainty to bridge the scales and understand the emergence of function. With a view towards applications in the life sciences, we discuss the state of the art of combining machine learning and multiscale modeling, identify applications and opportunities, raise open questions, and address potential challenges and limitations. We anticipate that it will stimulate discussion within the community of computational mechanics and reach out to other disciplines including mathematics, statistics, computer science, artificial intelligence, biomedicine, systems biology, and precision medicine to join forces towards creating robust and efficient models for biological systems.

machine learning, multiscale modeling, physics-bas↗

Operator learning for predicting multiscale bubble growth dynamics

We report simulating and predicting multiscale problems that couple multiple physics and dynamics across many orders of spatiotemporal scales is a great challenge that has not been investigated systematically by deep neural networks (DNNs). Herein, we develop a framework based on operator regression, the so-called deep operator network (DeepONet), with the long-term objective to simplify multiscale modeling by avoiding the fragile and time-consuming “hand-shaking” interface algorithms for stitching together heterogeneous descriptions of multiscale phenomena. To this end, as a first step, we investigate if a DeepONet can learn the dynamics of different scale regimes, one at the deterministic macroscale and the other at the stochastic microscale regime with inherent thermal fluctuations. Specifically, we test the effectiveness and accuracy of the DeepONet in predicting multirate bubble growth dynamics, which is described by a Rayleigh–Plesset (R–P) equation at the macroscale and modeled as a stochastic nucleation and cavitation process at the microscale by dissipative particle dynamics (DPD). First, we generate data using the R–P equation for multirate bubble growth dynamics caused by randomly time-varying liquid pressures drawn from Gaussian random fields (GRFs). Our results show that properly trained DeepONets can accurately predict the macroscale bubble growth dynamics and can outperform long short-term memory networks. We also demonstrate that the DeepONet can extrapolate accurately outside the input distribution using only very few new measurements. Subsequently, we train the DeepONet with DPD data corresponding to stochastic bubble growth dynamics. Although the DPD data are noisy and we only collect sparse data points on the trajectories, the trained DeepONet model is able to predict accurately the mean bubble dynamics for time-varying GRF pressures. Taken together, our findings demonstrate that DeepONets can be employed to unify the macroscale and microscale models of the multirate bubble growth problem, hence providing new insight into the role of operator regression via DNNs in tackling realistic multiscale problems and in simplifying modeling with heterogeneous descriptions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ECLEIRS: Exact conservation law embedded identification of reduced states for parameterized nonlinear conservation laws from sparse and noisy data

Multi-query applications such as parameter estimation, uncertainty quantification and design optimization for parameterized partial differential equation (PDE) systems are expensive. While reduced/latent state dynamics approaches for parameterized PDEs offer a viable alternative, these approaches rely on high-quality data and struggle with highly sparse spatiotemporal noisy measurements typically obtained from experiments. Furthermore, there is no guarantee that these models satisfy governing physical conservation laws. In this article, we propose a reduced state dynamics approach, referred to as ECLEIRS, that embeds exact conservation in the solution and flux representation by utilizing a space-time divergence-free neural network formulation. We compare ECLEIRS with other reduced state dynamics approaches, those that do not enforce any physical constraints and those with physics-informed loss functions, for three shock-propagation problems: 1-D advection, 1-D Burgers and 2-D Euler equations. In conclusion, the numerical experiments conducted in this study demonstrate that ECLEIRS provides the most accurate prediction of dynamics for unseen parameters even in the presence of highly sparse and noisy data.

97 MATHEMATICS AND COMPUTING↗

Quantum block encoding for one-pair semiseparable matrices

Quantum block encoding (QBE) is a crucial step in the development of most quantum algorithms, as it provides an embedding of a given matrix into a suitable larger unitary matrix. Historically, the development of efficient techniques for QBE has mostly focused on sparse matrices; less effort has been devoted to data-sparse (e.g., rank-structured) matrices. In this work we examine a particular case of rank structure, namely, one-pair semiseparable matrices. We present a new block encoding approach that relies on a suitable factorization of the given matrix as the product of triangular and diagonal factors. To encode the matrix, the algorithm needs $2\log(N)+7$ ancillary qubits. Assuming that the data input oracles can be implemented with polylogarithmic depth, or that a QRAM input model is available, our proposed method requires $\mathcal{O}({\rm polylog} (N))$ time and has an error of $\mathcal{O}(N^2)$, where $N$ is the matrix size.

Antonioli, Giacomo [Pisa U.; CERN] (ORCID:00090000↗

Integrating Applied Energy and BER Smart Data Capabilities to Develop a DOE Data Fabric for Energy-Water R&D

Focal Area(s): 1) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization, unsupervised learning (including deep learning), and hardware-related efforts involving AI (e.g., edge computing). Science Challenge: DOE R&D, including DOE’s Basic Energy Research (BER)’s Environmental Systems Science Division (EESSD) program and DOE’s applied energy research (AER) programs (EERE, FE, and NE) are producers and consumers of Earth systems datasets. This white paper focuses on the first topic area from the call in relation to how crosscutting resources and innovations from DOE’s EESSD and AER can be brought to bear to mutual benefit and more efficient energy-water, Earth system data resources through improved. The overarching challenge posed by this call focuses on how DOE can directly leverage artificial intelligence (AI) to engineer a substantial (paradigm-changing) improvement in Earth System Predictability? While stemming from DOE BER’s EESSD program, this is a challenge that is faced and also being addressed by DOE’s AER programs. Over the past decade plus, FE, EERE, and NE programs have made important strides towards addressing this need. These strides are in many ways highly complementary to EESSD’s MODEX efforts. Energy water systems spanning metocean to groundwater to surface water systems all are data driven whether for basic energy or applied energy. These are remote, multi-variate, complex natural, and in many cases engineered, systems. Key needs and challenges of both EESSD and AER include developing data-focused tools to enhance data search and discovery to fill in knowledge gaps (address sparse data challenge), and rapidly transform datasets, including disparate and multi-source data. Leveraging DOE on-premise computing (HPC, exascale) infrastructure supports the computing-intensive algorithms required to execute these data acquisition and transformation processes to derive enriched knowledge and data, driving AI/ML and big data analytics for these systems. The opportunity lies in combining BER and AER efforts to provide a more robust, advanced, efficient and complete computing data fabric to address energy-water data acquisition and assimilation needs which currently pose significant impediments to AI/ML predictions and research.

54 ENVIRONMENTAL SCIENCES↗

Predicting fault slip via transfer learning

Abstract Data-driven machine-learning for predicting instantaneous and future fault-slip in laboratory experiments has recently progressed markedly, primarily due to large training data sets. In Earth however, earthquake interevent times range from 10’s-100’s of years and geophysical data typically exist for only a portion of an earthquake cycle. Sparse data presents a serious challenge to training machine learning models for predicting fault slip in Earth. Here we describe a transfer learning approach using numerical simulations to train a convolutional encoder-decoder that predicts fault-slip behavior in laboratory experiments. The model learns a mapping between acoustic emission and fault friction histories from numerical simulations, and generalizes to produce accurate predictions of laboratory fault friction. Notably, the predictions improve by further training the model latent space using only a portion of data from a single laboratory earthquake-cycle. The transfer learning results elucidate the potential of using models trained on numerical simulations and fine-tuned with small geophysical data sets for potential applications to faults in Earth.

58 GEOSCIENCES↗

Machine-Learning of Nonlocal Kernels for Anomalous Subsurface Transport from Breakthrough Curves

Anomalous behavior is ubiquitous in subsurface solute transport due to the presence of high degrees of heterogeneity at different scales in the media. Although fractional models have been extensively used to describe the anomalous transport in various subsurface applications, their application is hindered by computational challenges. Simpler nonlocal models characterized by integrable kernels and finite interaction length represent a computationally feasible alternative to fractional models; yet, the informed choice of their kernel functions still remains an open problem. We propose a general data-driven framework for the discovery of optimal kernels on the basis of very small and sparse data sets in the context of anomalous subsurface transport. Using spatially sparse breakthrough curves recovered from fine-scale particle-density simulations, we learn the best coarse-scale nonlocal model using a nonlocal operator regression technique. Predictions of the breakthrough curves obtained using the optimal nonlocal model show good agreement with fine-scale simulation results even at locations and time intervals different from the ones used to train the kernel, confirming the excellent generalization properties of the proposed algorithm. A comparison with trained classical models and with black-box deep neural networks confirms the superiority of the predictive capability of the proposed model.

97 MATHEMATICS AND COMPUTING↗