Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “sparse data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Informing forest carbon inventories under the Paris Agreement using ground-based forest monitoring data

Human interactions with forests have shaped Earth's climate for millennia and will continue to do so as we target net-zero emission goals. Accurately characterizing these climate impacts requires making reliable forest carbon data available for forest monitoring and planning. Here, we develop a semi-automated process for submitting forest carbon measurements from the largest relevant scientific database to the International Panel on Climate Change's Emission Factor Database, which currently has sparse forest carbon data. Building this bridge from scientific research to international policy is an important step towards managing forests in a net-zero motivated future. Humans have been influencing Earth's climate via transformative impacts on forests for millennia, and forests are now recognized as critical to climate change mitigation under the Paris Agreement. The efficacy of climate change mitigation planning and reporting depends on quality data on forest carbon (C) stocks and changes. The Emission Factor Database (EFDB) of the International Panel on Climate Change (IPCC) is intended to be a definitive source for such data, but needs comprehensive and well-documented data to be so. To facilitate submission of forest C estimates from scientific studies to EFDB, we develop and document a process for semi-automated data submission from the Global Forest C database (ForC v4.0), which is the largest compilation of ground-based forest C estimates. We then assess the data currently available through ForC and provide recommendations for improving forest data collection, analysis, and reporting. As of September 2024, ForC contained ~19,286 records potentially relevant to EFDB, 1068 of which had been submitted and posted to EFDB. These represented 19% of the total EFDB records for forest land. Records were unevenly distributed across variables and geographic regions. ForC records (37%) reviewed could not be submitted because the original publication lacked required information. In the future, ground-based forest C estimates should target gaps in the record, and studies should ensure that they report all information necessary for inclusion in EFDB. Given that climate change is rapidly impacting the world's forests, timely reporting of recent estimates will be critical to accurate forest C inventories.

54 ENVIRONMENTAL SCIENCES↗

Predicting Open Quantum Dynamics with Data-Informed Quantum-Classical Dynamics

We introduce a data-informed quantum-classical dynamics (DIQCD) approach for predicting the evolution of an open quantum system. The equation of motion in DIQCD is a Lindblad equation with a flexible, time-dependent Hamiltonian that can be optimized to fit sparse and noisy data from local observations of an extensive open quantum system. We demonstrate the accuracy and efficiency of DIQCD for both experimental and simulated quantum devices. We show that DIQCD can predict entanglement dynamics of ultracold molecules (calcium fluoride) in optical tweezer arrays. DIQCD also successfully predicts carrier mobility in organic semiconductors (rubrene) with accuracy comparable to nearly exact numerical methods.

Lindblad equation↗

Exploring the Whole Set of Accurate Sparse Interpretable Models

In data science applications, there are often many models that fit the data well. This phenomenon was called the Rashomon Effect by Leo Breiman. The set of good models is called the Rashomon Set, and the goal of this project is to locate, store, and study the Rashomon sets for classes of interpretable models, including decision trees and generalized additive models.

97 MATHEMATICS AND COMPUTING↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗

Grid Topology Discovery Algorithm Evaluation of Suitability for Utility Deployment (CRADA 606 Final Report)

This work presents the results of a field-informed demonstration aimed at evaluating the practical suitability of a topology discovery algorithm for utility environments. We demonstrated an algorithm that uses a graph-theory-informed state estimation approach for model selection. In collaboration with Survalent and Peninsula Light Co., the algorithm was applied to real feeder models and field measurements from supervisory control and data acquisition (SCADA) and advanced metering infrastructure (AMI) systems to identify the operational topology of a power distribution system. The demonstration assessed the algorithm’s performance under realistic data conditions, including sparse and noisy measurements, and examined its ability to identify the most likely network configurations. The results confirmed that the approach can effectively narrow down feasible topologies, providing operators with improved situational awareness of network status. Key lessons learned emphasize the need for systematic data validation and strategic sensor placement to enhance observability. These insights inform future deployment strategies and guide refinements for broader adoption in utility operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A ModEx Framework for Watershed Subsurface Investigation With Limited Geophysical Data Using Machine Learning and Hydrologic Modeling

Abstract Subsurface heterogeneity influences watershed hydrology strongly but remains difficult to characterize at catchment scales with sparse and costly field data. Geophysical surveys such as electromagnetic induction (EMI) provide local spatial subsurface images yet scaling them to watershed scales and converting EMI‐derived resistivity into hydraulic properties remains a challenge. We present a Model–Experiment (ModEx) framework that integrates limited EMI data with machine learning (ML) and hydrologic modeling to improve process representation and guide field investigations. Sparse EMI surveys were scaled to the catchment scale using a Random Forest model, and the resulting resistivity fields were combined with nearby borehole constraints to parameterize a hydrologic model. The EMI‐informed hydrological simulations improved predictions of streamflow sustained by subsurface flow and shallow saturation patterns. By combining EMI data and ML with hydrologic modeling, the ModEx framework guides future subsurface surveys, providing a transferable and efficient strategy for data–model integration across diverse watersheds. Plain Language Summary Mapping the underground network of soil and rock that controls water is essential for predicting floods and droughts, but seeing underground is difficult and expensive. We cannot drill everywhere, so scientists use geophysical tools to scan broad areas. There are two key challenges: these geophysical scans are often sparse across the whole watershed, and the geophysical data is hard to translate into water‐related properties. We used artificial intelligence to solve these problems. We taught a computer to find patterns linking the limited geophysical data to the land surface properties. This allowed it to fill in the gaps and create a complete, useful subsurface map for the entire watershed. This new map improves hydrologic simulations, leading to more accurate predictions of water movement in the watershed. It also helps scientists build better models with less data and generates a priority map showing where to measure next, making future investigations more efficient. Key Points Limited EMI scaled with ML improves catchment‐scale subsurface parameterization for hydrologic models The framework integrates hydrologic modeling with limited geophysical data to support subsurface investigation design ModEx framework offers a transferable data–model integration strategy that quantifies and reduces uncertainty guiding watershed studies

Chen, Hang↗

Machine Learning Eliminates Reanalysis Warm Bias and Reveals Weaker Winter Surface Cooling Over Arctic Sea Ice

The surface energy budget governs Arctic sea-ice growth/melt, yet observations are sparse, and reanalysis data sets suffer from systematic biases. Here, we train a neural network with observational data to bias-correct hourly ERA5 fluxes over Arctic ice-covered regions (≥70°N; sea-ice concentration >80%) for 1994–2024. Training data cover two full seasonal cycles and different sea-ice regimes. The neural network reduces RMSE for net shortwave radiation by ∼40%, downward longwave radiation by ∼16% and the total surface energy budget by ∼55%, eliminating the wintertime warm bias of ∼4 K in ERA5. Wintertime surface cooling is reduced by ∼50%, yielding thermodynamic ice-growth estimates of ∼80–120 cm, consistent with SMOS–CryoSat satellite thickness increases and in contrast to the 150–200 cm growth implied by ERA5. Our bias-corrected data capture the observed clear/cloudy states of the winter boundary layer and can be used to study Arctic climatology, evaluate climate models and drive sea-ice-ocean models.

Hossain, Akil [Alfred Wegener Institute for Polar ↗

Efficient learning of accurate surrogates for simulations of complex systems

Machine learning methods are increasingly deployed to construct surrogate models for complex physical systems at a reduced computational cost. However, the predictive capability of these surrogates degrades in the presence of noisy, sparse or dynamic data. Here, we introduce an online learning method empowered by optimizer-driven sampling that has two advantages over current approaches: it ensures that all local extrema (including endpoints) of the model response surface are included in the training data, and it employs a continuous validation and update process in which surrogates undergo retraining when their performance falls below a validity threshold. We find, using benchmark functions, that optimizer-directed sampling generally outperforms traditional sampling methods in terms of accuracy around local extrema even when the scoring metric is biased towards assessing overall accuracy. Finally, the application to dense nuclear matter demonstrates that highly accurate surrogates for a nuclear equation-of-state model can be reliably autogenerated from expensive calculations using few model evaluations.

79 ASTRONOMY AND ASTROPHYSICS↗

Bayesian Gaussian process inference for neutron spin echo measurement

Neutron spin echo (NSE) spectroscopy provides unique access to microscopic dynamics, but its application is often constrained by low neutron flux, long acquisition times, and significant noise. Here, we present a Bayesian inference approach based on Gaussian process regression (GPR) to reconstruct high-quality spin echo signals from sparse and noisy data by exploiting correlations in reciprocal space. Benchmarks on synthetic datasets and validation with experimental NSE measurements of dendrimers show that GPR suppresses noise, interpolates missing intensity values, and accommodates irregular observations. The method improves accuracy, shortens acquisition times, and enables high-throughput and real-time studies. Beyond NSE, the framework is broadly applicable to other low signal-to-noise ratio scattering techniques, thereby extending the scope of neutron spectroscopy.

Tung, Chi-Huan [Oak Ridge National Laboratory (ORN↗

Measurement of the mean excitation energy of liquid argon

The mean excitation energy (I-value) of liquid argon is a critical input for energy estimation in neutrino oscillation experiments. It is measured to be 199 +4 -5 eV using the range of 402.2 MeV protons from the Fermilab Linac. This compares to the author's recent evaluation of (197 ± 7) eV based on a combination of an oscillator strength distribution analysis, gaseous argon range measurements, sparse stopping power data on solid argon, and an extrapolation of data on the effect of phase from other substances. Using all sources of information, we recommend a value of (198 ± 4) eV for liquid argon, which is significantly higher than 188 eV, from ICRU-37's gaseous argon evaluation, commonly used in Monte Carlo codes such as Geant4.

47 OTHER INSTRUMENTATION↗

Evaluation of Hypotheses for the Fate of Historical Discharges from Hanford Site BC Cribs and Trenches

A preliminary review of declassified reports and historical documents from the 1950s and 1960s indicated they may contain valuable information regarding potential impacts to groundwater from liquid waste discharged to the 216-BC Cribs and Trenches during the late 1950s. Transport of contaminants in the vadose zone beneath the BC Cribs area is simulated by a vadose zone model developed for the Composite Analysis (DOE/RL-2019-52, Composite Analysis for Low-Level Waste Disposal in the Hanford Site Central Plateau ). Results from this model indicate there has been no substantial breakthrough of contaminants (particularly technetium-99) to groundwater from the vadose zone up to the present day. However, there is uncertainty in the anisotropy and other model input parameters as well as acknowledged sparseness of sampling data for this area. In this study, historical documents and other information were reviewed to evaluate the possibility that contaminant breakthrough to groundwater occurred during or shortly after operation of the BC Cribs sites, and whether refinements of the vadose zone model parametrization should be considered.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Generating synthetic signaling networks for in silico modeling studies

Predictive models of signaling pathways have proven to be difficult to develop. Reasons include the uncertainty in the number of species, the complexity in species’ interactions, and the sparseness and uncertainty in experimental data. Traditional approaches to developing mechanistic models rely on collecting experimental data and fitting a single model to that data. This approach works for simple systems but has proven unreliable for complex systems such as biological signaling networks. For example, uncertainty and sparseness of the data often result in overfitted models that have little predictive value beyond recapitulating the experimental data itself. Thus, there is a need to develop new approaches to create predictive mechanistic models of complex systems. However, to determine the effectiveness of any new algorithm, a baseline model is needed to test its performance. To meet this need, we developed a method for generating artificial synthetic networks that are reasonably realistic and thus can be treated as ground truth models. These synthetic models can then be used to generate synthetic data for developing and testing algorithms designed to recover the underlying network topology and associated parameters. Here, we describe a simple approach for generating synthetic signaling networks that can be used for this purpose.

42 ENGINEERING↗

Computing Sparse Tensor Decompositions via Chapel and C++/MPI Interoperability without Intermediate I/O

We extend an existing approach for efficient use of shared mapped memory across Chapel and C++ for graph data stored as 1-D arrays to sparse tensor data stored using a combination of 2-D and 1-D arrays. We describe the specific extensions that provide use of shared mapped memory tensor data for a particular C++ tensor decomposition tool called GentenMPI. We then demonstrate our approach on several real-world datasets, providing timing results that illustrate minimal overhead incurred using this approach. Finally, we extend our work to improve memory usage and provide convenient random access to sparse shared mapped memory tensor elements in Chapel, while still being capable of leveraging high performance implementations of tensor algorithms in C++.

97 MATHEMATICS AND COMPUTING↗

Entropy Stable Conservative Flux Form Neural Networks

We propose an entropy-stable conservative flux form neural network (CFN) to predict the dynamics of unknown governing conservation laws. The design of the network is based on the entropy-stable, second-order, and non-oscillatory Kurganov-Tadmor (KT) scheme. The proposed entropy-stable CFN, hereafter referred to as ESCFN, uses slope limiting as a denoising mechanism, ensuring accurate predictions in both noisy and sparse observation environments, as well as in both smooth and discontinuous regions. Importantly, our method is designed to predict long term dynamics of the unknown conservation law exclusively from a short temporal window of observed data, that is, without oracle knowledge of the PDE or later-time solution profiles. Numerical experiments demonstrate that the ESCFN achieves both stability and conservation while maintaining accuracy over extended time domains, and successfully predicts shock propagation speeds in long-term simulations. Furthermore, it is also robust to both noisy and sparse data environments.

Hyperbolic conservation laws↗

High-throughput spin-bath characterization of spin defects in semiconductors

Detailed knowledge of the local environments of spin defects in semiconductors, such as nitrogenvacancy (NV) centers in diamond or divacancies in silicon carbide, is crucial for optimizing control and entanglement protocols in quantum sensing and information applications. However, at present a direct experimental characterization of individual defect environments is not scalable, as conventional spin-bath measurements are time consuming and difficult to automate. Achieving high-throughput characterization requires short experiments to probe the spin bath. However, with fewer and noisier measurements, the inverse problem of recovering spin-bath properties from measured data becomes ill posed, with multiple spin baths having a high likelihood of yielding the same data. In this work, we present a set of computational tools to resolve the ill-posed inverse problem of recovering the atomic positions and hyperfine couplings of random nuclei surrounding spin defects from sparse, noisy experimental coherence data, which can be obtained in hours. Here, we use a trans-dimensional Bayesian approach that incorporates ab initio data to yield full posterior distributions over nuclear spin environments, enabling robust recovery from limited data. We also provide practical tools and guidelines to determine the limits of detectability for hyperfine couplings under specific dynamical decoupling sequences and sampling conditions. In addition, we demonstrate how the tools developed here, in combination with ab initio simulations of spin baths, can guide the design of efficient experimental protocols for application-specific high-throughput screening. To showcase the utility of our approach, we apply it to design fast dynamical decoupling experiments to characterize the spin baths often individual NV centers in diamond. While the primary focus is on accelerating spin-bath characterization of spin defects, this Bayesian approach also lays the foundation for digital-twin studies of spin defects, where a virtual model of the spin-defect system evolves in real time with ongoing experimental measurements. Together, the set of tools we designed and applied paves the way for scalable deployment of spin defects in semiconductors for quantum sensing and information applications.

Bayesian methods↗

G2Aero Database of Airfoils - Curated Airfoils

This dataset contains a curated set of 19,164 airfoil shapes from various applications and the data-driven design space of separable shape tensors (PGA space), which can be used as a parameter space for machine-learning applications focused on airfoil shapes. We constructed the airfoil dataset in two main stages. First, we identified 13 baseline airfoils from the NREL 5MW and IEA 15MW reference wind turbines. We reparameterized these shapes using least-squares fits of 8-order CST parametrizations, which involve 18 coefficients. By uniformly perturbing all 18 CST coefficients by +/-20% around each baseline airfoil, we generated 1,000 unique airfoils. Each airfoil was sampled with 1,001 shape landmarks whose x-coordinates followed a cosine distribution along the chord. This process resulted in a total of 13,000 airfoil shapes, each with 1,001 landmarks. In the second phase, we gathered additional airfoils from the extensive BigFoil database, which consolidates data from sources such as the University of Illinois Urbana-Champaign (UIUC) airfoil database, the JavaFoil database, the NACA-TR-824 database, and others. We undertook a thorough pre-processing step to filter out shapes with sparse, noisy, or incomplete data. We also removed airfoils with sharp leading edge and those exceeding our threshold for trailing edge thickness. Additionally, we thinned out the collection of NACA airfoils-- parametric sweeps of NACA airfoils with increasing thickness and camber present in BigFoil database-- by selecting every fourth step in the parameter sweeps. Finally, we regularized the airfoils by reparametrizing them with an 8-order CST parametrization (with 1,001 shape landmarks with x coordinated following cosine distribution along the chord) and removing airfoils with high reconstruction errors. This data pre-processing resulted in a set of 6,164 airfoils. In total, our curated airfoil dataset comprises 19,164 airfoils, each with 1,001 landmarks, and is stored in the curated_airfoils.npz file. Using this curated airfoil dataset, we utilized the separable shape tensors framework to develop a data-driven parameterization of airfoils based on principal geodesic analysis (PGA) of separable shape tensors. This PGA space is provided in PGAspace.npz file.

airfoils↗