Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments

There is an increasing aspiration to utilize machine learning (ML) for various tasks of relevance to national security. ML models have thus far been mostly applied to tasks and domains that, while impactful, have sufficient volume of data. For predictive tasks of national security relevance, ML models of great capacity (ability to approximate nonlinear trends in input-output maps) are often needed to capture the complex underlying physics. However, scientific problems of relevance to national security are often accompanied by various sources of sparse and/or incomplete data, including experiments and simulations, across different regimes of operation, of varying degrees of fidelity, and include noise with different characteristics and/or intensity. State-of-the-art ML models, despite exhibiting superior performance on the task and domain they were trained on, may suffer detrimental loss in performance in such sparse data environments. This report summarizes the results of the Laboratory Directed Research and Development project entitled Trust-Enhancing Probabilistic Transfer Learning for Sparse and Noisy Data Environments. The objective of the project was to develop a new transfer learning (TL) framework that aims to adaptively blend the data across different sources in tackling one task of interest, resulting in enhanced trustworthiness of ML models for mission- and safety-critical systems. The proposed framework determines when it is worth applying TL and how much knowledge is to be transferred, despite uncontrollable uncertainties. The framework accomplishes this by leveraging concepts and techniques from the fields of Bayesian inverse modeling and uncertainty quantification, relying on strong mathematical foundations of probability and measure theories to devise new uncertainty-aware TL workflows.

97 MATHEMATICS AND COMPUTING↗

The importance of cycle-by-cycle data in performing rapid battery technology development and validation

Lithium-ion battery (LiB) technology is playing a crucial role in transforming the predominantly fossil fuel-based transportation and stationary storage sectors to achieve a low-carbon economy. Rapid innovation in the LiB materials to electrode to cell design is happening to satisfy the performance, life, and safety metrics required by those myriads of applications. Lately, advanced analytics, such as machine-learning or artificial intelligence (ML/AI) techniques, are being used more frequently to aid in expedited LiB technology development, performance validation, and life prediction. The success of these techniques often relies on a large volume of well-defined and high-quality battery test data. On the other hand, most battery developers and research and development (R&D) communities are still following a classical approach to develop batteries, which is running calendar- and/or cycle-aging tests, performing reference performance tests (RPTs), and conducting post-mortem analyses periodically without paying attention to the wealth of data often not collected during the calendar or cycle life aging tests. This sparse data collection approach is time- and resource-intensive, requiring data capture and evaluation of months to years of RPT data to diagnose accurate battery state of performance, health, and safety. Even so, the underlying aging modes and mechanisms can be missed. If collected properly, battery test data during cycling or calendaring can be efficiently combined with ML/AI techniques to create powerful tools in the rapid diagnosis of battery state of performance, health, and safety along with insights into underlying aging modes and mechanisms. In this report, we discuss the importance of effective cycle-by-cycle (CBC) data collection with example case studies. Within a reasonable timeframe, RPT data are often inadequate in capturing many of the crucial battery aging dynamics, which often predominantly show up in CBC test data. Finally, we also show examples of ML/AI techniques that use CBC data in rapid diagnosis and projection of LiB state of health (SOH) to motivate the scientific community in collecting and using CBC data to facilitate expeditious technology development and validation.

25 ENERGY STORAGE↗

Stochastic modeling and statistical calibration with model error and scarce data

This paper introduces a procedure to assess the predictive accuracy of stochastic models subject to model error and sparse data. Model error is introduced as uncertainty on the coefficients of appropriate polynomial chaos expansions (PCE). The error associated with finite sample size allows us to conceive of these coefficients as statistics of the data that we describe as random variables whose influence on output quantities of interest is evaluated through the extended polynomial chaos expansion (EPCE). A Bayesian data assimilation scheme is introduced to update these expansions by considering the resulting nested chaos expansion as a hierarchical probabilistic model. Stochastic models of quantities of interest (QoI) are thus constructed and efficiently evaluated. Here, the Metropolis–Hastings Markov chain Monte Carlo procedure is used to sample the posterior. Two illustrative analytical and numerical problems are used to demonstrate the proposed approach.

Bayesian inference↗

Volumetric Rendering on Wavelet-Based Adaptive Grid

Numerical modeling of physical phenomena frequently involves processes across a wide range of spatial and temporal scales. In the last two decades, the advancements in wavelet-based numerical methodologies to solve partial differential equations, combined with the unique properties of wavelet analysis to resolve localized structures of the solution on dynamically adaptive computational meshes, make it feasible to perform large-scale numerical simulations of a variety of physical systems on a dynamically adaptive computational mesh that changes both in space and time. Volumetric visualization of the solution is an essential part of scientific computing, yet the existing volumetric visualization techniques do not take full advantage of multi-resolution wavelet analysis and are not fully tailored for visualization of a compressed solution on the wavelet-based adaptive computational mesh. Our objective is to explore the alternatives for the visualization of time-dependent data on space-time varying adaptive mesh using volume rendering while capitalizing on the available sparse data representation. Two alternative formulations are explored. The first one is based on volumetric ray casting of multi-scale datasets in wavelet space. Rather than working with the wavelets at the finest possible resolution, a partial inverse wavelet transform is performed as a preprocessing step to obtain scaling functions on a uniform grid at a user-prescribed resolution. As a result, a solution in physical space is represented by a superposition of scaling functions on a coarse regular grid and wavelets on an adaptive mesh. An efficient and accurate ray casting algorithm is based just on these coarse scaling functions. Additional details are added during the ray tracing by taking an appropriate number of wavelets into account based on support overlap with the interpolation point, wavelet coefficient magnitude, and other characteristics, such as opacity accumulation (front to back ordering) and deviation from frontal viewing direction. The second approach is based on complementing of wavelet-based adaptive mesh to the traditional Adaptive Mesh Refinement (AMR) mesh. Both algorithms are illustrated and compared to the existing volume visualization software for Rayleigh-Benard thermal convection and electron density data sets in terms of rendering time and visual quality for different data compression of both wavelet-based and AMR adaptive meshes.

Vezolainen, Alexei V.↗

Multi-head physics-informed neural networks for learning functional priors and uncertainty quantification

In numerous applications, the integration of prior knowledge and historical information is essential, particularly for tasks requiring the solution of ordinary or partial differential equations (ODEs/PDEs) in data-sparse or noisy environments. For instance, achieving accurate solutions to time-dependent PDEs with limited initial condition measurements necessitates an effective strategy for embedding prior knowledge. Hard-parameter sharing architectures in neural networks (NNs) have demonstrated success in both traditional and scientific machine learning domains, facilitating the learning of informative representations. Here, in this study, we introduce a novel, yet efficient, method to enhance physics-informed neural networks (PINNs) by incorporating a multi-head structure that enables the learning of functional priors from both empirical data and governing physical laws. This prior information can then be used to address data sparsity and high-level noise in solving ODE/PDE problems with uncertainty quantification (UQ). The approach, termed Multi-Head PINN (MH-PINN), consists of a shared body NN and multiple head NNs, each corresponding to an individual PINN instance. Our framework for functional prior learning is carried out in two stages: (1) training the MH-PINNs to develop a shared body NN alongside multiple head NNs, and (2) employing these trained head NNs to estimate a prior distribution through a normalizing flow-based density estimator. The learned functional prior can then be applied as a regularization mechanism in deterministic contexts or as an informative prior within a Bayesian inference framework, aiding in the resolution of subsequent ODE/PDE tasks. We evaluate the efficacy of MH-PINNs across five benchmark problems, including a high-dimensional parametric PDE, all characterized by data sparsity or substantial noise levels. Our findings reveal that MH-PINNs deliver accurate solutions and robust UQ, demonstrating adaptability across a range of complex and challenging scenarios.

Bayesian inference↗

Non-linear relationships between daily temperature extremes and US agricultural yields uncovered by global gridded meteorological datasets

Global agricultural commodity markets are highly integrated among major producers. Prices are driven by aggregate supply rather than what happens in individual countries in isolation. Furthermore, estimating the effects of weather-induced shocks on production, trade patterns and prices hence requires a globally representative weather data set. Recently, two data sets that provide daily or hourly records, GMFD and ERA5-Land, became available. Starting with the US, a data rich region, we formally test whether these global data sets are as good as more fine-scaled country-specific data in explaining yields and whether they estimate similar response functions. While GMFD and ERA5-Land have lower predictive skill for US corn and soybeans yields than the fine-scaled PRISM data, they still correctly uncover the underlying non-linear temperature relationship. All specifications using daily temperature extremes under any of the weather data sets outperform models that use a quadratic in average temperature. Correctly capturing the effect of daily extremes has a larger effect than the choice of weather data. In a second step, focusing on Sub Saharan Africa, a data sparse region, we confirm that GMFD and ERA5-Land have superior predictive power to CRU, a global weather data set previously employed for modeling climate effects in the region.

54 ENVIRONMENTAL SCIENCES↗

Suppressing simulation bias in multi-modal data using transfer learning

Abstract Many problems in science and engineering require making predictions based on few observations. To build a robust predictive model, these sparse data may need to be augmented with simulated data, especially when the design space is multi-dimensional. Simulations, however, often suffer from an inherent bias. Estimation of this bias may be poorly constrained not only because of data sparsity, but also because traditional predictive models fit only one type of observed outputs, such as scalars or images, instead of all available output data modalities, which might have been acquired and simulated at great cost. To break this limitation and open up the path for multi-modal calibration, we propose to combine a novel, transfer learning technique for suppressing the bias with recent developments in deep learning, which allow building predictive models with multi-modal outputs. First, we train an initial neural network model on simulated data to learn important correlations between different output modalities and between simulation inputs and outputs. Then, the model is partially retrained, or transfer learned, to fit the experiments; a method that has never been implemented in this type of architecture. Using fewer than 10 inertial confinement fusion experiments for training, transfer learning systematically improves the simulation predictions while a simple output calibration, which we design as a baseline, makes the predictions worse. We also offer extensive cross-validation with real and carefully designed synthetic data. The method described in this paper can be applied to a wide range of problems that require transferring knowledge from simulations to the domain of experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Markov chain Monte Carlo (MCMC) Bayesian inference approach to analyze apparent activation barriers and reaction orders from microreactor data

Statistical analysis of steady-state catalytic kinetic data is often limited by data sparsity due to the slow pace at which the data is collected. Data sparsity and limitations in statistical analysis make it difficult to differentiate between mechanistic models and catalytic sites. A Bayesian inference tool is reported for catalysis researchers to estimate error in the determination of reaction orders from steady state microreactor data. The benefits of a Bayesian inference approach are discussed, as an alternative to the more common frequentist approach. The approach incorporates prior knowledge of the system and the data collected to form an error estimate on reaction orders. We investigated the effects of three distinct data treatments—individual fitting of trials, pooled analysis, and constrained regression methods—on the precision and uncertainty of reaction order determinations. To assess the robustness of our findings, we conducted sensitivity analyses to evaluate the influence of Bayesian parameters on uncertainty estimation. Additionally, we utilized synthetic data to illustrate how data quality impacts the precision of uncertainty assessments. We show Bayesian analysis can obtain a more precise estimation of error with a sparse data set than a frequentist analysis. Finally, this work provides strong evidence that the adoption of Bayesian analysis of kinetic data may help researchers make more precise arguments as to the strength of their evidence for a particular mechanistic hypothesis, or in comparing across different catalysts.

42 ENGINEERING↗

A scalable framework for efficient coupling of thermal and microstructural simulations in additive manufacturing

Predicting microstructure evolution in metal additive manufacturing (AM) is important for process optimization, but spatiotemporal scale disparities between thermal transport and microstructure evolution create significant challenges for efficient data transfer between simulation codes. To address this, we present Stork, a scalable framework for coupling thermal and microstructural simulations. Stork uses a sparse data representation to identify and store active solidification sub-volumes, enabling highly parallel quad-linear interpolation from coarse thermal grids to fine microstructure grids without large intermediate storage. We demonstrate the framework by coupling the semi-analytic heat transfer code 3DThesis with the time-parallel cellular automata code Toucan. This approach achieves over two orders of magnitude reduction in data generation time and file size compared to prior workflows. Numerical studies show that quad-linear interpolation preserves grain morphology and crystallographic texture in laser powder bed fusion (LPBF) simulations for coarsening ratios up to 16. Overall, Stork provides a scalable pathway for high-throughput, component-scale AM simulations on modern high-performance computing systems.

36 MATERIALS SCIENCE↗

Automated identification of local contamination in remote atmospheric composition time series

Abstract. Atmospheric observations in remote locations offer a possibility of exploring trace gas and particle concentrations in pristine environments. However, data from remote areas are often contaminated by pollution from local sources. Detecting this contamination is thus a central and frequently encountered issue. Consequently, many different methods exist today to identify local contamination in atmospheric composition measurement time series, but no single method has been widely accepted. In this study, we present a new method to identify primary pollution in remote atmospheric datasets, e.g., from ship campaigns or stations with a low background signal compared to the contaminated signal. The pollution detection algorithm (PDA) identifies and flags periods of polluted data in five steps. The first and most important step identifies polluted periods based on the derivative (time derivative) of a concentration over time. If this derivative exceeds a given threshold, data are flagged as polluted. Further pollution identification steps are a simple concentration threshold filter, a neighboring points filter (optional), a median, and a sparse data filter (optional). The PDA only relies on the target dataset itself and is independent of ancillary datasets such as meteorological variables. All parameters of each step are adjustable so that the PDA can be “tuned” to be more or less stringent (e.g., flag more or fewer data points as contaminated). The PDA was developed and tested with a particle number concentration dataset collected during the Multidisciplinary drifting Observatory for the Study of Arctic Climate (MOSAiC) expedition in the central Arctic. Using strict settings, we identified 62 % of the data as influenced by local contamination. Using a second independent particle number concentration dataset also collected during MOSAiC, we evaluated the performance of the PDA against the same dataset cleaned by visual inspection. The two methods agreed in 94 % of the cases. Additionally, the PDA was successfully applied to a trace gas dataset (CO2), also collected during MOSAiC, and to another particle number concentration dataset, collected at the high-altitude background station Jungfraujoch, Switzerland. Thus, the PDA proves to be a useful and flexible tool to identify periods affected by local contamination in atmospheric composition datasets without the need for ancillary measurements. It is best applied to data representing primary pollution. The user-friendly and open-access code enables reproducible application to a wide suite of different datasets. It is available at https://doi.org/10.5281/zenodo.5761101 (Beck et al., 2021).

54 ENVIRONMENTAL SCIENCES↗

Frictionless knowledge injection for few-shot learning

Cutting-edge machine learning methods often require large volumes of curated training data, precluding their use in national security problems with rare events in massive datasets. We present a method for incorporating abstract knowledge into models tailored for sparse data. A subject matter expert defines salient concepts using data examples, which are encoded in the model’s embedding space. Models are then trained to respect these concepts. This method enables knowledge injection, yielding effective models with limited labeled data and the ability to assess model sensitivity for subject matter expertise across the nonproliferation mission space, as demonstrated with Raman spectra analysis.

Stomps, Jordan [ORNL] (ORCID:0000000178114479)↗

Multi-year assessment of the impact of ship-borne radiosonde observations on polar WRF forecasts in the Arctic

Abstract To compensate for the lack of conventional observations over the Arctic Ocean, ship-borne radiosonde observations have been regularly carried out during summer Arctic expeditions and the observed data have been broadcast via the global telecommunication system since 2017. With these data obtained over the data-sparse Arctic Ocean, observing system experiments were carried out using a polar-optimized version of the Weather Research and Forecasting (WRF) model and the WRF Data Assimilation (WRFDA) system to investigate their effects on analyses and forecasts over the Arctic. The results of verification against reanalysis data reveal: (1) DA effects on analyses and forecasts; (2) the reason for the year-to-year variability of DA effects; and (3) the possible role of upper-level potential vorticity in delayed DA effects. The overall assimilation effects of the extra data on the analyses and forecasts over the Arctic are positive. Initially, the DA effects are the most apparent in the temperature variables in the middle/lower troposphere, which spread to the wind variables in the upper troposphere. The effects decrease with time but reappear after approximately 120 h, even in the 240-h forecasts. The effects on forecasts vary depending on the proximity of the radiosonde observation locations to the high synoptic variability. The upper-level potential vorticity is known to play an important role in the development of Arctic cyclones, and it is suggested as a possible explanation for the delayed DA effects after about 120 h.

Geology↗

Inference of bipolar neutrino flavor oscillations near a core-collapse supernova based on multiple measurements at Earth

Neutrinos in compact-object environments, such as core-collapse supernovae, can experience various kinds of collective effects in flavor space, engendered by neutrino-neutrino interactions. These include “bipolar” collective oscillations, which are exhibited by neutrino ensembles where different flavors dominate at different energies. Considering the importance of neutrinos in the dynamics and nucleosynthesis in these environments, it is desirable to ascertain whether an Earth-based detection could contain signatures of bipolar oscillations that occurred within a supernova envelope. To that end, we, in this study, continue examining a cost-function formulation of statistical data assimilation (SDA) to infer solutions to a small-scale model of neutrino flavor transformation. SDA is an inference paradigm designed to optimize a model with sparse data. Our model consists of two monoenergetic neutrino beams with different energies emanating from a source and coherently interacting with each other and with a matter background, with radially varying interaction strengths. We attempt to infer flavor transformation histories of these beams using simulated measurements of the flavor content at locations “in vacuum” (that is, far from the source), which could in principle correspond to Earth-based detectors. Within the scope of this small-scale model, we found that: (i) based on such measurements, the SDA procedure is able to infer whether bipolar oscillations had occurred within the protoneutron star envelope, and (ii) if the measurements sample the full amplitude of the neutrino oscillations in vacuum, then the amplitude of the prior bipolar oscillations is well predicted. This result intimates that the inference paradigm can well complement numerical integration codes, via its ability to infer flavor evolution at physically inaccessible locations.

79 ASTRONOMY AND ASTROPHYSICS↗

Towards A Geo-Data Science Method for Assessing Rare Earth Element and Critical Mineral Occurrences in Coal and Other Sedimentary Systems

While preliminary analyses of data from open-source resources (Ekmann, 2012) show promising concentrations of REE in individual coal samples from a number of sites and basins in the U.S., other sparse data for REE in domestic coal-related strata suggest that many occurrences are low, “subeconomic” concentrations. At present, there is no method for systematically assessing potential sedimentary occurrences of REE. However, the geologic processes responsible for REE occurrences in coal-related strata are systematic; the unpredictability of REE resources in coal-related strata is due to poorly quantified spatial resource trends and the lack of an exploration method tailored to these resources. Thus, there exists a need for a systematic assessment approach that incorporates knowledge of geological variation in the mechanisms of REE enrichment within coal basins to help minimize geologic uncertainty and reduce commercial exploration risk.

01 COAL, LIGNITE, AND PEAT↗

Generation of Data-Driven Expected Energy Models for Photovoltaic Systems

Although unique expected energy models can be generated for a given photovoltaic (PV) site, a standardized model is also needed to facilitate performance comparisons across fleets. Current standardized expected energy models for PV work well with sparse data, but they have demonstrated significant over-estimations, which impacts accurate diagnoses of field operations and maintenance issues. This research addresses this issue by using machine learning to develop a data-driven expected energy model that can more accurately generate inferences for energy production of PV systems. Irradiance and system capacity information was used from 172 sites across the United States to train a series of models using Lasso linear regression. The trained models generally perform better than the commonly used expected energy model from international standard (IEC 61724-1), with the two highest performing models ranging in model complexity from a third-order polynomial with 10 parameters (Radj2 = 0.994) to a simpler, second-order polynomial with 4 parameters (Radj2=0.993), the latter of which is subject to further evaluation. Subsequently, the trained models provide a more robust basis for identifying potential energy anomalies for operations and maintenance activities as well as informing planning-related financial assessments. We conclude with directions for future research, such as using splines to improve model continuity and better capture systems with low (≤1000 kW DC) capacity.

14 SOLAR ENERGY↗

Model Choice Metrics to Optimize Profile-QSAR Performance

Predicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. Here, this method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the model’s ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement.

Biological and medical sciences, Computer science↗

Reply to Comment by Peterie Et Al. on “Accelerated Fill‐Up of the Arbuckle Group Aquifer and Links to U.S. Midcontinent Seismicity”

Abstract Peterie et al. question one observation in our paper: associating pressure increases to injection volumes at distances of up to 25 km from an injection well. In this reply, we show that the comment misunderstands our analysis and the evidence that led to this conclusion. We also show that gauge‐depth‐corrected pressures, used by the authors to produce statewide pressure maps, are discrepant with the static fluid level data, provided in our original compilation and analysis. The discrepancies are a result of the pressure correction method employed, which naïvely substitutes formation pressure for bottomhole pressure to calculate wellbore fluid density. Their linearly interpolated pressure maps, based on sparse data, contain interpolation and extrapolation artifacts that contradict injection trends in the state, the Theis solution, and the superposition principle. We reiterate that pressure and static fluid level increases in Class I wells existed prior to 2013, most notably in central Kansas, where recent earthquakes are cited in the comment as evidence of a pressure plume emanating from the Kansas‐Oklahoma border, 90 km away. We show that the space‐time pattern of seismicity in this area is inconsistent with a northward propagating pressure plume and, instead, seismicity appears to be centered on and near a cluster of high‐rate injection wells, two of which are among the highest rate wells in the state. These observations, along with recent M4 + earthquakes during continued decreases in wastewater injection in southern Kansas and northern Oklahoma, question the usefulness of the comment for understanding and managing societally significant earthquakes.

Ansari, Esmail↗

Solving Inverse Stochastic Problems from Discrete Particle Observations Using the Fokker--Planck Equation and Physics-Informed Neural Networks

The Fokker--Planck (FP) equation governing the evolution of the probability density function (PDF) is applicable to many disciplines, but it requires specification of the coefficients for each case, which can be functions of space-time and not just constants and hence require the development of a data-driven modeling approach. When the data available is directly on the PDF, there exist methods for inverse problems that can be employed to infer the coefficients and thus determine the FP equation and subsequently obtain its solution. Herein, we address a more realistic scenario, where only sparse data are given on the particles' positions at a few time instants, which are not sufficient to accurately construct directly the PDF even at those times from existing methods, e.g., kernel estimation algorithms. To this end, we develop a general framework based on physics-informed neural networks (PINNs) that introduces a new loss function using the Kullback--Leibler divergence to connect the stochastic samples with the FP equation to simultaneously learn the equation and infer the multidimensional PDF at all times. In particular, we consider two types of inverse problems, type I, where the FP equation is known but the initial PDF is unknown, and type II, in which, in addition to the unknown initial PDF, the drift and diffusion terms are also unknown. In both cases, we investigate problems with either Brownian or Lévy noise or a combination of both. Here, we demonstrate the new PINN framework in detail in the one-dimensional (1D) case, but we also provide results for up to five dimensions demonstrating that we can infer both the FP equation and dynamics simultaneously at all times with high accuracy using only very few discrete observations of the particles.

97 MATHEMATICS AND COMPUTING↗