Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Sparse Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Bayesian State-Space Modeling Framework for Understanding and Predicting Golden Eagle Movements Using Telemetry Data

Predicting raptor movements through a wind power plant under given atmospheric and topographical conditions is a crucial first step in the overall goal of quantifying the risk of turbine-related collisions and mortalities. Extracting behavioral traits of golden eagles (Aquila chrysaetos) from telemetry data requires the fusion of noisy and sparse movement data (location, heading, velocity) with a stochastic mathematical representation of the eagles' decision-making processes. In this study, we framed this problem in a Bayesian state-space framework where both observations and decision-making are assumed to be stochastic processes connected through hidden states (mode of flight, intent), and the unknown model parameters are assumed to be random variables that are calibrated using the available telemetry data. This framework allowed for rigorous consideration of underlying uncertainties while allowing for both data and prior biological knowledge to contribute to a probabilistic and predictive agent-based movement model. We implemented and applied the Bayesian framework to understand movement behavior of 23 GPS-tagged golden eagles travelling in the western US for years 2019 and 2020. Our preliminary findings show that the Bayesian state-space framework provides a robust inverse modeling apparatus to decode eagle behavioral characteristics from telemetry data. This study was primarily aimed at verifying and validating the framework with selected golden eagle tracks (both long- and short-ranged), with future research aimed at extending the framework to include multi-mode flight, consideration of atmospheric data and uplift mechanisms, eagle-to-eagle interaction, and eagle-to-turbine interaction.

Bayesian modeling↗

Machine learning-based inversion for acoustic impedance with large synthetic training data: Workflow and data characterization

Where wells are sparse or training data are difficult to label with high-quality wireline-derived impedance logs, machine learning (ML)-based inversion of acoustic impedance typically depends on small training data sets, leading to biased prediction. We have advanced a novel workflow that applies large synthetic seismic training data to reduce facies-related bias. Using a geologically realistic model as the truth model, we randomly select sparse seed wells to perform sequential Gaussian simulation (SGS) for impedance models of the same geometry and simulate facies variability. We implement random forest regression on 30 features extracted from the synthetic volume. We observe that more seed wells tend to reduce facies-induced bias by sampling more types of facies, resulting in a better prediction. We then focus on the responses of SGS models to facies changes, the number of seed wells necessary for a useful synthetic model, and how much a synthetic model can help ML-based inversion. Here, we observe that the SGS synthetic training model outperforms well-direct training in general. For modeled clastic shore-zone systems in Miocene Gulf of Mexico, two or more seed wells are necessary for a significant reduction of root-mean-square error and outliners, and improvement of facies imaging. In a field-data test, we apply a similar workflow to quantitatively predict acoustic impedance, which is then converted to a sand-volume map at a high-frequency sequence (10–100 m), revealing detailed facies and sandstone patterns. Such results are valuable in many geologic and engineering applications, such as hydrocarbon and CO 2 reservoir prospecting, reserve estimation, simulation, etc.

3D seismic↗

Estimation of hydraulic conductivity in a watershed using sparse multi-source data via Gaussian process regression and Bayesian experimental design

Enhanced water management systems depend on accurate estimation of subsurface hydraulic properties. However, geologic formations can vary significantly, so information from a single source (e.g., widely spaced boreholes) is insufficient in characterizing subsurface aquifer properties. Therefore, multiple sources of information are needed to complement the hydrogeology understanding of a region. Here, this study presents a numerical framework in which information from different measurement sources is combined to characterize the 3D random field in a multi-fidelity prediction model. Coupled with the model, a Bayesian experimental design was used to determine the best future sampling locations. The Upper Sangamon watershed in east-central Illinois was selected as the case study site, where the multi-fidelity Gaussian process model was used to estimate the hydraulic conductivity in the region of interest. Multi-source observation data were obtained from electrical resistivity and borehole pumping tests. The accuracy of the model prediction is dependent on the locations and the distribution of both high- and low-fidelity data. Furthermore, the multi-fidelity model was compared with the single-fidelity model. The uncertainties and confidence in the measurements and parameter estimates were quantified and used to design future cycles of data collection to further improve the confidence intervals.

54 ENVIRONMENTAL SCIENCES↗

Data-Driven Closures and Assimilation for Stiff Multiscale Random Dynamics

Here, we introduce a data-driven and physics-informed framework for propagating uncertainty in stiff, multiscale random ordinary differential equations (RODEs) driven by correlated (colored) noise. Unlike systems subjected to Gaussian white noise, a deterministic equation for the joint probability density function (PDF) of RODE state variables does not exist in closed form. Moreover, such an equation would require as many phase-space variables as there are states in the RODE system. To alleviate this curse of dimensionality, we instead derive exact, albeit unclosed, reduced-order PDF (RoPDF) equations for low-dimensional observables/quantities of interest. The unclosed terms take the form of state-dependent conditional expectations, which are directly estimated from data at sparse observation times. However, for systems exhibiting stiff, multiscale dynamics, data sparsity introduces regression discrepancies that compound during RoPDF evolution. This is overcome by introducing a kinetic-like defect term to the RoPDF equation, which is learned by assimilating in sparse, low-fidelity RoPDF estimates. Two assimilation methods are considered, namely nudging and deep neural networks, which are successfully tested against Monte Carlo simulations.

97 MATHEMATICS AND COMPUTING↗

GeoThermalCloud: Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale that has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify highvalue data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies.

58 GEOSCIENCES↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cloud Fusion of Big Data and Multi-Physics Models using Machine Learning for Discovery, Exploration, and Development of Hidden Geothermal Resources

The primary goals of this project are identifying hidden geothermal resources in the USA and designing profitable enhanced geothermal systems (EGS). Many non-obvious processes and parameters could characterize geothermal resources and could control the ultimate energy potential of geothermal fields. Diverse datasets (e.g., geology, geochemistry, geophysics, satellite, airborne geophysics) are available to help characterize geothermal resources, but this data is sparse and multi-scale. This has hindered attempts to leverage the datasets for geothermal exploration and profitable EGS design. Recent advancements in machine learning (ML) give promise to overcome these issues. Modern ML methods and tools can (1) analyze large datasets, (2) assimilate model ensembles that include a multitude of inputs and outputs, (3) process sparse datasets, (4) perform transfer learning between sites with different data quality, (5) extract hidden geothermal signatures from field and simulation data, (6) label geothermal resources and processes, (7) identify high-value data acquisition targets, and (8) guide geothermal exploration and production by selecting optimal exploration, production, and drilling strategies. In this work, we implement ML-based geothermal exploration and an enhanced geothermal systems (EGS) design tool to achieve the above goals. Our exploration tool is GeoThermalCloud (GTC) EGS design tool is GeoDT-ML. GTC (github.com/SmartTensors/GeoThermalCloud.jl) utilizes a LANL unsupervised ML platform called SmartTensors (https://tensors.lanl.gov/) to automate data analyses and interpretations by extracting hidden signatures to identify geothermal prospects. It enables the identification of critical measurements needed to identify geothermal resource signatures. GeoDT-ML (github.com/SmartTensors/GeoThermalCloud.jl/tree/master/) adds coupling to GeoDT (https://github.com/GeoDesignTool/GeoDT.git) for stochastic EGS design optimization and performance prediction. GeoDT-ML leverages recent advances in deep learning and high-performance computing. Contributors to this effort include LANL, PNNL, Google, Stanford, and Julia Computing.

15 GEOTHERMAL ENERGY↗

Applying 3D Geologic Modeling Workflows to the Argillite Reference Case (Rev. 1)

The objective of this short report is to document the application of our 3D geologic modeling workflow to an argillite (shale) host rock. Over the past four years, our team at Los Alamos National Laboratory has developed a geologic modeling workflow that can be applied to generic alluvial basins such as those found in the western United States. In “frontier” or “exploratory” basins where data are sparse, the first steps are to collect, evaluate and integrate available subsurface data into conceptual geologic models. Those models form the basis for constructing the geologic framework model, a 3D geocellular model ideally constrained by seismic and borehole data. To date we have constructed our models using “synthetic” well data derived from conceptual models, without the prospect of validating our workflow using “real” subsurface data. We were tasked to investigate whether our workflow designed for alluvial basin sediments could be applied to other potential repository host rocks. This task also provided the opportunity to work with high-quality subsurface data collected specifically for siting and evaluating a nuclear waste repository. Nagra, the Swiss governmental agency responsible for the disposal of the nation’s radioactive waste, generously provided us with data from two deep boreholes drilled through their argillaceous target formation. The aim of our proof-of-concept demonstration is to evaluate whether geostatistical methods offer a viable approach to property modeling in argillaceous rocks. Nagra provided us with the well data on the condition that we maintain confidentiality with all transferred information and results. Fortunately, Nagra posts numerous technical reports on its public website that describe the subsurface geology in great detail. All of the information and illustrations in this report related to the Swiss repository enterprise are taken from the Nagra public website.

58 GEOSCIENCES↗

Artificial Judgement Assistance from teXt (AJAX): Applying Open Domain Question Answering to Nuclear Non-proliferation Analysis

Nuclear non-proliferation analysis is complex and subjective, as the data is sparse, and examples are rare and diverse. While analysing non-proliferation data, it is often desired that the findings be completely auditable such that any claim or assertion can be sourced directly to the reference material from which it was derived. Currently this is accomplished by analysts thoroughly documenting underlying assumptions and clearly referencing details to source documents. This is a labour-intensive and time-consuming process that can be difficult to scale with geometrically increasing quantities of data. In this work, we describe an approach to leverage bi-directional language models for nuclear non-proliferation analysis. It has been shown recently that these models not only capture language syntax but also some of the relational knowledge present in the training data. We have devised a unique Salt and Pepper strategy for testing the knowledge present in the language models, while also introducing auditability function in our pipeline. We demonstrate that fine-tuning the bi-directional language models on domain specific corpus improves their ability to answer domain-specific factoid questions. Our hope is that the results presented in this paper will further the natural language processing (NLP) field by introducing the ability to audit the answers provided by the language models to bring forward the source of said knowledge.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Machine learning and shallow groundwater chemistry to identify geothermal prospects in the Great Basin, USA

This study discovers various geothermal prospects in the Great Basin, USA based on shallow groundwater chemical (geochemical) data. The geochemical data are expected to include hidden (latent) information that is a proxy for geothermal prospectivity. We processed the sparse geochemical data in the Great Basin at 14,341 locations including 18 attributes. Next, a non-negative matrix factorization with customized k-means clustering is applied to the geochemical data matrix that automatically finds three hidden geothermal signatures representing modestly, moderately, and highly confident geothermal prospects. The algorithm also evaluated the probability of occurrence of these types of resources through the studied region. There is a consistency between regional geothermal prospectivity as estimated by our ML methodology and the traditional play fairway analysis conducted over a portion of the study area. We also identify the dominant data attributes associated with each signature. Finally, our ML analyses allow us to reconstruct attributes from sparse into continuous over the study domain. The predicted continuous attributes can be used for future detailed geothermal explorations in the Great Basin.

15 GEOTHERMAL ENERGY↗

A Bayesian model for multivariate discrete data using spatial and expert information with application to inferring building attributes

When modeling sparsely observed multivariate data, strong prior information elicited from experts can be used to bolster predictive accuracy and counteract sampling bias. Similarly, modeling autocorrelation in space can help make use of co-occurrence patterns present in many types of spatial data. To make use of both expert prior information and spatial structure, we propose a novel graphical model for a spatial Bayesian network developed specifically to address challenges in inferring the attributes of buildings from geographically sparse observational data. This model is implemented as the sum of a spatial multivariate Gaussian random field and a tabular conditional probability function in real-valued space prior to projection onto the probability simplex. This modeling form is especially suitable for the usage of prior information in the form of sets of atomic rules obtained from experts. To perform inference with missing data, we implement a Markov chain Monte Carlo scheme composed of alternating steps of Gibbs sampling of missing entries and Hamiltonian Monte Carlo for model parameters. A case study in building attribution is presented to highlight the advantages and limitations of this approach.

97 MATHEMATICS AND COMPUTING↗

Physics-informed neural networks for identification of material properties using standing waves

A metallic structure in its initial stage of failure involves plastic deformation or environmental degradation that changes the elastic modulus and density. This work presents the detection of change in wave velocity (a function of elastic modulus and density) as a system identification problem. A physics-informed neural network (PINN) is proposed to solve the system identification problem. The PINN takes the spatial coordinates of scanning locations and time as inputs and provides the displacement and wave velocity as outputs. The governing partial differential equation of standing waves in a rod is incorporated into the neural network as physics in the form of a loss function. The wave velocity vector is randomly initiated. During the training of the network, physics is used to determine and update the wave velocity target vector from the network’s displacement predictions. The measured data, comprising sparse displacement response on the rod structure, are used to train the PINN. The wave velocity at the sparse locations on the rod is learned from the predicted displacements during the training. Using the predictions of the trained network, the response of free vibration or material property variation can be reconstructed at unscanned locations on the structure to obtain high-resolution maps for full-field imaging to detect and localize the changes caused by plastic deformation. The PINN’s sparse scanning and simultaneous prediction capability during training can lead to high scanning and data-processing speeds. This capability yields a nondestructive evaluation system that can predict the presence of degraded material locations as the structural vibrations are scanned and processed in real time.

Rathod, Vivek↗

Encoding nonlinear and unsteady aerodynamics of limit cycle oscillations using nonlinear sparse Bayesian learning

This article investigates the applicability of a recently proposed, nonlinear sparse Bayesian learning (NSBL) algorithm to identify and estimate the complex aerodynamics of limit cycle oscillations. NSBL provides a semi-analytical framework for determining the data-optimal sparse model nested within a (potentially) over-parameterized model. This is particularly relevant to nonlinear dynamical systems where modelling approaches involve the use of physics-based and data-driven components. In such cases, the data-driven components, where analytical descriptions of the physical processes are not readily available, are often prone to overfitting, meaning that the empirical aspects of these models will often involve the calibration of an unnecessarily large number of parameters. While an overparameterized model may fit the observed data well, such models may be inadequate for making predictions in regimes that are different from those wherein the data were recorded. In view of this, it is desirable to not only calibrate the model parameters, but also identify the optimal compromise between data fit and model complexity. In this article, we exhibit the optimal model discovery for an aeroelastic system wherein the structural dynamics are well-known and described by a differential equation model, coupled with a semi-empirical aerodynamic model for laminar separation flutter, resulting in low-amplitude limit cycle oscillations (LCO). To illustrate the performance of the algorithm, in this article, we use synthetic data and demonstrate the ability of the algorithm to correctly rediscover the optimal model and model parameters, given a known data-generating model. The synthetic data are generated from a forward simulation of a known differential equation model with parameters selected so as to mimic the dynamics observed in wind-tunnel experiments. Subsequently, we demonstrate the performance of the algorithm for model selection using noisy LCO data from wind tunnel experiments. As there is no ground truth available for the experimental data case, we provide a comparison between NSBL and Bayesian model selection to validate the results, and demonstrate the use of NSBL as an efficient alternative to traditional methods.

97 MATHEMATICS AND COMPUTING↗

Transformer-powered surrogates close the ICF simulation-experiment gap with extremely limited data

Abstract Recent advances in machine learning, specifically transformer architecture, have led to significant advancements in commercial domains. These powerful models have demonstrated superior capability to learn complex relationships and often generalize better to new data and problems. This paper presents a novel transformer-powered approach for enhancing prediction accuracy in multi-modal output scenarios, where sparse experimental data is supplemented with simulation data. The proposed approach integrates transformer-based architecture with a novel graph-based hyper-parameter optimization technique. The resulting system not only effectively reduces simulation bias, but also achieves superior prediction accuracy compared to the prior method. We demonstrate the efficacy of our approach on inertial confinement fusion experiments, where only 10 shots of real-world data are available, as well as synthetic versions of these experiments.

97 MATHEMATICS AND COMPUTING↗

Optimal sensor placement for reconstructing wind pressure field around buildings using compressed sensing

Deciding how to optimally deploy sensors in a large, complex, and spatially extended structure is critical to ensure that the surface pressure field is accurately captured for subsequent analysis and design. In some cases, reconstruction of missing data is required in downstream tasks such as the development of digital twins. Here, this paper presents a data-driven sparse sensor selection algorithm, aiming to provide the most information contents for reconstructing aerodynamic characteristics of wind pressures over tall building structures parsimoniously. The algorithm first fits a set of basis functions to the training data, then applies a computationally efficient QR algorithm that ranks existing pressure sensors in order of importance based on the state reconstruction to this tailored basis. The findings of this study show that the proposed algorithm successfully re- constructs the aerodynamic characteristics of tall buildings from sparse measurement locations, generating stable and optimal solutions across a range of conditions. As a result, this study serves as a promising first step toward leveraging the success of data-driven and machine learning algorithms to supplement traditional genetic algorithms currently used in wind engineering.

42 ENGINEERING↗

Stochastic AC optimal power flow: A data-driven approach

There is an emerging need for efficient solutions to stochastic AC Optimal Power Flow (AC-OPF) to ensure optimal and reliable grid operations in the presence of increasing demand and generation uncertainty. Herein this paper presents a highly scalable data-driven algorithm for stochastic AC-OPF that has extremely low sample requirement. The novelty behind the algorithm’s performance involves an iterative scenario design approach that merges information regarding constraint violations in the system with data-driven sparse regression. Compared to conventional methods with random scenario sampling, our approach is able to provide feasible operating points for realistic systems with much lower sample requirements. Furthermore, multiple sub-tasks in our approach can be easily paralleled and based on historical data to enhance its performance and application. We demonstrate the computational improvements of our approach through simulations on different test cases in the IEEE PES PGLib-OPF benchmark library.

42 ENGINEERING↗

Active and Transfer Learning of High-Dimensional Neural Network Potentials for Transition Metals

Classical molecular dynamics (MD) simulations represent a very popular and powerful tool for materials modeling and design. The predictive power of MD hinges on the ability of the interatomic potential to capture the underlying physics and chemistry. There have been decades of seminal work on developing interatomic potentials, albeit with a focus predominantly on capturing the properties of bulk materials. Such physics-based models, while extensively deployed for predicting the dynamics and properties of nanoscale systems over the past two decades, tend to perform poorly in predicting nanoscale potential energy surfaces (PESs) when compared to high-fidelity first-principles calculations. These limitations stem from the lack of flexibility in such models, which rely on a predefined functional form. Machine learning (ML) models and approaches have emerged as a viable alternative to capture the diverse size-dependent cluster geometries, nanoscale dynamics, and the complex nanoscale PESs, without sacrificing the bulk properties. Here, in this study, we introduce an ML workflow that combines transfer and active learning strategies to develop high-dimensional neural networks (NNs) for capturing the cluster and bulk properties for several different transition metals with applications in catalysis, microelectronics, and energy storage, to name a few. Our NN first learns the bulk PES from the high-quality physics-based models in literature and subsequently augments this learning via retraining with a higher-fidelity first-principles training data set to concurrently capture both the nanoscale and bulk PES. Our workflow departs from status-quo in its ability to learn from a sparsely sampled data set that nonetheless covers a diverse range of cluster configurations from near-equilibrium to highly nonequilibrium as well as learning strategies that iteratively improve the fingerprinting depending on model fidelity. All the developed models are rigorously tested against an extensive first-principles data set of energies and forces of cluster configurations as well as several properties of bulk configurations for 10 different transition metals. Our approach is material agnostic and provides a methodology to transfer and build upon the learnings from decades of seminal work in molecular simulations on to a new generation of ML-trained potentials to accelerate materials discovery and design.

36 MATERIALS SCIENCE↗