Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Loosely Conditioned Emulation of Global Climate Models With Generative Adversarial Networks

Climate models encapsulate our best understanding of the Earth system, allowing research to be conducted on its future under alternative assumptions of how human-driven climate forces are going to evolve. An important application of climate models is to provide metrics of mean and extreme climate changes, particularly under these alternative future scenarios, as these quantities drive the impacts of climate on society and natural systems. Because of the need to explore a wide range of alternative scenarios and other sources of uncertainties in a computationally efficient manner, climate models can only take us so far, as they require significant computational resources, especially when attempting to characterize extreme events, which are rare and thus demand long and numerous simulations in order to accurately represent their changing statistics. Here we use deep learning in a proof of concept that lays the foundation for emulating global climate model output for different scenarios. We train two "loosely conditioned" Generative Adversarial Networks (GANs) that emulate daily precipitation output from a fully coupled Earth system model: one GAN modeling Fall-Winter behavior and the other Spring-Summer. Our GANs are trained to produce spatiotemporal samples: 32 days of precipitation over a 64x128 regular grid discretizing the globe. We evaluate the generator with a set of related performance metrics based upon KL divergence, and find the generated samples to be nearly as well matched to the test data as the validation data is to test. We also find the generated samples to accurately estimate the mean number of dry days and mean longest dry spell in the 32 day samples. Our trained GANs can rapidly generate numerous realizations at a vastly reduced computational expense, compared to large ensembles of climate models, which greatly aids in estimating the statistics of extreme events.

climate emulation, extreme climate, impacts, machi↗

Data augmentation for disruption prediction via robust surrogate models

The goal of this work is to generate large statistically representative data sets to train machine learning models for disruption prediction provided by data from few existing discharges. Such a comprehensive training database is important to achieve satisfying and reliable prediction results in artificial neural network classifiers. Here, we aim for a robust augmentation of the training database for multivariate time series data using Student t process regression. We apply Student t process regression in a state space formulation via Bayesian filtering to tackle challenges imposed by outliers and noise in the training data set and to reduce the computational complexity. Thus, the method can also be used if the time resolution is high. We use an uncorrelated model for each dimension and impose correlations afterwards via colouring transformations. We demonstrate the efficacy of our approach on plasma diagnostics data of three different disruption classes from the DIII-D tokamak. To evaluate if the distribution of the generated data is similar to the training data, we additionally perform statistical analyses using methods from time series analysis, descriptive statistics and classic machine learning clustering algorithms.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Interpretable Machine Learning for Molecular Biosignatures: a Novel Single-Sample Feature Importance Method That Is Sensitive To Statistical Interactions

Isotope ratio mass spectrometry (IRMS) of volatiles (e.g., CO 2 ) promises to be a powerful tool for potential biosignature detection for future missions to ocean worlds (OW) such as Europa and Enceladus. Machine learning (ML) methods for IRMS data could enable science autonomy by onboard prediction of seawater chemistry and biosignature presence. However, ML models are likely to be complex and involve statistical interactions between features (variables), which can make predictions seem opaque and enigmatic. For ML predictions as significant as extraterrestrial biosignatures, we must place extraordinary confidence in models. It is therefore essential that these models make interpretable predictions (i.e., human-understandable) and include false-prediction diagnostics. We achieve high accuracy and interpretability in ML biosignature and seawater chemistry models for OW through a nearest-neighbors feature selection tool that detects statistical interactions between predictors, constructs interaction networks for visualization of selected features working together to make a prediction, and reports single-sample feature importance scores for false-detection diagnostics. Here we develop a novel single-sample nearest-neighbors projected distance regression(ssNPDR) feature selection method that improves upon existing single-sample algorithms through the inclusion of statistical interactions while providing false-prediction diagnostics for ML models.

geochemistry↗

Predictive Battery Lifetime Modeling at NREL [Slides]

Battery lifetime models are used to extrapolate data from accelerated aging tests to simulate degradation in real-world applications such as electric vehicles and battery energy storage systems. Methods developed at NREL utilize both expert domain-knowledge and machine-learning to identify models, using statistical methods such as cross-validation and bootstrap resampling to interrogate model performance and quantify uncertainty. These models can be utilized in systems level simulations to predict battery performance or technoeconomic models to estimate the lifetime cost of battery systems.

25 ENERGY STORAGE↗

Accelerating phase-field-based microstructure evolution predictions via surrogate models trained by machine learning methods

Abstract The phase-field method is a powerful and versatile computational approach for modeling the evolution of microstructures and associated properties for a wide variety of physical, chemical, and biological systems. However, existing high-fidelity phase-field models are inherently computationally expensive, requiring high-performance computing resources and sophisticated numerical integration schemes to achieve a useful degree of accuracy. In this paper, we present a computationally inexpensive, accurate, data-driven surrogate model that directly learns the microstructural evolution of targeted systems by combining phase-field and history-dependent machine-learning techniques. We integrate a statistically representative, low-dimensional description of the microstructure, obtained directly from phase-field simulations, with either a time-series multivariate adaptive regression splines autoregressive algorithm or a long short-term memory neural network. The neural-network-trained surrogate model shows the best performance and accurately predicts the nonlinear microstructure evolution of a two-phase mixture during spinodal decomposition in seconds, without the need for “on-the-fly” solutions of the phase-field equations of motion. We also show that the predictions from our machine-learned surrogate model can be fed directly as an input into a classical high-fidelity phase-field model in order to accelerate the high-fidelity phase-field simulations by leaping in time. Such machine-learned phase-field framework opens a promising path forward to use accelerated phase-field simulations for discovering, understanding, and predicting processing–microstructure–performance relationships.

36 MATERIALS SCIENCE↗

ORION: Operational FoRecastIng Of INduced Seismicity

The Operational Forecasting of Induced Seismicity toolkit “ORION” (ORION) is an open-source, observation-based ensemble forecasting toolkit which is geared towards helping operators understand the seismic hazard (i.e., probabilistic assessment of the magnitude and frequency of induced seismic events) at a site. ORION analyzes how the seismic hazard evolves during injection and suggests possible mitigation strategies to employ if an earthquake that exceeds certain threshold is observed. Through its ensemble modeling approach, ORION leverages the benefits of statistical-, physics-, and machine learning-based forecasting methodologies, while reducing the impact of each model’s respective limitations. The ORION toolkit consists of an easy-to-use GUI interface that affords a user as much or as little interaction as desired. Advanced capabilities allow the user to upload local, high-precision earthquake catalogs, projected injection profiles and/or spatiotemporal estimates of pressure/stress, and to tune various model parameters. ORION will then provide a spatial and temporal ensemble forecast of seismicity defined as the probability of exceedance of a given earthquake magnitude over a forecast period. Additionally, ORION will provide probability distribution of the statistically derived maximum possible earthquake magnitude that may be expected. Finally, ORION will provide suggested operational management strategies (e.g. reduce injection volumes at specific wells) based on the level of hazard.

NRAP,NRAP Tools,ORION,Operational Forcasting of In↗

Evaluating proxies for the drivers of natural gas productivity using machine-learning models

We report the extensive development of unconventional reservoirs using horizontal drilling and multistage hydraulic fracturing has generated large volumes of reservoir characterization and production data. The analysis of this abundant data using statistical methods and advanced machine-learning (ML) techniques can provide data-driven insights into well performance. Most predictive modeling studies have focused on the impact that different well completion and stimulation strategies have on well production but have not fully exploited the available in situ rock property data to determine its role in reservoir productivity. We have used machine-learning techniques to rank rock mechanical properties, microseismic attributes, and stimulation parameters in the order of their significance for predicting natural gas production from an unconventional reservoir. The data for this study came from a hydraulically fractured well in the Marcellus Shale in Monongalia County, West Virginia. The data classes included measurements aggregated by well completion stage that included (1) gas production, (2) well-log-derived measurements including bulk density, elastic moduli, shear impedance, compressional impedance, brittleness, and gamma measurements, (3) microseismic attributes, (4) long-period long-duration (LPLD) event counts, (5) fracture counts, and (6) stimulation parameters that included the fluid injection volume and average pumping pressure. To identify observable proxies for the drivers of gas production, we evaluated five commonly used ML approaches including multivariate adaptive regression spline, Gaussian mixture model, random forest, gradient boosting, and neural network. We selected five variables including LPLD event count, seismogenic b-value, hydraulic diffusivity, cumulative moment, and fluid volume as the features most likely to impact gas productivity at the stage level in the study area. The data-driven selection of these parameters for their importance in determining gas production can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs. Plain language summary: We use machine-learning methods and data-driven selection of reservoir parameters to rank and better understand their importance in determining gas production, which can help reservoir engineers design more effective hydraulic-fracture treatments in the Marcellus Shale and other similar unconventional reservoirs.

58 GEOSCIENCES↗

Spatiotemporal Learning in Power Modules: Wavelet-Enhanced Forecasting of Thermomechanical Degradation

Detecting internal defects in power electronics packages is critical for their performance and reliability, especially under extreme operating conditions, as these defects can lead to catastrophic failure if not properly addressed. Confocal scanning acoustic microscopy (C-SAM) plays a key role in the nondestructive evaluation of bond layer degradation within a power electronics package by detecting defects such as delamination, voids, and cracks. However, accurately quantifying and predicting these defects from C-SAM images remains a significant challenge due to the low noise-to-signal ratio, which typically arises from both imaging process and bond patterns itself. In this paper, we explore machine learning strategies for processing C-SAM images and providing predictive models of defect growth. We use C-SAM images of sintered copper and sintered silver samples, which are obtained under accelerated thermal experiments, as the representative dataset for our study. We investigate the effect of Fourier transforms and wavelet transforms on these datasets to remove high-frequency noise and address noise across multiple scales with histogram equalization to enhance the contrast and improve the visibility of defects. As a result, defect boundaries can be clearly distinguished, enabling more accurate tracking of their growth over time. We then employ different time-series forecasting algorithms on the denoised images to formulate an image-based lifetime prediction model. Statistical models and deep-learning techniques are trained on images obtained in the early stages of thermal shock, and defect growth in the later stages is predicted. Our work serves as a preliminary attempt to improve the accuracy of lifetime prediction models of power electronics packages, which is critical under extreme operating environments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Probabilistic-learning-based stochastic surrogate model from small incomplete datasets for nonlinear dynamical systems

We consider a high-dimensional nonlinear computational model of a dynamical system, parameterized by a vector-valued control parameter, in the presence of uncertainties represented by an uncontrolled parameter modeled by a vector-valued random variable, and possibly with stochastic excitation. The objective is to construct a statistical surrogate model where the input is any deterministic value of the control parameter, and the output is a vector-valued observation of the computational model, which is a random vector whose probability measure is updated using a target dataset. To construct this statistical surrogate model, the stochastic response of the computational model must be built, which is a vector-valued time-discretized stochastic process in high dimension, depending on the control parameter. It is assumed that the computational cost of a single evaluation of the deterministic model is high. For the probabilistic updating, we consider a subset of the components of the observation of the computational model, defined as the “identification observation” of the computational model, for which a small target dataset is available. Therefore, the target dataset is associated with partial observability, corresponding to an incomplete data case. Given a prior probability model of the random control and uncontrolled parameters, a training dataset is constructed, consisting of realizations of the random triplet composed of the stochastic response, the random identification observation, and the random control parameter. Since the computational cost of a single evaluation of the deterministic model is assumed to be large, the training dataset is also of small size. The main challenges in this problem are the high dimensionality, partial observability leading to incomplete data in the target dataset for the identification observation of the computational model (which is not sufficient to identify the computational stochastic responses), and the availability of a small training dataset. To address these challenges, we propose a methodology based on statistical methods for constructing necessary reduced representations, direct probabilistic learning under constraints using probabilistic learning on manifolds (PLoM) constrained by the target dataset, and the use of a weak formulation of the Fourier transform of probability measures. Statistical conditioning is also employed to explore the learned dataset. The constructed predictive statistical surrogate model can be implemented in the context of online computation. Here, we apply this approach to a problem of nonlinear stochastic dynamics in high dimensions within the framework of deformable solids mechanics.

Engineering↗

The Statistical Emulators of GGCMI Phase 2: Responses of Year-to-Year Variation of Crop Yield to CO 2 , Temperature, Water and Nitrogen Perturbations

Understanding the impact of climate change on year-to-year variation of crop yield is critical to global food stability and security. While crop model emulators are believed to be lightweight tools to replace the models, few emulators have been developed to capture such interannual variation of crop yield in response to climate variability. In this study, we developed a statistical emulator with a machine learning algorithm to reproduce the response of year-to-year variation of four crop yields to CO 2 (C), temperature (T), water (W), and nitrogen (N) perturbations defined in the Global Gridded Crop Model Intercomparison Project (GGCMI) phase 2. The emulators were able to explain more than 52 % of the variance of simulated yield and performed well in capturing the year-to-year variation of global average and gridded crop yield over current croplands in the baseline. With the changes in CO 2 –temperature–water–nitrogen (CTWN) perturbations, the emulators could reproduce the year-to-year variation of crop yield well over most current cropland. The variation of R and the mean absolute error was small under the single CTWN perturbations and dual-factor perturbations. These emulators thus provide statistical response surfaces of yield, including both its mean and interannual variability, to climate factors. They could facilitate spatiotemporal downscaling of crop model simulation, projecting the changes in crop yield variability in the future and serving as a lightweight tool for multi-model ensemble simulation. The emulators enhanced the flexibility of crop yield estimates and expanded the application of large-ensemble simulations of crop yield under climate change.

Climate change↗

Personalized Tucker Decomposition: Modeling Commonality and Peculiarity on Tensor Data

In this paper, we propose a personalized Tucker decomposition (perTucker) to address the limitations of traditional tensor decomposition methods in capturing heterogeneity across different datasets. perTucker decomposes tensor data into shared global components and personalized local components. We introduce an order orthogonality assumption and develop a proximal gradient regularized block coordinate descent algorithm guaranteed to converge to a stationary point. The unique and common representations learned by perTucker reveal intrinsic statistical patterns in data and provide valuable information for a wide range of downstream analytics, including anomaly detection, source classification, and clustering. We demonstrate perTucker’s effectiveness through a simulation study and two case studies on solar flare detection and tonnage signal classification.

14 SOLAR ENERGY↗

A review of computing-based automated fault detection and diagnosis of heating, ventilation and air conditioning systems

We report faults in Heating, Ventilation, and Air Conditioning (HVAC) systems of buildings result in significant energy waste in building operation. With fast-growing sensing data availability and advancement in computing, computational modeling has demonstrated strong capability to detect and diagnose HVAC system faults, hence, ensuring efficient building operation. This paper comprehensively reviews the state-of-the-art computing-based fault detection and diagnosis (FDD) for HVAC systems. Overall, the reviewed computing-based FDD methods are classified as two major approaches: knowledge-based and data-driven approaches. We then identify multiple important topics, including data availability, training data size, data quality, approach generality, capability, interpretability, and required modeling efforts, along with corresponding metrics to summarize the most updated FDD development. Generally, the knowledge-based approaches are further divided as physics-based modeling, Diagnostic Bayesian Network, and performance indicator-based methods while data-driven approaches include supervised learning, unsupervised learning, and regression and statistics-based methods. State-of-the-art FDD development, remaining challenges, and future research directions are further discussed to push forward FDD in practice. Availability of fault data, capability of existing methods to deal with complex fault situations (such as simultaneous faults), modeling interpretability for data-driven methods, and required engineering efforts for physics-based methods are identified as remaining challenges in FDD development. Improving modeling fidelity and reducing modeling efforts are essential for applying physics-based methods in real buildings. Meanwhile, addressing fault data availability, increasing algorithm adaptability, and handling multiple faults are essential to further enhance the applicability of data-driven FDD approaches.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

OperonSEQer: A set of machine-learning algorithms with threshold voting for detection of operon pairs using short-read RNA-sequencing data

Operon prediction in prokaryotes is critical not only for understanding the regulation of endogenous gene expression, but also for exogenous targeting of genes using newly developed tools such as CRISPR-based gene modulation. A number of methods have used transcriptomics data to predict operons, based on the premise that contiguous genes in an operon will be expressed at similar levels. While promising results have been observed using these methods, most of them do not address uncertainty caused by technical variability between experiments, which is especially relevant when the amount of data available is small. In addition, many existing methods do not provide the flexibility to determine the stringency with which genes should be evaluated for being in an operon pair. We present OperonSEQer, a set of machine learning algorithms that uses the statistic and p-value from a non-parametric analysis of variance test (Kruskal-Wallis) to determine the likelihood that two adjacent genes are expressed from the same RNA molecule. We implement a voting system to allow users to choose the stringency of operon calls depending on whether your priority is high recall or high specificity. In addition, we provide the code so that users can retrain the algorithm and re-establish hyperparameters based on any data they choose, allowing for this method to be expanded as additional data is generated. We show that our approach detects operon pairs that are missed by current methods by comparing our predictions to publicly available long-read sequencing data. OperonSEQer therefore improves on existing methods in terms of accuracy, flexibility, and adaptability.

59 BASIC BIOLOGICAL SCIENCES↗

Curating a Standardized Dataset for Statistical Biosignature Classification

In recent years, machine learning has been explored as a toolkit for planetary science and operations [Helbert, Azari]. Machine learning has been used to improve our understanding of possible biosignatures and mineral signatures to improve science return on future missions [Warren-Rhodes, Cleaves].

Biosignatures↗

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]↗

Multiscale Flow for robust and optimal cosmological analysis

We propose Multiscale Flow, a generative Normalizing Flow that creates samples and models the field-level likelihood of two-dimensional cosmological data such as weak lensing. Multiscale Flow uses hierarchical decomposition of cosmological fields via a wavelet basis and then models different wavelet components separately as Normalizing Flows. The log-likelihood of the original cosmological field can be recovered by summing over the log-likelihood of each wavelet term. This decomposition allows us to separate the information from different scales and identify distribution shifts in the data such as unknown scale-dependent systematics. The resulting likelihood analysis can not only identify these types of systematics, but can also be made optimal, in the sense that the Multiscale Flow can learn the full likelihood at the field without any dimensionality reduction. We apply Multiscale Flow to weak lensing mock datasets for cosmological inference and show that it significantly outperforms traditional summary statistics such as power spectrum and peak counts, as well as machine learning–based summary statistics such as scattering transform and convolutional neural networks. We further show that Multiscale Flow is able to identify distribution shifts not in the training data such as baryonic effects. Finally, we demonstrate that Multiscale Flow can be used to generate realistic samples of weak lensing data.

79 ASTRONOMY AND ASTROPHYSICS↗

Denitrification and the challenge of scaling microsite knowledge to the globe

Here, our knowledge of microbial processes—who is responsible for what, the rates at which they occur, and the substrates consumed and products produced—is imperfect for many if not most taxa, but even less is known about how microsite processes scale to the ecosystem and thence the globe. In both natural and managed environments, scaling links fundamental knowledge to application and also allows for global assessments of the importance of microbial processes. But rarely is scaling straightforward: More often than not, process rates in situ are distributed in a highly skewed fashion, under the influence of multiple interacting controls, and thus often difficult to sample, quantify, and predict. To date, quantitative models of many important processes fail to capture daily, seasonal, and annual fluxes with the precision needed to effect meaningful management outcomes. Nitrogen cycle processes are a case in point, and denitrification is a prime example. Statistical models based on machine learning can improve predictability and identify the best environmental predictors but are—by themselves—insufficient for revealing process-level knowledge gaps or predicting outcomes under novel environmental conditions. Hybrid models that incorporate well-calibrated process models as predictors for machine learning algorithms can provide both improved understanding and more reliable forecasts under environmental conditions not yet experienced. Incorporating trait-based models into such efforts promises to improve predictions and understanding still further, but much more development is needed.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-informed machine learning assisted uncertainty quantification for the corrosion of dissimilar material joints

Jointing techniques like the Self-Piercing Riveting (SPR), Resistance Spot Welding (RSW) and Rivet-Weld (RW) joints are used for mass production of dissimilar material joints due to their high performance, short cycle time, and adaptability. However, the service life and safety usage of these joints can be largely impacted by the galvanic corrosion due to the difference in equilibrium potentials between the metals with the presence of electrolyte. Here, in this paper, we focus on Al-Fe galvanic corrosion and develop physics-informed machine learning based surrogate model for statistical corrosion analysis, which enables the reliability analysis of dissimilar material joints under corrosion environment. In this study, a physics-based finite element (FE) corrosion model has been developed to simulate the galvanic corrosion between a Fe cathode and an Al anode. Geometric and environmental factors including crevice gap, roughness of anode, conductivity, and the temperature of the electrolyte are investigated. Further, a thorough Uncertainty Quantification (UQ) analysis is conducted for the overall corrosion behavior of the Fe-Al joints. It is found that the electrolyte conductivity has the largest effects on the material loss and needs to be managed closely for better corrosion control. This will help in designing and manufacturing joints with improved corrosion performance.

42 ENGINEERING↗