Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Statistically-informed deep learning for gravitational wave parameter estimation

We introduce deep learning models to estimate the masses of the binary components of black hole mergers, $(m_1,m_2)$, and three astrophysical properties of the post-merger compact remnant, namely, the final spin, $a_\mathrm f$, and the frequency and damping time of the ringdown oscillations of the fundamental $\ell = m = 2$ bar mode, $(\omega_\mathrm R, \omega_\mathrm I)$. Our neural networks combine a modified WaveNet architecture with contrastive learning and normalizing flow. We validate these models against a Gaussian conjugate prior family whose posterior distribution is described by a closed analytical expression. Upon confirming that our models produce statistically consistent results, we used them to estimate the astrophysical parameters $(m_1,m_2, a_\mathrm f, \omega_\mathrm R, \omega_\mathrm I)$ of five binary black holes: GW150914, GW170104, GW170814, GW190521 and GW190630. We use PyCBC Inference to directly compare traditional Bayesian methodologies for parameter estimation with our deep learning based posterior distributions. Our results show that our neural network models predict posterior distributions that encode physical correlations, and that our data-driven median results and 90% confidence intervals are similar to those produced with gravitational wave Bayesian analyses. This methodology requires a single V100 NVIDIA GPU to produce median values and posterior distributions within two milliseconds for each event. Furthermore, this neural network, and a tutorial for its use, are available at the Data and Learning Hub for Science.

79 ASTRONOMY AND ASTROPHYSICS↗

Systematic quark/gluon identification with ratios of likelihoods

Discriminating between quark- and gluon-initiated jets has long been a central focus of jet substructure, leading to the introduction of numerous observables and calculations to high perturbative accuracy. At the same time, there have been many attempts to fully exploit the jet radiation pattern using tools from statistics and machine learning. We propose a new approach that combines a deep analytic understanding of jet substructure with the optimality promised by machine learning and statistics. After specifying an approximation to the full emission phase space, we show how to construct the optimal observable for a given classification task. This procedure is demonstrated for the case of quark and gluons jets, where we show how to systematically capture sub-eikonal corrections in the splitting functions, and prove that linear combinations of weighted multiplicity is the optimal observable. In addition to providing a new and powerful framework for systematically improving jet substructure observables, we demonstrate the performance of several quark versus gluon jet tagging observables in parton-level Monte Carlo simulations, and find that they perform at or near the level of a deep neural network classifier. Combined with the rapid recent progress in the development of higher order parton showers, we believe that our approach provides a basis for systematically exploiting subleading effects in jet substructure analyses at the Large Hadron Collider (LHC) and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Learning stochastic dynamics with statistics-informed neural network

We introduce a machine-learning framework named statistics-informed neural network (SINN) for learning stochastic dynamics from data. This new architecture was theoretically inspired by a universal approximation theorem for stochastic systems, which we introduce in this paper, and the projection-operator formalism for stochastic modeling. Here, we devise mechanisms for training the neural network model to reproduce the correct statistical behavior of a target stochastic process. Numerical simulation results demonstrate that a well-trained SINN can reliably approximate both Markovian and non-Markovian stochastic dynamics. We demonstrate the applicability of SINN to coarse-graining problems and the modeling of transition dynamics. Furthermore, we show that the obtained reduced-order model can be trained on temporally coarse-grained data and hence is well suited for rare-event simulations.

97 MATHEMATICS AND COMPUTING↗

Interface learning in fluid dynamics: Statistical inference of closures within micro–macro-coupling models

Many complex multiphysics systems in fluid dynamics involve using solvers with varied levels of approximations in different regions of the computational domain to resolve multiple spatiotemporal scales present in the flow. The accuracy of the solution is governed by how the information is exchanged between these solvers at the interface and several methods have been devised for such coupling problems. In this article, we construct a data-driven model by spatially coupling a microscale lattice Boltzmann method (LBM) solver and macroscale finite difference method (FDM) solver for reaction-diffusion systems. The coupling between the micro-macro solvers has one to many mapping at the interface leading to the interface closure problem, and we propose a statistical inference method based on neural networks to learn this closure relation. The performance of the proposed framework in a bifidelity setting partitioned between the FDM and LBM domain shows its promise for complex systems where analytical relations between micro-macro solvers are not available.

42 ENGINEERING↗

Unsupervised Clustering and Supervised Regression Learning to Select High Temperature Oxidation-Resistant Materials

High temperature oxidation and corrosion degradation mechanisms dictate the lifetime of materials critical to energy production. The combination of modeling and experimental approaches such as machine learning (ML) and data analytics, with sufficient experimental data, can accelerate the development of new materials while limiting its cost. In the present work, ML will be applied to two high temperature oxidation data libraries (Oak Ridge National Laboratory and National Air and Space Administration) that comprised of about 5000 mass change sample datasheets for a variety of materials and temperatures in dry air and air + 10 % H2O. A python code was developed to prepare the data for machine learning by collecting and formatting oxidation rate constants, alloy compositions and environment of exposure into a single data frame. Scikit-learn library and Statistics and Machine Learning Toolbox within MathWorks were then used to perform unsupervised clustering and supervised regression learning. The impact of dataset distribution on the performance of the developed ML models was evaluated. Potential strategies to improve the predictions and enhance extrapolative capability of the previously trained model were investigated.

Romedenne, Marie [ORNL] (ORCID:0000000317936561)↗

Investigating the influence of particle distribution on force and torque statistics using hierarchical machine learning

An accurate representation of hydrodynamic force and torque experienced by every particle in a distribution can be obtained from particle resolved (PR) simulations. These unique quantities are influenced by the deterministic position of surrounding particles. However, systems simulated with this methodology are typically limited to particles due to the involved computational cost. This resource requirement is a major bottleneck in analyzing the effect of variations in particle distribution. Here, this article attempts to address this bottleneck by availing relatively inexpensive deep learning models. The surrogate models that we employ in this article use a physics‐based hierarchical framework and symmetry‐preserving neural networks to achieve robustness with limited training data. This article first performs additional generalizability tests on PR data of distinct distributions that are not involved in the training process. The models are then deployed on several different particle distributions. Impact of clustering and structure on the observed statistics are investigated.

42 ENGINEERING↗

Decoding defect statistics from diffractograms via machine learning

Abstract Diffraction techniques can powerfully and nondestructively probe materials while maintaining high resolution in both space and time. Unfortunately, these characterizations have been limited and sometimes even erroneous due to the difficulty of decoding the desired material information from features of the diffractograms. Currently, these features are identified non-comprehensively via human intuition, so the resulting models can only predict a subset of the available structural information. In the present work we show (i) how to compute machine-identified features that fully summarize a diffractogram and (ii) how to employ machine learning to reliably connect these features to an expanded set of structural statistics. To exemplify this framework, we assessed virtual electron diffractograms generated from atomistic simulations of irradiated copper. When based on machine-identified features rather than human-identified features, our machine-learning model not only predicted one-point statistics (i.e. density) but also a two-point statistic (i.e. spatial distribution) of the defect population. Hence, this work demonstrates that machine-learning models that input machine-identified features significantly advance the state of the art for accurately and robustly decoding diffractograms.

36 MATERIALS SCIENCE↗

Jet classification using high-level features from anatomy of top jets

Recent advancements in deep learning models have significantly enhanced jet classification performance by analyzing low-level features (LLFs). However, this approach often leads to less interpretable models, emphasizing the need to understand the decision-making process and to identify the high-level features (HLFs) crucial for explaining jet classification. To address this, we consider the top jet tagging problems and introduce an analysis model (AM) that analyzes selected HLFs designed to capture important features of top jets. Our AM mainly consists of the following three modules: a relation network analyzing two-point energy correlations, mathematical morphology and Minkowski functionals for generalizing jet constituent multiplicities, and a recursive neural network analyzing subjet constituent multiplicity to enhance sensitivity to subjet color charges. We demonstrate that our AM achieves performance comparable to the Particle Transformer (ParT) while requiring fewer computational resources in a comparison of top jet tagging using jets simulated at the hadronic calorimeter angular resolution scale. Furthermore, as a more constrained architecture than ParT, the AM exhibits smaller training uncertainties because of the bias-variance tradeoff. We also compare the information content of AM and ParT by decorrelating the features already learned by AM. Lastly, we briefly comment on the results of AM with finer angular resolution inputs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Special Issue: Geostatistics and Machine Learning

Abstract Recent years have seen a steady growth in the number of papers that apply machine learning methods to problems in the earth sciences. Although they have different origins, machine learning and geostatistics share concepts and methods. For example, the kriging formalism can be cast in the machine learning framework of Gaussian process regression. Machine learning, with its focus on algorithms and ability to seek, identify, and exploit hidden structures in big data sets, is providing new tools for exploration and prediction in the earth sciences. Geostatistics, on the other hand, offers interpretable models of spatial (and spatiotemporal) dependence. This special issue on Geostatistics and Machine Learning aims to investigate applications of machine learning methods as well as hybrid approaches combining machine learning and geostatistics which advance our understanding and predictive ability of spatial processes.

58 GEOSCIENCES↗

Predicting industrial building energy consumption with statistical and machine-learning models informed by physical system parameters

The industrial sector consumes about one-third of global energy, making them a frequent target for energy use reduction. Variation in energy usage is observed with weather conditions, as space conditioning needs to change seasonally, and with production, energy-using equipment is directly tied to production rate. Previous models were based on engineering analyses of equipment and relied on site-specific details. Others consisted of single-variable regressors that did not capture all contributions to energy consumption. Further, new modeling techniques could be applied to rectify these weaknesses. Applying data from 45 different manufacturing plants obtained from industrial energy audits, a supervised machine-learning model is developed to create a general predictor for industrial building energy consumption. The model uses features of air enthalpy, solar radiation, and wind speed to predict weather-dependency; motor, steam, and compressed air system parameters to capture support equipment contributions; and operating schedule, production rate, number of employees, and floor area to determine production-dependency. Results showed that a model that used a linear regressor over a transformed feature space could outperform a support vector machine and utilize features more representative of physical systems. Using informed parameters to build a reliable predictor will more accurately characterize a manufacturing facility's energy savings opportunities.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Emerging opportunities for hybrid perovskite solar cells using machine learning

While there are several bottlenecks in hybrid organic–inorganic perovskite (HOIP) solar cell production steps, including composition screening, fabrication, material stability, and device performance, machine learning approaches have begun to tackle each of these issues in recent years. Different algorithms have successfully been adopted to solve the unique problems at each step of HOIP development. Specifically, high-throughput experimentation produces vast amount of training data required to effectively implement machine learning methods. Here, we present an overview of machine learning models, including linear regression, neural networks, deep learning, and statistical forecasting. Experimental examples from the literature, where machine learning is applied to HOIP composition screening, thin film fabrication, thin film characterization, and full device testing, are discussed. These paradigms give insights into the future of HOIP solar cell research. As databases expand and computational power improves, increasingly accurate predictions of the HOIP behavior are becoming possible.

Hering, Abigail R. (ORCID:0000000270806953)↗

On reading Youden: Learning about the practice of statistics and applied statistical research from a master applied statistician

From reading William John “Jack” Youden’s books and articles, Youden (1900-1971), an analytical chemist, becomes an applied statistician by the time he joins the National Bureau of Standards (NBS) in 1948. Here, this article traces his transition from chemist to applied statistician and what his body of work mostly at NBS (1948-1965) demonstrates about his practice of statistics and the role that applied statistical research plays in it. There is much we can learn from a master applied statistician.

42 ENGINEERING↗