Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Accurate and efficient parameterization of an atomic cluster expansion (ACE) potential for ammonia under extreme conditions

We present a machine learning interatomic potential for ammonia designed to capture its complex multiphase behavior, including both molecular and superionic phases. The potential is based on the atomic cluster expansion (ACE) formulation and has been parameterized to facilitate high-fidelity molecular dynamics simulations of ammonia under extreme conditions, for pressures up to 100 GPa and for temperatures above 500 K and up to 6000 K. A diverse range of configurations was generated through high-quality ab initio molecular dynamics simulations, covering insulating and superionic ice phases, liquid ammonia, molecular nitrogen (N 2 ) and hydrogen (H 2 ), and metastable compounds that form upon dissociation, including $NH^{+}_{4}$, $H^{+}_{3}$, N 2 H 4 , and N 3 H. We demonstrate that the ammonia ACE potential accurately reproduces experimental and density functional theory predicted isotherms and Hugoniots. Crucially, the potential is able to capture the intricate phase behavior of ammonia, including the transition from insulating molecular fluid to the superionic phase. This work provides a robust interatomic potential that can be used for large-scale, accurate simulations of ammonia under extreme thermodynamic conditions, offering a powerful tool for investigating its behavior in various phases and applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

An interactive machine learning platform for analyzing multi-particle coincidence data from cold target recoil ion momentum spectroscopy

We present SCULPT (Supervised Clustering and Uncovering Latent Patterns with Training), a comprehensive software platform for analyzing tabulated high-dimensional multi-particle coincidence data from Cold Target Recoil Ion Momentum Spectroscopy (COLTRIMS) experiments. The software addresses critical challenges in modern momentum spectroscopy by integrating advanced machine learning techniques with physics-informed analysis in an interactive web-based environment. SCULPT implements uniform manifold approximation and projection for non-linear dimensionality reduction to reveal correlations in high-dimensional data. We also discuss potential extensions to deep autoencoders for feature learning and genetic programming for automated discovery of physically meaningful observables. A novel adaptive confidence scoring system provides quantitative reliability assessments by evaluating user-selected clustering quality metrics with predefined weights that reflect each metric’s robustness. The platform features configurable molecular profiles for different experimental systems, interactive visualization with selection tools, and comprehensive data filtering capabilities. Utilizing a subset of SCULPT’s capabilities, we analyze photo-double-ionization data measured using the COLTRIMS method for three-body dissociation of the D 2 O molecule, revealing distinct fragmentation channels and their correlations with physics parameters. The software’s modular architecture and web-based implementation make it accessible to the broader atomic and molecular physics community, significantly reducing the time required for complex multi-dimensional analyses. This opens the door to finding and isolating rare events exhibiting non-linear correlations on the fly during experimental measurements, which can help steer exploration and improve the efficiency of experiments.

Artificial neural networks↗

A Deep Generative Model for Non-Intrusive Identification of EV Charging Profiles

The proliferation of electric vehicles (EVs) brings environmental benefits and technical challenges to power grids. An identification algorithm which can accurately extract individual EV charging profiles out of widely available smart meter measurements has attracted great interests. This paper proposes a non-intrusive identification framework for EV charging profile extraction, which is driven by deep generative models (DGM). First, the proposed DGM is designed as a representation layer embedded into the Markov process and used to model the joint probability distribution of available time-series data. A novel contribution is to approximate posterior distributions by neural networks whose parameters are obtained by variational inference and supervised learning. Second, the EV charging status is inferred from the DGM via dynamic programming. Lastly, the desired EV charging profile can be reconstructed by the rated power of EV models and inferred status. Compared with the benchmark Hidden Markov Models, the proposed framework can better handle noise in data with less computational complexity and better overall accuracy performances with smaller recall. The proposed framework is validated by numerical experiments on the Pecan Street dataset.

33 ADVANCED PROPULSION SYSTEMS↗

Opportunities in AI/ML for the Rubin LSST Dark Energy Science Collaboration

The Vera C. Rubin Observatory's Legacy Survey of Space and Time (LSST) will produce unprecedented volumes of heterogeneous astronomical data (images, catalogs, and alerts) that challenge traditional analysis pipelines. The LSST Dark Energy Science Collaboration (DESC) aims to derive robust constraints on dark energy and dark matter from these data, requiring methods that are statistically powerful, scalable, and operationally reliable. Artificial intelligence and machine learning (AI/ML) are already embedded across DESC science workflows, from photometric redshifts and transient classification to weak lensing inference and cosmological simulations. Yet their utility for precision cosmology hinges on trustworthy uncertainty quantification, robustness to covariate shift and model misspecification, and reproducible integration within scientific pipelines. This white paper surveys the current landscape of AI/ML across DESC's primary cosmological probes and cross-cutting analyses, revealing that the same core methodologies and fundamental challenges recur across disparate science cases. Since progress on these cross-cutting challenges would benefit multiple probes simultaneously, we identify key methodological research priorities, including Bayesian inference at scale, physics-informed methods, validation frameworks, and active learning for discovery. With an eye on emerging techniques, we also explore the potential of the latest foundation model methodologies and LLM-driven agentic AI systems to reshape DESC workflows, provided their deployment is coupled with rigorous evaluation and governance. Finally, we discuss critical software, computing, data infrastructure, and human capital requirements for the successful deployment of these new methodologies, and consider associated risks and opportunities for broader coordination with external actors.

Aubourg, Eric [APC, Paris] (ORCID:000000025592023X↗

Bayesian Approach to the Joint Inversion of Gravity and Magnetic Data, with Application to the Ismenius Area of Mars

This viewgraph presentation reviews a Bayesian approach to the inversion of gravity and magnetic data with specific application to the Ismenius Area of Mars. Many inverse problems encountered in geophysics and planetary science are well known to be non-unique (i.e. inversion of gravity the density structure of a body). In hopes of reducing the non-uniqueness of solutions, there has been interest in the joint analysis of data. An example is the joint inversion of gravity and magnetic data, with the assumption that the same physical anomalies generate both the observed magnetic and gravitational anomalies. In this talk, we formulate the joint analysis of different types of data in a Bayesian framework and apply the formalism to the inference of the density and remanent magnetization structure for a local region in the Ismenius area of Mars. The Bayesian approach allows prior information or constraints in the solutions to be incorporated in the inversion, with the "best" solutions those whose forward predictions most closely match the data while remaining consistent with assumed constraints. The application of this framework to the inversion of gravity and magnetic data on Mars reveals two typical challenges - the forward predictions of the data have a linear dependence on some of the quantities of interest, and non-linear dependence on others (termed the "linear" and "non-linear" variables, respectively). For observations with Gaussian noise, a Bayesian approach to inversion for "linear" variables reduces to a linear filtering problem, with an explicitly computable "error" matrix. However, for models whose forward predictions have non-linear dependencies, inference is no longer given by such a simple linear problem, and moreover, the uncertainty in the solution is no longer completely specified by a computable "error matrix". It is therefore important to develop methods for sampling from the full Bayesian posterior to provide a complete and statistically consistent picture of model uncertainty, and what has been learned from observations. We will discuss advanced numerical techniques, including Monte Carlo Markov

data analysis↗

Evaluation of Classifier Complexity for Delay Tolerant Network Routing

The growing popularity of small cost effective satellites (SmallSats, CubeSats, etc.) creates the potential for a variety of new science applications involving multiple nodes functioning together or independently to achieve a task, such as swarms and constellations. As this technology develops and is deployed for missions in Low Earth Orbit and beyond, the use of delay tolerant networking (DTN) techniques may improve communication capabilities within the network. In this paper, a network hierarchy is developed from heterogeneous networks of SmallSats, surface vehicles, relay satellites and ground stations which form an integrated network. There is a tradeoff between complexity, flexibility, and scalability of user defined schedules versus autonomous routing as the number of nodes in the network increases. To address these issues, this work proposes a machine learning classifier based on DTN routing metrics. A framework is developed which will allow for the use of several categories of machine learning algorithms (decision tree, random forest and deep learning) to be applied to a dataset of historical network statistics, which allows for the evaluation of algorithm complexity versus performance to be explored. We develop the emulation of a hierarchical network, consisting of tens of nodes which form a cognitive network architecture. CORE (Common Open Research Emulator) is used to emulate the network using bundle protocol and DTN IP neighbor discovery.

Dudukovich, Rachel↗

NASA GeneLab Multi-study Visualization Portal

NASA GeneLab has helped advance the field of Space Biology by providing a public repository where researchers can store, share, analyze and visualize the results of space flight related omics experiments. The GeneLab data visualization portal allows any user, regardless of bioinformatics knowledge or access to computational resources, to interact with the experimental data, draw their own conclusions, and gain insights about the effects of space on living systems. These tools help democratize scientific research and foster the NASA Open Science initiative. The new multi-study feature of the GeneLab visualization platform allows users to mine study metadata from RNA sequencing (RNA-seq) experiments to identify samples of interest by filtering datasets based on organism, tissue, assay technology type, and/or factor. Once samples are selected from multiple datasets, users can combine and normalize the sample data, then utilize the visualization displays, including Principal Component Analysis (PCA) plots, to assess sample distributions. Finally, users can perform differential gene expression analysis on the combined data and visualize the results through PCA plots, Volcano plots, Pair plots, Heatmap, Ideogram and Gene Set Enrichment Analysis. All user-generated results and visualizations will be available for download. Here, we present a biological study using samples from multiple GeneLab RNA-seq datasets and analyzed using the multi-study visualization platform to demonstrate inter- and intra-study variability, as well as commonly differentially expressed genes between spaceflight and ground control conditions across datasets. This new feature opens a wide range of possibilities and opportunities for further development including combining other assay technology types and integration with batch effect correction techniques and machine learning applications. Overall, this tool allows users to increase the statistical power of individual experiments, validate hypothesis, identify patterns, and opens the door to new and exciting research.

space biology↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗

Contrasting Time-Frequency Representations for Unknown Waveform Detection

Identifying unseen electromagnetic waveforms is critical for many applications, like interference management, electronic warfare and spectrum management. Traditionally this is done using statistical methods for anomaly detection, which has evolved to deep learning models for identifying the unseen data, formally termed as open set recognition. Some prior methods use a generative model to emulate open set data, which face challenges in generating synthetic samples for open set while simultaneously selecting an optimal discriminator for accurate classification. To alleviate this issue, we propose a discriminative model that effectively combines time and frequency domain features of communication signals for accurate predictions. We further introduce a cosine similarity loss that makes the domain specific features unique to enhance the prediction rate. Additionally, our model avoids generic feature vectors by extracting class-specific features during training, resulting in improved class representation. The experiment results show that this combined feature approach with cosine loss outperforms single-domain models and improves accuracy by 10% over models without cosine loss.

99 - GENERAL AND MISCELLANEOUS↗

Machine Learning Inference of Random Medium Properties

Earth materials are heterogeneous across a range of spatial scales, but the resolvability of small structures is limited by sparse data coverage, noise, bandlimitedness, and other difficulties. In practice, heterogeneities below a certain size cannot be recovered from seismic data except through statistical medium descriptions, which even then can be difficult to uniquely determine. To improve the characterization of such heterogeneities, we develop a novel supervised machine learning (ML) model that provides insight about the recoverability of statistical medium properties from elastic waveform data and succeeds despite cycle-skipping and other challenges well known from elastic waveform inversion. We demonstrate the approach using random media generated by superimposing self-affine random variations on homogeneous and layered background structures. After training on sparsely-recorded, high-frequency waveforms from hundreds of different random medium realizations, we show the ability of our ML model to recover correlation lengths and other statistical properties of interest to near-surface and crustal seismology, among other fields. For frequency passbands and spatial offsets encountered in seismology, Gaussian correlation lengths and the amplitude of the random variations relative to the background model are recovered even in challenging scenarios involving unknown medium parameters, complex crustal structures, and low signal-to-noise ratio. In comparison, von Kármán correlation lengths, which are related to larger-wavelength variations of the medium than Gaussian correlation lengths, are not as well recovered. These results provide one of the first and most systematic investigations of the recoverability of statistical properties of heterogeneities below the resolution limit of deterministic seismic tomography, and suggest practical ML strategies for high-frequency waveform seismology.

58 GEOSCIENCES↗

Improving statistical precision in Monte Carlo samples with negative weights via reweighting and uncertainty quantification

High statistical precision is critical for Monte Carlo (MC) samples in high energy physics and is degraded by negatively weighted events. This paper investigates a procedure to learn the relationship between the negative and positive weight distributions of any sample, allowing the reduction of statistical uncertainty by reweighting kinematically equivalent events with the same sign. A robust uncertainty quantification method is required for the practical application of such method. Two methods for the estimation of the reweighting uncertainty are developed: one at the event and another one at the final observable level. The latter method is strongly favored. The gains in statistical precision are then quantified. The method is demonstrated on Sherpa vector boson plus jets samples when using all generated events and when restricted to the signal region of a mock analysis. It is demonstrated to significantly reduce stochastic behavior in sparse MC samples while decreasing the overall uncertainty with a sufficiently well-known reweighting function.

Monte Carlo methods↗

A Machine Learning Assisted Development of a Model for the Populations of Convective and Stratiform Clouds

Abstract Traditional parameterizations of the interaction between convection and the environment have relied on an assumption that the slowly varying large‐scale environment is in statistical equilibrium with a large number of small and short‐lived convective clouds. They fail to capture nonequilibrium transitions such as the diurnal cycle and the formation of mesoscale convective systems as well as observed precipitation statistics and extremes. Informed by analysis of radar observations, cloud‐permitting model simulation, theory, and machine learning, this work presents a new stochastic cloud population dynamics model for characterizing the interactions between convective and stratiform clouds, with the goal of informing the representation of these interactions in global climate models. Fifteen wet seasons of precipitating cloud observations by a C‐band radar at Darwin, Australia are fed into a machine learning algorithm to obtain transition functions that close a set of coupled equations relating large‐scale forcing, mass flux, the convective cell size distribution, and the stratiform area. Under realistic large‐scale forcing, the derived transition functions show that, on the one hand, interactions with stratiform clouds act to dampen the variability in the size and number of convective cells and therefore in the convective mass flux. On the other, for a given convective area fraction, a larger number of smaller cells is more favorable for the growth of stratiform area than a smaller number of larger cells. The combination of these two factors gives rise to solutions with a few convective cells embedded in a large stratiform area, reminiscent of mesoscale convective systems.

54 ENVIRONMENTAL SCIENCES↗

Variance Preserving Spectral Subsampling

Generating statistically faithful short-duration gamma-ray spectra from a single long measurement is essential in nuclear safeguards, supporting tasks such as algorithm development and machine-learning applications, especially when list-mode data are unavailable. Existing subsampling methods often distort the statistical characteristics of genuine short-duration measurements, leading to biased or unreliable analytical outcomes and thereby undermining downstream tasks. In this work, we compare five subsampling approaches using a benchmark set of 156 genuine replicate spectra collected with a high-purity germanium detector. We evaluate each method with respect to run-to-run variance, channel-to-channel variance, and preservation of total counts (losslessness). Across a wide range of subsampling ratios, only binomial subsampling without replacement consistently reproduces the statistical properties of genuine short-duration spectra, maintaining proper dispersion even in sparse spectral regions and perfectly preserving total counts. These results provide a mathematically principled and practically validated framework for generating synthetically shortened spectra when true short-duration measurements are unavailable.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

EFIT‐AI: Machine Learning and Artificial Intelligence Assisted Equilibrium Reconstruction for Tokamak Experiments and Burning Plasmas (Final Report)

The EFIT-AI project is creating a modern advanced equilibrium reconstruction code suitable for tokamak experiments of burning plasmas. EFIT [1,2] was the first and is the most extensively used equilibrium reconstruction code in the world. This project builds on the production-level experience and adds key elements as follows. 1. A Model Order Reduction (MOR) version of the two-dimensional (2D) Grad-Shafranov equation solver (EFIT-MORNN) using physics-informed neural networks. 2. Improved optimization and data analysis capabilities using a Bayesian framework enhanced with machine learning. 3. A MOR version of the three-dimensional (3D) perturbed equilibrium reconstruction tool.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Lithium-ion battery physics and statistics-based state of health model

A pseudo-2d model using COMSOL Multiphysics® software is developed to simulate performance and performance degradation of Li-ion batteries consisting of layered and olivine cathodes with graphite anode when subjected to peak shaving grid service. Multiple degradation pathways are considered, including solid electrolyte interphase (SEI) formation and breakdown at the anode, cathode dissolution and its synergistic effect on SEI formation at the anode. The model is validated by simulating commercial cylindrical cell performance. A global model is developed to simulate performance across all chemistries, along with individual chemistry models using global model parameters as initial values. There is good agreement between these models for various optimization parameters such as SEI equilibrium potential, cathode dissolution exchange current density, solvent diffusivity in the SEI and SEI ionic conductivity. To circumvent time constraints related to the COMSOL model, a 0d global model is developed which fits data well and provides more clarity on differences in cathode dissolution exchange current density. Again, good agreement for various optimization parameters is obtained among the COMSOL global & individual chemistry models and the 0-d model. The lessons learned from the physics-based model is used to develop a top down statistics-based model using current, voltage and anode volumetric change per mole lithium intercalated, along with their interactions as degradation predictors. This model predicts out of sample degradation for multiple grid services and electric vehicle drive cycle with high accuracy and provides the pathway to develop an efficient battery management system combining machine learning and findings from physics-based computationally intensive algorithms.

Crawford, Aladsair J.↗

Learning to simulate high energy particle collisions from unlabeled data

In many scientific fields which rely on statistical inference, simulations are often used to map from theoretical models to experimental data, allowing scientists to test model predictions against experimental results. Experimental data is often reconstructed from indirect measurements causing the aggregate transformation from theoretical models to experimental data to be poorly-described analytically. Instead, numerical simulations are used at great computational cost. We introduce Optimal-Transport-based Unfolding and Simulation (OTUS), a fast simulator based on unsupervised machine-learning that is capable of predicting experimental data from theoretical models. Without the aid of current simulation information, OTUS trains a probabilistic autoencoder to transform directly between theoretical models and experimental data. Identifying the probabilistic autoencoder’s latent space with the space of theoretical models causes the decoder network to become a fast, predictive simulator with the potential to replace current, computationally-costly simulators. Here, we provide proof-of-principle results on two particle physics examples, Z-boson and top-quark decays, but stress that OTUS can be widely applied to other fields.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Mathematics: The Tao of Data Science

The two pieces, "Ten Research Challenge Areas in Data Science" by Jeannette M. Wing and “Challenges and Opportunities in Statistics and Data Science: Ten Research Areas” by Xuming He and Xihong Lin, provide an impressively complete list of data science challenges from luminaries in the field of data science. They have done an extraordinary job, so this response offers a complementary viewpoint from a mathematical perspective and evangelizes advanced mathematics as a key tool for meeting the challenges they have laid out. Notably, we pick up the themes of scientific understanding of machine learning and deep learning, computational considerations such as cloud computing and scalability, balancing computational and statistical considerations, and inference with limited data. We propose that mathematics is an important key to establishing rigor in the field of data science and as such has an essential role to play in its future.

97 MATHEMATICS AND COMPUTING↗

Acoustic Energy Release During the Laboratory Seismic Cycle: Insights on Laboratory Earthquake Precursors and Prediction

Abstract Machine learning can predict the timing and magnitude of laboratory earthquakes using statistics of acoustic emissions. The evolution of acoustic energy is critical for lab earthquake prediction; however, the connections between acoustic energy and fault zone processes leading to failure are poorly understood. Here, we document in detail the temporal evolution of acoustic energy during the laboratory seismic cycle. We report on friction experiments for a range of shearing velocities, normal stresses, and granular particle sizes. Acoustic emission data are recorded continuously throughout shear using broadband piezo‐ceramic sensors. The coseismic acoustic energy release scales directly with stress drop and is consistent with concepts of frictional contact mechanics and time‐dependent fault healing. Experiments conducted with larger grains (10.5 μm) show that the temporal evolution of acoustic energy scales directly with fault slip rate. In particular, the acoustic energy is low when the fault is locked and increases to a maximum during coseismic failure. Data from traditional slide‐hold‐slide friction tests confirm that acoustic energy release is closely linked to fault slip rate. Furthermore, variations in the true contact area of fault zone particles play a key role in the generation of acoustic energy. Our data show that acoustic radiation is related primarily to breaking/sliding of frictional contact junctions, which suggests that machine learning‐based laboratory earthquake prediction derives from frictional weakening processes that begin very early in the seismic cycle and well before macroscopic failure.

58 GEOSCIENCES↗