Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

OperonSEQer: A set of machine-learning algorithms with threshold voting for detection of operon pairs using short-read RNA-sequencing data

Operon prediction in prokaryotes is critical not only for understanding the regulation of endogenous gene expression, but also for exogenous targeting of genes using newly developed tools such as CRISPR-based gene modulation. A number of methods have used transcriptomics data to predict operons, based on the premise that contiguous genes in an operon will be expressed at similar levels. While promising results have been observed using these methods, most of them do not address uncertainty caused by technical variability between experiments, which is especially relevant when the amount of data available is small. In addition, many existing methods do not provide the flexibility to determine the stringency with which genes should be evaluated for being in an operon pair. We present OperonSEQer, a set of machine learning algorithms that uses the statistic and p-value from a non-parametric analysis of variance test (Kruskal-Wallis) to determine the likelihood that two adjacent genes are expressed from the same RNA molecule. We implement a voting system to allow users to choose the stringency of operon calls depending on whether your priority is high recall or high specificity. In addition, we provide the code so that users can retrain the algorithm and re-establish hyperparameters based on any data they choose, allowing for this method to be expanded as additional data is generated. We show that our approach detects operon pairs that are missed by current methods by comparing our predictions to publicly available long-read sequencing data. OperonSEQer therefore improves on existing methods in terms of accuracy, flexibility, and adaptability.

59 BASIC BIOLOGICAL SCIENCES↗

Code for the manuscript "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Mode

We disclose a python/pytorch implementation of the physics-informed machine learning algorithm described in "Lagrangian Attention Tensor Networks for Velocity Gradient Statistical Modeling", LA-UR-24-30678. Direct numerical simulation (DNS) of ubiquitous turbulence phenomena is computationally infeasible for realistic flows. As a result, reduced modeling for turbulent flows aim to reduce the number of resolved scales while retaining accurate representations of the small-scale physics. The dynamics of the velocity gradient tensor (VGT) is a key ingredient in reduced or subgrid turbulence models. The evolution equation for the VGT involves nonlocal terms, requiring closure modeling. This implementation of the novel methodology of Lagrangian Attention Tensor Networks (LATN), utilizes a structured representation of the history of the VGT to inform a physics-informed machine learning algorithm. This addition of structured memory terms is shown to outperform previous models when trained and evaluated on DNS data.

Livescu, Daniel [LANL]↗

Multiscale Flow for robust and optimal cosmological analysis

We propose Multiscale Flow, a generative Normalizing Flow that creates samples and models the field-level likelihood of two-dimensional cosmological data such as weak lensing. Multiscale Flow uses hierarchical decomposition of cosmological fields via a wavelet basis and then models different wavelet components separately as Normalizing Flows. The log-likelihood of the original cosmological field can be recovered by summing over the log-likelihood of each wavelet term. This decomposition allows us to separate the information from different scales and identify distribution shifts in the data such as unknown scale-dependent systematics. The resulting likelihood analysis can not only identify these types of systematics, but can also be made optimal, in the sense that the Multiscale Flow can learn the full likelihood at the field without any dimensionality reduction. We apply Multiscale Flow to weak lensing mock datasets for cosmological inference and show that it significantly outperforms traditional summary statistics such as power spectrum and peak counts, as well as machine learning–based summary statistics such as scattering transform and convolutional neural networks. We further show that Multiscale Flow is able to identify distribution shifts not in the training data such as baryonic effects. Finally, we demonstrate that Multiscale Flow can be used to generate realistic samples of weak lensing data.

79 ASTRONOMY AND ASTROPHYSICS↗

Denitrification and the challenge of scaling microsite knowledge to the globe

Here, our knowledge of microbial processes—who is responsible for what, the rates at which they occur, and the substrates consumed and products produced—is imperfect for many if not most taxa, but even less is known about how microsite processes scale to the ecosystem and thence the globe. In both natural and managed environments, scaling links fundamental knowledge to application and also allows for global assessments of the importance of microbial processes. But rarely is scaling straightforward: More often than not, process rates in situ are distributed in a highly skewed fashion, under the influence of multiple interacting controls, and thus often difficult to sample, quantify, and predict. To date, quantitative models of many important processes fail to capture daily, seasonal, and annual fluxes with the precision needed to effect meaningful management outcomes. Nitrogen cycle processes are a case in point, and denitrification is a prime example. Statistical models based on machine learning can improve predictability and identify the best environmental predictors but are—by themselves—insufficient for revealing process-level knowledge gaps or predicting outcomes under novel environmental conditions. Hybrid models that incorporate well-calibrated process models as predictors for machine learning algorithms can provide both improved understanding and more reliable forecasts under environmental conditions not yet experienced. Incorporating trait-based models into such efforts promises to improve predictions and understanding still further, but much more development is needed.

59 BASIC BIOLOGICAL SCIENCES↗

Physics-informed machine learning assisted uncertainty quantification for the corrosion of dissimilar material joints

Jointing techniques like the Self-Piercing Riveting (SPR), Resistance Spot Welding (RSW) and Rivet-Weld (RW) joints are used for mass production of dissimilar material joints due to their high performance, short cycle time, and adaptability. However, the service life and safety usage of these joints can be largely impacted by the galvanic corrosion due to the difference in equilibrium potentials between the metals with the presence of electrolyte. Here, in this paper, we focus on Al-Fe galvanic corrosion and develop physics-informed machine learning based surrogate model for statistical corrosion analysis, which enables the reliability analysis of dissimilar material joints under corrosion environment. In this study, a physics-based finite element (FE) corrosion model has been developed to simulate the galvanic corrosion between a Fe cathode and an Al anode. Geometric and environmental factors including crevice gap, roughness of anode, conductivity, and the temperature of the electrolyte are investigated. Further, a thorough Uncertainty Quantification (UQ) analysis is conducted for the overall corrosion behavior of the Fe-Al joints. It is found that the electrolyte conductivity has the largest effects on the material loss and needs to be managed closely for better corrosion control. This will help in designing and manufacturing joints with improved corrosion performance.

42 ENGINEERING↗

Adaptation Strategies Strongly Reduce the Future Impacts of Climate Change on Simulated Crop Yields

Abstract Simulations of crop yield due to climate change vary widely between models, locations, species, management strategies, and Representative Concentration Pathways (RCPs). To understand how climate and adaptation affects yield change, we developed a meta‐model based on 8703 site‐level process‐model simulations of yield with different future adaptation strategies and climate scenarios for maize, rice, wheat and soybean. We tested 10 statistical models, including some machine learning models, to predict the percentage change in projected future yield relative to the baseline period (2000–2010) as a function of explanatory variables related to adaptation strategy and climate change. We used the best model to produce global maps of yield change for the RCP4.5 scenario and identify the most influential variables affecting yield change using Shapley additive explanations. For most locations, adaptation was the most influential factor determining the projected yield change for maize, rice and wheat. Without adaptation under RCP4.5, all crops are expected to experience average global yield losses of 6%–21%. Adaptation alleviates this average projected loss by 1–13 percentage points. Maize was most responsive to adaptive practices with a projected mean yield loss of −21% [range across locations: −63%, +3.7%] without adaptation and −7.5% [range: −46%, +13%] with adaptation. For maize and rice, irrigation method and cultivar choice were the adaptation types predicted to most prevent large yield losses, respectively. When adaptation practices are applied, some areas are predicted to experience yield gains, especially at northern high latitudes. These results reveal the critical importance of implementing adequate adaptation strategies to mitigate the impact of climate change on crop yields.

54 ENVIRONMENTAL SCIENCES↗

Regional variability in the environmental controls of precipitation regimes in the tropics

The environmental factors that control precipitation regimes in Boreal winter rainfall in the tropics and their regional variabilities are examined using a simple statistical analysis and a machine learning model. Radar-derived precipitation from field campaigns at Darwin Australia, Manaus Brazil, and the Equatorial Indian Ocean, along with the corresponding large-scale environmental variables from ERA5 reanalysis, are used. The dependence of marginal distributions of frequencies of these regimes on five environmental variables are calculated. The variables are hourly column integrated precipitable water (PW), convective available potential energy (CAPE), convective inhibition (CIN), and lower and upper tropospheric wind shear. The simple machine learning model that predicts the probability of transition from suppressed to an active regime as a function of the environmental variables is designed and optimized for a potential application as a trigger function for convection parameterizations. To the first order the analysis shows an abrupt increase in the probability of an active regime near PW > 60 mm and CIN of <100 J kg -1 . The key differences in the frequencies of active regimes among the regions are found to be related to the fact that over Darwin there is strong variability in PW while CIN is generally low. Over Amazon, on the other hand, both PW and CIN are quite variable. Over the DYNAMO domain the comparatively frequent low PW is compensated for by consistently low CIN resulting in a moderate frequency of an active regime.

54 ENVIRONMENTAL SCIENCES↗

Cross-Layered Distributed Data-Driven Framework for Enhanced Smart Grid Cyber-Physical Security

Smart Grid (SG) research and development has drawn much attention from academia, industry and government due to the great impact it will have on society, economics and the environment. Securing the SG is a considerably significant challenge due the increased dependency on communication networks to assist in physical process control, exposing them to various cyber-threats. In addition to attacks that change measurement values using False Data Injection (FDI) techniques, attacks on the communication network may disrupt the power system's real-time operation by intercepting messages, or by flooding the communication channels with unnecessary data. Addressing these attacks requires a cross-layer approach. In this paper a cross-layered strategy is presented, called Cross-Layer Ensemble CorrDet with Adaptive Statistics(CECD-AS), which integrates the detection of faulty SG measurement data as well as inconsistent network inter-arrival times and transmission delays for more reliable and accurate anomaly detection and attack interpretation. Numerical results show that CECD-AS can detect multiple False Data Injections, Denial of Service (DoS) and Man In The Middle (MITM) attacks with a high F1-score compared to current approaches that only use SG measurement data for detection such as the traditional physics-based State Estimation, Ensemble CorrDet with Adaptive Statistics strategy and other machine learning classification-based detection schemes.

cyber-physical security↗

Thinking Bayesian for plasma physicists

Bayesian statistics offers a powerful technique for plasma physicists to infer knowledge from the heterogeneous data types encountered. To explain this power, a simple example, Gaussian Process Regression, and the application of Bayesian statistics to inverse problems are explained. The likelihood is the key distribution because it contains the data model, or theoretic predictions, of the desired quantities. By using prior knowledge, the distribution of the inferred quantities of interest based on the data given can be inferred. Because it is a distribution of inferred quantities given the data and not a single prediction, uncertainty quantification is a natural consequence of Bayesian statistics. The benefits of machine learning in developing surrogate models for solving inverse problems are discussed, as well as progress in quantitatively understanding the errors that such a model introduces.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A low-complexity non-intrusive approach to predict the energy demand of buildings over short-term horizons

Reliable, non-intrusive, short-term (of up to 12 hours ahead) prediction of a building's energy demand is a critical component of intelligent energy management applications. A number of such approaches have been proposed over time, utilizing various statistical and, more recently, machine learning techniques, such as decision trees, neural networks and support vector machines. Importantly, all of these works barely outperform simple seasonal auto-regressive integrated moving average models, while their complexity is significantly higher. Here, we propose a novel low-complexity non-intrusive approach that improves the predictive accuracy of the state-of-the-art by up to ~10%. The backbone of our approach is a K-nearest neighbours search method, that exploits the demand pattern of the most similar historical days, and incorporates appropriate time-series pre-processing and easing. In the context of this work, we evaluate our approach against state-of-the-art methods and provide insights on their performance.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Lack of clear standards and usable comparisons of downscaled climate projections pose a roadblock for US climate discovery and adaptation

Abstract The release of global climate projections coupled with the demand for local-resolution climate-forced meteorology has prompted many research groups to downscale these projections using various statistical, dynamical, and current machine learning techniques. Such downscaled datasets are being used to plan infrastructure and other community needs over the coming decades. Faced with roughly a dozen available US downscaled datasets, many practitioners ask, ‘What are the relevant differences between datasets?’ This work highlights the difficulty of comparing downscaled datasets and illustrates ways in which datasets differ even when using identical climate model input data. We show that substantial variability in precipitation projections arises from downscaling alone and that the downscaled dataset agreement varies depending on global climate projection. This analysis emphasizes the need for greater coordination and movement toward rigorous benchmarking of downscaling strategies within the downscaling research community, à la the land-modeling community, to better quantify downscaling dataset differences, strengths, and weaknesses for practitioners.

Hartke, Samantha H. (ORCID:0000000202394723)↗

Predicting fusion ignition at the National Ignition Facility with physics-informed deep learning

Here, an inertial confinement fusion experiment, carried out at the National Ignition Facility, has achieved ignition by generating fusion energy exceeding the laser energy that drove the experiment. Prior to the experiment, a generative machine learning model that combines radiation hydrodynamics simulations, deep learning, experimental data, and Bayesian statistics was used to predict, with a probability greater than 70%, that ignition was the most likely outcome for this shot.

Spears, Brian K. [Lawrence Livermore National Labo↗

Laser Powder Bed Fusion Microstructure Surrogate Model

SAND2025-11467O The Laser Powder Bed Fusion (LPBF) Microstructure Surrogate Model is a machine-learning-based tool. It predicts statistics of microstructures that are produced by the LPBF additive manufacturing process. It includes a series of codes for training, testing, and analyzing the model as well as utility scripts for handling data. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Moser, Daniel [Sandia National Lab. (SNL-CA), Live↗

Experimental Aging and Lifetime Prediction in Grid Applications for Large-Format Commercial Li-Ion Batteries

Due to the growth of electric vehicle and stationary energy storage markets, the production and use of lithium-ion batteries has grown exponentially in recent years. For many of these applications, large-format lithium-ion batteries are being utilized, as large cells have less inactive material relative to their energy capacity and require fewer electrical connections to assemble into packs. And especially for stationary energy storage systems, where energy delivered is the only revenue source, the economics of these battery systems is highly dependent on cell lifetime. However, testing of large-format lithium-ion batteries is time consuming and requires high current channels and large testing chambers, making information on the performance of commercial, large-format lithium-ion batteries hard to come by. Here, accelerated aging test data from four commercial large-format lithium-ion batteries is reported. These batteries span both NMC-Gr and LFP-Gr cell chemistries, pouch and prismatic formats, and a range of cell designs with varying power capabilities. Accelerated aging test results are analyzed to examine both cell performance, in terms of efficiency and thermal response under load, as well as cell lifetime. Cell thermal response is characterized by measuring temperature during cycle aging, which is used to calculated a normalized thermal resistance value that may help estimate both cell cooling needs or to help extrapolate aging test results to different thermal environments. Cell lifetime is evaluated qualitatively, considering simply the average calendar and cycle life across a range of conditions, as well as quantitatively, using statistical modeling and machine-learning methods to identify predictive aging models from the accelerated aging data. These predictive aging models are then used to investigate cell sensitivities to stressors, such as cycling temperature, voltage window, and C-rate, as well as to predict cell lifetime in various stationary storage applications. Results from this work show that cell lifetime and sensitivity to aging conditions varies substantially across commercial cells, necessitating testing for specific cell formats to make quantitative lifetime predictions. That being said, all commercial cells tested here are predicted to reach at least 10-year lifetimes for stationary storage applications. Based on the aging test results and modeling, some cells are expected to be relatively insensitive to temperature and use-case, making them suited for simple use cases with little or no thermal management and simple controls, while the lifetime of other cells could be extended to 20+ years if operated with thermal management and degradation-aware controls.

battery↗

System Engineers and Decisions: It?s All about Knowledge

In order to guarantee that a system meets adequate levels of reliability and availability, system performances are continuously monitored and analyzed thanks to the technological advancements driving the Industry 4.0 revolution. An Industry 4.0 approach is typically based on advanced statistical, big data mining, machine learning, and internet-of-things methods designed to detect anomalies in the behavior of system, detect the most likely failure modes, and provide indications to system engineers on when maintenance activities should be performed before system performance are deemed unacceptable (which can be generated by diagnostic and prognostic methods). However, these analyses, which are designed to automatize and increase the efficacy of the system maintenance program, require large amount of data which can come in various forms: numeric, textual, images, sounds etc. Such data constitutes the historic knowledge benchmark to track system performances and support system engineer decisions. Here we claim that data is not sufficient to support this kind of analyses when applied to systems characterized by complex architectures and behaviors. Robust system engineer decisions require the ability to understand the system operational context that lies behind the observed data elements. In this respect, system models are in fact necessary to “put data in context” and capture relationships between data elements. Industry 4.0 methods require in fact contextual knowledge as a basis upon which hypotheses can be generated and assumptions tested. In our view, for complex systems, model-based system engineering (MBSE) models can afford this contextual knowledge, as they are typically used to describe systems architecture and dynamic behaviors. System knowledge is here intended as the blending of collected data and system architecture which takes the form of a “knowledge graph”. A knowledge graph is a database which consists of a large set of nodes (in our case an entity can be either a data or an MBSE element) which are linked to each other. The types of nodes and links follow a pre-defined topology, sometimes also refers as an ontology, that is designed to fit the actual decisions that needs to be performed. We show here how a knowledge graph can be defined to support system engineer maintenance decisions and how the same graph can be built based on system MBSE models and pre-processed data from numeric (through anomaly detections and diagnostic methods) and textual elements (through technical language processing TLP).

97 - MATHEMATICS AND COMPUTING↗

Cluster characterization in atom probe tomography: Machine learning using multiple summary functions

In this work, we develop a machine learning-based method to characterize intracluster concentration (ρ c ), background concentration (ρ b ), clustering radius (r̄), and radius dispersity (δ r ) in simulated atom probe tomography data using multiple spatial statistics summary functions to train a Bayesian regularized neural network. Here, we build upon previous work that utilized Ripley’s K-function by incorporating additional features from nearest-neighbor spatial statistics summary functions to better characterize concentration-based metrics. The addition of nearest-neighbor based features allows for highly accurate estimates of ρ c and ρ b , both with 90% of the predictions within 4.0% of the real value; the root-mean-square errors are reduced by 81.5% and 92.8% from predictions using only K-function based features, respectively. Additionally, including these nearest-neighbor based features improves the ability to differentiate between r̄ and δ r .

36 MATERIALS SCIENCE↗

Deep Learning Parameterization of Vertical Wind Velocity Variability via Constrained Adversarial Training

Atmospheric models with typical resolution in the tenths of kilometers cannot resolve the dynamics of air parcel ascent, which varies on scales ranging from tens to hundreds of meters. Small-scale wind fluctuations are thus characterized by a subgrid distribution of vertical wind velocity W with standard deviation σ W . The parameterization of σ W is fundamental to the representation of aerosol–cloud interactions, yet it is poorly constrained. Using a novel deep learning technique, this work develops a new parameterization for σ W merging data from global storm-resolving model simulations, high-frequency retrievals of W , and climate reanalysis products. The parameterization reproduces the observed statistics of σ W and leverages learned physical relations from the model simulations to guide extrapolation beyond the observed domain. Incorporating observational data during the training phase was found to be critical for its performance. The parameterization can be applied online within large-scale atmospheric models, or offline using output from weather forecasting and reanalysis products.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning-Based Anomaly Detection for PMT Data Quality Monitoring in the SBN and DUNE

Maintaining high-quality detector data is essential for achieving the scientific objectives of the Short-Baseline Neutrino (SBN) Program at Fermilab. Current data quality monitoring (DQM) procedures rely primarily on threshold-based metrics and manual inspection of detector monitoring plots, making the detection of subtle or gradually developing anomalies both time-consuming and dependent on expert interpretation. This project developed and evaluated a machine-learning workflow for automatically identifying anomalous photomultiplier tube (PMT) channels in the Short-Baseline Near Detector (SBND) using optical-hit amplitude data. A Python-based analysis program was developed to process ROOT files, extract statistical features describing individual PMT amplitude distributions, and generate feature vectors for anomaly detection. These features were used to train an Isolation Forest model using data representing normal detector operation. The trained model was subsequently applied to independent detector runs to identify channels exhibiting statistically unusual behavior relative to the learned reference response. To support expert interpretation, the workflow generated complementary diagnostic products, including anomaly score distributions, normalized amplitude comparisons, decision-tree visualizations, and principal component analysis (PCA) projections. This project demonstrated the feasibility of integrating unsupervised machine learning into detector data-quality monitoring and developed a complete workflow for automated PMT performance assessment to aid expert-driven review. Beyond its technical contributions, the VFP appointment fostered a research collaboration between Aurora University and Fermilab and provided direct workforce development benefits by training the visiting faculty member in detector-scale machine-learning methods that are now being incorporated into undergraduate coursework and research. The methodology developed here provides a foundation for future applications to ProtoDUNE and other liquid argon time projection chamber (LArTPC) detectors, contributing to ongoing efforts to improve detector reliability, reduce manual monitoring requirements, and enable scalable data quality monitoring for future large-scale neutrino experiments, including the Deep Underground Neutrino Experiment (DUNE).

Colón Santana, Juan A. [Unlisted, US, IL]↗