Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Hazard Detection Detector Cards

This report presents a comprehensive summary of five advanced anomaly detection tools developed and deployed by Oak Ridge National Laboratory in support of the VA’s Health Information Technology modernization. These detectors—Order Path Tracker, Trend Watcher, Pain Pointer, Performance Monitor, and Patient Record Flag Detector—leverage statistical and machine learning methods to monitor workflow disruptions, detect anomalies in care sequences and volumes, identify bottlenecks, and track system-level performance metrics across VistA and Millennium systems. All detectors have been integrated into the Health Data Analytics Platform (HDAP), with most having completed deployment and testing using live data from targeted stations in cardiology and oncology domains. This work enhances VA’s capacity for proactive system surveillance, promotes patient safety, and informs data-driven operational improvements across the EHR ecosystem.

97 MATHEMATICS AND COMPUTING↗

An Overview of Electric Vehicle Load Modeling Strategies for Grid Integration Studies

The adoption of electric vehicles (EVs) has emerged as a solution to reduce greenhouse gas emissions in the transportation sector, which has motivated the implementation of public policies to promote their use in several countries. However, the high adoption of EVs poses challenges for the electricity sector, as it would imply an increase in energy demand and possible impacts on the power quality (PQ) of the power grid. Therefore, it is important to conduct EV integration studies in the power grid to determine the amount that can be incorporated without causing problems and identify the areas of the power sector that will require reinforcements. Accurate EV load patterns are required for this type of study that, through mathematical modeling, reflect both the dynamic behavior and the factors that influence the decision to recharge EVs. This article aims to present an overview of EVs, examine the different factors considered in the literature for modeling EV load patterns, and review modeling methods. EV load modeling methods are classified into deterministic, statistical, and machine learning. The article shows that each modeling method has its advantages, disadvantages, and data requirements, ranging from simple load modeling to more accurate models requiring large datasets.

Computer Science↗

Data from: Understanding the biogeochemical and spatial drivers of methane and carbon dioxide fluxes in a large temperate reservoir

This dataset contains spatially resolved measurements of CO₂ and CH₄ fluxes and associated environmental variables collected across 200 sites in Douglas Reservoir (Tennessee, USA) between July 29-August 2, 2024. Measurements include diffusive fluxes of CO₂ and CH₄, CH₄ ebullition, and biogeochemical and spatial variables such as dissolved oxygen, temperature, conductivity, chlorophyll-a, pH, water depth, and distance from the dam. Sampling was conducted using a spatially balanced design to capture longitudinal and depth-related gradients throughout the reservoir. The dataset is structured to support analyses of spatial variability, flux pathway comparisons, and modeling approaches (e.g., spatial statistics and machine learning) aimed at understanding controls on reservoir CO₂ and CH₄ fluxes and improving upscaling to whole-reservoir and regional estimates.

Neeper, Jamie [ORNL] (ORCID:0009000842342101)↗

Assessing Current and Future Infrastructure Hazards

Project Objective Execute intelligent analytics via an advanced analytical framework, to assess the current state of offshore infrastructure, evaluate infrastructure life, and identify technologies to reduce infrastructure hazards, costs, and extend infrastructure life. Approach &amp; Results thus Far• Build comprehensive dataset• Perform data-driven analytics to evaluate infrastructure integrity<p> 1. Remaining lifespan</p><p> 2. Likelihood of future risk</p><p>• Apply data-driven advanced spatial, statistical, and Machine Learning (ML) models to quantify existing infrastructure integrity</p><p>• Release data and models through a smart, online platform hosted by Energy Data eXchange (EDX)</p>

Romeo, Lucy F.↗

Utility-scale Building Type Assignment Using Smart Meter Data

United States building energy use accounted for 40% of total energy use, 74% of peak demand, and $412 billion in 2019. Building energy modeling allows researchers to simulate building physics, gain insights into possible energy/demand saving opportunities, and assess cost-effective resilience amidst climate change. Many building features needed to create building energy models are readily available such as 2D footprints and LiDAR (height). A critical feature that is not generally obtainable is the building type. In partnership with a utility, a years worth of real-world, 15-minute electrical use data has been examined. The smart meter data is compared to 97 different prototype building energy models to assign building type. Real-world considerations including data preparation, quality assurance, and handling of missing values for advanced metering infrastructure data are addressed. Euclidean distance for pattern-matching of energy use, dynamic time warping, and time-window statistics with machine learning are compared for determining building type from measured electricity use.

Bass, Brett↗

Power Electronics Materials and Bonded Interfaces - Reliability and Lifetime

Advanced packaging technologies are currently being designed and developed by the power electronics industry however, the maximum operating temperature is still limited to 175 degrees Celsius for the silicon carbide devices. Bonded materials such as sintered copper and polymeric materials are potential candidates for high temperature operation, but it is critical to characterize and evaluate its reliability under harsh operating conditions. In this project, we discuss the results of the accelerated experiments conducted on sintered copper and polymeric materials. Additionally, a novel framework to develop the lifetime prediction model of bonded interfaces through employing statistical and machine learning models are described. In this task, scanning acoustic microscope images of bonded interfaces obtained under thermal cycling experiments are used as the data.

ENGINEERING↗

A Causal Approach to Model Validation and Calibration

This poster presents a novel method for validation and verification that focuses on identifying causal relationships between data elements, moving beyond traditional statistical and machine learning approaches. These methods employ causal discovery techniques to reveal the underlying mechanisms of data generation. The research utilizes structural causal models and directed acyclic graphs to depict causal relationships. This approach assists in achieving alignment between simulation models and reality.

97 MATHEMATICS AND COMPUTING↗

A Computational Workflow of Elucidating Viral Impact on Mediating Microbial Response to In-situ Experimental Warming: Bridging microbial modeling to carbon and mineral modeling

Viruses are abundant in soils and shape microbial communities in ways that can potentially influence ecosystem processes, yet their contributions to carbon cycling and mineral transformations remain poorly understood. Here we present a multi-phase framework that links virus-host interactions to soil biogeochemistry by combining ecological simulations, genome- and community-scale metabolic modeling, and statistical and machine-learning analyses. We first calibrated microbial abundance profiles under explicit infection scenarios to capture how viral pressure alters community structure, then explored alternative interaction strategies, including kill-the-winner, piggyback-the-winner, and mixed lytic-lysogenic modes, through forward simulations. These ecological shifts were translated into metabolic consequences using exchange fluxes summarized into biologically meaningful categories, while integrated statistical and machine-learning screens elevated subtle but consistent signals. Application of this framework revealed that viral infections shift the balance between organic and inorganic fluxes, redirecting metabolism from diffuse organic transformations toward inorganic pools such as protons and CO 2 , directly linking viral regulation to respiration and soil carbon balance. The roll-up analysis also isolated perturbations in critical mineral ions, including magnesium, manganese, zinc, and copper, which serve as essential enzymatic cofactors. In piggyback-the-winner scenarios, uptake of these ions was strongly suppressed. Contrasting viral strategies produced distinct community structures and metabolic outcomes, from broad suppression under kill-the-winner dynamics to dramatic redistributions under high-lytic and high-gain lysogenic regimes that collapsed vulnerable microbial populations while promoting opportunists. Together, these results provide a tractable path to trace viral perturbations from host abundance shifts to metabolic flux adjustments and ecosystem-scale processes, offering a practical way to include viruses in earth system models.

54 ENVIRONMENTAL SCIENCES↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Scalability Testing Approach for Internet of Things for Manufacturing SQL and NoSQL Database Latency and Throughput

The proliferation of low-cost sensors and industrial data solutions has continued to push the frontier of manufacturing technology. Machine learning and other advanced statistical techniques stand to provide tremendous advantages in production capabilities, optimization, monitoring, and efficiency. The tremendous volume of data gathered continues to grow, and the methods for storing the data are critical underpinnings for advancing manufacturing technology. This work aims to investigate the ramifications and design tradeoffs within a decoupled architecture of two prominent database management systems (DBMS): sql and NoSQL. A representative comparison is carried out with Amazon Web Services (AWS) DynamoDB and AWS Aurora MySQL. The technologies and accompanying design constraints are investigated, and a side-by-side comparison is carried out through high-fidelity industrial data simulated load tests using metrics from a major US manufacturer. The results support the use of simulated client load testing for comparing the latency of database management systems as a system scales up from the prototype stage into production. As a result of complex query support, MySQL is favored for higher-order insights, while NoSQL can reduce system latency for known access patterns at the expense of integrated query flexibility. Here, by reviewing this work, a manufacturer can observe that the use of high-fidelity load testing can reveal tradeoffs in IoTfM write/ingestion performance in terms of latency that are not observable through prototype-scale testing of commercially available cloud DB solutions.

AWS↗

A superconducting quantum simulator based on a photonic-bandgap metamaterial

Synthesizing many-body quantum systems with various ranges of interactions facilitates the study of quantum chaotic dynamics. Such extended interaction range can be enabled by using nonlocal degrees of freedom such as photonic modes in an otherwise locally connected structure. Here, we present a superconducting quantum simulator in which qubits are connected through an extensible photonic-bandgap metamaterial, thus realizing a one-dimensional Bose-Hubbard model with tunable hopping range and on-site interaction. Using individual site control and readout, we characterize the statistics of measurement outcomes from many-body quench dynamics, which enables in situ Hamiltonian learning. Further, the outcome statistics reveal the effect of increased hopping range, showing the predicted crossover from integrability to ergodicity. Our work enables the study of emergent randomness from chaotic many-body evolution and, more broadly, expands the accessible Hamiltonians for quantum simulation using superconducting circuits.

Science & Technology - Other Topics↗

Out-of-Distribution Detection and Radiological Data Monitoring Using Statistical Process Control

Abstract Machine learning (ML) models often fail with data that deviates from their training distribution. This is a significant concern for ML-enabled devices as data drift may lead to unexpected performance. This work introduces a new framework for out of distribution (OOD) detection and data drift monitoring that combines ML and geometric methods with statistical process control (SPC). We investigated different design choices, including methods for extracting feature representations and drift quantification for OOD detection in individual images and as an approach for input data monitoring. We evaluated the framework for both identifying OOD images and demonstrating the ability to detect shifts in data streams over time. We demonstrated a proof-of-concept via the following tasks: 1) differentiating axial vs. non-axial CT images, 2) differentiating CXR vs. other radiographic imaging modalities, and 3) differentiating adult CXR vs. pediatric CXR. For the identification of individual OOD images, our framework achieved high sensitivity in detecting OOD inputs: 0.980 in CT, 0.984 in CXR, and 0.854 in pediatric CXR. Our framework is also adept at monitoring data streams and identifying the time a drift occurred. In our simulations tracking drift over time, it effectively detected a shift from CXR to non-CXR instantly, a transition from axial to non-axial CT within few days, and a drift from adult to pediatric CXRs within a day—all while maintaining a low false positive rate. Through additional experiments, we demonstrate the framework is modality-agnostic and independent from the underlying model structure, making it highly customizable for specific applications and broadly applicable across different imaging modalities and deployed ML models.

Zamzmi, Ghada↗

Transient anisotropic kernel for probabilistic learning on manifolds

PLoM (Probabilistic Learning on Manifolds) is a method introduced in 2016 for handling small training datasets by projecting an Itô equation from a stochastic dissipative Hamiltonian dynamical system, acting as the MCMC generator, for which the KDE-estimated probability measure with the training dataset is the invariant measure. PLoM performs a projection on a reduced-order vector basis related to the training dataset, using the diffusion maps (DMAPS) basis constructed with a time-independent isotropic kernel. In this paper, we propose a new ISDE projection vector basis built from a transient anisotropic kernel, providing an alternative to the DMAPS basis to improve statistical surrogates for stochastic manifolds with heterogeneous data. The construction ensures that for times near the initial time, the DMAPS basis coincides with the transient basis. For larger times, the differences between the two bases are characterized by the angle of their spanned vector subspaces. The optimal instant yielding the optimal transient basis is determined using an estimation of mutual information from Information Theory, which is normalized by the entropy estimation to account for the effects of the number of realizations used in the estimations. Consequently, this new vector basis better represents statistical dependencies in the learned probability measure for any dimension. Three applications with varying levels of statistical complexity and data heterogeneity validate the proposed theory, showing that the transient anisotropic kernel improves the learned probability measure.

Diffusion maps↗

DEPRECATED AI-Batt-OS (Autonomous Identification of Battery Life Models - Open Source) [SWR 21-17]

DEPRECATED. This repository was archived by the owner on Jun 30, 2026. It is now read-only. Open source implementation of some of the methods utilized by AI-Batt, a battery lifetime modeling and analysis toolkit provided by the National Laboratory of the Rockies (NLR). This software demonstrates the use of bi-level optimization and symbolic regression techniques to semi-autonomously identify algebraic models predicting the capacity fade of lithium-ion batteries during calendar aging. Modeling the degradation of batteries is a complex task, due to the difficulty in separating the time-dependent and time-independent factors impacting cell level degradation, across multiple data series with different numbers of measurements and/or data quality. Bi-level optimization enables model parameters to be optimized to either the entire data set or to individual data series, allowing statistical disambiguation of global behaviors (data series independent) and local behaviors (data series dependent). Symbolic regression is used to automatically search for optimal low-dimesional models predicting the variation of locally optimized parameters versus time-independent experimental variables from millions of possible models, resulting in a more accurate and repeatable model identification process than is possible by a manual search. The provided tools also implement cross-validation and bootstrap resampling schemes, empowering statistical model comparison/selection and quantification of model uncertainties. An example script replicates the results from the manuscript "Challenging Practices of Algebraic Battery Life Models through Statistical Validation and Model Identification via Machine-Learning", submitted to ECS. All code is written in MATLAB. Requires the Statistics and Machine Learning Toolbox. Contact Dr. Paul Gasper at Paul.Gasper@nlr.gov for any questions.

Gasper, Paul↗

Physics guided machine learning using simplified theories

Recent applications of machine learning, in particular deep learning, motivate the need to address the generalizability of the statistical inference approaches in physical sciences. In this Letter, we introduce a modular physics guided machine learning framework to improve the accuracy of such data-driven predictive engines. The chief idea in our approach is to augment the knowledge of the simplified theories with the underlying learning process. To emphasize their physical importance, our architecture consists of adding certain features at intermediate layers rather than in the input layer. To demonstrate our approach, we select a canonical airfoil aerodynamic problem with the enhancement of the potential flow theory. We include the features obtained by a panel method that can be computed efficiently for an unseen configuration in our training procedure. By addressing the generalizability concerns, our results suggest that the proposed feature enhancement approach can be effectively used in many scientific machine learning applications, especially for the systems where we can use a theoretical, empirical, or simplified model to guide the learning module.

42 ENGINEERING↗