Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “statistical learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

An Overview of Electric Vehicle Load Modeling Strategies for Grid Integration Studies

The adoption of electric vehicles (EVs) has emerged as a solution to reduce greenhouse gas emissions in the transportation sector, which has motivated the implementation of public policies to promote their use in several countries. However, the high adoption of EVs poses challenges for the electricity sector, as it would imply an increase in energy demand and possible impacts on the power quality (PQ) of the power grid. Therefore, it is important to conduct EV integration studies in the power grid to determine the amount that can be incorporated without causing problems and identify the areas of the power sector that will require reinforcements. Accurate EV load patterns are required for this type of study that, through mathematical modeling, reflect both the dynamic behavior and the factors that influence the decision to recharge EVs. This article aims to present an overview of EVs, examine the different factors considered in the literature for modeling EV load patterns, and review modeling methods. EV load modeling methods are classified into deterministic, statistical, and machine learning. The article shows that each modeling method has its advantages, disadvantages, and data requirements, ranging from simple load modeling to more accurate models requiring large datasets.

Computer Science↗

Data from: Understanding the biogeochemical and spatial drivers of methane and carbon dioxide fluxes in a large temperate reservoir

This dataset contains spatially resolved measurements of CO₂ and CH₄ fluxes and associated environmental variables collected across 200 sites in Douglas Reservoir (Tennessee, USA) between July 29-August 2, 2024. Measurements include diffusive fluxes of CO₂ and CH₄, CH₄ ebullition, and biogeochemical and spatial variables such as dissolved oxygen, temperature, conductivity, chlorophyll-a, pH, water depth, and distance from the dam. Sampling was conducted using a spatially balanced design to capture longitudinal and depth-related gradients throughout the reservoir. The dataset is structured to support analyses of spatial variability, flux pathway comparisons, and modeling approaches (e.g., spatial statistics and machine learning) aimed at understanding controls on reservoir CO₂ and CH₄ fluxes and improving upscaling to whole-reservoir and regional estimates.

Neeper, Jamie [ORNL] (ORCID:0009000842342101)↗

Data exploration systems for databases

Data exploration systems apply machine learning techniques, multivariate statistical methods, information theory, and database theory to databases to identify significant relationships among the data and summarize information. The result of applying data exploration systems should be a better understanding of the structure of the data and a perspective of the data enabling an analyst to form hypotheses for interpreting the data. This paper argues that data exploration systems need a minimum amount of domain knowledge to guide both the statistical strategy and the interpretation of the resulting patterns discovered by these systems.

Greene, Richard J.↗

Ceramic processing: Experimental design and optimization

The objectives of this paper are to: (1) gain insight into the processing of ceramics and how green processing can affect the properties of ceramics; (2) investigate the technique of slip casting; (3) learn how heat treatment and temperature contribute to density, strength, and effects of under and over firing to ceramic properties; (4) experience some of the problems inherent in testing brittle materials and learn about the statistical nature of the strength of ceramics; (5) investigate orthogonal arrays as tools to examine the effect of many experimental parameters using a minimum number of experiments; (6) recognize appropriate uses for clay based ceramics; and (7) measure several different properties important to ceramic use and optimize them for a given application.

Weiser, Martin W.↗

Using Historical Data to Automatically Identify Air-Traffic Control Behavior

This project seeks to develop statistical-based machine learning models to characterize the types of errors present when using current systems to predict future aircraft states. These models will be data-driven - based on large quantities of historical data. Once these models are developed, they will be used to infer situations in the historical data where an air-traffic controller intervened on an aircraft's route, even when there is no direct recording of this action.

trajectory generation↗

Multispectral Imagery Research and Applications

The NASA Short-term Prediction Research and Transition (SPoRT) Center developed techniques to improve the quality and interpretation of multispectral imagery derived from NASA/NOAA geostationary satellites using statistical and machine learning approaches. A physically-based machine learning approach, DustTracker-AI, was developed to overcome the problem of night-time dust detection and to augment dust analysis with satellite products such as the Dust RGB. Additionally, an updated limb-correction and intercalibration methodology for short-wave, near-infrared, thermal infrared, and water vapor bands was developed for the purpose of developing a suite of high-quality RGB imagery that can be used at high viewing angles and across the constellation of geostationary sensors. This presentation will briefly highlight the techniques developed to detect dust in difficult night-time scenes and improve the quality and interpretation of multispectral imagery.

Emily Berndt↗

Assessing Current and Future Infrastructure Hazards

Project Objective Execute intelligent analytics via an advanced analytical framework, to assess the current state of offshore infrastructure, evaluate infrastructure life, and identify technologies to reduce infrastructure hazards, costs, and extend infrastructure life. Approach &amp; Results thus Far• Build comprehensive dataset• Perform data-driven analytics to evaluate infrastructure integrity<p> 1. Remaining lifespan</p><p> 2. Likelihood of future risk</p><p>• Apply data-driven advanced spatial, statistical, and Machine Learning (ML) models to quantify existing infrastructure integrity</p><p>• Release data and models through a smart, online platform hosted by Energy Data eXchange (EDX)</p>

Romeo, Lucy F.↗

Utility-scale Building Type Assignment Using Smart Meter Data

United States building energy use accounted for 40% of total energy use, 74% of peak demand, and $412 billion in 2019. Building energy modeling allows researchers to simulate building physics, gain insights into possible energy/demand saving opportunities, and assess cost-effective resilience amidst climate change. Many building features needed to create building energy models are readily available such as 2D footprints and LiDAR (height). A critical feature that is not generally obtainable is the building type. In partnership with a utility, a years worth of real-world, 15-minute electrical use data has been examined. The smart meter data is compared to 97 different prototype building energy models to assign building type. Real-world considerations including data preparation, quality assurance, and handling of missing values for advanced metering infrastructure data are addressed. Euclidean distance for pattern-matching of energy use, dynamic time warping, and time-window statistics with machine learning are compared for determining building type from measured electricity use.

Bass, Brett↗

Power Electronics Materials and Bonded Interfaces - Reliability and Lifetime

Advanced packaging technologies are currently being designed and developed by the power electronics industry however, the maximum operating temperature is still limited to 175 degrees Celsius for the silicon carbide devices. Bonded materials such as sintered copper and polymeric materials are potential candidates for high temperature operation, but it is critical to characterize and evaluate its reliability under harsh operating conditions. In this project, we discuss the results of the accelerated experiments conducted on sintered copper and polymeric materials. Additionally, a novel framework to develop the lifetime prediction model of bonded interfaces through employing statistical and machine learning models are described. In this task, scanning acoustic microscope images of bonded interfaces obtained under thermal cycling experiments are used as the data.

ENGINEERING↗

A Causal Approach to Model Validation and Calibration

This poster presents a novel method for validation and verification that focuses on identifying causal relationships between data elements, moving beyond traditional statistical and machine learning approaches. These methods employ causal discovery techniques to reveal the underlying mechanisms of data generation. The research utilizes structural causal models and directed acyclic graphs to depict causal relationships. This approach assists in achieving alignment between simulation models and reality.

97 MATHEMATICS AND COMPUTING↗

A Computational Workflow of Elucidating Viral Impact on Mediating Microbial Response to In-situ Experimental Warming: Bridging microbial modeling to carbon and mineral modeling

Viruses are abundant in soils and shape microbial communities in ways that can potentially influence ecosystem processes, yet their contributions to carbon cycling and mineral transformations remain poorly understood. Here we present a multi-phase framework that links virus-host interactions to soil biogeochemistry by combining ecological simulations, genome- and community-scale metabolic modeling, and statistical and machine-learning analyses. We first calibrated microbial abundance profiles under explicit infection scenarios to capture how viral pressure alters community structure, then explored alternative interaction strategies, including kill-the-winner, piggyback-the-winner, and mixed lytic-lysogenic modes, through forward simulations. These ecological shifts were translated into metabolic consequences using exchange fluxes summarized into biologically meaningful categories, while integrated statistical and machine-learning screens elevated subtle but consistent signals. Application of this framework revealed that viral infections shift the balance between organic and inorganic fluxes, redirecting metabolism from diffuse organic transformations toward inorganic pools such as protons and CO 2 , directly linking viral regulation to respiration and soil carbon balance. The roll-up analysis also isolated perturbations in critical mineral ions, including magnesium, manganese, zinc, and copper, which serve as essential enzymatic cofactors. In piggyback-the-winner scenarios, uptake of these ions was strongly suppressed. Contrasting viral strategies produced distinct community structures and metabolic outcomes, from broad suppression under kill-the-winner dynamics to dramatic redistributions under high-lytic and high-gain lysogenic regimes that collapsed vulnerable microbial populations while promoting opportunists. Together, these results provide a tractable path to trace viral perturbations from host abundance shifts to metabolic flux adjustments and ecosystem-scale processes, offering a practical way to include viruses in earth system models.

54 ENVIRONMENTAL SCIENCES↗

Next-Cycle Optimal Dilute Combustion Control via Online Learning of Cycle-to-Cycle Variability Using Kernel Density Estimators

Dilute combustion using exhaust gas recirculation (EGR) presents a cost-effective method for increasing the efficiency of spark-ignition (SI) engines. However, the maximum amount of EGR that can be used at a given condition is limited by a rapid increment of cycle-to-cycle variability (CCV). This study describes a methodology to design a model-based stochastic optimal controller to adjust the cycle-to-cycle fuel injection quantity in order to reduce CCV and further extend the dilute limit. Given the complexity and chaotic nature of combustion events, the controller was enhanced with online learning in order to identify the statistical properties of combustion efficiency, which are needed to generate predictions for next-cycle events. This study showed that a kernel density estimator (KDE) can be used to learn the combustion properties in real time and can be incorporated into the feedback policy in order to calculate the optimal control command. Experimental results suggested that the dilute limit can be extended from 18.5% to 21% EGR fraction at an operating condition relevant for highway cruising. Additionally, the proposed controller can achieve a large CCV reduction with less fuel enrichment compared to previous methods, overall contributing to an increase in 0.2% indicated fuel conversion efficiency.

33 ADVANCED PROPULSION SYSTEMS↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Scalability Testing Approach for Internet of Things for Manufacturing SQL and NoSQL Database Latency and Throughput

The proliferation of low-cost sensors and industrial data solutions has continued to push the frontier of manufacturing technology. Machine learning and other advanced statistical techniques stand to provide tremendous advantages in production capabilities, optimization, monitoring, and efficiency. The tremendous volume of data gathered continues to grow, and the methods for storing the data are critical underpinnings for advancing manufacturing technology. This work aims to investigate the ramifications and design tradeoffs within a decoupled architecture of two prominent database management systems (DBMS): sql and NoSQL. A representative comparison is carried out with Amazon Web Services (AWS) DynamoDB and AWS Aurora MySQL. The technologies and accompanying design constraints are investigated, and a side-by-side comparison is carried out through high-fidelity industrial data simulated load tests using metrics from a major US manufacturer. The results support the use of simulated client load testing for comparing the latency of database management systems as a system scales up from the prototype stage into production. As a result of complex query support, MySQL is favored for higher-order insights, while NoSQL can reduce system latency for known access patterns at the expense of integrated query flexibility. Here, by reviewing this work, a manufacturer can observe that the use of high-fidelity load testing can reveal tradeoffs in IoTfM write/ingestion performance in terms of latency that are not observable through prototype-scale testing of commercially available cloud DB solutions.

AWS↗

A superconducting quantum simulator based on a photonic-bandgap metamaterial

Synthesizing many-body quantum systems with various ranges of interactions facilitates the study of quantum chaotic dynamics. Such extended interaction range can be enabled by using nonlocal degrees of freedom such as photonic modes in an otherwise locally connected structure. Here, we present a superconducting quantum simulator in which qubits are connected through an extensible photonic-bandgap metamaterial, thus realizing a one-dimensional Bose-Hubbard model with tunable hopping range and on-site interaction. Using individual site control and readout, we characterize the statistics of measurement outcomes from many-body quench dynamics, which enables in situ Hamiltonian learning. Further, the outcome statistics reveal the effect of increased hopping range, showing the predicted crossover from integrability to ergodicity. Our work enables the study of emergent randomness from chaotic many-body evolution and, more broadly, expands the accessible Hamiltonians for quantum simulation using superconducting circuits.

Science & Technology - Other Topics↗

Bayesian Statistics and Uncertainty Quantification for Safety Boundary Analysis in Complex Systems

The analysis of a safety-critical system often requires detailed knowledge of safe regions and their highdimensional non-linear boundaries. We present a statistical approach to iteratively detect and characterize the boundaries, which are provided as parameterized shape candidates. Using methods from uncertainty quantification and active learning, we incrementally construct a statistical model from only few simulation runs and obtain statistically sound estimates of the shape parameters for safety boundaries.

Active Learning↗