Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Scalability analysis of heavy-duty gas turbines using data-driven machine learning

With the increasing integration of variable renewable energy sources into power systems, the role of flexible power generation technologies like gas turbines (GT) in rapid grid balancing remains crucial. This sustained importance underscores the need for scaled and precise modeling of GT to ensure effective integration within evolving energy frameworks. While physics-driven GT models integrate thermodynamics, fluid dynamics, and combustion principles, they often rely on approximate mathematical representations to accommodate scaling that may not capture the actual complex dynamics for GTs and inertial effects associated to GTs with different ratings. In this study, a data-driven model is proposed using machine learning (ML) techniques to conduct GT scalability analysis and performance evaluation with high accuracy. The ML model, trained on data from various operating conditions and performance parameters, aims to uncover intricate relationships and patterns, resembling GT characteristics at different scales (ratings). The model is developed to capture complex system interaction and to adapt to changing operational scenarios at different capacities, providing valuable insights of power system dynamics. In this study, the real-time digital simulator platform was employed to generate training data for the ML model and assess its dynamic characteristics. The ultimate objective was to develop a detailed modeling framework based on governing equations and data-driven ML capable of predicting key performance indicators, in thermal systems such as GTs, including power output, speed, fuel consumption, and exhaust temperature under diverse operating conditions at different scales. The developed ML framework demonstrated high accuracy, with mean relative errors for GT power prediction, reference speed, exhaust temperature, and compressor pressure ratio (CPR) parameters consistently below 0.1% across typical load fluctuation scenarios. Maximum deviations were limited to approximately 0.5 K for exhaust temperature and 0.009 for CPR, underscoring the model’s ability to replicating dynamic GT behavior with high precision. The adaptability of the ML model enables its application across diverse operational conditions and its extension to other thermal systems. By leveraging advanced ML techniques, this study presents a robust and scalable modeling framework that enhances GT simulation precision, facilitating improved integration into evolving power systems.

24 POWER TRANSMISSION AND DISTRIBUTION

Aggregate data‐driven dynamic modeling of active distribution networks with DERs for voltage stability studies

Abstract Electric distribution networks increasingly host distributed energy resources based on power electronic converter (PEC) toward active distribution networks (ADN). Despite advances in computational capabilities, electromagnetic transient models are limited in scalability because of their reliance on exact data about the distribution system and each of its components. Similarly, the use of the DER_A model, which is intended to examine the combined dynamic behavior of many DERs, is limited by the difficulty in parameterization. There is a need for improved dynamic models of DERs for use in large power system simulations for stability analysis. This paper proposes an aggregate model‐free, data‐driven approach for deriving a dynamic partitioned model (DPM) of ADNs. Detailed residential distribution feeders were first developed, including PEC‐based DERs and composite load models (CMLDs), from which the aggregated DPM was derived. The performance was evaluated through various case studies and validated against the detailed ADN model and state‐of‐the‐art DER_A model with CMLD. The data‐driven DPM achieved a of over 90%, accurately representing the aggregated dynamic behavior of ADNs. Furthermore, the DPM significantly accelerated the simulation process with a computational speedup of 68 times compared to the detailed ADN and a 3.5 times speedup compared to the DER_A CMLD model.

42 ENGINEERING

The Double-edged Sword of Data-driven Super-Resolution: Adversarial Super-resolution Models

Data-driven super-resolution (SR) methods are often integrated into imaging pipelines as preprocessing steps to improve downstream tasks such as classification and detection. However, these SR models introduce a previously unexplored attack surface into imaging pipelines. In this paper, we present AdvSR, a framework demonstrating that adversarial behavior can be embedded directly into SR model weights during training, requiring no access to inputs at inference time. Unlike prior attacks that perturb inputs or rely on backdoor triggers, AdvSR operates entirely at the model level. By jointly optimizing for reconstruction quality and targeted adversarial outcomes, AdvSR produces models that appear benign under standard image quality metrics while inducing downstream misclassification. We evaluate AdvSR on three SR architectures (SRCNN, EDSR, SwinIR) paired with a YOLOv11 classifier and demonstrate that AdvSR models can achieve high attack success rates with minimal quality degradation. These findings highlight a new model-level threat for imaging pipelines, with implications for how practitioners source and validate models in safety-critical applications.

Sullivan, Haley [ORNL] (ORCID:0000000274069217)

Data-Driven Model for Photovoltaic Generation: Comparison with Physical Models Using a Microgrid in Puerto Rico

Photovoltaic (PV) generation is a critical component of microgrids, but its accurate modeling is challenging due to the complex and dynamic interactions between solar irradiance, temperature, and PV system installation. This paper develops a multilayer perceptron (MLP) model that inputs solar irradiance and temperature to estimate the PV generation, and it compares the proposed data-driven model’s performance to two well-known physical models: the single-diode model and the inverter model. The results demonstrate that all the models can reach high levels of accuracy. However, the MLP model outperforms the physical models on average by 4.5 to 6.6 percent in R squared scores and 220 to 290 Watts in RMSE scores, and it does not require physical system parameters. Moreover, the data-driven model can overcome the limitations of the lack of real-time PV generation data.

R pesante colón, Marcos

Data-driven upper bounds and event attribution for unprecedented heatwaves

The last decade has seen numerous record-shattering heatwaves in all corners of the globe. In the aftermath of these devastating events, there is interest in identifying worst-case thresholds or upper bounds that quantify just how hot temperatures can become. Generalized Extreme Value theory provides a data-driven estimate of extreme thresholds; however, upper bounds may be exceeded by future events, which undermines attribution and planning for heatwave impacts. Here, we show how the occurrence and relative probability of observed yet unprecedented events that exceed a priori upper bound estimates, so-called “impossible” temperatures, has changed over time. We find that many unprecedented events are actually within data-driven upper bounds, but only when using modern spatial statistical methods. Furthermore, there are clear connections between anthropogenic forcing and the “impossibility” of the most extreme temperatures. Robust understanding of heatwave thresholds provides critical information about future record-breaking events and how their extremity relates to historical measurements.

54 ENVIRONMENTAL SCIENCES

On the Prediction of Aerosol-Cloud Interactions Within a Data-Driven Framework

Aerosol-cloud interactions (ACI) pose the largest uncertainty for climate projection. Among many challenges of understanding ACI, the question of whether ACI can be deterministically predicted has not been explicitly answered. Here we attempt to answer this question by predicting cloud droplet number concentration N c from aerosol number concentration N a and ambient conditions using a data-driven framework. We use aerosol properties, vertical velocity fluctuations, and meteorological states from the ACTIVATE field observations (2020–2022) as predictors to estimate N c . We show that the campaign-wide N c can be successfully predicted using machine learning models despite the strongly nonlinear and multi-scale nature of ACI. However, the observation-trained machine learning model fails to predict N c in individual cases while it successfully predicts N c of randomly selected data points that cover a broad spatiotemporal scale. This suggests that, within a data-driven framework, the N c prediction is uncertain at fine spatiotemporal scales.

54 ENVIRONMENTAL SCIENCES

Data-Driven Mapping of the Cesium Cadmium Bromide Phase Space Utilizing a Soft-Chemistry Approach

Soft-chemistry techniques provide a versatile approach to synthesizing inorganic materials under mild conditions, enabling access to compositions and structures that are challenging to achieve through traditional thermodynamically driven solid-state methods. However, these solution-based routes often result in phase competition, requiring precise control over reaction conditions to achieve selective product formation. While one-variable-at-a-time (OVAT) approaches have traditionally been used for phase selection, data-driven strategies are emerging as more efficient methods for navigating complex synthetic spaces. Ternary metal halides, such as cesium cadmium bromides (Cs–Cd–Br), are of growing interest due to their potential in wide and ultrawide band gap applications. Unlike the well-studied cesium lead halide phases, the compositional diversity and solution-based synthesis of ternary Cs–Cd–Br phases remain largely unexplored. This study systematically investigates the synthetic phase space of the Cs–Cd–Br system by constructing a data-driven phase map. Using a common set of precursors and a standardized experimental procedure, we successfully synthesize all four known Cs–Cd–Br phases—CsCdBr 3 , Cs 2 CdBr 4 , Cs 3 CdBr 5 , and Cs 7 Cd 3 Br 13 —each exhibiting distinct structures, morphologies, and optical properties. Our findings highlight the potential of soft-chemistry methods for expanding the library of ternary metal halides and provide key insights into the thermodynamic and kinetic factors governing phase formation.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Data-driven global ocean modeling for seasonal to decadal prediction

Accurate modeling of ocean dynamics is crucial for enhancing our understanding of complex ocean circulation processes, predicting climate variability, and tackling challenges posed by climate change. Although great efforts have been made to improve traditional numerical models, predicting global ocean variability over multiyear scales remains challenging. Here, we propose ORCA-DL (Oceanic Reliable foreCAst via Deep Learning), a data-driven three-dimensional ocean model for seasonal to decadal prediction of global ocean dynamics. ORCA-DL accurately simulates the three-dimensional structure of global ocean dynamics with high physical consistency and outperforms state-of-the-art numerical models in capturing extreme events, including El Niño–Southern Oscillation and upper ocean heat waves. Moreover, ORCA-DL stably emulates ocean dynamics at decadal timescales, demonstrating its potential even for skillful decadal predictions and climate projections. Our results demonstrate the high potential of data-driven models for providing efficient and accurate global ocean modeling and prediction.

Science & Technology - Other Topics

A Data-Driven Framework for Predicting the Sorting and Screening Performance of an Integrated Biomass Feedstock Preprocessing System

The characteristics of mechanically sorted and screened lignocellulosic biomass, such as the mass contents of corn stover anatomical fractions (leaves, husks, stalks, cobs, etc.), can be used to calculate the intermediate feedstock quality attributes “yield” and “purity” that indicate the conversion efficiency of biocrude. No prior study has investigated the correlations from the characteristics of raw biomass and preprocessing unit operation parameters to those intermediate feedstock quality attributes. This work presents a data-driven framework for assessing and predicting the intermediate feedstock quality attributes in an integrated biomass feedstock preprocessing system. Our study used corn stover as a typical type of herbaceous biomass because of its abundance in the U.S. It began with data acquisition of moisture content, particle size distribution, and anatomical fractions of the materials after each unit operation in the system. The objective of this preprocessing system is to minimize husks and leaves and maximizing cobs and stalks by mechanically separating the materials into three streams via disc screen and air separator. Prototype neural network models were then developed to evaluate the feasibility of predicting process outcomes based on measurable parameters. It is found that incorporating physical constraints into these prediction models significantly enhances the accuracy of the predicted yield and purity against the ground truth data. The experimental data and model predictions indicate that decreasing throughput increases purity, while higher throughput results in lower purity. Finally, an optimization problem was introduced to search optimal combinations of feed material properties and preprocessing unit operation parameters, as the intermediate feedstock quality attributes – yield and purity, appeared to be competing factors. The study also suggests the continual need to improve the data-driven framework’s predictability by incorporating more accurate physical models to describe the dynamics in the preprocessing units such as the air separator.

09 - BIOMASS FUELS

Data-driven multi-element substitution of TiFe alloys for tunable thermodynamics and enhanced activation behaviour for hydrogen storage

Due to their high volumetric hydrogen storage capacity under moderate storage conditions, TiFe alloys have been widely investigated as candidates for practical solid-state hydrogen storage. Partially substituting Ti or Fe sites can improve the key characteristics of TiFe alloys, such as the first hydrogen absorption step (activation) and the equilibrium hydrogen pressure (thermodynamic properties). However, the selection of substitution elements has heavily relied on intuition and trial-and-error. Also, conventional substitution strategies have mainly focused on single-element substitution within the TiFe alloy, limiting the design space and tunability for target applications. Here, to address this limitation, we report a multi-element substitution strategy motivated by an efficient, data-driven machine learning (ML) approach combined with corroborating density functional theory (DFT) calculations. Our models successfully predict experimentally measured hydride stability in five selected alloys using only compositional descriptors. Most importantly, the multi-element substitution leads to enhanced activation properties compared to pure TiFe, achieving near room-temperature activation behaviour. This work provides a method for on-demand tuning of hydrogen storage and activation properties, which may have broad implications for data-driven discovery of energy storage materials.

Cho, YongJun [Korea Advanced Institute Science and

Solar Forecasting, Net Load Forecasting, and Data-Driven Distributed Solar Visibility Prizes (Final Technical Report)

The American-Made Solar Forecasting Prize, Net Load Forecasting Prize, and Data-Driven Distribution (3D) Solar Visibility Prize is a multimillion-dollar prize competition designed to energize U.S. solar innovation through a series of contests that accelerate the entrepreneurial process from years to months. The activities incentivized by these three prizes will support the governmentwide approach to increase American energy dominance by promoting innovation and early deployment of energy technologies, resulting in wider adoption, which is critical for secure, affordable, and reliable solar energy.

14 SOLAR ENERGY

Examples of Mission-driven Data Science from Jefferson Lab and ACES

This presentation details mission-driven data science initiatives at Jefferson Lab and the Joint Institute for Advanced Computing on Environmental Studies (ACES). JLab, a U.S. Department of Energy Office of Science national laboratory, operates the Continuous Electron Beam Accelerator Facility (CEBAF), and is the lead institute for the new High Performance Data Facility (HPDF) Hub. The Joint Institute for ACES brings together interdisciplinary teams in health informatics, climate modeling, computer science, and physics to address environmental challenges, including flood modeling. The Hampton Roads region, particularly Norfolk and Virginia Beach, faces increasing flood risks, motivating the need for rapid, reliable, and risk-aware decision support. ACES’s flooding work has a focus on uncertainty quantification (UQ) and machine learning (ML) for coastal flood management. The work is motivated by the increasing vulnerability of communities such as Norfolk and Virginia Beach, Virginia, to frequent coastal flooding events, and the need for rapid, reliable decision support. The research develops computationally efficient ML surrogate models to forecast water levels and flooding risk. A central theme is the quantification and calibration of predictive uncertainty, especially for out-of-distribution (OOD) scenarios, using techniques such as Monte Carlo Dropout, Deep Ensembles, Gaussian Processes, and Deep Quantile Regression (DQR). The study demonstrates that distance-aware UQ is critical for reliable scientific AI, particularly in high-dimensional, safety-critical, and real-time applications.

McSpadden, Diana [Thomas Jefferson National Accele

Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”

This page contains the datasets and code #O5097 Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”. Data literature-derived cultivation experiments for polyhydroxybutyrate (PHB) production in Synechocystis sp. PCC 6803 and were used for ML model development, interpretation, and experimental validation.

Lalonde, Jessica N. [Los Alamos National Laborator

ZENN: A thermodynamics-inspired computational framework for heterogeneous data–driven modeling

Traditional entropy-based methods—such as cross-entropy loss in classification problems—have long been essential tools for representing the information uncertainty and physical disorder in data and for developing artificial intelligence algorithms. However, the rapid growth of data across various domains has introduced new challenges, particularly the integration of heterogeneous datasets with intrinsic disparities. To address this, we introduce a zentropy-enhanced neural network (ZENN), extending zentropy theory into the data science domain via intrinsic entropy, enabling more effective learning from heterogeneous data sources. ZENN simultaneously learns both energy and intrinsic entropy components, capturing the underlying structure of multisource data. To support this, we redesign the neural network architecture to better reflect the intrinsic properties and variability inherent in diverse datasets. We demonstrate the effectiveness of ZENN on classification tasks and energy landscape reconstructions, showing its superior generalization capabilities and robustness-particularly in predicting high-order derivatives. In image and text classification tasks, ZENN demonstrates superior generalization by introducing a learnable temperature variable that models latent multisource heterogeneity, allowing it to surpass state-of-the-art models on CIFAR-10/100, BBC News, and AG News. As a practical application in materials science, we employ ZENN to reconstruct the Helmholtz energy landscape of Fe3Pt using data generated from density functional theory and capture key material behaviors, including negative thermal expansion and the critical point in the temperature–pressure space. Overall, this work presents a zentropy-grounded framework for data-driven machine learning, positioning ZENN as a versatile and robust approach for scientific problems involving complex, heterogeneous datasets.

36 MATERIALS SCIENCE

Next-Generation Materials Design: Quantum Mechanics and Data-Driven Modeling

The future of materials design is rapidly advancing through the combination of quantum mechanics and data-driven modeling. These approaches integrate quantum principles with advanced data analysis, enabling precise insights into material behavior. This talk will highlight recent progress in using these methods for computational design, particularly in high-entropy alloy catalysts, emphasizing the role of hierarchical machine-learning architectures for accurate predictions. Additionally, I will discuss our work on developing machine learning interatomic potentials (MLPs) for single-element metals, metal oxides, and alloys under extreme conditions, focusing on melting behavior and phase properties at high temperatures and pressures. We have also refined our MLP models to capture dynamic surface interactions, such as CO2 and CO adsorption on MgO, using both static and molecular dynamics simulations. These models maintain high accuracy while significantly reducing computational costs compared to first-principles calculations. By enabling efficient and accurate simulations, this work supports broader community adoption, optimizes datasets for materials discovery, and extends the accessible time, size, and environmental conditions beyond the limits of experiments and traditional simulations.

machine learning

Developing an oxidation materials ontology for data-driven materials design

Materials data is complex, and managing and storing materials data for use and reuse is a common challenge. An ontology-based data management framework can address these challenges through encoding data attributes and relationships in a flexible way. This presentation discusses the creation of an ontology for alloy oxidation test data and reviews the logic, structure and interoperability of the ontology.

advanced alloy development

Data-Driven Closures and Assimilation for Stiff Multiscale Random Dynamics

Here, we introduce a data-driven and physics-informed framework for propagating uncertainty in stiff, multiscale random ordinary differential equations (RODEs) driven by correlated (colored) noise. Unlike systems subjected to Gaussian white noise, a deterministic equation for the joint probability density function (PDF) of RODE state variables does not exist in closed form. Moreover, such an equation would require as many phase-space variables as there are states in the RODE system. To alleviate this curse of dimensionality, we instead derive exact, albeit unclosed, reduced-order PDF (RoPDF) equations for low-dimensional observables/quantities of interest. The unclosed terms take the form of state-dependent conditional expectations, which are directly estimated from data at sparse observation times. However, for systems exhibiting stiff, multiscale dynamics, data sparsity introduces regression discrepancies that compound during RoPDF evolution. This is overcome by introducing a kinetic-like defect term to the RoPDF equation, which is learned by assimilating in sparse, low-fidelity RoPDF estimates. Two assimilation methods are considered, namely nudging and deep neural networks, which are successfully tested against Monte Carlo simulations.

97 MATHEMATICS AND COMPUTING