Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data-driven modeling”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Roadmap and Benchmarking: Privacy in Federated Load Forecasting

Data-driven techniques for energy demand forecasting continue to emerge with promising impacts on distribution grid planning. However, the development of robust and generalizable machine learning models requires that representative high quality training data are available. Distributed energy resources have begun to embed intelligence, gathering large amounts of data on customer demand, behavior, and household devices that are connected to the grid. Though utilities aggregate meter-level demand data for load shaping, demand response, outage management, reliability planning, and billing applications, there lies an inherent privacy concern in sharing consumption data that may identify individual consumer behavioral patterns. Hence, while sharing the data is crucial, the private sensitive customer data must be safeguarded from being exposed or manipulated. In this study, we propose a roadmap for implementing a based privacy preserving framework to support the advancement of data-driven analytics in data-sensitive distributed energy resources environments. The roadmap incorporates federated learning–a distributed training framework, differential privacy–a statistical framework that provides guarantees to safeguard the leakage of sensitive data, secure multiparty computation and homomorphic encryption– techniques for encrypting model gradients and applying secure aggregation on the server. Moreover, we perform baseline experiments on the federated short-term load forecasting (STLF) task using open-source residential load profile datasets, offering insights into the challenges of integrating differential privacy into federated learning.

Abebe, Waqwoya [Oak Ridge National Laboratory (ORN

Detecting thermodynamic phase transition via explainable machine learning of photoemission spectroscopy

Identifying thermodynamic signatures of electronic phases, such as superconductivity, is challenging in low-dimensional materials due to strong fluctuations and low probing volume. Spectroscopic methods are often used to identify new bulk phases, but their main measurable quantity—electronic energy gaps—is no longer an effective order parameter in low-dimensional and fluctuating systems. Combining angle-resolved photoemission with a domain-adversarial neural network, we report a data-driven method to identify thermodynamic phase transitions solely based on single-particle spectra. We demonstrate 97.6% accuracy in cuprate superconductor Bi 2 Sr 2 CaCu 2 O 8+δ with strong superconducting fluctuations. This model notably compensates for the scarcity of experimental data by leveraging virtually inexhaustible simulated data. Further, its explainability reveals the crucial role of in-gap spectral weight in detecting phase fluctuations and thermodynamic transitions. Our work pinpoints the spectroscopic signatures of fluctuating orders and enables using spectroscopy for machine-learning-assisted material discovery for low-dimensional and strong coupling systems.

2D materials

Carbon Utilization and Storage Partnership of the Western United States

This technical report documents research conducted under DOE Award No. DE-FE0031837 focused on evaluating the feasibility of carbon capture, utilization, and storage (CCUS) systems in the central and western United States. The project integrated geologic characterization, reservoir simulation, infrastructure modeling, and economic analysis to assess CO₂ storage potential near industrial sources and develop strategies for transport and sequestration. The work included subsurface modeling, risk assessment, monitoring and verification (MRV) planning, and evaluation of regulatory pathways such as EPA Underground Injection Control (UIC) Class VI permitting and IRS 45Q tax credit eligibility. Results demonstrate the viability of multiple storage approaches, including saline formations, enhanced coalbed methane recovery, and basalt mineralization, supported by data-driven workflows and regional analyses. The project also produced permitting templates, technology transfer activities, and stakeholder engagement efforts to support deployment readiness. These findings contribute to the development of scalable, economically viable CCUS systems and provide a repeatable framework for future carbon management projects.

20 FOSSIL-FUELED POWER PLANTS

State, Local, and Tribal Program

NLR's State, Local, and Tribal Program delivers customized, data-driven support that strengthens local energy systems - expanding access to America's abundant energy resources, reducing costs, and supporting energy reliability across the country. NLR's world-class staff use a wide variety of cutting-edge energy tools and capabilities to deliver robust modeling, validation, and deployment support to hundreds of communities annually.

29 ENERGY PLANNING, POLICY, AND ECONOMY

General Purpose Data-Driven System Monitoring for Space Operations

Modern space propulsion and exploration system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. The Inductive Monitoring System (IMS) is a data-driven system health monitoring software tool that has been successfully applied to several aerospace applications. IMS uses a data mining technique called clustering to analyze archived system data and characterize normal interactions between parameters. This characterization, or model, of nominal operation is stored in a knowledge base that can be used for real-time system monitoring or for analysis of archived events. Ongoing and developing IMS space operations applications include International Space Station flight control, satellite vehicle system health management, launch vehicle ground operations, and fleet supportability. As a common thread of discussion this paper will employ the evolution of the IMS data-driven technique as related to several Integrated Systems Health Management (ISHM) elements. Thematically, the projects listed will be used as case studies. The maturation of IMS via projects where it has been deployed, or is currently being integrated to aid in fault detection will be described. The paper will also explain how IMS can be used to complement a suite of other ISHM tools, providing initial fault detection support for diagnosis and recovery.

Satellites

General Purpose Data-Driven System Monitoring for Space Operations

Modern space propulsion and exploration system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. The Inductive Monitoring System (IMS) is a data-driven system health monitoring software tool that has been successfully applied to several aerospace applications. IMS uses a data mining technique called clustering to analyze archived system data and characterize normal interactions between parameters. This characterization, or model, of nominal operation is stored in a knowledge base that can be used for real-time system monitoring or for analysis of archived events. Ongoing and developing IMS space operations applications include International Space Station flight control, spacecraft vehicle system health management, launch vehicle ground operations, and fleet supportability. As a common thread of discussion this paper will employ the evolution of the IMS data-driven technique as related to several Integrated Systems Health Management (ISHM) elements. Thematically, the projects listed will be used as case studies. The maturation of IMS via projects where it has been deployed or is currently being integrated to aid in fault detection will be described. The paper will also explain how IMS can be used to complement a suite of other ISHM tools, providing initial fault detection support for diagnosis and recovery

Space Propulsion

First Measurement of Sub-GeV nu_mu Charged-Current Coherent Pion Production on Argon in MicroBooNE

Coherent pion production, characterized by a neutrino interacting with an entire nucleus without breaking it apart, results in a forward-going muon, pion, and a low-momentum recoil nucleus. This process provides a sensitive probe of neutrino-nucleus interactions and offers a potential standard candle for neutrino-flux normalization in neutrino-oscillation experiments. We present the first measurement of the flux-averaged charged-current coherent pion production cross section on argon nucleus using the MicroBooNE liquid argon time projection chamber. The analysis employs particle identification together with a data-driven background parameterization to isolate the coherent signal. This measurement uses the full MicroBooNE dataset collected from the Fermilab Booster Neutrino Beam, corresponding to an exposure of 1.26E10^21 protons on target and an average neutrino energy of approximately 0.8 GeV. The result provides the first constraint on charged-current coherent pion production on argon nucleus at sub-GeV energies and supplies important input for improving neutrino interaction modeling in current and future experiments such as DUNE.

Hussain, Adil [Kansas State U.] (ORCID:00000001621

AutoBEM: A scalable framework for nationwide building energy simulation and retrofit evaluation in the United States

This paper presents AutoBEM, an integrated, automated framework for nationwide building energy modeling and retrofit evaluation in the United States. Unlike prior UBEM platforms that either rely primarily on representative stock sampling or operate at city scale, AutoBEM automates the generation of building-resolved, physics-based EnergyPlus/OpenStudio simulation models at national scale using GIS-derived geometry, prototype-based assumptions, and standardized scalable workflows. Leveraging the Model America dataset and high-performance computing, AutoBEM generates and simulates energy models for 122.9 million buildings, representing 97.8% of the U.S. building stock. These models are being made publicly and freely available as the Model America v1.0 (MAv1) dataset. AutoBEM supports detailed, building-level assessments of energy consumption, CO2 emissions, and post-processed anthropogenic heat emissions (AHE), and evaluates 151 energy conservation measures (ECMs) using localized utility pricing and building characteristics. In addition, AutoBEM incorporates both typical and future climate conditions through integration with Typical Meteorological Year (TMY) and Future TMY (fTMY) weather data derived from IPCC scenarios. In a case study of Phoenix, Arizona, AutoBEM identified several high-efficiency HVAC upgrades and selected envelope measures with short modeled payback periods (1.5 years) for certain building types and standards. Simulations under future climate scenarios (SSP5–RCP8.5) project an 11.3% increase in electricity use and a 32% reduction in natural gas demand by 2100, underscoring the need for climate-adaptive retrofit planning. By enabling reproducible, bottom-up, and location-specific analysis at scale, AutoBEM provides a step toward a national digital twin of the built environment and supports data-driven screening and planning for decarbonization, resilience, and energy equity.

Li, Hang [ORNL] (ORCID:0000000306001920)

End-To-End Decentralized Transmission Line Protection in IBR-Dominated Weak Grids Using Interpretable Data-Driven Methods

Traditional transmission line protection relies on predictable synchronous-based fault signatures, which frequently fail under the non-standard, current-limited fault characteristics of Inverter-Based Resources (IBRs). This study investigates how to achieve secure, communication-free fault isolation in IBR-dominated weak grids without relying on opaque, computationally heavy "black-box" machine learning algorithms. To address this, we propose a novel, standalone, and inherently interpretable data-driven protection framework. Unlike centralized methods requiring multi-terminal communication, this decentralized approach relies solely on local measurements using a hierarchical linear-kernel Support Vector Machine (SVM). The methodology decomposes the protection task into four sequential stages that mimic traditional protection elements: fault detection and fault direction identification, fault type classification, zone classification, and location estimation. This multi-stage architecture allows for specialized feature engineering at each stage, combining high computational efficiency with logic traceability. The framework's end-to-end performance was validated via C-code and PSCAD/EMTDC co-simulation, utilizing a real-world utility network and an OEM black-box IBR model. The proposed relay achieves 97.2% overall accuracy and provides a reliable trip decision within a 2.5-cycle window. The results confirm 100% accuracy in fundamental fault detection, reliable zone selectivity across low to moderate fault resistances, and robust security against non-fault transients, proving its immediate viability for integration into commercial numerical relays.

24 POWER TRANSMISSION AND DISTRIBUTION

Instabilities and phase transitions in architected metamaterials: a gradient-enhanced continuum approach

Architected metamaterials such as foams and lattices exhibit a wide range of properties governed by microstructural instabilities and emerging phase transitions. Their macroscopic response–including energy dissipation during impact, large recoverable deformations, morphing between configurations, and auxetic behavior–remains difficult to capture with conventional continuum models, which often rely on discrete approaches that limit scalability. In this work, we propose a nonlocal continuum formulation that captures both stable and unstable responses of elastic architected metamaterials. The framework extends anisotropic hyperelasticity by introducing nonlocal variables and internal length scales reflective of microstructural features. Local polyconvex free-energy models are systematically augmented with two families of non-(poly)convex energies, enabling both metastable and bistable responses. Implementation in a finite element framework enables solution using a hybrid monolithic–staggered strategy. Simulations capture densification fronts, forward and reverse transitions, hysteresis loops, imperfection sensitivity, and globally coordinated auxetic modes. Overall, this framework provides a robust foundation for accelerated modeling of instability-driven phenomena in architected metamaterials, while enabling extensions to anisotropic, dissipative, and active systems as well as integration with data-driven and machine learning approaches.

42 ENGINEERING

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization

Reconstructing Quasar Spectra and Measuring the Lyα Forest with SpenderQ

Quasar spectra carry the imprint of foreground intergalactic medium (IGM) through absorption features. In particular, absorption caused by neutral hydrogen gas, the "Lyα forest," is a key spectroscopic tracer for cosmological analyses used to measure cosmic expansion and test physics beyond the standard model. Despite their importance, current methods for measuring Lyα absorption cannot directly derive the intrinsic quasar continuum and make strong assumptions on its shape, thus distorting the measured Lyα clustering. We present SpenderQ , a ML-based approach for directly reconstructing the intrinsic quasar spectra and measuring the Lyα forest from observations. SpenderQ uses the Spender spectrum autoencoder to learn a compact and redshift-invariant latent encoding of quasar spectra, combined with an iterative procedure to identify and mask absorption regions. To demonstrate its performance, we apply SpenderQ to 400,000 synthetic quasar spectra created to validate the Dark Energy Spectroscopic Instrument Year 1 Lyα cosmological analyses. SpenderQ accurately reconstructs the true intrinsic quasar spectra, including the broad Lyβ, Lyα, SiIV, CIV, and CIII emission lines. Redward of Lyα, SpenderQ provides percent-level reconstructions of the true quasar spectra. Blueward of Lyα, SpenderQ reconstructs the true spectra to < 5%. SpenderQ reproduces the shapes of individual quasar spectra more robustly than the current state-of-the-art. We, thus, expect it will significantly reduce biases in Lyα clustering measurements and enable studies of quasars and their physical properties. SpenderQ also provides informative latent variable encodings that can be used to, e.g., classify quasars with Broad Absorption Lines. Overall, SpenderQ provides a new data-driven approach for unbiased Lyα forest measurements in cosmological, quasar, and IGM studies.

Hahn, ChangHoon [Arizona U., Astron. Dept. - Stewa

Hybrid Data‐Driven Discovery of High‐Performance Silver Selenide‐Based Thermoelectric Composites

Optimizing material compositions often enhances thermoelectric performances. However, the large selection of possible base elements and dopants results in a vast composition design space that is too large to systematically search using solely domain knowledge. To address this challenge, a hybrid data-driven strategy that integrates Bayesian optimization (BO) and Gaussian process regression (GPR) is proposed to optimize the composition of five elements (Ag, Se, S, Cu, and Te) in AgSe-based thermoelectric materials. Data is collected from the literature to provide prior knowledge for the initial GPR model, which is updated by actively collected experimental data during the iteration between BO and experiments. Within seven iterations, the optimized AgSe-based materials prepared using a simple high-throughput ink mixing and blade coating method deliver a high power factor of 2100 µW m −1 K −2 , which is a 75% improvement from the baseline composite (nominal composition of Ag 2 Se 1 ). In conclusion, the success of this study provides opportunities to generalize the demonstrated active machine learning technique to accelerate the development and optimization of a wide range of material systems with reduced experimental trials.

36 MATERIALS SCIENCE

Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”

This page contains the datasets and code #O5097 Jupyter Notebook Code for “Data-Driven Insights to Accelerate Advanced Biomanufacturing”. Data literature-derived cultivation experiments for polyhydroxybutyrate (PHB) production in Synechocystis sp. PCC 6803 and were used for ML model development, interpretation, and experimental validation.

Lalonde, Jessica N. [Los Alamos National Laborator

Flux Cube Reconstruction from Slitless Spectroscopy

Slitless spectroscopy enables efficient, large-area surveys without target preselection, yet it faces challenges from source blending, higher noise, and lost spatial–spectral information. We present an advanced, nonparametric, data-driven algorithm that leverages multiple dispersion angles to reconstruct three-dimensional flux distributions, providing low-resolution integral field unit capabilities from slitless data. By treating each pixel as an independent element, our method naturally handles source confusion without requiring prior assumptions regarding redshifts, templates, or model libraries. We validate the algorithm using simulated Roman Space Telescope wide-field slitless spectroscopy images that are equivalent to what is expected from the High-Latitude Time-Domain Survey. First, we demonstrate that a host-galaxy model reconstructed from multiple dispersion angles can be used to accurately subtract host light from a transient, recovering a Type Ia supernova spectrum with minimal bias. Second, we showcase a high-fidelity flux-cube reconstruction of a complex galaxy, successfully measuring the redshift and recovering continuum, emission, and absorption features. This approach highlights the potential of multi-dispersion-angle slitless data to provide spatially resolved spectral information in a nonparametric way, which is traditionally accessible only with integral field spectroscopy, opening a new window into large, unbiased, and spatially resolved studies of galaxy evolution.

Griggio, M. [Space Telescope Science Institute, Ba

A Probabilistic Approach to Load Modeling for Central HVAC Systems in Large Commercial Buildings for Retrofit Decisions Under Uncertainty

Retrofitting central HVAC systems in large commercial buildings with advanced technologies like heat recovery chillers (HRCs) offers a significant opportunity to enhance energy efficiency. However, analyzing these retrofits is challenging with traditional whole-building simulation tools, which require intensive calibration and struggle to model innovative system configurations and controls. To overcome these limitations, this study proposes a load profilebased retrofit analysis framework that provides better decisions under uncertainty. The main focus of this paper is the development of a probabilistic load profile model that can be used in the framework by using exploratory data analysis (EDA) of measured building data to properly quantify its inherent variability. A non-parametric Gaussian Process (GP) model was employed to capture the time- and weather-dependent characteristics of the heating load while explicitly modeling its uncertainty. The model's effectiveness is demonstrated through strong predictive performance on unseen data and physically interpretable insights into load behavior. This data-driven, probabilistic load profile serves as a robust and flexible input for subsequent system simulations, enabling a more confident and statistically sound analysis of retrofit potential.

Ham, S W

Plant Bioengineering Atlas: A Knowledge Graph of Genes, DNA Constructs, and Plant Traits.

Plant bioengineering has generated tens of thousands of genotype-to-phenotype relationships, but this knowledge remains fragmented across narrative literature and difficult to use computationally. Inconsistent descriptions of DNA constructs, host species, and traits, including variable species names, omitted regulatory elements, and inconsistent gene symbols, impede data reuse, comparative analysis, and design-build-test-learn cycles. Here, we present the Plant Bioengineering Atlas, a literature-mined, ontology-grounded knowledge base assembled using an artificial intelligence (AI)-aided extraction pipeline. A large language model parsed open-access primary research articles to generate structured, provenance-anchored records of engineered genes, modification types, promoter-gene-terminator constructs, host species, target traits, and reported phenotypes, with every record traceable to its source. The current release contains 14,358 curated records encompassing 6,998 distinct genes across 436 plant species from 6,452 papers published between 2000 and 2026. Corpus analysis reveals that experiments are concentrated in a small group of model and crop species, disease and pathogen resistance is the most frequently engineered trait class, and constitutive regulatory parts (particularly the CaMV 35S promoter and NOS terminator) remain pervasive. Two in five records omit one or both flanking regulatory elements (i.e., promoter and terminator), while only 23.4% describe cassettes in which both elements resolve to named part classes, exposing a systematic reproducibility gap. We organize these data into a knowledge graph linking genes, constructs, species, and traits; provide access through an interactive web portal; and propose an AI-compatible documentation standard for AI-ready reporting. The Plant Bioengineering Atlas provides a foundation for data-driven hypothesis generation and AI-aided plant biodesign.

, Genes, DNA Constructs

BEAST: Expanding Sustainable Data Infrastructure for High-Enthalpy Facilities

Reproducible, data-driven thermal protection system (TPS) research requires that experimental records from high-enthalpy testing be consistently structured, traceable, and accessible across campaigns and institutions. In practice, however, arcjet and plasma facilities data remain largely fragmented: raw diagnostics are stored in ad hoc formats, material sample histories are disconnected from test conditions, and metadata standards are absent, precluding systematic cross-campaign analysis and long-term reuse. BEAST (Backend for Experiment Analysis, Storage, and Traceability) is an open-source, web-based platform that addresses these limitations by providing a unified, queryable infrastructure for high-enthalpy ground-test data [1]. First presented at the 15th Ablation Workshop [2], BEAST has since undergone significant development. The platform ingests and structures multi-channel time-series diagnostics, facility configurations, and material property records within a common provenance model, ensuring end-to-end traceability from raw sensor acquisition to reduced experimental quantities. A versioned material library links specimen identity and processing history to the specific runs in which each sample was tested. An integrated modeling workbench enables training and evaluation of regression models directly on archived experimental data, supporting condition interpolation and the construction of empirical material response databases. Beyond its original deployment at NASA Ames Research Center, BEAST has been designed to be facility-agnostic, with ongoing efforts to extend its adoption to other facilities. Its modular architecture accommodates heterogeneous diagnostic setups and facility types, and its future open-source distribution allows institutions to build on a common data standard rather than maintaining isolated, bespoke solutions. BEAST is further integrated within a broader ecosystem of companion tools: arcjetCV [3] extracts recession rates and shock standoff distances from high-speed video using computer vision, and miniSTARscan [4] provides sub-minute, portable photogrammetric surface reconstruction of test articles before and after exposure. All tools share a common data schema, enabling seamless ingestion of surface geometry, imagery, and time-series data into a single, coherent experimental record.

Database