Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hybrid data-driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

A framework for strategic discovery of credible neural network surrogate models under uncertainty

The widespread integration of deep neural networks in developing data-driven surrogate models for high-fidelity simulations of complex physical systems highlights the critical necessity for robust uncertainty quantification techniques and credibility assessment methodologies, ensuring the reliable deployment of surrogate models in consequential decision-making. Here, this study presents the Occam Plausibility Algorithm for surrogate models (OPAL-surrogate), providing a systematic framework to uncover predictive neural network-based surrogate models within the large space of potential models, including various neural network classes and choices of architecture and hyperparameters. The framework is grounded in hierarchical Bayesian inferences and employs model validation tests to evaluate the credibility and prediction reliability of the surrogate models under uncertainty. Leveraging these principles, OPAL-surrogate introduces a systematic and efficient strategy for balancing the trade-off between model complexity, accuracy, and prediction uncertainty. The effectiveness of OPAL-surrogate is demonstrated through two modeling problems, including the deformation of porous materials for building insulation and turbulent combustion flow for ablation of solid fuels within hybrid rocket motors.

42 ENGINEERING↗

Benchmarking Bayesian Optimization Frameworks and Acquisition Strategies for Materials Discovery and Autonomous Laboratories

Bayesian optimization (BO) can accelerate materials discovery by guiding expensive experiments toward the most promising processing conditions. We systematically compare five BO surrogate and framework combinations (Gaussian processes in Ax, Gaussian processes and Monte-Carlo neural networks in BayBE, random forests in Lolopy, and tree-structured Parzen (TPE) estimators in Hyperopt) on three benchmarks that mimic common materials design tasks (a discrete solid-electrolyte composition space, a hybrid discrete/continuous laminate-composite design problem solved with micromechanics modeling, and the continuous Ishigami analytic function which is a standard optimization benchmark). Each BO surrogate is paired with posterior mean, probability of improvement, and expected improvement acquisition functions and run for 100 trials from randomized initial samples with uniform random search providing a control. Across five random seeds per setting, BayBE’s Gaussian-process surrogate with expected improvement consistently reached ≥95 % of the known optimum in the fewest evaluations, while Lolopy’s random forest matched or exceeded GP performance on purely categorical or mixed spaces at a higher computational cost. Posterior mean alone often stagnated at local optima, underscoring the need for exploration, whereas probability and expected improvement balanced exploration and exploitation leading to better optimization in fewer trials. Execution times ranged from milliseconds for TPE to minutes for neural-network and random-forest surrogates. These results establish baseline expectations for BO in automated materials laboratories and highlight expected improvement with Gaussian processes as a reliable first choice, with random forests offering a strong alternative when categorical variables dominate. The benchmark suite and code are released to facilitate future surrogate, acquisition, and constraint-handling research in data-driven materials optimization.

Bayesian optimization↗

Online peak-aware energy scheduling with untrusted advice

This paper studies the online energy scheduling problem in a hybrid model where the cost of energy is proportional to both the volume and peak usage, and where energy can be either locally generated or drawn from the grid. Inspired by recent advances in online algorithms with Machine Learned (ML) advice, we develop parameterized deterministic and randomized algorithms for this problem such that the level of reliance on the advice can be adjusted by a trust parameter. We then analyze the performance of the proposed algorithms using two performance metrics: robustness that measures the competitive ratio as a function of the trust parameter when the advice is inaccurate, and consistency for competitive ratio when the advice is accurate. Since the competitive ratio is analyzed in two different regimes, we further investigate the Pareto optimality of the proposed algorithms. Our results show that the proposed deterministic algorithm is Pareto-optimal, in the sense that no other online deterministic algorithms can dominate the robustness and consistency of our algorithm. Furthermore, we show that the proposed randomized algorithm dominates the Pareto-optimal deterministic algorithm. Our large-scale empirical evaluations using real traces of energy demand, energy prices, and renewable energy generations highlight that the proposed algorithms outperform worst-case optimized algorithms and fully data-driven algorithms.

Lee, Russell↗

Stability-Constrained Learning for Frequency Regulation in Power Grids With Variable Inertia

The increasing penetration of converter-based renewable generation has resulted in faster frequency dynamics, and low and variable inertia. As a result, there is a need for frequency control methods that are able to stabilize a disturbance in the power system at timescales comparable to the fast converter dynamics. This paper proposes a combined linear and neural network controller for inverter-based primary frequency control that is stable at time-varying levels of inertia. We model the time-variance in inertia via a switched affine hybrid system model. We derive stability certificates for the proposed controller via a quadratic candidate Lyapunov function. We test the proposed control on a 12-bus 3-area test network, and compare its performance with a base case linear controller, optimized linear controller, and finite-horizon Linear Quadratic Regulator (LQR). Our proposed controller achieves faster mean settling time and over 50% reduction in average control cost across 100 inertia scenarios compared to the optimized linear controller. Unlike LQR which requires complete knowledge of the inertia trajectories and system dynamics over the entire control time horizon, our proposed controller is real-time tractable, and achieves comparable performance to LQR.

data-driven control↗

Numerical Investigation of Fluid Flow and Space Charge in Liquid Argon Time Projection Chamber (LArTPC) Detectors

Overview This project focused on developing a high-fidelity numerical framework to simulate the multiphysics environment within Liquid Argon Time Projection Chamber (LArTPC) detectors. The primary objective was to characterize the complex interplay between ion transport, background fluid dynamics, and electric field distortions—a critical factor for the calibration and sensitivity of next-generation High Energy Physics experiments, such as DUNE. Technical Achievements The research successfully yielded a hybrid numerical space-charge solver utilizing a Cell-Centered Finite Volume Method (FVM) for ion transport coupled with a Finite Element Method (FEM) for electric potential. Key accomplishments include: • Verification & Validation: The 3-D solver was rigorously verified against 1-D analytical solutions, demonstrating high numerical accuracy in predicting space-charge-induced field deviations. • Field Distortion Analysis: 3D simulations revealed that space charge effects introduce significant non-uniformities in the electric field. Critically, the research identified that background LAr flow velocities, when comparable to ion drift velocities, markedly exacerbate these distortions. • Technology Transfer: The resulting source code and comprehensive user manuals were successfully transferred to collaborators at Fermilab, providing a portable computational tool for the broader scientific community. Challenges and Future Directions While the space-charge solver achieved all performance metrics, the integrated fluid dynamics modeling encountered convergence challenges stemming from the extreme 200-fold disparity in length scales between the detector's 37 mm inlet pipes and the 8-meter global domain. To address this, the project has identified a clear technical pivot toward Hierarchical Geometric Adaptive Mesh Refinement (HG-AMR). By implementing an h-type refinement strategy with hanging nodes, future iterations of this solver will be capable of resolving localized high-gradient inlet flows without the prohibitive computational costs of regular grids. This advancement, combined with data-driven uncertainty quantification based on MicroBooNE-style calibration, will enable the precise modeling of detector responses in large-scale cryogenic environments where direct measurement remains difficult. Impact The computational tools developed under this award provide a foundation for enhancing the energy resolution and spatial reconstruction of noble liquid detectors. By bridging the gap between theoretical fluid dynamics and experimental field calibration, this work supports the DOE’s mission to advance the frontiers of neutrino physics and dark matter detection.

42 ENGINEERING↗

Data-driven analysis and prediction of wastewater treatment plant performance: Insights and forecasting for sustainable operations

Here this study presents a comprehensive performance and forecasting analysis of the As-Samra wastewater treatment plant (WWTP) in Jordan, with two main objectives. Firstly, a thorough evaluation of the plant's performance is conducted. The analysis involves independently assessing historical operational conditions, plant production, and their statistical correlations using various statistical techniques. The second objective focuses on developing a data-driven forecasting approach to predict the plant's production one month in advance, using multiple machine learning models. The results highlight the effectiveness of principal component analysis (PCA) in simplifying operational data, revealing distinct operational clusters, and identifying seasonal production patterns while showing correlations between operational conditions and overall power production. The support vector machine (SVM) forecasting model emerged as the top performer, showcasing the potential of a hybrid forecasting approach. The findings offer valuable perspectives for enhancing operational efficiency, refining production planning, and ultimately improving the environmental impact of the plant.

42 ENGINEERING↗

A systematic feature extraction and selection framework for data-driven whole-building automated fault detection and diagnostics in commercial buildings

In data-driven automated fault detection and diagnostics (AFDD) modeling for building energy systems, feature engineering is a critical process of extracting information from high-dimensional and noisy sensor measurement and turning it into informative and representative inputs or features for data-driven modeling. However, few studies specifically discuss the feature engineering, especially the interactions between feature extraction and feature selection in whole-building AFDD. We developed a systematic feature extraction and selection framework for whole-building AFDD. In this framework, features are aggressively extracted from raw sensor data using statistical feature extraction techniques with various window sizes and statistics. With many features extracted, a hybrid feature selection algorithm that combines the filter and wrapper method then selects the best feature set. The framework considers diversity in the duration of fault behavior among fault types in whole-building AFDD, thus achieving high model generalization. We implemented our developed framework in a virtual testbed calibrated with measured data from Oak Ridge National Laboratory's Flexible Research Platform designed to mimic the operation of a typical small commercial building. The AFDD model is trained by the simulation data generated from the virtual testbed. The results show that (1) the developed framework improves the generalization of the AFDD model by 10.7% compared with literature-reported feature extraction and selection methods and (2) features with diverse window sizes and statistics are selected, providing insight into physical systems beyond the current understanding of buildings and faults and improving the detection and diagnostics of multiple fault types.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Instabilities and phase transitions in architected metamaterials: a gradient-enhanced continuum approach

Architected metamaterials such as foams and lattices exhibit a wide range of properties governed by microstructural instabilities and emerging phase transitions. Their macroscopic response–including energy dissipation during impact, large recoverable deformations, morphing between configurations, and auxetic behavior–remains difficult to capture with conventional continuum models, which often rely on discrete approaches that limit scalability. In this work, we propose a nonlocal continuum formulation that captures both stable and unstable responses of elastic architected metamaterials. The framework extends anisotropic hyperelasticity by introducing nonlocal variables and internal length scales reflective of microstructural features. Local polyconvex free-energy models are systematically augmented with two families of non-(poly)convex energies, enabling both metastable and bistable responses. Implementation in a finite element framework enables solution using a hybrid monolithic–staggered strategy. Simulations capture densification fronts, forward and reverse transitions, hysteresis loops, imperfection sensitivity, and globally coordinated auxetic modes. Overall, this framework provides a robust foundation for accelerated modeling of instability-driven phenomena in architected metamaterials, while enabling extensions to anisotropic, dissipative, and active systems as well as integration with data-driven and machine learning approaches.

42 ENGINEERING↗

Sustainable materials acceleration platform reveals stable and efficient wide-bandgap metal halide perovskite alloys

The vast chemical space of emerging semiconductors, like metal halide perovskites, and their varied requirements for semiconductor applications have rendered trial-and-error environmentally unsustainable. Here, in this work, we demonstrate RoboMapper, a materials acceleration platform (MAP), that achieves 10-fold research acceleration by formulating and palletizing semiconductors on a chip, thereby allowing high-throughput (HT) measurements to generate quantitative structure-property relationships (QSPRs) considerably more efficiently and sustainably. We leverage the RoboMapper to construct QSPR maps for the mixed ion FA 1-y Cs y Pb(I 1-x Br x ) 3 halide perovskite in terms of structure, bandgap, and photostability with respect to its composition. We identify wide-bandgap alloys suitable for perovskite-Si hybrid tandem solar cells exhibiting a pure cubic perovskite phase with favorable defect chemistry while achieving superior stability at the target bandgap of ~1.7 eV. RoboMapper’s palletization strategy reduces environmental impacts of data generation in materials research by more than an order of magnitude, paving the way for sustainable data-driven materials research.

36 MATERIALS SCIENCE↗

Development of Steady-State and Dynamic Mass and Energy Constrained Neural Networks for Distributed Chemical Systems Using Noisy Transient Data

The paper presents the development of algorithms for mass and energy constrained neural network models that can exactly conserve the overall mass and energy of distributed chemical process systems, even though the noisy transient data used for optimal model training violate the same. In contrast to approximately satisfying mass and energy balance constraints of a system by soft penalization of objective function, algorithms have been developed for solving equality-constrained nonlinear optimization problems, thus providing the guarantee of exactly satisfying the system mass and energy conservation laws. For developing dynamic mass-energy constrained network models for distributed systems, hybrid series and parallel dynamic-static neural networks have been leveraged. The developed algorithms for solving both the training and forward problems are validated using both steady-state and dynamic data in the presence of various noise characteristics. The developed data-driven algorithms are flexible to exactly satisfy mass and energy balance constraints for dynamic chemical processes if the system holdup information is available. The proposed network structures and algorithms are applied to the development of data-driven lumped and distributed models of an adiabatic superheater/reheater system, a nonisothermal continuous stirred tank reactor, as well as an electrically heated plug-flow reactor system where one form of energy gets transformed to another. It has been observed that the mass-energy constrained neural networks yield a root mean squared error of <1% with respect to the system truth for the case studies evaluated in this work.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Informing the planning of rotating power outages in heat waves through data analytics of connected smart thermostats for residential buildings

Abstract With climate change, heat waves have become more frequent and intense. Rotating power outages happen when the power supply is unable to meet the cooling demand increase resulting from extreme high temperatures. Power outages during heat waves expose residents to high risks of overheating. In this study, we propose a novel data-driven inverse modelling approach to inform decision makers and grid operators on planning rotating power outages. We first infer the building thermal characteristics using the connected smart thermostat data, and used the estimated thermal dynamics to simulate the thermal resilience during a heat wave event. Our proposed method was tested for the California power outage in August 2020 by using the open source Ecobee Donate Your Data dataset. We found in California the power outage should not last more than two hours during heat waves to avoid overheating risks. Informing the residents in advance so they can prepare for it through pre-cooling is a simple but effective strategy to expand the acceptable power outage duration. In addition to assisting power outage planning, the proposed method can be used for other applications, such as to evaluate a building energy efficiency policy, to examine fuel poverty, and to estimate the load shifting potential of building stocks.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Search for photons above 10 18 eV by simultaneously measuring the atmospheric depth and the muon content of air showers at the Pierre Auger Observatory

The Pierre Auger Observatory is the most sensitive instrument to detect photons with energies above 1 0 17 eV . It measures extensive air showers generated by ultrahigh energy cosmic rays using a hybrid technique that exploits the combination of a fluorescence detector with a ground array of particle detectors. The signatures of a photon-induced air shower are a larger atmospheric depth of the shower maximum ( X max ) and a steeper lateral distribution function, along with a lower number of muons with respect to the bulk of hadron-induced cascades. In this work, a new analysis technique in the energy interval between 1 and 30 EeV ( 1 EeV = 1 0 18 eV ) has been developed by combining the fluorescence detector-based measurement of X max with the specific features of the surface detector signal through a parameter related to the air shower muon content, derived from the universality of the air shower development. No evidence of a statistically significant signal due to photon primaries was found using data collected in about 12 years of operation. Thus, upper bounds to the integral photon flux have been set using a detailed calculation of the detector exposure, in combination with a data-driven background estimation. The derived 95% confidence level upper limits are 0.0403, 0.01113, 0.0035, 0.0023, and 0.0021 km − 2 sr − 1 yr − 1 above 1, 2, 3, 5, and 10 EeV, respectively, leading to the most stringent upper limits on the photon flux in the EeV range. Compared with past results, the upper limits were improved by about 40% for the lowest energy threshold and by a factor 3 above 3 EeV, where no candidates were found and the expected background is negligible. The presented limits can be used to probe the assumptions on chemical composition of ultrahigh energy cosmic rays and allow for the constraint of the mass and lifetime phase space of super-heavy dark matter particles. Published by the American Physical Society 2024

79 ASTRONOMY AND ASTROPHYSICS↗

Facilitating Staging-based Unstructured Mesh Processing to Support Hybrid In-Situ Workflows

In-situ and in-transit processing alleviate the gap between the computing and I/O capabilities by scheduling data analytics close to the data source. Hybrid in-situ processing splits data analytics into two stages: the data processing that runs in-situ aims to extract regions of interest, which are then transferred to staging services for further in-transit analytics. To facilitate this type of hybrid in-situ processing, the data staging service needs to support complex intermediate data representations generated/consumed by the in-situ tasks. Unstructured (or irregular) mesh is one such derived data representation that is typically used and bridges simulation data and analytics. However, how staging services efficiently support unstructured mesh transfer and processing remains to be explored. This paper investigates design options for transferring and processing unstructured mesh data using staging services. Using polygonal mesh data as an example, we show that hybrid in-situ workflows with staging-based unstructured mesh processing can effectively support hybrid in-situ workflows, and can significantly decrease data movement overheads.

data-driven↗

Data-Driven Day-Ahead PV Estimation Using Autoencoder-LSTM and Persistence Model

Inherent variability in photovoltaic (PV) and associated impacts on power systems is a challenging problem for both the PV owners and the grid operators. Existing statistical and machine learning algorithms typically work well for weather conditions similar to historical data. Furthermore, uncertain weather conditions pose a great challenge to the estimation accuracy of the estimation models. With the enhanced integration of intelligent electronic devices and the realization of associated automation in the power grid, renewable energy data is becoming more accessible, which can be utilized by deep learning models and improve the PV power generation estimation accuracy. In this paper, a hybrid deep learning model driven by external weather data is proposed to do day-ahead PV output forecasting at 15-minute-interval. The proposed model is motivated by the recent advancement of Long-Short-Term-Memory (LSTM) networks and AutoEncoder (AE), which estimates uncertainties in sequence while making the prediction for complex weather conditions. Meanwhile, the persistence model (PM) is used to predict continuous sunny weather conditions. The forecasting result is validated with data from multiple locations

42 ENGINEERING↗

A Digital Twin Framework Utilizing Machine Learning for Robust Predictive Maintenance: Enhancing Tire Health Monitoring

We introduce a novel digital twin (DT) framework for the predictive maintenance of long-term physical systems. Using monitoring tire health as an application, we show how the DT framework can be used to enhance automotive safety and efficiency, and how the technical challenges can be overcome using a three-step approach. First, to manage the data complexity over a long operation span, we employ data reduction techniques to concisely represent physical tires using historical performance and usage data. Relying on these data, for fast real-time prediction, we train a transformer-based model offline on our concise dataset to predict future tire health over time, represented as remaining casing potential (RCP). Based on our architecture, our model quantifies both epistemic and aleatoric uncertainties, providing reliable confidence intervals around predicted RCP. Second, to incorporate real-time data, we update the predictive model in the DT framework, ensuring its accuracy throughout its lifespan with the aid of hybrid modeling and the use of the discrepancy function. Third, to assist decision-making in predictive maintenance, we implement a tire state decision algorithm, which strategically determines the optimal timing for tire replacement based on RCP forecasted by our transformer model. This approach ensures that our DT accurately predicts system health, continually refines its digital representation, and supports predictive maintenance decisions. Furthermore, our framework effectively embodies a physical system, leveraging big data and machine learning (ML) for predictive maintenance, model updates, and decision-making.

advanced computing infrastructure↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Developing Data-Driven Synthetic Infrastructure Models for Resilience Analysis

Research on infrastructure resilience has produced promising methods to simulate and optimize complex networks to improve performance. However, restrictions on sharing infrastructure models and the steep cost of developing and maintaining infrastructure models presents a roadblock to adoption. To overcome this limitation, this research focuses on methods to create data-driven infrastructure models that will help improve infrastructure resilience and security. The analysis couples incomplete utility data, geospatial data, machine learning, and synthetic network generation methods to rapidly develop and update infrastructure models. The methods are validated using realistic utility models and site-specific data, with a focus on Puerto Rico due to its unique infrastructure challenges and available data. This research highlights promising opportunities for the use of synthetic network generation and machine learning to create infrastructure models when very little data is available. Results demonstrate that hybrid methods, which combine sparse utility data with synthetic models, can enhance model accuracy, and machine learning can predict model attributes using training data from other models. However, the complexity of infrastructure systems means that even minor changes in network connectivity can significantly impact simulation results. Resilience analysis using synthetic infrastructure models shows that while some system behaviors are preserved, the magnitude of disruptions may not be accurately represented, indicating the need for more research and validation before using synthetic models for critical infrastructure investment decisions. The framework outlined in this report represents a significant advance to infrastructure model development and could be applied to additional domains and sites. Future research will continue to streamline and validate methods to help reduce roadblocks to resilience analysis.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Dynamic Control of Sodium Cold Trap Purification Temperature Using LSTM System Identification

This study investigates the dynamic regulation of the sodium cold trap purification temperature at Argonne National Laboratory’s liquid sodium test facility, employing long short-term memory (LSTM) system identification techniques. The investigation introduces an innovative hybrid approach by integrating model predictive control (MPC) based on first principles dynamic models with a multi-step time–frequency LSTM model in predicting the temperature profiles of a sodium cold trap purification system. The long short-term memory–model predictive controller (LSTM-MPC) model employs a sliding window scheme to gather training samples for multi-step prediction, leveraging historical data to construct predictive models that capture the non-linearities of the complex system dynamics without explicitly modeling the underlying physical processes. The performance of the LSTM-MPC and MPC were evaluated through simulation experiments, where both models were assessed on their capacity to maintain the cold trap temperature within predefined set-points while minimizing deviations and overshoots. Results obtained show how the data-driven LSTM-MPC model demonstrates stability and adaptability. In contrast, the traditional MPC model exhibits irregularities, particularly evident as overshoots around set-point limits, which can potentially compromise its effectiveness over long prediction time intervals. The findings obtained offer valuable insights into integrating data-driven techniques for enhancing real-time monitoring systems.

LSTM-MPC↗