Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Data driven”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Effects of Optimisation Parameters on Data-Driven Magnetofrictional Modelling of Active Regions

Context . The solar magnetic field plays an essential role in the formation, evolution, and dynamics of large-scale eruptive structures in the corona. The estimation of the coronal magnetic field, the ultimate driver of space weather, particularly in the ‘low’ and ‘middle’ corona, is presently limited due to practical difficulties. Data-driven time-dependent magnetofrictional modelling (TMFM) of active region magnetic fields has been proven to be a useful tool to study the corona. The input to the model is the photospheric electric field that is inverted from a time series of the photospheric magnetic field. Constraining the complete electric field, that is, including the non-inductive component, is critical for capturing the eruption dynamics. We present a detailed study of the effects of optimisation of the non-inductive electric field on the TMFM of AR 12473. Aims . We aim to study the effects of varying the non-inductive electric field on the data-driven coronal simulations, for two alternative parametrisations. By varying parameters controlling the strength of the non-inductive electric field, we wish to explore the changes in flux rope formation and their early evolution and other parameters, for instance, axial flux and magnetic field magnitude. Methods . We used the high temporal and spatial resolution cadence vector magnetograms from the Helioseismic and Magnetic Imager (HMI) on board the Solar Dynamics Observatory (SDO). The non-inductive electric field component in the photosphere is critical for energising and introducing twist to the coronal magnetic field, thereby allowing unstable configurations to be formed. We estimated this component using an approach based on optimising the injection of magnetic energy. Results . Our data show that flux ropes are formed in all of the simulations except for those with the lower values of these optimised parameters. However, the flux rope formation, evolution and eruption time varies depending on the values of the optimisation parameters. The flux rope is formed and has overall similar evolution and properties with a large range of non-inductive electric fields needed to determine the non-inductive electric field component that is critical for energising and introducing twist to the coronal magnetic field. Conclusions . This study shows that irrespective of non-inductive electric field values, flux ropes are formed and erupted, which indicates that data-driven TMFM can be used to estimate flux rope properties early in their evolution without needing to employ a lengthy optimisation process.

A. Kumari

Data-Driven Voltage Regulation of Distribution Grid Using Nonlinear Autoregressive Model with Exogenous Inputs (NARX)

This article proposes data-driven control via a nonlinear autoregressive model with exogenous inputs (NARX) for real-time voltage regulation of a modified feeder using reactive power sources. Traditional voltage control strategies rely on rule-based heuristics or optimization techniques, which often require detailed system models and extensive computational resources. The NARX-based controller learns system dynamics from historical data and predicts optimal reactive power dispatch in real-time for voltage correction. The proposed approach is evaluated on a power system feeder model under varying load and network conditions. Simulation results demonstrate that the NARX-based controller achieves improved voltage regulation, offering higher adaptability to system fluctuations. This study highlights the potential of data-driven control for enhancing the reliability of power distribution networks.

Donge, Vrushabh [ORNL] (ORCID:0000000306062803)

Data Imbalance, Uncertainty Quantification, and Transfer Learning in Data‐Driven Parameterizations: Lessons From the Emulation of Gravity Wave Momentum Transport in WACCM

Abstract Neural networks (NNs) are increasingly used for data‐driven subgrid‐scale parameterizations in weather and climate models. While NNs are powerful tools for learning complex non‐linear relationships from data, there are several challenges in using them for parameterizations. Three of these challenges are (a) data imbalance related to learning rare, often large‐amplitude, samples; (b) uncertainty quantification (UQ) of the predictions to provide an accuracy indicator; and (c) generalization to other climates, for example, those with different radiative forcings. Here, we examine the performance of methods for addressing these challenges using NN‐based emulators of the Whole Atmosphere Community Climate Model (WACCM) physics‐based gravity wave (GW) parameterizations as a test case. WACCM has complex, state‐of‐the‐art parameterizations for orography‐, convection‐, and front‐driven GWs. Convection‐ and orography‐driven GWs have significant data imbalance due to the absence of convection or orography in most grid points. We address data imbalance using resampling and/or weighted loss functions, enabling the successful emulation of parameterizations for all three sources. We demonstrate that three UQ methods (Bayesian NNs, variational auto‐encoders, and dropouts) provide ensemble spreads that correspond to accuracy during testing, offering criteria for identifying when an NN gives inaccurate predictions. Finally, we show that the accuracy of these NNs decreases for a warmer climate (4 × CO 2 ). However, their performance is significantly improved by applying transfer learning, for example, re‐training only one layer using ∼1% new data from the warmer climate. The findings of this study offer insights for developing reliable and generalizable data‐driven parameterizations for various processes, including (but not limited to) GWs.

54 ENVIRONMENTAL SCIENCES

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis

A Data-driven approach to Core Power distribution reconstruction in a Nuclear Reactor

This report presents the initial development of a data-driven approach for reconstructing the core power distribution in a nuclear reactor (power shape synthesis) using ex-core sensors. Traditional techniques rely on deploying a large number of detectors throughout the reactor core. However, this approach is not feasible for innovative reactor concepts like Advanced Reactors and Microreactors. First, the tight lattice pitch, designed to maximize power density, limits the space available for sensors. Secondly, the harsh operating conditions are not compatible with commercially available detectors. The method proposed in this work integrates high-fidelity modeling with data-driven techniques to accurately reconstruct power distribution across various reactor types, thereby reducing the reliance on in-core sensors. Purdue University Reactor One (PUR-1) was selected as the test case. The CAD model representing the latest configuration of the PUR-1 core was imported into the OpenMC simulation framework, and the model was built. Additionally, the previously developed MCNP6 model was updated. The two models were assessed against the data collected during an experimental campaign conducted in July 2024. Thirty gold foils were placed in three Irradiation Assemblies in PUR-1 core. Using the measured activity of the irradiated foils, the neutron flux at different core locations was estimated.

22 GENERAL STUDIES OF NUCLEAR REACTORS

MAD 3 (Material Data Driven Design) User Manual (v1.01)

MAD 3 (Material Data Driven Design) is a novel and unique software solution that provides initial plastic anisotropy of polycrystalline metals using crystallographic texture information, developed at Sandia National Laboratories. In this document, we describe the structure and functionality of the current MAD 3 software (v1.01).

36 MATERIALS SCIENCE

Development of Data-Driven Models for Performance Prediction and Chemical Dosing of a Full-Scale Controlled Phosphorus Precipitation Reactor

This study evaluated the use of data-driven models to improve control of a struvite precipitation reactor that removes phosphorus from wastewater while producing a fertilizer product. The researchers developed predictive models for influent orthophosphate concentration, effluent orthophosphate concentration, and phosphorus removal using operational data from a full-scale MagPrex™ reactor at a water resource recovery facility in Denver, Colorado. Model predictions were used to recommend magnesium chloride dosing adjustments needed to achieve a target effluent phosphorus concentration. Several machine learning approaches were tested, with ridge regression providing the best predictions for influent orthophosphate concentration and phosphorus removal, and XGBoost providing the best predictions for effluent orthophosphate concentration. Simulation results indicated that the decision-support approach could correctly identify dosing adjustments in most cases and reduce chemical use. Full-scale implementation achieved lower accuracy due to changing operating conditions and limited historical data in some operating ranges. Here, the results demonstrate the potential of data-driven tools to support phosphorus recovery process control while also identifying practical limitations that affect deployment in full-scale systems.

42 ENGINEERING

Data Driven UAM Flight Energy Consumption Prediction and Risk Assessment

With the current technological advancements revolutionizing the concept of Urban Air Mobility (UAM) and package delivery, there is also, a concurrent need to quantify the operational safety of these vehicles in terms of their associated risk. Conducting safe flight operations is critical for UAM vehicles which are electrically Vertical Takeoff and Landing (eVTOL) vehicles, to operate in current Air traffic control. In this paper, a data-driven method for UAM vehicle energy consumption prediction and risk quantification with conditional value-at-risk based on energy consumption distribution is presented. Significant factors affecting energy consumption, such as density altitude, aircraft design, airspeed, and collision avoidance algorithms, are considered in the data-driven based energy consumption prediction of different eVTOL

Data-driven

Data‐driven variational method for discrepancy modeling: Dynamics with small‐strain nonlinear elasticity and viscoelasticity

Abstract The effective inclusion of a priori knowledge when embedding known data in physics‐based models of dynamical systems can ensure that the reconstructed model respects physical principles, while simultaneously improving the accuracy of the solution in the previously unseen regions of state space. This paper presents a physics‐constrained data‐driven discrepancy modeling method that variationally embeds known data in the modeling framework. The hierarchical structure of the method yields fine scale variational equations that facilitate the derivation of residuals which are comprised of the first‐principles theory and sensor‐based data from the dynamical system. The embedding of the sensor data via residual terms leads to discrepancy‐informed closure models that yield a method which is driven not only by boundary and initial conditions, but also by measurements that are taken at only a few observation points in the target system. Specifically, the data‐embedding term serves as residual‐based least‐squares loss function, thus retaining variational consistency. Another important relation arises from the interpretation of the stabilization tensor as a kernel function, thereby incorporating a priori knowledge of the problem and adding computational intelligence to the modeling framework. Numerical test cases show that when known data is taken into account, the data driven variational (DDV) method can correctly predict the system response in the presence of several types of discrepancies. Specifically, the damped solution and correct energy time histories are recovered by including known data in the undamped situation. Morlet wavelet analyses reveal that the surrogate problem with embedded data recovers the fundamental frequency band of the target system. The enhanced stability and accuracy of the DDV method is manifested via reconstructed displacement and velocity fields that yield time histories of strain and kinetic energies which match the target systems. The proposed DDV method also serves as a procedure for restoring eigenvalues and eigenvectors of a deficient dynamical system when known data is taken into account, as shown in the numerical test cases presented here.

Masud, Arif

MBX V1.2: Accelerating Data-Driven Many-Body Molecular Dynamics Simulations

The MBX software provides an advanced platform for molecular dynamics simulations, leveraging state-of-the-art MB-pol and MB-nrg data-driven many-body potential energy functions. Developed over the past decade, these potential energy functions integrate physics-based and machine-learned many-body terms trained on electronic structure data calculated at the "gold standard" coupled-cluster level of theory. Recent advancements in MBX have focused on optimizing its performance, resulting in the release of MBX v1.2. While the inherently many-body nature of MB-pol and MB-nrg ensures high accuracy, it poses computational challenges. MBX v1.2 addresses these challenges with significant performance improvements, including enhanced parallelism that fully harnesses the power of modern multicore CPUs. In conclusion, these advancements enable simulations on nanosecond time scales for condensed-phase systems, significantly expanding the scope of high-accuracy, predictive simulations of complex molecular systems powered by data-driven many-body potential energy functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Rare Lepton Decays and Differentiable Hadronization Models - From Signatures of New Physics to Data-driven Event Generation

This dissertation is partitioned into two parts: phenomenological studies focused on rare lepton decays as probes of heavy and light new physics, and the development of differentiable, data-driven hadronization models. Part I develops the phenomenology of new physics signatures stemming from rare charged lepton flavor violating decays probed by experiments at the intensity frontier. These include interactions mediated by both high-scale effective operators and light new physics, manifesting in multi-lepton final states ($\mu \to 5e$), elastic nuclear transitions ($\mu \to e$ conversion), baryon-number-violating muon capture, and time-dependent signals from ultralight dark matter ($\mu \to e \phi, \tau \to \ell \phi$). Part II develops two distinct strategies for advancing differentiable and data-driven hadronization models. One involves comprehensive reweighting frameworks for hadronization that enable efficient uncertainty estimation, facilitate parameter tuning, and interface naturally with differentiable programming paradigms. The other introduces machine-learning-based methods for extracting microscopic fragmentation dynamics directly from macroscopic observables through the deformation of existing models -- effectively providing solutions to the inverse problem of hadronization. Altogether, these studies advance the interpretability, flexibility, and precision of theoretical predictions for both high-intensity and high-energy experiments.

Menzo, Tony [Cincinnati U.] (ORCID:000000022013457

Data-Driven Digital Twin for Reliability Assessment of DC/DC Buck Converter

In commercial applications, the operation of DC/DC converters significantly impacts overall system performance and long-term reliability. This study introduces a data-driven digital twin (DT) approach for estimating critical degradation parameters of DC/DC BUCK converter under steady-state condition. Initially, a circuit-level MATLAB/Simulink digital model (DM C ) is refined against a hardware prototype’s switching model dataset using offline particle swarm optimization. The optimized digital model’s steady-state response is then verified with its average model response while varying the duty and load. Subsequently, degradation profiles are imposed on the inductor, capacitor, MOSFET in the DMC. A large dataset is generated from this model, allowing training, validation, and testing of machine learning (ML) models for component health regression tasks. The proposed method employs random forest ML models, achieving impressive regression results with a squared R value as high as 0.99978 and a root mean square error of 4.2× 10 –6 . The method is further validated on a medium power level DC/DC BUCK prototype with varying load conditions, and includes the analysis of MOSFET’s on-resistance under degradation conditions. This data-driven DT method shows promise for identifying parasitic degradation and ohmic loss parameters, enhancing converter reliability assessments in a non-invasive, generalized, and computationally efficient manner.

14 SOLAR ENERGY

Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?

Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties, including Monte Carlo dropout, deep ensemble, Bayesian neural networks, Gaussian Processes and conformal prediction. Lastly, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Efficient data-driven regression for reduced-order modeling of spatial pattern formation

We present an efficient data-driven regression approach for constructing reduced-order models (ROMs) of reaction-diffusion systems exhibiting pattern formation. The ROMs are learned non-intrusively from available training data of physically accurate numerical simulations. The method can be applied to general nonlinear systems through the use of polynomial model form, while not requiring knowledge of the underlying physical model, governing equations, or numerical solvers. The process of learning ROMs is posed as a low-cost least-squares problem in a reduced-order subspace identified via Proper Orthogonal Decomposition (POD). Numerical experiments on classical pattern-forming systems–including the Schnakenberg and Mimura–Tsujikawa models–demonstrate that higher-order surrogate models significantly improve prediction accuracy while maintaining low computational cost. The proposed method provides a flexible, non-intrusive model reduction framework, well suited for the analysis of complex spatio-temporal pattern formation phenomena.

Data-driven modeling

General Purpose Data-Driven Online System Health Monitoring with Applications to Space Operations

Modern space transportation and ground support system designs are becoming increasingly sophisticated and complex. Determining the health state of these systems using traditional parameter limit checking, or model-based or rule-based methods is becoming more difficult as the number of sensors and component interactions grows. Data-driven monitoring techniques have been developed to address these issues by analyzing system operations data to automatically characterize normal system behavior. System health can be monitored by comparing real-time operating data with these nominal characterizations, providing detection of anomalous data signatures indicative of system faults, failures, or precursors of significant failures. The Inductive Monitoring System (IMS) is a general purpose, data-driven system health monitoring software tool that has been successfully applied to several aerospace applications and is under evaluation for anomaly detection in vehicle and ground equipment for next generation launch systems. After an introduction to IMS application development, we discuss these NASA online monitoring applications, including the integration of IMS with complementary model-based and rule-based methods. Although the examples presented in this paper are from space operations applications, IMS is a general-purpose health-monitoring tool that is also applicable to power generation and transmission system monitoring.

Iverson, David L.

Navigation Sensor Technology Assessment Capability for Data-Driven Systems Analysis

The capability to assess the mission performance of novel navigation technologies from a systems-level approach and to provide quantitative results is crucial. With growing interest and innovations from NASA’s commercial partners, data-driven results that quantify the technological impact will guide research developments and facilitate stakeholder decision-making while requiring less time and resources. This paper presents a method to evaluate various sensor combinations and their impact on the overall system performance using an existing six degrees-of-freedom, physics-based engineering simulation for a government reference lunar lander as a testbed. Selected navigation technologies over various technology readiness levels, including Inertial Measurement Units, Navigation Doppler Lidar, and radar altimeter-radar velocimeter, were studied in this paper. Results using this testbed to perform sensitivity studies and to provide quantitative assessments are reported. One key finding is that there are diminishing returns for reducing sensor errors. The most influential sensor parameter to landing success are vehicle configuration and mission dependent. Results from this method could be used by stakeholders to make data-driven systems-based decisions and by technology developers as guidance for the specific parameter improvements that will have the most significant impact on mission success. This capability could be further used to assess alternative scenarios should one type of technology become unavailable or to re-assess initial assumptions and identify appropriate requirements from a system perspective.

Esther Lee

Crossing the Finish Line: Integration of Data-Driven Process Control for Maximization of Energy and Resource Efficiency in Advanced Water Resource Recovery Facilities

Improvements in process monitoring and control at water resource recovery facilities (WRRFs) could result in reductions in electricity consumption, chemical inputs, and greenhouse gas emissions, as well as improved energy recovery. Many current WRRF data collection, monitoring, and control approaches use 20th century process monitoring and control systems, which require large design safety factors to ensure reliability in the absence of more advanced, precise controls. Implementation of more modern data-driven control tools could lead to more efficient operations that provide intrinsic reliability with better overall process performance at full-scale. This project (1) developed and demonstrated data-driven process controls at full-scale facilities for five promising WRRF process technologies that provide whole-plant approaches and offer substantial energy and resource recovery benefits, and (2) created a Machine Learning (ML) Toolkit and an implementation guide of new process control approaches that walks users through each step of the ML workflow and illustrates the steps through case study examples.

54 ENVIRONMENTAL SCIENCES

Data-Driven Invertible Neural Surrogates of Atmospheric Transmission

We present Data-Driven Invertible Neural Surrogates of Atmospheric transmission, or DINSAT. DINSAT is a novel framework for inferring an atmospheric transmission profile from a spectral scene. This framework leverages a lightweight, physics-based simulator that is automatically tuned -- by virtue of autodifferentiation and differentiable programming -- to construct a surrogate atmospheric profile to model the observed data. The framework has utility in (i) performing atmospheric correction, (ii) recasting spectral data between various modalities (e.g. radiance and reflectance at the surface and at the sensor), and (iii) inferring atmospheric transmission profiles, such as absorbing bands and their relative magnitudes. We demonstrate the utility of these methods by performing a canonical atmospheric correction task for the purposes of further analysis - in this case, target detection within a scene.

Koch, James V.