Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data driven model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23

Geographically aware estimates of remotely sensed water properties for Chesapeake Bay

This software contains the code for a data-driven model to produce high-resolution (< 1 km2) estimates of temperature, salinity, and turbidity over decadal time scales as required by next-generation Earth Systems Models. Our model fits daily MODIS Aqua reflectance data to surface observations (< 1 m depth) from 2000-2021 in Chesapeake Bay, USA.

Stachelek, Jemma↗

Duration of Contagion of Foot-And-Mouth Disease Virus in Infected Live Pigs and Carcasses

Data-driven modeling of incursions of high-consequence, transboundary pathogens of animals is a critical component of veterinary preparedness. However, simplifying assumptions and excessive use of proxy measures to compensate for gaps in available data may compromise modeled outcomes. The current investigation was prospectively designed to address two major gaps in current knowledge of foot-and-mouth disease virus (FMDV) pathogenesis in pigs: the end (duration) of the infectious period and the viability of FMDV in decaying carcasses. By serial exposure of sentinel groups of pigs to the same group of donor pigs infected by FMDV A24 Cruzeiro, it was demonstrated that infected pigs transmitted disease at 10 days post infection (dpi), but not at 15 dpi. Assuming a latent period of 1 day, this would result in a conservative estimate of an infectious duration of 9 days, which is considerably longer than suggested by a previous report from an experiment performed in cattle. Airborne contagion was diminished within two days of removal of infected pigs from isolation rooms. FMDV in muscle was inactivated within 7 days in carcasses stored at 4°C. By contrast, FMDV infectivity in vesicle epithelium harvested from intact carcasses stored under similar conditions remained remarkably high until the study termination at 11 weeks post mortem. The output from this study consists of experimentally determined data on contagion associated with FMDV-infected pigs. This information may be utilized to update parameterization of models used for foot-and-mouth disease outbreak simulations involving areas of substantial pig production.

59 BASIC BIOLOGICAL SCIENCES↗

Investigating boosted decision trees as a guide for inertial confinement fusion design

Inertial confined fusion experiments at the National Ignition Facility have recently entered a new regime approaching ignition. Improved modeling and exploration of the experimental parameter space were essential to deepening our understanding of the mechanisms that degrade and amplify the neutron yield. The growing prevalence of machine learning in fusion studies opens a new avenue for investigation. Here in this paper, we have applied the Gradient-Boosted Decision Tree machine-learning architecture to further explore the parameter space and find correlations with the neutron yield, a key performance indicator. We find reasonable agreement between the measured and predicted yield, with a mean absolute percentage error on a randomly assigned test set of 35.5%. This model finds the characteristics of the laser pulse to be the most influential in prediction, as well as the hohlraum laser entrance hole diameter and an enhanced capsule fabrication technique. We used the trained model to scan over the design space of experiments from three different campaigns to evaluate the potential of this technique to provide design changes that could improve the resulting neutron yield. While these data-driven model cannot predict ignition without examples of ignited shots in the training set, it can be used to indicate that an unseen shot design will at least be in the upper range of previously observed neutron yields.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Multi-Scale Computational Platform for Predictive Modeling of Corrosion in Al-Steel Joints (Final Report)

The research team proposed to develop innovative multi-scale models to predict corrosion and the resulting mechanical performances in aluminum-steel joints. The methods of joining considered are resistance spot welding, self-piercing riveting, and rivet-welding, all suitable for mass production applications. The multi-scale models integrate high throughput first-principle calculations based on density functional theory (DFT), high throughput calculation of phase diagrams (CALPHAD) modeling, and finite element method (FEM) simulations. These models are to be validated through laboratory experiments. Furthermore, the models are available as open source so as to enable scientists and engineers in the community to adapt and contribute to the development and application. The approaches rely on the research team’s extensive experience on the prediction of properties of individual phases at finite temperatures and variable compositions through DFT calculations, and our broad expertise on dissimilar material joining and their corrosion. The proposed computational framework enables high throughput computations for improved predictions of corrosion and the associated mechanical performance in dissimilar material joints, resulting in significant reduction in computational time needed by the current state-of-the-art methods. With the participation of researchers from three universities, an auto manufacturer, two manufacturing technology/equipment suppliers, and a software developer/vendor, the interdisciplinary research team applies the technical development on both phase-based modeling and laboratory experiments into the automobile body joining processes for validation and technology demonstration. The global cost of corrosion was estimated at about 3.4% of the global GDP in 2013. By using available corrosion control practices, it is estimated a saving between 15-35% of the cost of corrosion. In the U.S., more than $276 billion is spent repairing corrosion damage. Prediction of the corrosion and its impact on performance of the dissimilar material joints is critical for reducing the massive number of the current corrosion-based recalls for automobiles. Thus, the project goal is to develop models to enable predictive maintenance and end-of-life planning of multi-metal joints with risk of corrosion under different conditions such as exposure to high temperatures in summer and salt solutions in winter, quantified through its pH. An academia-industry consortium led by the University of Michigan and including Pennsylvania State University, University of Illinois Urbana-Champaign, University of Georgia, General Motors Company, Livermore Software Technology Corporation, and Optimal Process Technologies, LLC. created multi-scale models for prediction of corrosion in aluminum-steel joint structures such of them used in vehicle subassemblies – chassis and transmission systems. Starting from the first principle calculations, the team developed mathematical and data-driven models to predict the metallic components, which are formed during joining of two metals, for example aluminum and steel - a lightweight multilateral system which is currently used in more than 60% car bodies. These models were used for simulating chemical reactions that are happening when the joining metallic components are exposed to high temperatures and different pH values. The team was able to predict how the corrosion installs on the metallic components and how they lead to a sudden failure of components in cars. Newly developed machine learning algorithms combining Science, Technology, Engineering and Math disciplines, advanced finite element simulation and experimental validations have been integrated in a platform for prediction of the corrosion evolution and prediction the failure of joints under mechanical loadings and fatigue. Moreover, based on machine learning and inverse analysis, the team proposed solutions for designing new metallic alloys less susceptible to corrosion when joining multi-material assembles. An average of 4% error compared with experiments was achieved for the most common joints that are used in vehicle subassemblies.

36 MATERIALS SCIENCE↗

A physics-constrained neural ordinary differential equations approach for robust learning of stiff chemical kinetics

The high computational cost associated with solving for detailed chemistry poses a significant challenge for predictive computational fluid dynamics (CFD) simulations of turbulent reacting flows. While deep learning techniques have been explored to develop faster surrogate models, they often fail to integrate reliably with CFD solvers. This instability arises because traditional deep learning approaches optimize for training error without ensuring compatibility with ordinary differential equation (ODE) solvers, resulting in accumulation of errors over time. Recently, neuralODE (NODE) based approaches have been shown to be a promising technique to emulate and accelerate detailed chemistry computations. Here, in the present work, we extend this NODE framework for stiff chemical kinetics by incorporating mass conservation constraints directly into the loss function during training. This ensures that the total mass as well as the individual elemental species masses are conserved in an a-posteriori manner. Proof-of-concept studies are performed with the novel physics-constrained NODE (PC-NODE) approach for homogeneous autoignition of hydrogen-air mixture over a range of composition and thermodynamic conditions. It is demonstrated that the PC-NODE framework not only improves the physical consistency of the resulting data-driven model with respect to mass conservation criteria, but also improves training efficiency. PC-NODE is shown to achieve 2–100× speedup relative to the hydrogen-air detailed chemical mechanism depending on the type of the ODE solver (implicit or explicit) used during autoregressive inference tests. Lastly, a-posteriori studies are performed wherein the trained PC-NODE model is coupled with a CFD solver. It is shown that higher accuracy is achieved with PC-NODE relative to the purely data-driven NODE approach. Moreover, PC-NODE also exhibits robustness and generalizability to unseen initial conditions from within (interpolative capability) as well as outside (extrapolative capability) the training regime.

computational combustion↗

Morphological analysis of the polarized synchrotron emission with WMAP and Planck

The bright polarized synchrotron emission, away from the Galactic plane, originates mostly from filamentary structures. We implement a filament finder algorithm which allows the detection of bright elongated structures in polarized intensity maps. We analyse the sky at 23 and 30 GHz as observed respectively by WMAP and Planck. We identify 19 filaments, 13 of which have been previously observed. For each filament, we study the polarization fraction, finding values typically larger than for the areas outside the filaments, excluding the Galactic plane, and a fraction of about 30% is reached in two filaments. We study the polarization spectral indices of the filaments, and find a spectral index consistent with the values found in previous analysis (about -3.1) for more diffuse regions. Decomposing the polarization signals into the E and B families, we find that most of the filaments are detected in P E , but not in P B . We then focus on understanding the statistical properties of the diffuse regions of the synchrotron emission at 23 GHz. Using Minkowski functionals and tensors, we analyse the non-Gaussianity and statistical isotropy of the polarized intensity maps. For a sky coverage corresponding to 80% of the fainter emission, and on scales smaller than 6 degrees (ℓ > 30), the deviations from Gaussianity and isotropy are significantly higher than 3σ. The level of deviation decreases for smaller scales, however, it remains significantly high for the lowest analised scale (~ 1.5°). When 60% sky coverage is analysed, we find that the deviations never exceed 3σ. Finally, we present a simple data-driven model to generate non-Gaussian and anisotropic simulations of the synchrotron polarized emission. The simulations are fitted in order to match the spectral and statistical properties of the faintest 80% sky coverage of the data maps.

79 ASTRONOMY AND ASTROPHYSICS↗

Clustering of red sequence galaxies in the fourth data release of the Kilo-Degree Survey

We present a sample of luminous red sequence galaxies as the basis for a study of the large-scale structure in the fourth data release of the Kilo-Degree Survey. The selected galaxies are defined by a red sequence template, in the form of a data-driven model of the colour-magnitude relation conditioned on redshift. In this work, the red sequence template was built using the broad-band optical+near infrared photometry of KiDS-VIKING and the overlapping spectroscopic data sets. The selection process involved estimating the red sequence redshifts, assessing the purity of the sample and estimating the underlying redshift distributions of redshift bins. After performing the selection, we mitigated the impact of survey properties on the observed number density of galaxies by assigning photometric weights to the galaxies. We measured the angular two-point correlation function of the red galaxies in four redshift bins and constrain the large-scale bias of our red sequence sample assuming a fixed ΛCDM cosmology. We find consistent linear biases for two luminosity-threshold samples (‘dense’ and ‘luminous’). We find that our constraints are well characterised by the passive evolution model.

79 ASTRONOMY AND ASTROPHYSICS↗

In Situ Infrared Spectroscopy of a Plasma Jet and Data-Driven Solution of Multi-Scale Plasma Chemistry Problems

Multi-scale problems are commonly known in many scientific and engineering fields where microscopic behaviors are coupled with macroscopic processes. This is also an unsolved problem in low-temperature plasma chemistry where hundreds of chemical species are involved in thousands of chemical reactions. To address this problem, a physics-informed data-driven modeling is developed to solve such a multi-scale problem using the experimental Fourier-transform infrared spectroscopy (FTIR) measurements of several species’ concentrations. The modeling based on modern machine learning techniques provides concentrations of other relevant species along with the electron temperature and gas temperature at the location of FTIR measurements. For example, the concentrations O, OH, and H 2 O 2 play key roles in plasma-based cancer therapy. This approach overcomes the multi-scale difficulties of microscopic low-temperature plasma chemistry coupling with macroscopic gas flow and allows the acquisition of a full picture of output species concentrations. Presented here for the helium-air jet at atmospheric pressure, the ML-based modeling can be used to describe and possibly control multiscale systems using partial experimental data sets.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Daylight simulation workflows incorporating measured bidirectional scattering distribution functions

Daylight predictions of architectural spaces depend on good estimates of light transfer through skylights, windows and other fenestration systems. For clear glazing and painted surfaces, parametric transmission and reflection models have proven adequate, but there are many cases where light-scattering, semi-specular shading and daylighting materials defy simple characterization. Something as commonplace as fabric roller shades and venetian blinds may turn daylight prediction into guesswork, and numerous advanced systems on the market tuned specifically to enhance daylight are not sufficiently characterized to distinguish their performance. In this paper, we describe new tools available to handle novel and specialized fabrics, materials, and devices using data-driven modelling of bi-directional scattering distribution functions (BSDFs). These representations are usually tabulated at constant or adjustable angular resolution for efficient point-in-time and annual daylight simulations. We describe a variety of BSDF simulation workflows, including some of the tools and methods that make advanced analysis possible, and highlight some of the current challenges. We conclude with a discussion of future work and how such data might be created and shared worldwide.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

1.2.4.404 - Data-Driven Approach for Hydropower Plant Controller Prototyping Using Remote Hardware in the Loop (DR-HIL)

Real-time prototyping of hydropower plant controls is important for reducing the cost and the risk of field deployment. This project will 1) collect design and operational data from actual hydro plants and 2) use a physics-informed machine learning approach for real-time emulation of hydropower plants, including hydro turbine and hydrodynamics. The data-driven models will be interfaced with digital real-time simulation at NREL’s Flatirons campus for hardware-in-the-loop (HIL) testing of the governor hardware device or controller-HIL (CHIL). The proposed approach will also establish the connectivity based remote CHIL testing capability using real-time data streams from an actual hydro plant. This integrated hydro-plant emulation with CHIL will be used to prototype hydro-governor controls and eventually provide an opportunity to test hydropower integrated with various technologies (e.g. conventional and renewable generation, energy conversion, etc.) as HIL.

controls prototyping↗

Toward prediction of turbulent atmospheric flows over propagating oceanic waves via machine-learning augmented large-eddy simulation

Wind-wave interactions have important effects on the energy harvesting of offshore wind farms. High-fidelity large-eddy simulation (LES) is a powerful approach for investigating wind-wave interactions in turbulent oceanic environments. Due to the large scale of the flow domain and the high grid resolution required to resolve multi-scale flow motions, however, brute-force LES of wind-wave interactions is computationally very expensive. We propose augmenting brute-force LES via machine-learning data-driven modeling (ML-LES) to dramatically reduce the computational time required to obtain converged turbulence statistics when the brute-force approach is employed. Namely, we employ a convolutional neural network (CNN) autoencoder trained and validated with LES data sets to develop a highly efficient ML-LES approach for computing turbulence statistics from just a few snapshots of instantaneous LES flow fields. Further, our results demonstrate the accuracy and efficiency of ML-LES in predicting the mean velocity, velocity fluctuations, and turbulence kinetic energy in highly stretched computational grid systems required to carry out simulations in real-life oceanic environments.

42 ENGINEERING↗

A database of global storm surge reconstructions

Storm surges are among the deadliest coastal hazards and understanding how they have been affected by climate change and variability in the past is crucial to prepare for the future. However, tide gauge records are often too short to assess trends and perform robust statistical analyses. Here we use a data-driven modeling framework to simulate daily maximum surge values at 882 tide gauge locations across the globe. We use five different atmospheric reanalysis products for the storm surge reconstruction, the longest one going as far back as 1836. The data that we generate can be used, for example, for long-term trend analyses of the storm surge climate and identification of regions where changes in the intensity and/or frequency of storms surges have occurred in the past. It also provides a better basis for robust extreme value analysis, especially for tide gauges where observational records are short. The data are made available for public use through an interactive web-map as well as a public data repository.

54 ENVIRONMENTAL SCIENCES↗

Application of artificial intelligence methods in the international roughness index prediction of rigid and composite pavements: a systematic review

The International Roughness Index (IRI) is a widely adopted metric for quantifying pavement roughness, directly influencing vehicle safety, ride comfort, and overall roadway performance. In recent years, the use of Machine Learning (ML) models for IRI prediction has gained momentum, with the goal of improving the allocation of maintenance and rehabilitation resources by enabling accurate assessments of pavement conditions. Most prior reviews, however, have concentrated on flexible pavements, leaving a notable gap regarding rigid and composite pavements. To address this gap, the present study conducts a systematic review of Artificial Intelligence (AI) methods applied to IRI prediction for rigid and composite pavements. Literature published between 2004 and 2025 is synthesized to highlight prevailing trends, methodological contributions, and directions for future research. Particular attention is given to the types of models employed, the datasets used for training and validation, and the role of input variables and data-processing strategies. Across the included studies, ensemble learning methods (especially gradient boosting variants such as XGBoost), artificial neural networks, and hybrid architectures frequently achieved high predictive skill, with several models reporting test-set coefficients of determination approaching 0.9–0.96, indicating strong potential for capturing the influence of traffic, pavement structure, and climatic factors. Since these results are obtained from heterogeneous datasets and evaluation protocols, they are interpreted qualitatively rather than as strict cross-study rankings. Analysis of input variables revealed that pavement age and initial IRI were included in 91% (21 of 23) and 78% (18 of 23) of studies, respectively. Climatic variables such as the freezing index appeared in 57% (13 of 23), while traffic-related factors were considered in 65% (15 of 23). The findings underscore the importance of standardized, high-quality datasets, such as those from the Long-Term Pavement Performance (LTPP) program, along with data consistency, model interpretability, computational efficiency, and replicability in enhancing IRI prediction. Future research should focus on incorporating input variable selection techniques to identify the most influential predictors, thereby improving accuracy and robustness. Integrating these approaches with advanced non-linear data-driven models, coupled with robust hyperparameter optimization, holds considerable promise for strengthening the reliability of IRI prediction and supporting resilient pavement management strategies.

42 ENGINEERING↗

Position Papers for Inverse Methods for Complex Systems under Uncertainty Workshop

The ability to solve inverse problems – inferring unknown parameters, structures, or states of a system from observed data – is essential for advancing scientific discovery and innovation capabilities for the DOE mission. Basic research needs and challenges are particularly acute in emerging areas such as the interactive, data-driven, modeling and simulation of digital twins; decision support for experiments at DOE scientific user facilities; and for other complex systems and workflows. Inverse problems are at the heart of understanding and controlling complex systems due to factors such as observational data with varying modalities and fidelities, inherent uncertainties in physical measurements and numerical models, and the computational demands of rapid and high-fidelity simulations. The convergence of recent scientific computing trends – scientific machine learning, artificial intelligence, and computing advances such as exascale computing – is creating unprecedented opportunities. These advancements offer the potential to revolutionize how we approach inverse problems to extract actionable insights with the required level of accuracy and computational efficiency. This workshop and the Call for Position Papers are vital steps in bringing together experts to collectively explore and identify the new computational and mathematical directions needed in inverse methods for complex systems under uncertainty.

97 MATHEMATICS AND COMPUTING↗

A smart alarm for particle accelerator beamline operations

We present the initial results of a proof-of-concept 'smart alarm' for the Continuous Electron Beam Accelerator Facility injector beamline at Jefferson Lab. To minimize machine downtime and improve operational efficiency, an autonomous alarm system able to identify and diagnose unusual machine states is needed. Our approach leverages a trained neural network capable of alerting operators (a) when an anomalous condition exists in the beamline and (b) identifying the element setting that is the root cause. The tool is based on an inverse model that maps beamline readings (diagnostic readbacks) to settings (beamline attributes operators can modify). The model takes as input readings from the machine and computes machine settings which are compared to control setpoints. Instances where predictions differ from setpoints by a user-defined threshold are flagged as anomalous. Given data corresponding to 354 anomalous injector configurations, the model can narrow the root cause of an anomalous condition to three potential candidates with 94.6% accuracy. Furthermore, compared to the current method of identifying anomalous conditions which raises an alarm when machine parameters drift outside their normal tolerances, the data-driven model can identify 83% more anomalous conditions.

43 PARTICLE ACCELERATORS↗

Cyber-Resilient Automatic Generation Control for Systems of AC Microgrids

In this paper we propose a co-design of the secondary frequency regulation in systems of AC microgrids and its cyber security solutions. We term the secondary frequency regulator a Micro-Automatic Generation Control (μ AGC) for highlighting its same functionality as the AGC in bulk power systems. We identify sensory challenges and cyber threats facing the μ AGC. To address the sensory challenges, we introduce a new microgrid model by exploiting the rank-one deficiency property of microgrid dynamics. This model is used to pose an optimal μ AGC control problem that is easily implemented, because it does not require fast frequency measurements. An end-to-end cyber security solution to the False Data Injection (FDI) attack detection and mitigation is developed for the proposed μ AGC. The front-end barrier of applying off-the-shelf algorithms for cyber attack detection is removed by introducing a data-driven modeling approach. Finally, we propose an observer-based corrective control for an islanded microgrid and a collaborative mitigation scheme in systems of AC microgrids. We demonstrate a collaborative role of systems of microgrids during cyber attacks. Furthermore, the performance of the proposed cyber-resilient μ AGC is tested in a system of two networked microgrids.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Reconstruction of the Density Power Spectrum from Quasar Spectra using Machine Learning

We describe a novel end-to-end approach using Machine Learning to reconstruct the power spectrum of cosmological density perturbations at high redshift from observed quasar spectra. State-of-the-art cosmological simulations of structure formation are used to generate a large synthetic dataset of line-of-sight absorption spectra paired with 1-dimensional fluid quantities along the same line-of-sight, such as the total density of matter and the density of neutral atomic hydrogen. With this dataset, we build a series of data-driven models to predict the power spectrum of total matter density. We are able to produce models which yield reconstruction to accuracy of about 1% for wavelengths $k \leq 2 h Mpc^{-1}$, while the error increases at larger $k$. We show the size of data sample required to reach a particular error rate, giving a sense of how much data is necessary to reach a desired accuracy. This work provides a foundation for developing methods to analyse very large upcoming datasets with the next-generation observational facilities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Accelerating computational fluid dynamics simulation of post-combustion carbon capture modeling with MeshGraphNets

Packed columns are commonly used in post-combustion processes to capture CO 2 emissions by providing enhanced contact area between a CO 2 -laden gas and CO 2 -absorbing solvent. To study and optimize solvent-based post-combustion carbon capture systems (CCSs), computational fluid dynamics (CFD) can be used to model the liquid–gas countercurrent flow hydrodynamics in these columns and derive key determinants of CO 2 -capture efficiency. However, the large design space of these systems hinders the application of CFD for design optimization due to its high computational cost. In contrast, data-driven modeling approaches can produce fast surrogates to study large-scale physics problems. We build our surrogates using MeshGraphNets (MGN), a graph neural network framework that efficiently learns and produces mesh-based simulations. We apply MGN to a random packed column modeled with over 160K graph nodes and a design space consisting of three key input parameters: solvent surface tension, inlet velocity, and contact angle. Our models can adapt to a wide range of these parameters and accurately predict the complex interactions within the system at rates over 1700 times faster than CFD, affirming its practicality in downstream design optimization tasks. This underscores the robustness and versatility of MGN in modeling complex fluid dynamics for large-scale CCS analyses.

97 MATHEMATICS AND COMPUTING↗