Regression-based projection for learning Mori-Zwanzig operators
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Wildfires have shown increasing trends in both frequency and severity across the contiguous United States (CONUS). However, process-based fire models have difficulties in accurately simulating the burned area over the CONUS due to a simplification of the physical process and cannot capture the interplay among fire, ignition, climate, and human activities. The deficiency of burned area simulation deteriorates the description of fire impact on energy balance, water budget, and carbon fluxes in the Earth system models (ESMs). Alternatively, fire models based on machine learning (ML), which capture statistical relationships between the burned area and environmental factors, have shown promising burned area predictions and corresponding fire impact simulation. We develop a hybrid framework (ELM2.1-XGBFire1.0) that integrates an eXtreme Gradient Boosting (XGBoost) wildfire model with the Energy Exascale Earth System Model (E3SM) land model (ELM) version 2.1. A Fortran–C–Python deep learning bridge is adapted to support online communication between ELM and the ML fire model. Specifically, the burned area predicted by the ML-based wildfire model is directly passed to ELM to adjust the carbon pool and vegetation dynamics after disturbance, which are then used as predictors in the ML-based fire model in the next time step. Evaluated against the historical burned area from Global Fire Emissions Database 5 from 2001–2019, the ELM2.1-XGBFire1.0 outperforms process-based fire models in terms of spatial distribution and seasonal variations. The ELM2.1-XGBFire1.0 has proven to be a new tool for studying vegetation–fire interactions and, more importantly, enables seamless exploration of climate–fire feedback, working as an active component of E3SM.
Software Health Management (SWHM) must make sure that the software always remains in safe and healthy regions of the state space. Boundaries between healthy and unhealthy regions are important for the detection of violations and health management.In this position paper, we present a framework, which employs techniques from Bayesian statistical modeling and active learning to efficiently characterize health boundaries in high-dimensional spaces. We will discuss, how this framework supports SWHM during design time and during operation of learning/adapting software systems
Near-real-time (NRT) streamflow data are critical importance for timely water resources management. Here, we developed an open-source tool, FlowStats, for NRT streamflow analysis and visualization in Germany, based on NRT meteorological data from the German Weather Service and simulated streamflow from a long short-term memory neural network (LSTM). The LSTM model achieved very good overall performance, median NSE of 0.80 for the test period across 1,479 catchments. FlowStats provides options for deriving various streamflow statistics, from normal and abnormal streamflow detection to drought and flood analyses. An example analysis from FlowStats revealed widespread below-normal to extreme low-flow conditions across Germany from March to May 2025, which weakened from June to September 2025. Drought analysis for September 2025 highlighted severe to extreme drought conditions in northwestern Germany, while flood classifications indicated that high-flow events occurred in southwestern Germany. FlowStats can be used for various hydrological assessments to support water resources management.
With the need to maintain long-term reliable energy using nuclear power plants, there is an underlying demand to ensure that the maintenance of plant components and systems is also done in an efficient and cost-effective manner. One way to achieve this is by moving from time-based maintenance to condition-based maintenance. The research presented in this paper focuses on applying statistical and machine-learning-based methods to capture anomalies within data for fault detection to further develop into condition monitoring. This paper focuses on system data for a circulating water system (CWS) of a pressurized heavy-water reactor for detecting anomalies. The different methodologies used for detecting and capturing anomalies in the CWS data are matrix profile, density-based spatial clustering of applications with noise (DBSCAN), and support vector machines (SVMs). Matrix profile and DBSCAN are used to distinguish between normal data and anomalous data. This paper presents a hybrid method using DBSCAN and SVM when a portion of the data is used for DBSCAN to generate clusters. This portion of data is then used to train the SVM along with the clusters generated by DBSCAN as output. SVM is then tested on unseen data as a predictive tool, which can work in real time to categorize data points as either normal or anomalous. This paper presents results that show the high accuracies of DBSCAN and SVM in capturing anomalies within the data for a CWS for fault detection. Thus, the maintenance plan would be focused on component condition rather than a time-based schedule by switching to an automated system to identify and predict faults within a CWS.
The ability of feed-forward neural net architectures to learn continuous-valued mappings in the presence of noise is demonstrated in relation to parameter identification and real-time adaptive control applications. Factors and parameters influencing the learning performance of such nets in the presence of noise are identified. Their effects are discussed through a computer simulation of the Back-Error-Propagation algorithm by taking the example of the cart-pole system controlled by a nonlinear control law. Adequate sampling of the state space is found to be essential for canceling the effect of the statistical fluctuations and allowing learning to take place.
Deep learning has been successfully applied to the segmentation of 3D Computed Tomography (CT) scans. Establishing the credibility of these segmentations requires uncertainty quantification (UQ) to identify untrustworthy predictions. Recent UQ architectures include Monte Carlo dropout networks (MCDNs), which approximate deep Gaussian processes, and Bayesian neural networks (BNNs), which learn the distribution of the weight space. BNNs are advantageous over MCDNs for UQ but are thought to be computationally infeasible in high dimension, and neither architecture has produced interpretable geometric uncertainty maps. We propose a novel 3D Bayesian convolutional neural network (BCNN), the first deep learning method which generates statistically credible geometric uncertainty maps and scales for application to 3D data. We present experimental results on CT scans of graphite electrodes and laser-welded metals and show that our BCNN outperforms an MCDN in recent uncertainty metrics. The geometric uncertainty maps generated by our BCNN capture distributions of sigmoid values that are interpretable as confidence intervals, critical for applications that rely on deep learning for high-consequence decisions.
Deep Neural Networks (DNNs) are considered to be key components in many autonomous systems. Applications range from vision-based obstacle avoidance to intelligent/learning control and planning. Safety-critical applications as found in the aerospace domain require that the behavior of the DNN is validated and tested rigorously for safety of the autonomous system (AUS). In this paper, we present a framework to support testing of DNNs and the analysis of the network structure. Our framework employs techniques from statistical modeling and active learning to effectively generate test cases for DNN safety testing and performance analysis. We will present results of a case study on a physics-based Deep recurrent residual neural network (DR-RNN), which has been trained to emulate the aerodynamics behavior of a fixed-wing aircraft.
This study explores the intersection of digitalization and privacy within the energy sector, focusing on the emerging challenges and opportunities presented by integrating Distributed Energy Resources (DERs) and advanced metering infrastructure. The need for robust digital privacy measures has become crucial as the energy industry evolves towards a more decentralized, digitalized, and decarbonized future. This study delves into four cutting-edge privacy-preserving technologies—Homomorphic Encryption (HE), Secure Multiparty Computation (SMPC), Differential Privacy (DP), and Federated Learning (FL)—each offering unique solutions to safeguard consumer data by increasing digital connectivity and data exchange. Through a detailed examination of these methods, the study explains how each technology operates, its applications within the energy sector, and the specific privacy challenges it addresses. Homomorphic Encryption allows for secure computations on encrypted data, enabling data analysis without compromising privacy. Secure Multiparty Computation enables collaborative data analysis across different entities while protecting the confidentiality of the inputs. Differential Privacy introduces randomness into the assembled data set, preventing the identification of individual records in statistical databases. Lastly, Federated Learning offers a paradigm shift in data analysis, where machine learning models are trained at the edge, minimizing the centralization of sensitive data. The research underscores the significance of implementing these privacy-enhancing technologies to comply with strict data protection regulations, foster consumer trust, and enhance the security of the energy infrastructure. By providing a comprehensive overview of these methodologies and their practical implications for the energy sector, this study aims to contribute to the ongoing discourse on digital privacy, offering insights into how the energy industry can navigate the complexities of data privacy in the digital age.
Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.
To accurately model radiative fluxes at the surface and within the atmosphere, we need to know both vertical and horizontal structures of cloudiness. While MODIS provides accurate information on cloud horizontal structure, it has limited ability to estimate cloud vertical structure. ICESat/GLAS on the other hand, provides the vertical distribution and internal structure of clouds as deep as the laser beam can penetrate and return a signal. Having different orbits, MODIS and GLAS provide few collocated measurements; hence a statistical approach is needed to learn about 3D cloud structures from the two instruments. In the presentation, we show the results of the statistical analysis of vertical and horizontal structure of cloudiness using GLAS and MODIS cloud top(s) data acquired in October-November 2003. We revisit the (H1, C1) plot, previously used for analyzing cloud liquid water data, and illustrate cloud structure for single and multiple-layer clouds.
We outline a machine learning strategy for quantitively determining the conformation of AB-type diblock copolymers with excluded volume effects using small angle scattering. Complemented by computer simulations, a correlation matrix connecting conformations of different copolymers according to their scattering features is established on the mathematical framework of a Gaussian process, a multivariate extension of the familiar univariate Gaussian distribution. We show that the relevant conformational characteristics of copolymers can be probabilistically inferred from their coherent scattering cross sections without any restriction imposed by model assumptions. This work not only facilitates the quantitative structural analysis of copolymer solutions but also provides the reliable benchmarking for the related theoretical development of scattering functions.
Aqueous phase heterogeneous catalysis is important to various industrial processes, including biomass conversion, Fischer–Tropsch synthesis, and electrocatalysis. Accurate calculation of solvation thermodynamic properties is essential for modeling the performance of catalysts for these processes. Explicit solvation methods employing multiscale modeling, e.g., involving density functional theory and molecular dynamics have emerged for this purpose. Although accurate, these methods are computationally intensive. This study introduces machine learning (ML) models to predict solvation thermodynamics for adsorbates on a Pt(111) surface, aiming to enhance computational efficiency without compromising accuracy. In particular, ML models are developed using a combination of molecular descriptors and fingerprints and trained on previously published water–adsorbate interaction energies, energies of solvation, and free energies of solvation of adsorbates bound to Pt(111). These models achieve root mean square error values of 0.09 eV for interaction energies, 0.04 eV for energies of solvation, and 0.06 eV for free energies of solvation, demonstrating accuracy within the standard error of multiscale modeling. Feature importance analysis reveals that hydrogen bonding, van der Waals interactions, and solvent density, together with the properties of the adsorbate, are critical factors influencing solvation thermodynamics. Furthermore, these findings suggest that ML models can provide rapid and reliable predictions of solvation properties. This approach not only reduces computational costs but also offers insights into the solvation characteristics of adsorbates at Pt(111)–water interfaces.
Machine learning interatomic potentials (MLIPs) have greatly enhanced molecular dynamics (MD) simulations, achieving near-first-principles accuracy in thermal conductivity studies. In this work, we reveal that this accuracy, observed in BAs and diamond at sub-Debye temperatures, stems from an accidental error cancelation: classical statistics overestimates specific heat while underestimating phonon lifetimes, balancing out in thermal conductivity predictions. However, this balance is disrupted when isotopes are introduced, leading MLIP-based MD to significantly underpredict thermal conductivity compared to experiments and quantum statistics-based Boltzmann transport equation. This discrepancy arises not from classical statistics affecting phonon–isotope scattering rates but from its impact on the interplay between phonon–isotope and phonon–phonon scattering in the normal scattering-dominated BAs and diamond. In conclusion, this work underscores the limitations of MLIP-based MD for thermal conductivity studies at sub-Debye temperatures.
The global energy system is undergoing significant changes, including a shift in energy generating technologies to more renewable energy sources. However, the dependence of renewable energy sources on local environmental conditions could also increase disruptions in service through exposures to compound, extreme weather events. By fusing three diverse datasets (operations and maintenance tickets, weather data, and production data), this analysis presents a novel methodology to identify and evaluate performance impacts arising from extreme weather events across diverse geographical regions. Text analysis of maintenance tickets identified snow, hurricanes, and storms as the leading extreme weather events affecting photovoltaic plants in the United States. Statistical techniques and machine learning were then implemented to identify the magnitude and variability of these extreme weather impacts on site performance. Impacts varied between event and non-event days, with snow events causing the greatest reductions in performance (54.5%), followed by hurricanes (12.6%) and storms (1.1%). Machine learning analysis identified key features in determining if a day is categorized as low performing, such as low irradiance, geographic location, weather features, and site size. The analysis improves our understanding of compound, extreme weather event impacts on photovoltaic systems, which can inform planning activities, especially as the industry continues to expand into new geographic and climatic regions around the world.
Quantifying hydrologic exchange fluxes (HEFs) at the stream-groundwater interface and their residence time distributions (RTDs) in the subsurface are important for managing the water quality and ecosystem health in dynamic river corridors. However, direct simulating high-spatial resolution HEFs and RTDs can be time-consuming, especially for watershed-scale modeling. Efficient surrogate models linking RTDs to hydromorphic units (HUs) can be alternatives for simulating RTDs in large-scale models. A common concern of these surrogate models, though, is the transferability of the relationship between the RTDs and HUs from one river corridor to another. To address this issue, this work evaluates the HEFs and resulting RTD-HU relationships for two 10-km long river corridors along the Columbia River leveraging a one-way coupled three-dimensional transient surface-subsurface water transport modeling framework we previously developed. Applying such a framework at the two river corridors with similar HUs allows for quantitative comparisons of HEFs and RTDs using both statistical tests and machine learning classification models. Finally, our comparison shows that the similarity and transferability of the RTD-HU relationship is very low for the two investigated river sections, which suggests that devising a general algorithm to estimate RTDs based solely on surface water hydrodynamics and short-distance river channel topography data, as well as HU classification, might be nearly impossible.
Polygenic risk scores (PRS) estimate the genetic risk of an individual for a complex disease based on many genetic variants across the whole genome. Here, we compared a series of computational models for estimation of breast cancer PRS. A deep neural network (DNN) was found to outperform alternative machine learning techniques and established statistical algorithms, including BLUP, BayesA, and LDpred. In the test cohort with 50% prevalence, the Area Under the receiver operating characteristic Curve (AUC) were 67.4% for DNN, 64.2% for BLUP, 64.5% for BayesA, and 62.4% for LDpred. BLUP, BayesA, and LPpred all generated PRS that followed a normal distribution in the case population. However, the PRS generated by DNN in the case population followed a bimodal distribution composed of two normal distributions with distinctly different means. This suggests that DNN was able to separate the case population into a high-genetic-risk case subpopulation with an average PRS significantly higher than the control population and a normal-genetic-risk case subpopulation with an average PRS similar to the control population. This allowed DNN to achieve 18.8% recall at 90% precision in the test cohort with 50% prevalence, which can be extrapolated to 65.4% recall at 20% precision in a general population with 12% prevalence. Interpretation of the DNN model identified salient variants that were assigned insignificant p values by association studies, but were important for DNN prediction. These variants may be associated with the phenotype through nonlinear relationships.