Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interpretable models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Improving multiwell petrophysical interpretation from well logs via machine learning and statistical models

Well-log interpretation estimates in situ rock properties along well trajectory, such as porosity, water saturation, and permeability, to support reserve-volume estimation, production forecasts, and decision making in reservoir development. However, due to measurement errors, variability of well logs caused by multiple measurement vendors, different borehole tools, and nonuniform drilling/borehole conditions, estimations of rock properties with original well logs without proper preprocessing may not be accurate, especially in the context of multiwell estimation. Well-log normalization techniques such as two-point scaling and mean-variance normalization are commonly used to improve the robustness of multiwell rock-property estimation. However, these techniques do not consider the correlation between well logs and require subjective knowledge for their effective implementation. To reduce uncertainties and processing time associated with multiwell rock-property estimation from well logs, we develop discriminative adversarial (DA) and linear constraint models for well-log normalization and rock-property estimation. The DA neural network model developed for well-log normalization and interpretation can perform linear and nonlinear well-log normalization while considering the joint distribution of each well log and rock properties. However, the linear constraint model uses an ensemble of predictions from linear models to constrain well-log normalization and rock-property estimation. We also develop a divergence-based type well identification method to select type (training) wells for a test well based on the statistical similarity of associated well-log distributions instead of the interwell distance. We apply the DA model to perform well-log normalization and prediction of permeability for the Seminole San Andres Unit carbonate reservoir. Compared with the permeability predicted with the classical machine learning model without well-log normalization and models with two-point scaling normalization, the DA model yields the most accurate permeability prediction by decreasing the mean-squared error of permeability prediction by 20%–50%.

Geochemistry & Geophysics↗

Disturbance and Response Model (DRM)

The DRM is a framework to interpret ecosystem process model output for 3D fire behavior model input and interpret the 3D fire behavior model output for ecosystem process model inpu

Atchley, Adam↗

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models↗

Symbolic diagnostics to interpret and analyze neural network models

Embedded machine-learned models (EMLMs) have the promise to improve the predictive accuracy of engineering simulators in environments of national interest. EMLMs often comprise complex input-output maps (e.g., neural networks), which make them unamenable to rigorous analysis and generally difficult to interpret. In the face of decades of theory, this lack of interpretability is a significant barrier to building confidence in these models. This work outlines an approach to interpret EMLMs using sparse polynomial regression for comparison with theoretical understanding. To do so, we build on the concept of Locally Interpretable Model-agnostic Explanations (LIME) using physics-informed clustering, prototype selection, and library construction. While general, we demonstrate our method on tensor-basis neural networks used in Reynolds-Averaged Navier-Stokes simulations of hypersonic fluid flows. Results are presented for a simulated toy model and for direct numerical simulations (DNS) of turbulent flows over a flat plate.

97 MATHEMATICS AND COMPUTING↗

Temperature dependence of nitrate-reducing Fe(II) oxidation by Acidovorax strain BoFeN1 – evaluating the role of enzymatic vs. abiotic Fe(II) oxidation by nitrite

ABSTRACT Fe(II) oxidation coupled to nitrate reduction is a widely observed metabolism. However, to what extent the observed Fe(II) oxidation is driven enzymatically or abiotically by metabolically produced nitrite remains puzzling. To distinguish between biotic and abiotic reactions, we cultivated the mixotrophic nitrate-reducing Fe(II)-oxidizing Acidovorax strain BoFeN1 over a wide range of temperatures and compared it to abiotic Fe(II) oxidation by nitrite at temperatures up to 60°C. The collected experimental data were subsequently analyzed through biogeochemical modeling. At 5°C, BoFeN1 cultures consumed acetate and reduced nitrate but did not significantly oxidize Fe(II). Abiotic Fe(II) oxidation by nitrite at different temperatures showed an Arrhenius-type behavior with an activation energy of 80±7 kJ/mol. Above 40°C, the kinetics of Fe(II) oxidation were abiotically driven, whereas at 30°C, where BoFeN1 can actively metabolize, the model-based interpretation strongly suggested that an enzymatic pathway was responsible for a large fraction (ca. 62%) of the oxidation. This result was reproduced even when no additional carbon source was present. Our results show that at below 30°C, i.e. at temperatures representing most natural environments, biological Fe(II) oxidation was largely responsible for overall Fe(II) oxidation, while abiotic Fe(II) oxidation by nitrite played a less important role.

Microbiology↗

Theory and modeling of molecular modes in the NMR relaxation of fluids

Traditional theories of the nuclear magnetic resonance (NMR) autocorrelation function for intra-molecular dipole pairs assume a single-exponential decay, yet the calculated autocorrelation of realistic systems displays a rich, multi-exponential behavior, resulting in anomalous NMR relaxation dispersion (i.e., frequency dependence). We develop an approach to model and interpret the multi-exponential intra-molecular autocorrelation using simple, physical models within a rigorous statistical mechanical development that encompasses both rotational diffusion and translational diffusion in the same framework. Here, we recast the problem of evaluating the autocorrelation in terms of averaging over a diffusion propagator whose evolution is described by a Fokker–Planck equation. The time-independent part admits an eigenfunction expansion, allowing us to write the propagator as a sum over modes. Each mode has a spatial part that depends on the specified eigenfunction and a temporal part that depends on the corresponding eigenvalue (i.e., correlation time) with a simple, exponential decay. The spatial part is a probability distribution of the dipole pair, analogous to the stationary states of a quantum harmonic oscillator. Drawing inspiration from the idea of inherent structures in liquids, we interpret each of the spatial contributions as a specific molecular mode. These modes can be used to model and predict the NMR dipole–dipole relaxation dispersion of fluids by incorporating phenomena on the molecular level. We validate our statistical mechanical description of the distribution in molecular modes with molecular dynamics simulations interpreted without any relaxation models or adjustable parameters: the most important poles in the Padé–Laplace transform of the simulated autocorrelation agree with the eigenvalues predicted by the theory

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Interpretable, Extensible Linear Regression Model for Electron Density Prediction

This code allows us to generate a model to predict ground state electron densities directly from atomic structure thereby bypassing the need to perform direct DFT calculations. It is based on many body correlation descriptors defined in https://arxiv.org/abs/2004.14442 (IM release numbers: LLNL-JRNL- 808791 and LLNL-JRNL-840698). The code reads in electron densities, in CHGCAR format used by the DFT code VASP, calculates many-body correlation descriptors based on atoms surrounding a grid point and then uses linear regression to obtain the fitting parameters.

Kumar, Shashikant↗

When ancient numerical demons meet physics-informed machine learning: adjoint-based gradients for implicit differentiable modeling

Recent advances in differentiable modeling, a genre of physics-informed machine learning that trains neural networks (NNs) together with process-based equations, have shown promise in enhancing hydrological models' accuracy, interpretability, and knowledge-discovery potential. Current differentiable models are efficient for NN-based parameter regionalization, but the simple explicit numerical schemes paired with sequential calculations (operator splitting) can incur numerical errors whose impacts on models' representation power and learned parameters are not clear. Implicit schemes, however, cannot rely on automatic differentiation to calculate gradients due to potential issues of gradient vanishing and memory demand. Here we propose a “discretize-then-optimize” adjoint method to enable differentiable implicit numerical schemes for the first time for large-scale hydrological modeling. The adjoint model demonstrates comprehensively improved performance, with Kling–Gupta efficiency coefficients, peak-flow and low-flow metrics, and evapotranspiration that moderately surpass the already-competitive explicit model. Therefore, the previous sequential-calculation approach had a detrimental impact on the model's ability to represent hydrological dynamics. Furthermore, with a structural update that describes capillary rise, the adjoint model can better describe baseflow in arid regions and also produce low flows that outperform even pure machine learning methods such as long short-term memory networks. The adjoint model rectified some parameter distortions but did not alter spatial parameter distributions, demonstrating the robustness of regionalized parameterization. Despite higher computational expenses and modest improvements, the adjoint model's success removes the barrier for complex implicit schemes to enrich differentiable modeling in hydrology.

58 GEOSCIENCES↗

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue↗

The interpretation of temperature and salinity variables in numerical ocean model output and the calculation of heat fluxes and heat content

Abstract. The international Thermodynamic Equation of Seawater 2010 (TEOS-10) defined the enthalpy and entropy of seawater, thus enabling the global ocean heat content to be calculated as the volume integral of the product of in situ density, ρ, and potential enthalpy, h0 (with reference sea pressure of 0 dbar). In terms of Conservative Temperature, Θ, ocean heat content is the volume integral of ρcp0Θ, where cp0 is a constant “isobaric heat capacity”. However, many ocean models in the Coupled Model Intercomparison Project Phase 6 (CMIP6) as well as all models that contributed to earlier phases, such as CMIP5, CMIP3, CMIP2, and CMIP1, used EOS-80 (Equation of State – 1980) rather than the updated TEOS-10, so the question arises of how the salinity and temperature variables in these models should be physically interpreted, with a particular focus on comparison to TEOS-10-compliant observations. In this article we address how heat content, surface heat fluxes, and the meridional heat transport are best calculated using output from these models and how these quantities should be compared with those calculated from corresponding observations. We conclude that even though a model uses the EOS-80, which expects potential temperature as its input temperature, the most appropriate interpretation of the model's temperature variable is actually Conservative Temperature. This perhaps unexpected interpretation is needed to ensure that the air–sea heat flux that leaves and arrives in atmosphere and sea ice models is the same as that which arrives in and leaves the ocean model. We also show that the salinity variable carried by present TEOS-10-based models is Preformed Salinity, while the salinity variable of EOS-80-based models is also proportional to Preformed Salinity. These interpretations of the salinity and temperature variables in ocean models are an update on the comprehensive Griffies et al. (2016) paper that discusses the interpretation of many aspects of coupled Earth system models.

54 ENVIRONMENTAL SCIENCES↗

A review of computing-based automated fault detection and diagnosis of heating, ventilation and air conditioning systems

We report faults in Heating, Ventilation, and Air Conditioning (HVAC) systems of buildings result in significant energy waste in building operation. With fast-growing sensing data availability and advancement in computing, computational modeling has demonstrated strong capability to detect and diagnose HVAC system faults, hence, ensuring efficient building operation. This paper comprehensively reviews the state-of-the-art computing-based fault detection and diagnosis (FDD) for HVAC systems. Overall, the reviewed computing-based FDD methods are classified as two major approaches: knowledge-based and data-driven approaches. We then identify multiple important topics, including data availability, training data size, data quality, approach generality, capability, interpretability, and required modeling efforts, along with corresponding metrics to summarize the most updated FDD development. Generally, the knowledge-based approaches are further divided as physics-based modeling, Diagnostic Bayesian Network, and performance indicator-based methods while data-driven approaches include supervised learning, unsupervised learning, and regression and statistics-based methods. State-of-the-art FDD development, remaining challenges, and future research directions are further discussed to push forward FDD in practice. Availability of fault data, capability of existing methods to deal with complex fault situations (such as simultaneous faults), modeling interpretability for data-driven methods, and required engineering efforts for physics-based methods are identified as remaining challenges in FDD development. Improving modeling fidelity and reducing modeling efforts are essential for applying physics-based methods in real buildings. Meanwhile, addressing fault data availability, increasing algorithm adaptability, and handling multiple faults are essential to further enhance the applicability of data-driven FDD approaches.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental Soil Warming Impacts Soil Moisture and Plant Water Stress and Thereby Ecosystem Carbon Dynamics

Experimental soil heating experiments have found a consistent increase in soil-surface CO 2 emissions ( F s ), but inconsistent soil organic carbon (SOC) responses. Interpretation of heating effects is complicated by spatial heterogeneity and soil moisture, nitrogen availability, and microbial and plant responses. Here we applied a mechanistic ecosystem model to interpret heating impacts on a California forest subjected to 1 m deep, 4°C heating. The model accurately simulated control-plot CO 2 fluxes, SOC stocks, fine root biomass, soil moisture, and soil temperature, and the observed increases in F s and decreases in fine root biomass. We show that a complex suite of interactions can lead to a consistent increase in F s (~17%) over the 5-year study period, with very small changes in SOC stocks (<1%). Modeled increases in leaf water stress from soil drying reduced GPP and NPP. The resulting reduction in leaf and fine root allocation increased fine root litter inputs to the soil and reduced root exudation. Soil heating led to about a 50% larger increase in root autotrophic respiration than in heterotrophic respiration, with the heating effect on both these fluxes decreasing over the simulation period. Increased heterotrophic respiration led to increased soil N availability and plant N uptake. These heating responses are mechanistically linked, of magnitudes that can affect ecosystem dynamics, and long-term observations of them are rarely made. Therefore, we conclude that a coupled observational and mechanistic modeling framework is needed to interpret manipulation experiments, and to improve projections of climate change impacts on terrestrial ecosystem carbon dynamics.

54 ENVIRONMENTAL SCIENCES↗

Role of Neutrals Versus Transport in Determining the Pedestal Density Structure: Final Technical Report

In fusion devices the plasma density plays a crucial role in determining the fusion reaction rate and has a direct impact on the fusion gain of a given device. This density is in general regulated by the particle sources and transport near the plasma edge, which give rise to an edge density pedestal. When predicting the performance of future devices, this density pedestal is often prescribed, rather than predicted, due to a lack of models which allow confident extrapolation. This project aims to advance these models through the focused validation of theoretical models related to the transport of fueling neutral particles, and through interpretive transport modeling in present day fusion plasmas, in which the penetration of neutrals is altered to better simulate future reactor-like conditions. Achievements in theory and model validation under this project have advanced our understanding of how much of the edge density profile is set by transport versus direct ionization, enabling interesting projections to future burning plasma devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding the Drivers of Atlantic Multidecadal Variability using a Stochastic Model Hierarchy

The relative importance of ocean and atmospheric dynamics in generating Atlantic Multidecadal Variability (AMV) remains an open question. Comparisons between climate models with SLAB and fully-dynamic (FULL) ocean components are often used to explore this question, but cannot reveal how individual ocean processes generate these differences. We build a hierarchy of physically interpretable stochastic models to investigate the contribution of two upper-ocean processes to AMV: the role of seasonal variation and mixed-layer entrainment. This interpretability arises from the stochastic model’s simplified representation of sea surface temperature (SST), considering only the local upper ocean response to white-noise atmospheric forcing and its impact on surface heat exchange. We focus on understanding differences between SLAB and FULL non-eddy resolving pre-industrial control simulations of the Community Earth System Model 1 (CESM), and estimate the stochastic model parameters from each respective simulation. Despite its simplicity, the stochastic model reproduces temporal characteristics of SST variability in the SPG, including reemergence, seasonal-to-interannual persistence and power spectra. Furthermore, unrealistically persistent SST of the CESM-SLAB ocean simulation is reproduced in the equivalent stochastic model configuration where the mixed-layer depth (MLD) is constant. The stochastic model also reveals that vertical entrainment primarily damps SST variability, thus explaining why SLAB exhibits larger SST variance than FULL. Here, the stochastic model driven by temporally stochastic, spatially coherent forcing patterns reproduces the canonical AMV pattern. However, the amplitude of low-frequency variability remains underestimated, suggesting a role for ocean dynamics beyond entrainment.

54 ENVIRONMENTAL SCIENCES↗