Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model interpretability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Interpretable, Extensible Linear Regression Model for Electron Density Prediction

This code allows us to generate a model to predict ground state electron densities directly from atomic structure thereby bypassing the need to perform direct DFT calculations. It is based on many body correlation descriptors defined in https://arxiv.org/abs/2004.14442 (IM release numbers: LLNL-JRNL- 808791 and LLNL-JRNL-840698). The code reads in electron densities, in CHGCAR format used by the DFT code VASP, calculates many-body correlation descriptors based on atoms surrounding a grid point and then uses linear regression to obtain the fitting parameters.

Kumar, Shashikant↗

When ancient numerical demons meet physics-informed machine learning: adjoint-based gradients for implicit differentiable modeling

Recent advances in differentiable modeling, a genre of physics-informed machine learning that trains neural networks (NNs) together with process-based equations, have shown promise in enhancing hydrological models' accuracy, interpretability, and knowledge-discovery potential. Current differentiable models are efficient for NN-based parameter regionalization, but the simple explicit numerical schemes paired with sequential calculations (operator splitting) can incur numerical errors whose impacts on models' representation power and learned parameters are not clear. Implicit schemes, however, cannot rely on automatic differentiation to calculate gradients due to potential issues of gradient vanishing and memory demand. Here we propose a “discretize-then-optimize” adjoint method to enable differentiable implicit numerical schemes for the first time for large-scale hydrological modeling. The adjoint model demonstrates comprehensively improved performance, with Kling–Gupta efficiency coefficients, peak-flow and low-flow metrics, and evapotranspiration that moderately surpass the already-competitive explicit model. Therefore, the previous sequential-calculation approach had a detrimental impact on the model's ability to represent hydrological dynamics. Furthermore, with a structural update that describes capillary rise, the adjoint model can better describe baseflow in arid regions and also produce low flows that outperform even pure machine learning methods such as long short-term memory networks. The adjoint model rectified some parameter distortions but did not alter spatial parameter distributions, demonstrating the robustness of regionalized parameterization. Despite higher computational expenses and modest improvements, the adjoint model's success removes the barrier for complex implicit schemes to enrich differentiable modeling in hydrology.

58 GEOSCIENCES↗

A Data-Driven Method for Modeling Creep-Fatigue Stress- Strain Behavior Using Neural ODEs

In this paper, we introduce a data-driven machine learning approach for modeling one-dimensional stress–strain behavior under cyclic loading, utilizing experimental data from the nickel-based Alloy 617. The study employs uniaxial creep–fatigue test data acquired under various loading histories and compares two distinct neural network-based ODE models. The first model, known as the black-box model, comprehensively describes the strain–stress relationship using a Neural ODE equation. To interpret this black-box model, we apply the Sparse Identification of Nonlinear Dynamical Systems (SINDy) technique, transforming the black-box model into an equation-based model using symbolic regression. The second model, the Neural flow rule model, incorporates Hooke’s Law for the linear elastic component, with the nonlinear part characterized by a Neural ODE. Both models are trained with experimental data to accurately reflect the observed stress–strain behavior. We conduct a detailed comparison with the standard Chaboche model, which includes three back stresses. Our results demonstrate that the neural network-based ODE models precisely capture the experimental creep–fatigue mechanical behavior, exceeding the standard Chaboche model’s accuracy. Furthermore, an interpretable model derived from the black-box neural ODE model through symbolic regression achieves accuracy comparable to the Chaboche model, enhancing its interpretability. The results highlight the potential of neural network-based ODE models to depict complex creep–fatigue behavior, eliminating the necessity for experts to define a specific, material-focused model form.

creep-fatigue↗

The interpretation of temperature and salinity variables in numerical ocean model output and the calculation of heat fluxes and heat content

Abstract. The international Thermodynamic Equation of Seawater 2010 (TEOS-10) defined the enthalpy and entropy of seawater, thus enabling the global ocean heat content to be calculated as the volume integral of the product of in situ density, ρ, and potential enthalpy, h0 (with reference sea pressure of 0 dbar). In terms of Conservative Temperature, Θ, ocean heat content is the volume integral of ρcp0Θ, where cp0 is a constant “isobaric heat capacity”. However, many ocean models in the Coupled Model Intercomparison Project Phase 6 (CMIP6) as well as all models that contributed to earlier phases, such as CMIP5, CMIP3, CMIP2, and CMIP1, used EOS-80 (Equation of State – 1980) rather than the updated TEOS-10, so the question arises of how the salinity and temperature variables in these models should be physically interpreted, with a particular focus on comparison to TEOS-10-compliant observations. In this article we address how heat content, surface heat fluxes, and the meridional heat transport are best calculated using output from these models and how these quantities should be compared with those calculated from corresponding observations. We conclude that even though a model uses the EOS-80, which expects potential temperature as its input temperature, the most appropriate interpretation of the model's temperature variable is actually Conservative Temperature. This perhaps unexpected interpretation is needed to ensure that the air–sea heat flux that leaves and arrives in atmosphere and sea ice models is the same as that which arrives in and leaves the ocean model. We also show that the salinity variable carried by present TEOS-10-based models is Preformed Salinity, while the salinity variable of EOS-80-based models is also proportional to Preformed Salinity. These interpretations of the salinity and temperature variables in ocean models are an update on the comprehensive Griffies et al. (2016) paper that discusses the interpretation of many aspects of coupled Earth system models.

54 ENVIRONMENTAL SCIENCES↗

A review of computing-based automated fault detection and diagnosis of heating, ventilation and air conditioning systems

We report faults in Heating, Ventilation, and Air Conditioning (HVAC) systems of buildings result in significant energy waste in building operation. With fast-growing sensing data availability and advancement in computing, computational modeling has demonstrated strong capability to detect and diagnose HVAC system faults, hence, ensuring efficient building operation. This paper comprehensively reviews the state-of-the-art computing-based fault detection and diagnosis (FDD) for HVAC systems. Overall, the reviewed computing-based FDD methods are classified as two major approaches: knowledge-based and data-driven approaches. We then identify multiple important topics, including data availability, training data size, data quality, approach generality, capability, interpretability, and required modeling efforts, along with corresponding metrics to summarize the most updated FDD development. Generally, the knowledge-based approaches are further divided as physics-based modeling, Diagnostic Bayesian Network, and performance indicator-based methods while data-driven approaches include supervised learning, unsupervised learning, and regression and statistics-based methods. State-of-the-art FDD development, remaining challenges, and future research directions are further discussed to push forward FDD in practice. Availability of fault data, capability of existing methods to deal with complex fault situations (such as simultaneous faults), modeling interpretability for data-driven methods, and required engineering efforts for physics-based methods are identified as remaining challenges in FDD development. Improving modeling fidelity and reducing modeling efforts are essential for applying physics-based methods in real buildings. Meanwhile, addressing fault data availability, increasing algorithm adaptability, and handling multiple faults are essential to further enhance the applicability of data-driven FDD approaches.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Experimental Soil Warming Impacts Soil Moisture and Plant Water Stress and Thereby Ecosystem Carbon Dynamics

Experimental soil heating experiments have found a consistent increase in soil-surface CO 2 emissions ( F s ), but inconsistent soil organic carbon (SOC) responses. Interpretation of heating effects is complicated by spatial heterogeneity and soil moisture, nitrogen availability, and microbial and plant responses. Here we applied a mechanistic ecosystem model to interpret heating impacts on a California forest subjected to 1 m deep, 4°C heating. The model accurately simulated control-plot CO 2 fluxes, SOC stocks, fine root biomass, soil moisture, and soil temperature, and the observed increases in F s and decreases in fine root biomass. We show that a complex suite of interactions can lead to a consistent increase in F s (~17%) over the 5-year study period, with very small changes in SOC stocks (<1%). Modeled increases in leaf water stress from soil drying reduced GPP and NPP. The resulting reduction in leaf and fine root allocation increased fine root litter inputs to the soil and reduced root exudation. Soil heating led to about a 50% larger increase in root autotrophic respiration than in heterotrophic respiration, with the heating effect on both these fluxes decreasing over the simulation period. Increased heterotrophic respiration led to increased soil N availability and plant N uptake. These heating responses are mechanistically linked, of magnitudes that can affect ecosystem dynamics, and long-term observations of them are rarely made. Therefore, we conclude that a coupled observational and mechanistic modeling framework is needed to interpret manipulation experiments, and to improve projections of climate change impacts on terrestrial ecosystem carbon dynamics.

54 ENVIRONMENTAL SCIENCES↗

Role of Neutrals Versus Transport in Determining the Pedestal Density Structure: Final Technical Report

In fusion devices the plasma density plays a crucial role in determining the fusion reaction rate and has a direct impact on the fusion gain of a given device. This density is in general regulated by the particle sources and transport near the plasma edge, which give rise to an edge density pedestal. When predicting the performance of future devices, this density pedestal is often prescribed, rather than predicted, due to a lack of models which allow confident extrapolation. This project aims to advance these models through the focused validation of theoretical models related to the transport of fueling neutral particles, and through interpretive transport modeling in present day fusion plasmas, in which the penetration of neutrals is altered to better simulate future reactor-like conditions. Achievements in theory and model validation under this project have advanced our understanding of how much of the edge density profile is set by transport versus direct ionization, enabling interesting projections to future burning plasma devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding the Drivers of Atlantic Multidecadal Variability using a Stochastic Model Hierarchy

The relative importance of ocean and atmospheric dynamics in generating Atlantic Multidecadal Variability (AMV) remains an open question. Comparisons between climate models with SLAB and fully-dynamic (FULL) ocean components are often used to explore this question, but cannot reveal how individual ocean processes generate these differences. We build a hierarchy of physically interpretable stochastic models to investigate the contribution of two upper-ocean processes to AMV: the role of seasonal variation and mixed-layer entrainment. This interpretability arises from the stochastic model’s simplified representation of sea surface temperature (SST), considering only the local upper ocean response to white-noise atmospheric forcing and its impact on surface heat exchange. We focus on understanding differences between SLAB and FULL non-eddy resolving pre-industrial control simulations of the Community Earth System Model 1 (CESM), and estimate the stochastic model parameters from each respective simulation. Despite its simplicity, the stochastic model reproduces temporal characteristics of SST variability in the SPG, including reemergence, seasonal-to-interannual persistence and power spectra. Furthermore, unrealistically persistent SST of the CESM-SLAB ocean simulation is reproduced in the equivalent stochastic model configuration where the mixed-layer depth (MLD) is constant. The stochastic model also reveals that vertical entrainment primarily damps SST variability, thus explaining why SLAB exhibits larger SST variance than FULL. Here, the stochastic model driven by temporally stochastic, spatially coherent forcing patterns reproduces the canonical AMV pattern. However, the amplitude of low-frequency variability remains underestimated, suggesting a role for ocean dynamics beyond entrainment.

54 ENVIRONMENTAL SCIENCES↗

Assessing Metal Ion Assignment Accuracy in Protein Data Bank Models via Elemental Spectroscopy

Accurate representation of metal ions in macromolecular structures is critical for chemical interpretation, computational modeling, and machine-learning methods that rely on Protein Data Bank (PDB) entries. However, the elemental identity of metals modeled in crystallographic structures is often inferred indirectly and rarely validated experimentally. Here, we combine Particle Induced X-ray Emission (PIXE) and X-ray Fluorescence Spectroscopy (XRFS) to determine the elemental composition of protein samples used to generate 70 deposited metalloprotein crystal structures. By analyzing the original protein material employed for crystallization, but before the addition of crystallization buffer solutions, we assess whether the modeled metal ions in deposited structures are consistent with experimentally detectable elemental content. We find that in a majority of cases, the metals modeled in the corresponding PDB entries are inconsistent with the metals present in the protein samples before crystallization, or that additional metals are present but not represented in the structural models. Spectroscopic results were integrated with automated crystallographic validation metrics, including real-space Z-difference (RSZD) analysis and systematic rerefinement, to evaluate atomic-number mismatch at metal sites. PIXE and XRFS show strong agreement for dominant elemental signals and provide complementary, scalable approaches for identifying suspect metal assignments. This work does not address physiological or functional metalation but instead highlights a widespread data integrity issue in deposited macromolecular structures, PDB-wide. These results establish an experimentally corroborated link between elemental identity and crystallographic validation metrics, enabling the large-scale detection of chemically inconsistent annotations in structural databases used for computational modeling and machine learning.

Crystallization↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Searches for new phenomena in events with two leptons, jets, and missing transverse momentum in 139 fb –1 of √s = 13 TeV $pp$ collisions with the ATLAS detector

Searches for new phenomena inspired by supersymmetry in final states containing an e + e – or μ + μ – pair, jets, and missing transverse momentum are presented. These searches make use of proton–proton collision data with an integrated luminosity of 139 fb –1 , collected during 2015–2018 at a centre-of-mass energy √s = 13 TeV by the ATLAS detector at the Large Hadron Collider. Two searches target the pair production of charginos and neutralinos. One uses the recursive-jigsaw reconstruction technique to follow up on excesses observed in 36.1 fb –1 of data, and the other uses conventional event variables. The third search targets pair production of coloured supersymmetric particles (squarks or gluinos) decaying through the next-to-lightest neutralino ($\tilde{χ}$$^{0}_{2}$) via a slepton ($\tilde{ℓ}$) or Z boson into ℓ + ℓ – $\tilde{χ}$$^{0}_{1}$ , resulting in a kinematic endpoint or peak in the dilepton invariant mass spectrum. The data are found to be consistent with the Standard Model expectations. Results are interpreted using simplified models and exclude masses up to 900 GeV for electroweakinos, 1550 GeV for squarks, and 2250 GeV for gluinos.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Programmatic Advantages of Linear Equivalent Seismic Models

Underground explosions nonlinearly deform the surrounding earth material and can interact with the free surface to produce spall. However, at typical seismological observation distances the seismic wavefield can be accurately modeled using linear approximations. Although nonlinear algorithms can accurately simulate very near field ground motions, they are computationally expensive and potentially unnecessary for far field wave simulations. Conversely, linearized seismic wave propagation codes are orders of magnitude faster computationally and can accurately simulate the wavefield out to typical observational distances. Thus, devising a means of approximating a nonlinear source in terms of a linear equivalent source would be advantageous both for scenario modeling and for interpretation of seismic source models that are based on linear, far-field approximations. This allows fast linear seismic modeling that still incorporates many features of the nonlinear source mechanics built into the simulation results so that one can have many of the advantages of both types of simulations without the computational cost of the nonlinear computation. In this report we first show the computational advantage of using linear equivalent models, and then discuss how the near-source (within the nonlinear wavefield regime) environment affects linear source equivalents and how well we can fit seismic wavefields derived from nonlinear sources.

58 GEOSCIENCES↗

Overview of IMPACT Data Acquisition System and Data Reduction Process

This report documents the development of the data acquisition system (DAS) and data reduction methodologies for the Irradiated Material Property Accelerated Characterization Test (IMPACT) experiment at the Advanced Test Reactor (ATR). The IMPACT experiment is designed to enable in-pile measurement of thermal conductivity in metallic nuclear fuels, specifically U-10Zr, using an instrumented thermal conductivity probe. The DAS supports both passive temperature monitoring and active thermal interrogation of the probe through controlled AC and DC excitation. Significant modifications to laboratory-scale systems were required to accommodate the higher resistance paths associated with the in-pile application. Custom electronics and relay-controlled measurement sequencing were developed to enable the measurement and sufficient power delivery to the sensing region. A reduced-order, axisymmetric thermal model based on the thermal quadrupoles method is presented to support data interpretation. This model enables efficient evaluation of transient heat transfer behavior and facilitates solution of the inverse problem required to extract thermal properties from measured signals. Multiple boundary condition formulations are discussed to address varying experimental time scales and geometries. Additionally, machine learning techniques are introduced to support data reduction and improve confidence in inverse solutions. Convolutional neural networks are applied to identify the presence of gas gaps and other evolving geometric features that significantly impact thermal response during irradiation. These efforts contribute to the broader integration of digital twin frameworks and real-time modeling capabilities within the Advanced Fuels Campaign.

46 - INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AN↗

Search for supersymmetry in final states with two or three soft leptons and missing transverse momentum in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A search for supersymmetry in events with two or three low-momentum leptons and missing transverse momentum is performed. The search uses proton-proton collisions at $\sqrt{s}$ = 13 TeV collected in the three-year period 2016–2018 by the CMS experiment at the LHC and corresponding to an integrated luminosity of up to 137 fb -1 . The data are found to be in agreement with expectations from standard model processes. The results are interpreted in terms of electroweakino and top squark pair production with a small mass difference between the produced supersymmetric particles and the lightest neutralino. For the electroweakino interpretation, two simplified models are used, a wino-bino model and a higgsino model. Exclusion limits at 95% confidence level are set on $^{\sim0}_{χ2}/^{\sim±}_{χ1}$ masses up to 275 GeV for a mass difference of 10 GeV in the wino-bino case, and up to 205(150) GeV for a mass difference of 7.5 (3) GeV in the higgsino case. The results for the higgsino are further interpreted using a phenomenological minimal supersymmetric standard model, excluding the higgsino mass parameter μ up to 180 GeV with the bino mass parameter M 1 at 800 GeV. In the top squark interpretation, exclusion limits are set at top squark masses up to 540 GeV for four-body top squark decays and up to 480 GeV for chargino-mediated decays with a mass difference of 30 GeV.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

pyTCR: A tropical cyclone rainfall model for python

pyTCR is a climatology software package developed in the Python programming language. It integrates the capabilities of several legacy physical models and increases computational efficiency to allow rapid estimation of tropical cyclone (TC) rainfall consistent with the large-scale environment. Specifically, pyTCR implements a horizontally distributed and vertically integrated model [Zhu et al., 2013] for simulating rainfall driven by TCs. Along storm tracks, rainfall is estimated by computing the cross-boundary-layer, upward water vapor transport caused by different mechanisms including frictional convergence, vortex stretching, large-scale baroclinic effect (i.e., wind shear), topographic forcing, and radiative cooling [Lu et al., 2018]. The package provides essential functionalities for modeling and interpreting spatio-temporal TC rainfall data. pyTCR requires a limited number of model input parameters, making it a convenient and useful tool for analyzing rainfall mechanisms driven by TCs. To sample rare (most intense) rainfall events that are often of great societal interest, pyTCR adapts and leverages outputs from a statistical-dynamical TC downscaling model [Lin et al., 2023] capable of rapidly generating a large number of synthetic TCs given a certain climate. As a result, pyTCR significantly reduces computational effort and improves the efficiency in capturing extreme TC rainfall events at the tail of the distributions from limited datasets. Furthermore, the TC downscaling model is forced entirely by large-scale environmental conditions from reanalysis data or coupled General Circulation Models (GCMs), simplifying the projection of TC-induced rainfall and wind speed under future climate using pyTCR. Finally, pyTCR can be coupled with hydrological and wind models to assess risks associated with independent and compound events (e.g., storm surges and freshwater flooding).

54 ENVIRONMENTAL SCIENCES↗