Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model interpretability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

A Variable Eddington Factor Model for Thermal Radiative Transfer with Closure Based on Data-Driven Shape Function

Here, a new variable Eddington factor (VEF) model is presented for nonlinear problems of thermal radiative transfer (TRT). The VEF model is data-driven and acts on known (a-priori) radiation-diffusion solutions for material temperatures in the TRT problem. A linear auxiliary problem is constructed for the radiative transfer equation (RTE) whose emission source and opacities are evaluated at these known material temperatures. The solution to this RTE approximates the specific intensity distribution in phase-space and time. It is applied as a shape function to define the Eddington tensor for the presented VEF model. The shape function computed via the auxiliary RTE problem will capture some degree of transport effects within the TRT problem. The VEF moment equations closed with this approximate Eddington tensor will thus carry with them these captured transport effects. In this study, the temperature data comes from multigroup P 1 , P 1/3 , and flux-limited diffusion radiative transfer models. The proposed VEF model can be interpreted as a transport-corrected diffusion reduced-order model. Numerical results are presented on the Fleck-Cummings test problem which models a supersonic wavefront of radiation. The VEF model is shown to improve accuracy by 1–2 orders of magnitude compared to the considered radiation-diffusion model solutions to the TRT problem.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Method to account for natural fracture induced elastic anisotropy in geomechanical characterization of shale gas reservoirs

Shale has been usually recognized as a transverse isotropic (TI) medium in conventional geomechanical log interpretation due to its laminated nature. However, when natural fractures exist in the shale rock, additional elastic anisotropy is introduced, converting laminated Shale to an orthorhombic (OB) medium. Previous studies illustrate that neglecting the natural fracture induced anisotropy in shale geomechanical log interpretation could lead to inaccurate evaluations of elastic moduli and in-situ stresses. In this paper, a new method is developed to account for the natural fracture induced anisotropy in geomechanical log interpretation based upon the TI acoustic model developed by the author and a characterization technique of elastic wave anisotropy (Sayers, 1991). The new OB model incorporates the four acoustic log data inputs and five modeling constraints in a nonlinear optimization algorithm to solve for the nine independent stiffness coefficients of an OB rock, and further to solve for the geomechanical properties and in-situ stress profiles in an OB formation. The new method was validated with a Marcellus Gas Shale field case. Both the new OB model and the conventional TI model were applied to interpret the minimum horizontal stress profile for the same formation. By comparing the results, the OB model is more robust than the TI from two aspects. First, the average stress magnitude predicted by the OB model is closer to the one measured by the Diagnostic Fracture Injection Test (DFIT). Second, the OB model predicts a more obvious stress barrier between the lower Marcellus and upper Onondaga Limestone than the TI model does. Finally, the predicted stress barrier is consistent with the observation of the microseismic events of a horizontal well drilled and completed nearby, which reveals that no hydraulic fracture propagates downward through the bottom boundary of Marcellus Shale into the underlying Onondaga Limestone.

03 NATURAL GAS↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

Using machine learning with optical profilometry for GaN wafer screening

Abstract To improve the manufacturing process of GaN wafers, inexpensive wafer screening techniques are required to both provide feedback to the manufacturing process and prevent fabrication on low quality or defective wafers, thus reducing costs resulting from wasted processing effort. Many of the wafer scale characterization techniques—including optical profilometry—produce difficult to interpret results, while models using classical programming techniques require laborious translation of the human-generated data interpretation methodology. Alternatively, machine learning techniques are effective at producing such models if sufficient data is available. For this research project, we fabricated over 6000 vertical PiN GaN diodes across 10 wafers. Using low resolution wafer scale optical profilometry data taken before fabrication, we successfully trained four different machine learning models. All models predict device pass and fail with 70–75% accuracy, and the wafer yield can be predicted within 15% error on the majority of wafers.

36 MATERIALS SCIENCE↗

Global biomass supply modeling for long-run management of the climate system

Bioenergy is projected to have a prominent, valuable, and maybe essential, role in climate management. However, there is significant variation in projected bioenergy deployment results, as well as concerns about the potential environmental and social implications of supplying biomass. Bioenergy deployment projections are market equilibrium solutions from integrated modeling, yet little is known about the underlying modeling of the supply of biomass as a feedstock for energy use in these modeling frameworks. We undertake a novel diagnostic analysis with ten global models to elucidate, compare, and assess how biomass is supplied within the models used to inform long-run climate management. With experiments that isolate and reveal biomass supply modeling behavior and characteristics (costs, emissions, land use, market effects), we learn about biomass supply tendencies and differences. The insights provide a new level of modeling transparency and understanding of estimated global biomass supplies that informs evaluation of the potential for bioenergy in managing the climate and interpretation of integrated modeling. For each model, we characterize the potential distributions of global biomass supply across regions and feedstock types for increasing levels of quantity supplied, as well as some of the potential societal externalities of supplying biomass. We also evaluate the biomass supply implications of managing these externalities. Finally, we interpret biomass market results from integrated modeling in terms of our new understanding of biomass supply. Overall, we find little consensus between models on where biomass could be cost-effectively produced and the implications. We also reveal model specific biomass supply narratives, with results providing new insights into integrated modeling bioenergy outcomes and differences. The analysis finds that many integrated models are considering and managing emissions and land use externalities of supplying biomass and estimating that environmental and societal trade-offs in the form of land emissions, land conversion, and higher agricultural prices are cost-effective, and to some degree a reality of using biomass, to address climate change.

09 BIOMASS FUELS↗

A detailed study of interpretability of deep neural network based top taggers

Abstract Recent developments in the methods of explainable artificial intelligence (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input–output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton–proton collisions at the Large Hadron Collider. We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as neural activation pattern diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the particle flow interaction network model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

97 MATHEMATICS AND COMPUTING↗

Evidential Deep Learning: Enhancing Predictive Uncertainty Estimation for Earth System Science Applications

Abstract Robust quantification of predictive uncertainty is a critical addition needed for machine learning applied to weather and climate problems to improve the understanding of what is driving prediction sensitivity. Ensembles of machine learning models provide predictive uncertainty estimates in a conceptually simple way but require multiple models for training and prediction, increasing computational cost and latency. Parametric deep learning can estimate uncertainty with one model by predicting the parameters of a probability distribution but does not account for epistemic uncertainty. Evidential deep learning, a technique that extends parametric deep learning to higher-order distributions, can account for both aleatoric and epistemic uncertainties with one model. This study compares the uncertainty derived from evidential neural networks to that obtained from ensembles. Through applications of the classification of winter precipitation type and regression of surface-layer fluxes, we show evidential deep learning models attaining predictive accuracy rivaling standard methods while robustly quantifying both sources of uncertainty. We evaluate the uncertainty in terms of how well the predictions are calibrated and how well the uncertainty correlates with prediction error. Analyses of uncertainty in the context of the inputs reveal sensitivities to underlying meteorological processes, facilitating interpretation of the models. The conceptual simplicity, interpretability, and computational efficiency of evidential neural networks make them highly extensible, offering a promising approach for reliable and practical uncertainty quantification in Earth system science modeling. To encourage broader adoption of evidential deep learning, we have developed a new Python package, Machine Integration and Learning for Earth Systems (MILES) group Generalized Uncertainty for Earth System Science (GUESS) (MILES-GUESS) ( https://github.com/ai2es/miles-guess ), that enables users to train and evaluate both evidential and ensemble deep learning. Significance Statement This study demonstrates a new technique, evidential deep learning, for robust and computationally efficient uncertainty quantification in modeling the Earth system. The method integrates probabilistic principles into deep neural networks, enabling the estimation of both aleatoric uncertainty from noisy data and epistemic uncertainty from model limitations using a single model. Our analyses reveal how decomposing these uncertainties provides valuable insights into reliability, accuracy, and model shortcomings. We show that the approach can rival standard methods in classification and regression tasks within atmospheric science while offering practical advantages such as computational efficiency. With further advances, evidential networks have the potential to enhance risk assessment and decision-making across meteorology by improving uncertainty quantification, a longstanding challenge. This work establishes a strong foundation and motivation for the broader adoption of evidential learning, where properly quantifying uncertainties is critical yet lacking.

Schreck, John S.↗

Search for higgsinos in compressed mass spectra using low-momentum tracks in pp collisions at s=13 TeV with the ATLAS detector

This paper presents two searches for the electroweak production of higgsinos with compressed mass spectra using 140 fb−1 of s=13$$ \sqrt{s}=13 $$ TeV proton-proton collision data collected by the ATLAS experiment at the Large Hadron Collider. Events are required to feature an energetic jet, large missing transverse momentum, and at least one low-momentum charged particle that serves as a candidate higgsino decay product. In the first search, targeting higgsino mass splittings in the range of 0.3–1 GeV, the higgsinos are expected to predominantly decay into pions that are identified as low-momentum charged particles with large transverse impact parameters due to the long higgsino lifetime (cτ ≈ ?(0.1–10 mm)), and neural networks are used to discriminate between signal and background processes. The second search targets larger mass splittings in the range of 1–3 GeV, where the higgsinos are expected to decay promptly into low-momentum leptons, one of which is identified by dedicated low-momentum electron or muon taggers based on neural networks utilising tracking and calorimeter information. No significant excess above the Standard Model prediction is observed in either search and the results are interpreted within simplified models, to set lower limits on the masses of the higgsino-like charginos and neutralinos. Together, these searches exclude chargino masses below 126 GeV at 95% confidence level for mass splittings between the chargino and lightest neutralino in the range of 0.3–2 GeV. This represents the first ATLAS constraints in a portion of this parameter space and surpasses the limits previously set by other experiments.

Aad, G↗

Interpretable machine learning for knowledge generation in heterogeneous catalysis

Most applications of machine learning in heterogeneous catalysis thus far have used black-box models to predict computable physical properties (descriptors), such as adsorption or formation energies, that can be related to catalytic performance (that is, activity or stability). Here, extracting meaningful physical insights from these black-box models has proved challenging, as the internal logic of these black-box models is not readily interpretable due to their high degree of complexity. Interpretable machine learning methods that merge the predictive capacity of black-box models with the physical interpretability of physics-based models offer an alternative to black-box models. In this Perspective, we discuss the various interpretable machine learning methods available to catalysis researchers, highlight the potential of interpretable machine learning to accelerate hypothesis formation and knowledge generation, and outline critical challenges and opportunities for interpretable machine learning in heterogeneous catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Higgs Portal Vector Dark Matter Interpretation: Review of Effective Field Theory Approach and Ultraviolet Complete Models

In this report a review of the Higgs portal-vector dark matter interpretation of the spin-independent dark matter nu-cleon elastic scattering cross section is presented, where the invisible Higgs decay width measured at the LHC is used. Effective Field Theory and ultraviolet complete models are discussed. LHC interpretations show only the scalar and Majorana dark matter scenarios; we propose including interpretation for vec-tor dark matter in the EFT and UV completion theoretical framework. In addition, our studies suggest an extension of the LHC dark matter interpretations to the sub-GeV regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Multiscale graph neural network autoencoders for interpretable scientific machine learning

The goal of this work is to address two limitations in autoencoder-based models: latent space interpretability and compatibility with unstructured meshes. This is accomplished here with the development of a novel graph neural network (GNN) autoencoding architecture with demonstrations on complex fluid flow applications. To address the first goal of interpretability, the GNN autoencoder achieves reduction in the number nodes in the encoding stage through an adaptive graph reduction procedure. Further, this reduction procedure essentially amounts to flowfieldconditioned node sampling and sensor identification, and produces interpretable latent graph representations tailored to the flowfield reconstruction task in the form of so-called masked fields. These masked fields allow the user to (a) visualize where in physical space a given latent graph is active, and (b) interpret the time-evolution of the latent graph connectivity in accordance with the time-evolution of unsteady flow features (e.g. recirculation zones, shear layers) in the domain. To address the goal of unstructured mesh compatibility, the autoencoding architecture utilizes a series of multi-scale message passing (MMP) layers, each of which models information exchange among node neighborhoods at various lengthscales. The MMP layer, which augments standard single-scale message passing with learnable coarsening operations, allows the decoder to more efficiently reconstruct the flowfield from the identified regions in the masked fields. Analysis of latent graphs produced by the autoencoder for various model settings are conducted using unstructured snapshot data sourced from large-eddy simulations in a backward-facing step (BFS) flow configuration with an OpenFOAM-based flow solver at high Reynolds numbers.

97 MATHEMATICS AND COMPUTING↗

Integrating multi-modal remote sensing, deep learning, and attention mechanisms for yield prediction in plant breeding experiments

In both plant breeding and crop management, interpretability plays a crucial role in instilling trust in AI-driven approaches and enabling the provision of actionable insights. The primary objective of this research is to explore and evaluate the potential contributions of deep learning network architectures that employ stacked LSTM for end-of-season maize grain yield prediction. A secondary aim is to expand the capabilities of these networks by adapting them to better accommodate and leverage the multi-modality properties of remote sensing data. In this study, a multi-modal deep learning architecture that assimilates inputs from heterogeneous data streams, including high-resolution hyperspectral imagery, LiDAR point clouds, and environmental data, is proposed to forecast maize crop yields. The architecture includes attention mechanisms that assign varying levels of importance to different modalities and temporal features that, reflect the dynamics of plant growth and environmental interactions. The interpretability of the attention weights is investigated in multi-modal networks that seek to both improve predictions and attribute crop yield outcomes to genetic and environmental variables. This approach also contributes to increased interpretability of the model's predictions. The temporal attention weight distributions highlighted relevant factors and critical growth stages that contribute to the predictions. The results of this study affirm that the attention weights are consistent with recognized biological growth stages, thereby substantiating the network's capability to learn biologically interpretable features. Accuracies of the model's predictions of yield ranged from 0.82-0.93 R 2 ref in this genetics-focused study, further highlighting the potential of attention-based models. Further, this research facilitates understanding of how multi-modality remote sensing aligns with the physiological stages of maize. The proposed architecture shows promise in improving predictions and offering interpretable insights into the factors affecting maize crop yields, while demonstrating the impact of data collection by different modalities through the growing season. By identifying relevant factors and critical growth stages, the model's attention weights provide valuable information that can be used in both plant breeding and crop management. The consistency of attention weights with biological growth stages reinforces the potential of deep learning networks in agricultural applications, particularly in leveraging remote sensing data for yield prediction. To the best of our knowledge, this is the first study that investigates the use of hyperspectral and LiDAR UAV time series data for explaining/interpreting plant growth stages within deep learning networks and forecasting plot-level maize grain yield using late fusion modalities with attention mechanisms.

59 BASIC BIOLOGICAL SCIENCES↗

VERITAS : A density-functional theory-based multiband kinetic model for understanding x-ray spectroscopy of dense plasmas

X-ray spectroscopy has long been a powerful diagnostic tool for hot, dilute plasmas, providing insights into plasma conditions by measuring line shifts and broadenings of atomic transitions. The technique critically depends on the accuracy of atomic physics models used to interpret spectroscopic measurements for inferring plasma properties such as free-electron density and temperature. Over the past decades, the atomic and plasma physics communities have developed robust atomic physics models to account for various processes in hot, dilute classical plasmas. While these models have been successful in that regime, their applicability becomes uncertain when interpreting x-ray spectroscopy experiments of above-solid-density plasmas. Given that finite-temperature density-functional theory (DFT) offers a more accurate description of dense plasma environments, we present the development of a DFT-based multi-band kinetic model, VERITAS, designed to improve the interpretation of x-ray spectroscopic measurements in high-density plasmas produced by laser-driven spherical implosions. This work details the VERITAS model and its application to both time-integrated and time-resolved x-ray spectra from implosion experiments on OMEGA. The advantages and limitations of the VERITAS model will also be discussed, along with potential directions for advancing x-ray spectroscopy of dense and superdense plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Data-driven modeling of power generation for a coal power plant under cycling

Increased penetration of renewables for power generation has negatively impacted the dynamics of conventional fossil fuel-based power plants. The power plants operating on the base load are forced to cycle, to adjust to the fluctuating power demands. This results in an inefficient operation of the coal power plants, which leads up to higher operating losses. To overcome such operational challenge associated with cycling and to develop an optimal process control, this work analyzes a set of models for predicting power generation. Moreover, the power generation is intrinsically affected by the state of the power plant components, and therefore our model development also incorporates additional power plant process variables while forecasting the power generation. We present and compare multiple state-of-the-art forecasting data-driven methods for power generation to determine the most adequate and accurate model. We also develop an interpretable attention-based transformer model to explain the importance of process variables during training and forecasting. The trained deep neural network (DNN) LSTM model has good accuracy in predicting gross power generation under various prediction horizons with/without cycling events and outperforms the other models for long-term forecasting. The DNN memory-based models show significant superiority over other state-of-the-art machine learning models for short, medium and long range predictions. The transformer-based model with attention enhances the selection of historical data for multi-horizon forecasting, and also allows to interpret the significance of internal power plant components on the power generation. This newly gained insights can be used by operation engineers to anticipate and monitor the health of power plant equipment during high cycling periods.

01 COAL, LIGNITE, AND PEAT↗

End-to-end AI framework for interpretable prediction of molecular and crystal properties

We introduce an end-to-end computational framework that allows for hyperparameter optimization using the DeepHyper library, accelerated model training, and interpretable AI inference. The framework is based on state-of-the-art AI models including CGCNN, PhysNet, SchNet, MPNN, MPNN-transformer, and TorchMD-NET. We employ these AI models along with the benchmark QM9, hMOF, and MD17 datasets to showcase how the models can predict user-specified material properties within modern computing environments. We demonstrate transferable applications in the modeling of small molecules, inorganic crystals and nanoporous metal organic frameworks with a unified, standalone framework. We have deployed and tested this framework in the ThetaGPU supercomputer at the Argonne Leadership Computing Facility, and in the Delta supercomputer at the National Center for Supercomputing Applications to provide researchers with modern tools to conduct accelerated AI-driven discovery in leadership-class computing environments. We release these digital assets as open source scientific software in GitLab, and ready-to-use Jupyter notebooks in Google Colab.

36 MATERIALS SCIENCE↗

A structured framework for predicting sustainable aviation fuel properties using liquid-phase FTIR and machine learning

Sustainable aviation fuels have the potential to improve efficiency, reduce emissions, and enhance energy security. To help identify viable sustainable aviation fuels and accelerate research, machine learning models have been developed to predict relevant physicochemical properties. However, many models have limited applicability, leverage data from complex analytical techniques with confined spectral ranges, or use feature decomposition methods that offer limited interpretability. Using liquid-phase Fourier Transform Infrared (FTIR) spectra, this study presents a structured method for creating accurate and interpretable property prediction models for neat molecules, aviation fuels, and blends. Liquid FTIR spectra can be collected quickly and consistently, offering high reliability, sensitivity, and component specificity using less than 2 ml of sample. The method first decomposes FTIR spectra into fundamental building blocks using non-negative matrix factorization (NMF) to enable scientific analysis of FTIR spectra attributes and fuel properties. The NMF features are then used to create five ensemble models for predicting final boiling point, flash point, freezing point, density at 15°C, and kinematic viscosity at -20°C. All models were trained using experimental property data from neat molecules, aviation fuels, and blends. The models accurately predict key properties across a broad range of neat molecules and representative fuels and blends, while enabling interpretation of relationships between compositional elements, such as functional groups or chemical classes, and their resulting properties. This demonstrates strong potential to support sustainable aviation fuel research and development. The models and data are available on an interactive web tool.

Fourier transform infrared spectroscopy↗