Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “explainable neural networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Using Explainable Artificial Intelligence to Predict Perovskite Solar Cell Electrical Metastability from Operando Photoluminescence Images in Accelerated Stress Testing

Metal halide perovskite (MHP) solar cells exhibit a metastable response to bias governed by coupled ionic–electronic processes, complicating the conventional reciprocity relation between luminescence intensity and device open-circuit voltage (V oc ). This limits the use of luminescence as a diagnostic for device screening or accelerated stress testing, motivating new approaches that can interpret photoluminescence (PL) signals under nonequilibrium conditions. From the artificial intelligence perspective, we develop an explainable deep learning framework that integrates convolutional neural networks (CNN), long short-term memory (LSTM) layers, and an attention mechanism to learn spatiotemporal features from operando photoluminescence PL image sequences. The model achieves a mean absolute error of ±0.027 V in predicting open-circuit voltage transients and reduces extreme-tail errors by up to 78% compared to physics-based reciprocity calculations. Gradient-weighted Class Activation Mapping (Grad-CAM) provides interpretability by highlighting physically meaningful regions such as electrode edges and emergent defect features. From the engineering application perspective, this framework enables accurate, contactless prediction of device V oc and identification of degradation-relevant features during accelerated aging of perovskite solar cells. This approach demonstrates how explainable AI can enhance operando diagnostics and reliability analysis in photovoltaic devices under nonequilibrium conditions.

14 SOLAR ENERGY↗

Methods for Explainable Artificial Intelligence

We explored ways of quantifying information in a neural network. This can be used to determine the right size of a network or to infer the way in which a network is processing information. The first year and a half was somewhat exploratory while the last half of the project focused on approaches that seemed to show the most promise. The introduce a new way of computing explainable artificial intelligence (XAI) saliency maps that is several orders of magnitude faster than methods with similar fidelity We call it FastCAM. The method works be combining a Class Activation Map (CAM) method such as GradCAM with a forward activation map computed with a statistic we call SMOE Scale. The addition of the forward activation maps to CAM methods seems to always improve their fidelity. At the same time, computational overhead is not increased by very much. While Gradients with SmoothGrad scores better on some fidelity measures, it is overall not as good and requires more than 1500 times to compute. We demonstrate two completed applications of FastCAM on tasks outside of the LDRD at LLNL. The source code for FastCAM is currently being implanted into Captum, the official XAI toolkit for the popular deep learning toolkit PyTorch. The LDRD currently has 13 publications released to the public. 10 of them are journal length.

97 MATHEMATICS AND COMPUTING↗

Investigating the crust of neutron stars with neural-network quantum states

An accurate description of low-density nuclear matter is crucial for explaining the physics of neutron star crusts. In the density range between approximately 0.01 fm −3 and 0.1 fm −3 , matter transitions from neutron-rich nuclei to various higher-density pasta shapes, before ultimately reaching a uniform liquid. In this work, we introduce a variational Monte Carlo method based on a neural Pfaffian-Jastrow quantum state, which allows us to model the transition from the liquid phase to neutron-rich nuclei microscopically. At low densities, nuclear clusters dynamically emerge from the microscopic interactions among protons and neutrons, which we model based on pionless effective field theory. Our variational Monte Carlo approach represents a significant improvement over the state-of-the-art auxiliary-field diffusion Monte Carlo method, which is severely hindered by the fermion-sign problem in this low-density regime and cannot capture the onset of clusters. In addition to computing the energy per particle of symmetric nuclear matter and pure neutron matter, we analyze an intermediate isospin-asymmetry configuration to elucidate the formation of nuclear clusters. We also provide evidence that the presence of such nuclear clusters influences the amount of protons in the crust compared to protons in beta-equilibrated, neutrino-transparent matter.

Nuclear astrophysics↗

Preprocessing for Unintended Conducted Emissions Classification with ResNet

Characterization of Unintended Conducted Emissions (UCE) from electronic devices is important when diagnosing electromagnetic interference, performing nonintrusive load monitoring (NILM) of power systems, and monitoring electronic device health, among other applications. Prior work has demonstrated that UCE analysis can serve as a diagnostic tool for energy efficiency investigations and detailed load analysis. While explaining the feature selection of deep networks with certainty is often not fully comprehensive, or in other applications, quite lacking, additional tools/methods for further corroboration and confirmation can help further the understanding of the researcher. This is true especially in the subject application of the study in this paper. Often the focus of such efforts is the selected features themselves, and there is not as much understanding gained about the noise in the collected data. If selected feature and noise characteristics are known, it can be used to further shape the design of the deep network or associated preprocessing. This is additionally difficult when the available data are limited, as in the case which the authors investigated in this study. Here, the authors present a novel work (which is a proposed complementary portion of the overall solution to the deep network classification explainability problem for this application) by applying a systematic progression of preprocessing and a deep neural network (ResNet architecture) to classify UCE data obtained via current transformers. By using a methodical application of preprocessing techniques prior to a deep classifier, hypotheses can be produced concerning what features the deep network deems important relative to what it perceives as noise. For instance, it is hypothesized in this particular study as a result of execution of the proposed method and periodic inspection of the classifier output that the UCE spectral features are relatively close to each other or to the interferers, as systematically reducing the beta parameter of the Kaiser window produced progressively better classification performance, but only to a point, as going below the Beta of eight produced decreased classifier performance, as well as the hypothesis that further spectral feature resolution was not as important to the classifier as rejection of the leakage from a spectrally distant interference. This can be very important in unpredictable low-FNR applications, where knowing the difference between features and noise is difficult. As a side-benefit, much was learned regarding the best preprocessing to use with the selected deep network for the UCE collected from these low power consumer devices obtained via current transformers. Baseline rectangular windowed FFT preprocessing provided a 62% classification increase versus using raw samples. After performing a more optimal preprocessing, more than 90% classification accuracy was achieved across 18 low-power consumer devices for scenarios in which the in-band features-to-noise ratio (FNR) was very poor.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Neural-network quantum states for the nuclear many-body problem

A long-standing goal of nuclear theory is to explain how the structure and dynamics of atomic nuclei and neutron-star matter emerge from the underlying interactions among protons and neutrons. Achieving this goal requires solving the nuclear quantum many-body problem with high accuracy across a wide range of length scales and density regimes. In this review, we discuss how artificial neural network representations of the nuclear many-body wave function have significantly extended the capabilities of continuum quantum Monte Carlo methods. In particular, neural network quantum states enable calculations of larger systems than were previously accessible and provide a flexible framework for capturing phenomena that challenge conventional approaches, including the emergence of nuclear clusters and superfluid phases in dense matter. We highlight recent applications to finite nuclei, infinite nuclear and neutron matter, and dynamical processes relevant to lepton-nucleus and nucleus-nucleus scattering. We also discuss conceptual and methodological connections with condensed matter physics, emphasizing developments in neural network quantum states that bridge strongly correlated systems across disciplines. Together, these developments demonstrate how neural-network methods open new avenues toward unified and accurate descriptions of nuclear structure, matter, and reactions.

Lovato, Alessandro [Argonne; TIFPA-INFN, Trento; V↗

Understanding oxidation of Fe-Cr-Al alloys through explainable artificial intelligence

Abstract The oxidation resistance of FeCrAl based on alloying composition and oxidizing conditions is predicted using a combinatorial experimental and artificial intelligence approach. A neural network (NN) classification model was trained on the experimental FeCrAl dataset produced at GE Research. Furthermore, using the SHapley Additive exPlanations (SHAP) explainable artificial intelligence (XAI) tool, we explore how the NN can showcase further material insights that are unavailable directly from a black-box model. We report that high Al and Cr content forms protective oxide layer, while Mo in FeCrAl creates thick unprotective oxide scale that is vulnerable to spallation due to thermal expansion. Graphical abstract

Materials Science↗

How Captain Amerika uses neural networks to fight crime

Artificial neural network models can make amazing computations. These models are explained along with their application in problems associated with fighting crime. Specific problems addressed are identification of people using face recognition, speaker identification, and fingerprint and handwriting analysis (biometric authentication).

Rogers, Steven K.↗

Damage Detection Using Holography and Interferometry

This paper reviews classical approaches to damage detection using laser holography and interferometry. The paper then details the modern uses of electronic holography and neural-net-processed characteristic patterns to detect structural damage. The design of the neural networks and the preparation of the training sets are discussed. The use of a technique to optimize the training sets, called folding, is explained. Then a training procedure is detailed that uses the holography-measured vibration modes of the undamaged structures to impart damage-detection sensitivity to the neural networks. The inspections of an optical strain gauge mounting plate and an International Space Station cold plate are presented as examples.

Decker, Arthur J.↗

Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions

Models that predict brain responses to stimuli provide one measure of understanding of a sensory system and have many potential applications in science and engineering. Deep artificial neural networks have emerged as the leading such predictive models of the visual system but are less explored in audition. Prior work provided examples of audio-trained neural networks that produced good predictions of auditory cortical fMRI responses and exhibited correspondence between model stages and brain regions, but left it unclear whether these results generalize to other neural network models and, thus, how to further improve models in this domain. We evaluated model-brain correspondence for publicly available audio neural network models along with in-house models trained on 4 different tasks. Most tested models outpredicted standard spectromporal filter-bank models of auditory cortex and exhibited systematic model-brain correspondence: Middle stages best predicted primary auditory cortex, while deep stages best predicted non-primary cortex. However, some state-of-the-art models produced substantially worse brain predictions. Models trained to recognize speech in background noise produced better brain predictions than models trained to recognize speech in quiet, potentially because hearing in noise imposes constraints on biological auditory representations. The training task influenced the prediction quality for specific cortical tuning properties, with best overall predictions resulting from models trained on multiple tasks. The results generally support the promise of deep neural networks as models of audition, though they also indicate that current models do not explain auditory cortical responses in their entirety.

59 BASIC BIOLOGICAL SCIENCES↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Characterization of Extremes and Compound Impacts: Applications of Machine Learning and Interpretable Neural Networks

Focal Area: This white paper responds to Focal area III by exploring data fusion, learning and explainable AI methods in characterizing hydrological extremes and interconnections. It also addresses Focal area II by using probabilistic AI and ensemble ML for predicting extremes and compound extremes. Science Challenge: A key question associated with the integrated water (or hydrological) cycle grand challenge in the Earth and Environmental Systems Sciences Division (EESSD) strategic plan, is how the frequency and intensity of hydrological events will change. Prediction of the tail behavior (extremes) of the hydrological cycle is especially challenging, because of their stochasticity and low probability. These extreme events and their compound impacts have significant societal and economic consequences. It is anticipated for the next-generation Earth System models (ESMs), that model predictability of the water cycle will improve with increased resolution (e.g., regionally refined E3SM), advanced software and computational architectures, and improved model physics based on the data from ARM measurements and high-fidelity models. However, the challenges for predictability of low-probability high-impact extreme events will unlikely be alleviated with conventional modeling and data-driven approaches, as ESMs are calibrated largely for capturing the high-frequency mean climate states. Recent AI and ML applications have shown great potential in quantifying well-defined climate extremes (e.g., supervised learning of tropical cyclones/atmospheric rivers by ClimateNet1) but few efforts are dedicated to compound events, extreme drivers and uncertainty estimation. We envision the opportunity to develop and apply ML and interpretable AI methods extended on the existing efforts, specifically, for: (1) identification of compound extremes, (2) diagnosing drivers of extremes, (3) bias correction in extreme predictions and (4) probabilistic modeling of extremes.

54 ENVIRONMENTAL SCIENCES↗

Convolutional Neural Networks for the CHIPS Neutrino Detector R&D Project.

The CHerenkov detectors In mine PitS (Chips) neutrino detector R&D project aims to develop novel strategies and technologies for very large yet ‘cheap as chips’ water Cherenkov neutrino detectors. Via deployment in a body of water, use of commercially available components, and instrumentation coverage optimisation for the study of exclusively accelerator beam neutrinos, Chips will enable megaton scale detectors to become a reality at the cost of $200k-$300k per kt of sensitive mass. During the summer of 2019 a prototype Chips detector, Chips-5, was deployed into the Wentworth 2W disused mine pit in northern Minnesota, 7 mrad off the NuMI beam axis. A novel data acquisition system was introduced using cheap single-board computers and open-source software. This work presents a novel approach to water Cherenkov neutrino detector event reconstruction and classification. Three forms of a Convolutional Neural Network, a type of deep learning algorithm, have been trained to reject cosmic muon events, classify beam events, and estimate neutrino energies, all using only the raw detector event as input. When evaluated on the expected distribution of Chips-5 events, this new approach is shown to be robust and explainable as well as providing a significant performance increase over the standard likelihood-based reconstruction and simple neural network classification. Promisingly, the performance presented here is comparable to the more complex (and expensive) neutrino oscillation experiments within the field.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Semi-supervised Bayesian Low-shot Learning

Deep neural networks (NNs) typically outperform traditional machine learning (ML) approaches for complicated, non-linear tasks. It is expected that deep learning (DL) should offer superior performance for the important non-proliferation task of predicting explosive device configuration based upon observed optical signature, a task which human experts struggle with. However, supervised machine learning is difficult to apply in this mission space because most recorded signatures are not associated with the corresponding device description, or “truth labels.” This is challenging for NNs, which traditionally require many samples for strong performance. Semi-supervised learning (SSL), low-shot learning (LSL), and uncertainty quantification (UQ) for NNs are emerging approaches that could bridge the mission gaps of few labels and rare samples of importance. NN explainability techniques are important in gaining insight into the inferential feature importance of such a complex model. In this work, SSL, LSL, and UQ are merged into a single framework, a significant technical hurdle not previously demonstrated. Exponential Average Adversarial Training (EAAT) and Pairwise Neural Networks (PNNs) are chosen as the SSL and LSL methods of choice. Permutation feature importance (PFI) for functional data is used to provide explainability via the Variable importance Explainable Elastic Shape Analysis (VEESA) pipeline. A variety of uncertainty quantification approaches are explored: Bayesian Neural Networks (BNNs), ensemble methods, concrete dropout, and evidential deep learning. Two final approaches, one utilizing ensemble methods and one utilizing evidential learning, are constructed and compared using a well-quantified synthetic 2D dataset along with the DIRSIG Megascene.

97 MATHEMATICS AND COMPUTING↗

Artificial neural network implementation of a near-ideal error prediction controller

A theory has been developed at the University of Virginia which explains the effects of including an ideal predictor in the forward loop of a linear error-sampled system. It has been shown that the presence of this ideal predictor tends to stabilize the class of systems considered. A prediction controller is merely a system which anticipates a signal or part of a signal before it actually occurs. It is understood that an exact prediction controller is physically unrealizable. However, in systems where the input tends to be repetitive or limited, (i.e., not random) near ideal prediction is possible. In order for the controller to act as a stability compensator, the predictor must be designed in a way that allows it to learn the expected error response of the system. In this way, an unstable system will become stable by including the predicted error in the system transfer function. Previous and current prediction controller include pattern recognition developments and fast-time simulation which are applicable to the analysis of linear sampled data type systems. The use of pattern recognition techniques, along with a template matching scheme, has been proposed as one realizable type of near-ideal prediction. Since many, if not most, systems are repeatedly subjected to similar inputs, it was proposed that an adaptive mechanism be used to 'learn' the correct predicted error response. Once the system has learned the response of all the expected inputs, it is necessary only to recognize the type of input with a template matching mechanism and then to use the correct predicted error to drive the system. Suggested here is an alternate approach to the realization of a near-ideal error prediction controller, one designed using Neural Networks. Neural Networks are good at recognizing patterns such as system responses, and the back-propagation architecture makes use of a template matching scheme. In using this type of error prediction, it is assumed that the system error responses be known for a particular input and modeled plant. These responses are used in the error prediction controller. An analysis was done on the general dynamic behavior that results from including a digital error predictor in a control loop and these were compared to those including the near-ideal Neural Network error predictor. This analysis was done for a second and third order system.

Mcvey, Eugene S.↗

Evaluating the Trustworthiness of Explainable Artificial Intelligence (XAI) Methods Applied to Regression Predictions of Arctic Sea Ice Motion

Abstract Recent advances in explainable artificial intelligence (XAI) methods show promise for understanding predictions made by machine learning (ML) models. XAI explains how the input features are relevant or important for the model predictions. We train linear regression (LR) and convolutional neural network (CNN) models to make 1-day predictions of sea ice velocity in the Arctic from inputs of present-day wind velocity and previous-day ice velocity and concentration. We apply XAI methods to the CNN and compare explanations to variance explained by LR. We confirm the feasibility of using a novel XAI method [i.e., global layerwise relevance propagation (LRP)] to understand ML model predictions of sea ice motion by comparing it to established techniques. We investigate a suite of linear, perturbation-based, and propagation-based XAI methods in both local and global forms. Outputs from different explainability methods are generally consistent in showing that wind speed is the input feature with the highest contribution to ML predictions of ice motion, and we discuss inconsistencies in the spatial variability of the explanations. Additionally, we show that the CNN relies on both linear and nonlinear relationships between the inputs and uses nonlocal information to make predictions. LRP shows that wind speed over land is highly relevant for predicting ice motion offshore. This provides a framework to show how knowledge of environmental variables (i.e., wind) on land could be useful for predicting other properties (i.e., sea ice velocity) elsewhere. Significance Statement Explainable artificial intelligence (XAI) is useful for understanding predictions made by machine learning models. Our research establishes trustability in a novel implementation of an explainable AI method known as layerwise relevance propagation for Earth science applications. To do this, we provide a comparative evaluation of a suite of explainable AI methods applied to machine learning models that make 1-day predictions of Arctic sea ice velocity. We use explainable AI outputs to understand how the input features are used by the machine learning to predict ice motion. Additionally, we show that a convolutional neural network uses nonlinear and nonlocal information in making its predictions. We take advantage of the nonlocality to investigate the extent to which knowledge of wind on land is useful for predicting sea ice velocity elsewhere.

Hoffman, Lauren [Scripps Institution of Oceanograp↗

Measure Utility, Gain Trust: Practical Advice for XAI Researchers

Research into explanation of machine learning models, i.e. explainable AI (XAI), has seen a sympathetic exponential growth alongside deep artificial neural networks throughout the past decade. For historical reasons explanation and trust have been intertwined. However this focus on trust is too narrow, and has led the research community astray from tried and true empirical methods that lead to more defensible scientific knowledge about people and explanations. To address this, we contribute a practical path forward for researchers in the XAI field. We recommend researchers focus on the utility and impact of their explanations instead of trust. We outline five broad use cases where explanations are useful and, for each, we describe pseudo-experiments that rely on objective empirical measurements and falsifiable hypotheses. We believe that this experimental rigor is necessary to contribute to scientific knowledge in the field of XAI.

Davis, Brittany F.↗

Data Mining Methods Applied to Flight Operations Quality Assurance Data: A Comparison to Standard Statistical Methods

In a previous study, multiple regression techniques were applied to Flight Operations Quality Assurance-derived data to develop parsimonious model(s) for fuel consumption on the Boeing 757 airplane. The present study examined several data mining algorithms, including neural networks, on the fuel consumption problem and compared them to the multiple regression results obtained earlier. Using regression methods, parsimonious models were obtained that explained approximately 85% of the variation in fuel flow. In general data mining methods were more effective in predicting fuel consumption. Classification and Regression Tree methods reported correlation coefficients of .91 to .92, and General Linear Models and Multilayer Perceptron neural networks reported correlation coefficients of about .99. These data mining models show great promise for use in further examining large FOQA databases for operational and safety improvements.

Stolzer, Alan J.↗

How the Galaxy–Halo Connection Depends on Large-scale Environment

We investigate the connection between galaxies, dark matter halos, and their large-scale environments at z = 0 with Illustris TNG300 hydrodynamic simulation data. We predict stellar masses from subhalo properties to test two types of machine learning (ML) models: explainable boosting machines (EBMs) with simple galaxy environment features and E(3)-invariant graph neural networks (GNNs). The best-performing EBM models leverage spherically averaged overdensity features on 3 Mpc scales. Interpretations via SHapley Additive exPlanations also suggest that in the context of the TNG300 galaxy–halo connection, simple spherical overdensity on ∼3 Mpc scales is more important than cosmic web distance features measured using the DisPerSE algorithm. Meanwhile, a GNN with connectivity defined by a fixed linking length, L, outperforms the EBM models by a significant margin. As we increase the linking length scale, GNNs learn important environmental contributions up to the largest scales we probe (L = 10 Mpc). We conclude that 3 Mpc distance scales are most critical for describing the TNG galaxy–halo connection using the spherical overdensity parameterization, but that information on larger scales, which is not captured by simple environmental parameters or cosmic web features, can further augment these models. Our study highlights the benefits of using interpretable ML algorithms to explain models of astrophysical phenomena, and the power of using GNNs to flexibly learn complex relationships directly from data while imposing constraints from physical symmetries.

79 ASTRONOMY AND ASTROPHYSICS↗