Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Explainable deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Science Gateway for the Repeatable Analysis of Machine Learning Predicted Gravity Anomalies

In recent years, deep learning has become an increasingly popular alternative for modeling in geoscience applications due to its scalability and efficiency. However, the interpretability, compute, data volume, and hyperparameter tuning requirements of deep learning models make development and monitoring difficult. Furthermore, model explainability and communicating results obtained by these models to users or domain experts is a challenge, as domain experts in geoscience also need to have a deep understanding of how those models function in order to support their scientific works. Here, we describe a science gateway and machine learning pipeline for predicting gravity anomalies from geophysical data. The gateway, built on open-source technologies, provides a holistic view of the pipeline through interactive visualizations aimed at enabling efficient exploratory data analysis. The repeatability, reproducibility, and monitoring capabilities of this overall system allow us to iterate and analyze at scale. Using this pipeline and gateway, we can repeatedly produce accurate high-resolution gravity anomaly datasets. By describing the underlying technologies, implementation, and results, here we provide a foundation for the broader adoption of science gateways into cross-cutting geoscience and machine learning research projects as a means to improve the scientific discovery and collaboration in the geophysics and computational sciences community.

58 GEOSCIENCES↗

Applications of explainable artificial intelligence in renewable energy research

Researchers in renewable energy are applying deep learning (DL) to a variety of problems from diverse renewable energy domains, such as biofuels, wind, solar, power systems, buildings, vehicles, and transportation systems. Improvements in accuracy may be demonstrated using DL in laboratory settings. However, the lack of interpretability of DL models poses a practical limitation to their utility in advancing scientific knowledge and in the deployment of DL models in safety-critical energy systems. In this article, we discuss explainable artificial intelligence (XAI) as one pathway toward more interpretable DL models. We explore a brief timeline of U.S. national laboratory interest in XAI, an overview and taxonomy of methods in the field of XAI, and a selection of applications across renewable energy research domains. We conclude by highlighting pivotal areas where XAI can accelerate innovation in artificial intelligence for renewable energy research and other essential future directions.

97 MATHEMATICS AND COMPUTING↗

Interpretable Deep Learning for the Earth System with Fractal Nets

Focal Area 3: Explainable AI Our confidence in the projections made by Earth System Models (ESMs) depends on understanding them to be, in some important respects, faithful representations of the Earth system. Here we present an “explainable Artificial Intelligence (AI)” method that allows us to uncover the dynamical structure of the observed and modeled Earth system, discover hidden links across wide spatiotemporal scales, target model development efforts at poorly-represented dynamics, and optimize observed or modeled data collection to maximize predictive information. Science Challenge: Dynamical system science for the Earth system poses unique challenges given the large degree of internal climate variability. Thus, tools that help us understand how ESMs succeed and fail at representing these dynamics are crucial, particularly in relation to the observed system. Furthermore, the computational and memory constraints on ESM data output motivate in situ analysis of ESM dynamics, including automatic detection of dynamical shifts. Also, of key importance are procedures that leverage ESMs to optimize observational campaigns for improving process representation, reducing structural uncertainty and improving model skill.

54 ENVIRONMENTAL SCIENCES↗

Learning dynamical systems from data: An introduction to physics-guided deep learning

Modeling complex physical dynamics is a fundamental task in science and engineering. Traditional physics-based models are first-principled, explainable, and sample-efficient. However, they often rely on strong modeling assumptions and expensive numerical integration, requiring significant computational resources and domain expertise. While deep learning (DL) provides efficient alternatives for modeling complex dynamics, they require a large amount of labeled training data. Furthermore, its predictions may disobey the governing physical laws and are difficult to interpret. Physics-guided DL aims to integrate first-principled physical knowledge into data-driven methods. It has the best of both worlds and is well equipped to better solve scientific problems. Recently, this field has gained great progress and has drawn considerable interest across discipline Here, we introduce the framework of physics-guided DL with a special emphasis on learning dynamical systems. We describe the learning pipeline and categorize state-of-the-art methods under this framework. We also offer our perspectives on the open challenges and emerging opportunities.

97 MATHEMATICS AND COMPUTING↗

An Explainable Machine-Learning Model for Compensatory Reserve Measurement: Methods for Feature Selection and the Effects of Subject Variability

Tracking vital signs accurately is critical for triaging a patient and ensuring timely therapeutic intervention. The patient’s status is often clouded by compensatory mechanisms that can mask injury severity. The compensatory reserve measurement (CRM) is a triaging tool derived from an arterial waveform that has been shown to allow for earlier detection of hemorrhagic shock. However, the deep-learning artificial neural networks developed for its estimation do not explain how specific arterial waveform elements lead to predicting CRM due to the large number of parameters needed to tune these models. Alternatively, we investigate how classical machine-learning models driven by specific features extracted from the arterial waveform can be used to estimate CRM. More than 50 features were extracted from human arterial blood pressure data sets collected during simulated hypovolemic shock resulting from exposure to progressive levels of lower body negative pressure. A bagged decision tree design using the ten most significant features was selected as optimal for CRM estimation. This resulted in an average root mean squared error in all test data of 0.171, similar to the error for a deep-learning CRM algorithm at 0.159. By separating the dataset into sub-groups based on the severity of simulated hypovolemic shock withstood, large subject variability was observed, and the key features identified for these sub-groups differed. This methodology could allow for the identification of unique features and machine-learning models to differentiate individuals with good compensatory mechanisms against hypovolemia from those that might be poor compensators, leading to improved triage of trauma patients and ultimately enhancing military and emergency medicine.

60 APPLIED LIFE SCIENCES↗

A detailed study of interpretability of deep neural network based top taggers

Abstract Recent developments in the methods of explainable artificial intelligence (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input–output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton–proton collisions at the Large Hadron Collider. We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as neural activation pattern diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the particle flow interaction network model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

97 MATHEMATICS AND COMPUTING↗

A Search for Sterile-Neutrino-Based Muon Neutrino Disappearance Using the MicroBooNe Deep Learning Analysis

We describe a search for νµ disappearance using the MicroBooNE Deep Learning analysis 1µ1p selection. Presently, the unexplained MiniBooNE and LSND anomalies could be explained by a sterile neutrino impacting neutrino oscillations. Our analysis searches for the allowed parameter space that could describe such a sterile neutrino. We determine the allowed and excluded region of a 3+1 sterile-based muon neutrino disappearance model in MicroBooNE at 90% confidence. Our allowed region includes both the null model, and current global best fit model. Context for the underlying Deep Learning analysis is provided and several validation studies surrounding both the disappearance search, and originating 1µ1p selection are performed to strengthen confidence in the result. In addition, a next-generation deep learning tool for cosmic-ray-muon event discrimination is proposed and evaluated, demonstrating a removal of 70% of the remaining event background when added to current methods, under the cut criteria used.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Search for Sterile-Neutrino-Based Muon Neutrino Disappearance Using the MicroBooNe Deep Learning Analysis

We describe a search for νµ disappearance using the MicroBooNE Deep Learning analysis 1µ1p selection. Presently, the unexplained MiniBooNE and LSND anomalies could be explained by a sterile neutrino impacting neutrino oscillations. Our analysis searches for the allowed parameter space that could describe such a sterile neutrino. We determine the allowed and excluded region of a 3+1 sterile-based muon neutrino disappearance model in MicroBooNE at 90% confidence. Our allowed region includes both the null model, and current global best fit model. Context for the underlying Deep Learning analysis is provided and several validation studies surrounding both the disappearance search, and originating 1µ1p selection are performed to strengthen confidence in the result. In addition, a next-generation deep learning tool for cosmic-ray-muon event discrimination is proposed and evaluated, demonstrating a removal of 70% of the remaining event background when added to current methods, under the cut criteria used.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Making Corgis Important for Honeycomb Classification: Adversarial Attacks on Concept-based Explainability Tools

Methods for model explainability have become increasingly critical for testing the fairness and soundness of deep learning. Concept-based interpretability techniques, which use a small set of human-interpretable concept exemplars in order to measure the influence of a concept on a model's internal representation of input, are an important thread in this line of research. In this work we show that these explainability methods can suffer the same vulnerability to adversarial attacks as the models they are meant to analyze. We demonstrate this phenomenon on two well-known concept-based interpretability methods: TCAV and faceted feature visualization. We show that by leveraging the geometry of the problem and carefully perturbing the examples of the concept that is being investigated, we can radically change the output of the interpretability method. The attacks that we propose can either induce positive interpretations (polka dots are an important concept for a model when classifying zebras) or negative interpretations (stripes are not an important factor in identifying images of a zebra). Our work highlights the fact that in safety-critical applications, there is need for security around not only the machine learning pipeline but also the model interpretation process.

Brown, Davis R.↗

Enhanced Oblique Decision Tree Enabled Policy Extraction for Deep Reinforcement Learning in Power System Emergency Control

Deep reinforcement learning (DRL) algorithms have successfully solved many challenging problems in various power system control scenarios. However, their decision-making process is usually regarded as black-boxes. Furthermore, how DRL models interact with human intelligence remains an open problem. Thus, this paper proposes a policy extraction framework to extract a complex DRL model into an explainable policy. This framework includes three parts: 1) DRL training and data generation. We train an agent for a specific control task and generate data, which contains the control policy of the agent. 2) Policy extraction. We propose an information gain rate based weighted oblique decision tree (IGR-WODT) for DRL policy extraction. 3) Policy evaluation. We define three metrics to evaluate the performance of the proposed approach. A case study for the under-voltage load shedding problem shows that the IGR-WODT presents a performance enhancement compared with DRL, weighted oblique decision tree, and univariate decision tree. The proposed policy extraction method could provide an intuitive explanation of the neural network decision-making process to the dispatchers when making final decisions on power grid operation. Also, the resulted rule-based controller could replace the deep neural network-based controller in many field edge devices with limited computing resources, providing comparable performance.

deep reinforcement learning↗

Measure Utility, Gain Trust: Practical Advice for XAI Researchers

Research into explanation of machine learning models, i.e. explainable AI (XAI), has seen a sympathetic exponential growth alongside deep artificial neural networks throughout the past decade. For historical reasons explanation and trust have been intertwined. However this focus on trust is too narrow, and has led the research community astray from tried and true empirical methods that lead to more defensible scientific knowledge about people and explanations. To address this, we contribute a practical path forward for researchers in the XAI field. We recommend researchers focus on the utility and impact of their explanations instead of trust. We outline five broad use cases where explanations are useful and, for each, we describe pseudo-experiments that rely on objective empirical measurements and falsifiable hypotheses. We believe that this experimental rigor is necessary to contribute to scientific knowledge in the field of XAI.

Davis, Brittany F.↗

Refining water and carbon fluxes modeling in terrestrial ecosystems via plant hydraulics integration

Plant hydraulics substantially affects terrestrial water and carbon cycles by modulating water transport and carbon assimilation. Despite improved drought simulations in certain ecosystems through their integration into land surface models (LSMs), the broader application of plant hydraulics in diverse ecosystems and hydroclimates is still underexplored. Here, in this study, we implemented the recently developed Noah-Multiparameterization Land Surface Model (Noah-MP LSM) equipped with a plant hydraulics scheme (Noah-MP-PHS) across 40 FLUXNET sites globally. Employing the Shuffled Complex Evolution-University of Arizona (SCE-UA) auto-calibration algorithm, we optimized key plant hydraulics parameters for these sites spanning eight vegetation types in both arid and humid climates. Noah-MP-PHS significantly improves the simulation of evapotranspiration (ET) and gross primary production (GPP) by better representing atmospheric and soil water stress compared to traditional soil hydraulic schemes (SHSs, such as Noah and CLM). The augmented Noah-MP-PHS models reduce surface flux overestimation and underestimation, exhibiting an average increase of 0.14 and 0.15 in Kling-Gupta Efficiency (KGE) compared to Noah and CLM, respectively. The explicit consideration of plant capacitance in PHS reveals substantial deep-layer and nocturnal root water uptake especially under dry conditions. We employed eXplainable Machine learning (XML) to quantify the model’s relative sensitivity to newly introduced leaf-, stem and root-related parameters in PHS. The sensitivity analysis reveals a rise in root parameter importance and a decline in leaf and stem parameters as conditions shift from humid to arid. These findings indicate that as aridity states vary, the most influential parameters affecting surface fluxes variation may change in parameter calibration for PHS applications. Our findings underscore the importance of incorporating plant hydraulics into LSMs to enhance simulations of terrestrial water and carbon dynamics. These findings are crucial for understanding ecosystem responses to global climate changes and guide the broader application of PHS at larger scales.

54 ENVIRONMENTAL SCIENCES↗

Learning Global Proliferation Expertise Evolution Using AI-Driven Analytics and Public Information

Detecting and anticipating global proliferation expertise and capability evolution from unstructured, noisy, and incomplete public data streams is a highly desired, but extremely challenging task. Here, in this article, we present our pioneering data-driven approach to support the non-proliferation mission to detect and explain the evolution of proliferation expertise and capability development globally from terabytes of publicly available information (PAI), focusing on our knowledge extraction pipeline and descriptive analytics. We first discuss how we fuse nine open-source data streams, including multilingual data, to convert 4 TB of unstructured data to structured knowledge and encode dynamically evolving proliferation expertise representations—content and context graphs. For this, we rely on natural language processing (NLP) and deep learning (DL) models to perform information extraction, topic modeling, and distributed text representation (aka embedding) learning. We then present interactive, usable, and explainable descriptive analytics to refine domain knowledge and present it in a human-understandable form. Finally, we introduce future work avenues that will leverage our dynamic knowledge representations and descriptive analytics to enable predictive and prescriptive inferences to achieve real-time domain understanding and contextual reasoning about global proliferation expertise and capability evolution.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Super Resolving Unrolled Neural Networks for Remote Sensing

In remote sensing systems, the capabilities of the system are constrained by the complex interactions between size, weight, and power (SWAP) of potential designs. In electro-optical (EO) systems, examples of these critical parameters include the system’s sensitivity and resolution. Those parameters can be increased by ever larger optical apertures and focal planes but at the cost of more SWAP. Multi-image super resolution (MISR) techniques allow resolution to be enhanced via computation rather than more sophisticated optical hardware. These algorithms combine multiple images together into a single, higher resolution image, trading temporal resolution and computation for spatial resolution. Fielded MISR techniques, such as Drizzle, can require several hundred images to create a single super resolved image, implying reduced temporal resolution, increased data acquisition load, and limiting mission applications. Iterative techniques, such as model-based image reconstruction and compressive sensing, have been shown to create super resolved images using fewer images than Drizzle. They do this by posing an optimization problem that balances accuracy between a highly accurate physical model and an image model. In the case of super resolution, the physical model is defined by the relation between low resolution input images and the desired high resolution output image. The image model encodes some assumptions about the super resolved image. These assumptions are meant to suppress reconstruction artifacts that arise due to deterministic physical model error, stochastic measurement noise, and potential undersampling. In practice, the performance of iterative methods are limited by imaging models compatible with optimization. Deep learning-based methods can effectively learn image models of arbitrary complexity, but lack the theoretical explainability and robustness of iterative techniques. Consensus equilibrium (CE) generalizes the iterative techniques beyond optimization, enabling blackbox algorithms such as traditional and neural image denoisers to be used as the image model. CE-based approaches retain much of the explainability and robustness of iterative techniques while allowing the expressiveness of machine learning image models to be used. Additionally, by unrolling iterations of CE with an embedded image denoiser, the image denoiser can be further trained and specialized to the specific application with potentially higher quality reconstructions. Under this project, we demonstrated the feasibility of training an unrolled neural network based upon CE. While we didn’t train one, we showed that the CE process is differentiable and its gradient can be tractably computed. We also explored the usage of a variants of CE akin to generative neural works. Most importantly, we applied the CE framework to a number of problems including non-blind deconvolution, upsampling, single-image super resolution, MISR, event-based sensing, and saturated deconvolution. Our MISR prototype creates high quality reconstructions with an order of magnitude fewer images than previous approaches and, critically, produces these reconstructions fast enough for practical usage.

47 OTHER INSTRUMENTATION↗

Arm and shoulder muscle segmentation in axial MRI with UNet deep learning model

Quantifying individual upper-limb muscle volumes from MRI provides key insight into muscle-specific strength, deficits, and adaptations. Manual delineation is the gold standard but time‑intensive, and the performance of current deep learning approaches, particularly for small or anatomically complex muscles, remains incompletely characterized. We evaluated a state‑of‑the‑art deep learning framework across the entire upper limb and analyzed factors governing segmentation performance, with attention to the forearm. Three previously published MRI datasets (1.5 T, 3D GRE T1‑weighted; total n = 39) spanning young, middle‑aged, and older adults were curated and quality‑checked, including expert manual segmentations for 31 muscles. Following multiclass mask reconstruction, we trained three 3D nnU‑Net multiclass models matched to the muscle subsets present across datasets, using five‑fold cross‑validation and a composite Dice Similarity Coefficient (DSC) + cross entropy loss. Segmentation accuracy was assessed with DSC. Performance varied across muscles (mean DSC = 0.806 ± 0.098), ranging from 0.920 (Deltoid) to 0.461 (Extensor pollicis brevis). In uncertainty‑weighted regressions, muscle volume was positively associated with DSC (R2 = 0.36, p < 0.001), whereas training segmentation count and muscle orientation showed negligible associations (R2 ≤ 0.06). A weighted mixed‑effects model identified volume as the strongest evaluated predictor, explaining 23.9% of variance in DSC; orientation and training count each contributed <1%, leaving 61.5% unexplained. These results indicate that deep learning–based segmentation can accurately quantify muscle volume for many upper‑limb muscles but remains constrained for small, low‑contrast forearm muscles.

Gillespie, Samuel↗

Explainable multi-fidelity Bayesian neural network for distribution system state estimation

Distribution System State Estimation (DSSE) is frequently constrained by limited real-time measurements, the uncertainties introduced by distributed energy resources, and the presence of bad data. To address them, this paper proposes an enhanced Multi-Fidelity Bayesian Neural Network (MFBNN) DSSE approach. A low-fidelity layer based on a Deep Neural Network (DNN) is first pre-trained on pseudo-measurement data to learn fundamental state features. Subsequently, a high-fidelity Bayesian Neural Network (BNN) layer leverages limited but high-quality real-time measurements to refine these features, thereby achieving accurate DSSE. Additionally, the deep SHapley Additive exPlanation (SHAP) is developed to quantify the influence of measurement data on DSSE through dual perspectives of global feature importance and local nodal contributions, establishing a hierarchical explainability framework for machine learning-based DSSE. Comparative studies conducted on the IEEE 13-bus system and a real-world 2135-node system from Dominion Energy demonstrate that the proposed method excels in estimation accuracy, even under situations of high noise levels, bad data, and missing data. Further comparisons with Weighted Least Squares (WLS) and other machine learning-based DSSE approaches verify that the proposed framework offers higher accuracy, improved interpretability, and enhanced robustness.

Bad data↗

Semi-supervised Bayesian Low-shot Learning

Deep neural networks (NNs) typically outperform traditional machine learning (ML) approaches for complicated, non-linear tasks. It is expected that deep learning (DL) should offer superior performance for the important non-proliferation task of predicting explosive device configuration based upon observed optical signature, a task which human experts struggle with. However, supervised machine learning is difficult to apply in this mission space because most recorded signatures are not associated with the corresponding device description, or “truth labels.” This is challenging for NNs, which traditionally require many samples for strong performance. Semi-supervised learning (SSL), low-shot learning (LSL), and uncertainty quantification (UQ) for NNs are emerging approaches that could bridge the mission gaps of few labels and rare samples of importance. NN explainability techniques are important in gaining insight into the inferential feature importance of such a complex model. In this work, SSL, LSL, and UQ are merged into a single framework, a significant technical hurdle not previously demonstrated. Exponential Average Adversarial Training (EAAT) and Pairwise Neural Networks (PNNs) are chosen as the SSL and LSL methods of choice. Permutation feature importance (PFI) for functional data is used to provide explainability via the Variable importance Explainable Elastic Shape Analysis (VEESA) pipeline. A variety of uncertainty quantification approaches are explored: Bayesian Neural Networks (BNNs), ensemble methods, concrete dropout, and evidential deep learning. Two final approaches, one utilizing ensemble methods and one utilizing evidential learning, are constructed and compared using a well-quantified synthetic 2D dataset along with the DIRSIG Megascene.

97 MATHEMATICS AND COMPUTING↗