Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Explanability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Leveraging large language models to automate the identification of healthcare access barriers for veterans

Objective: To develop and evaluate an automated system for identifying healthcare barriers focusing on transportation issues in veterans’ clinical notes using large language models (LLMs) and to assess the impact of different prompting strategies on classification performance and explanation consistency. Methods: We developed a hybrid system combining pattern matching for templated notes with LLM analysis for free-text notes. Using 2000 manually annotated clinical notes, we compared four prompting strategies (dual-role short, dual-role long, analysis-first, analysis-only) across Mistral-7B and Llama-3.1 models. We evaluated classification performance using standard metrics and assessed explanation consistency through embedding similarity analysis. Results: The analysis-first strategy achieved superior performance, with Mistral-7B reaching an F1 score of 0.914, outperforming traditional machine learning approaches (GBM: 0.786, BERT: 0.811). LLMs demonstrated higher explanation consistency within models (mean cosine similarity 0.887–0.908) compared to cross-model similarities (0.767–0.872). Pattern matching successfully handled 6.7% of templated notes deterministically. Mistral-7B showed greater internal consistency but higher abstention rates compared to Llama-3.1. Conclusion: Requiring LLMs to analyze evidence before classification improves both accuracy and explanation consistency for identifying transportation barriers in clinical notes. This approach enables automated barrier detection at scale while providing clinically relevant explanations, supporting both population-level healthcare planning and individual patient care decisions.

Healthcare access barriers↗

ACES-GNN: can graph neural network learn to explain activity cliffs?

Graph Neural Networks (GNNs) have revolutionized molecular property prediction by leveraging graph-based representations, yet their opaque decision-making processes hinder broader adoption in drug discovery. This study introduces the Activity-Cliff-Explanation-Supervised GNN (ACES-GNN) framework, designed to simultaneously improve predictive accuracy and interpretability by integrating explanation supervision for activity cliffs (ACs) into GNN training. ACs, defined by structurally similar molecules with significant potency differences, pose challenges for traditional models due to their reliance on shared structural features. By aligning model attributions with chemist-friendly interpretations, the ACES-GNN framework bridges the gap between prediction and explanation. Validated across 30 pharmacological targets, ACES-GNN consistently enhances both predictive accuracy and attribution quality for ACs compared to unsupervised GNNs. Our results demonstrate a positive correlation between improved predictions and accurate explanations, offering a robust and adaptable framework to better understand and interpret ACs. This work underscores the potential of explanation-guided learning to advance interpretable artificial intelligence in molecular modeling and drug discovery.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Explaining machine-learning models for gamma-ray detection and identification

As more complex predictive models are used for gamma-ray spectral analysis, methods are needed to probe and understand their predictions and behavior. Recent work has begun to bring the latest techniques from the field of Explainable Artificial Intelligence (XAI) into the applications of gamma-ray spectroscopy, including the introduction of gradient-based methods like saliency mapping and Gradient-weighted Class Activation Mapping (Grad-CAM), and black box methods like Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). In addition, new sources of synthetic radiological data are becoming available, and these new data sets present opportunities to train models using more data than ever before. In this work, we use a neural network model trained on synthetic NaI(Tl) urban search data to compare some of these explanation methods and identify modifications that need to be applied to adapt the methods to gamma-ray spectral data. We find that the black box methods LIME and SHAP are especially accurate in their results, and recommend SHAP since it requires little hyperparameter tuning. We also propose and demonstrate a technique for generating counterfactual explanations using orthogonal projections of LIME and SHAP explanations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Explaining word embeddings with perfect fidelity: a case study in predicting research impact

The best-performing approaches for scholarly document quality prediction are based on embedding models. In addition to their performance when used in classifiers, embedding models can also provide predictions even for words that were not contained in the labelled training data for the classification model, which is important in the context of the ever-evolving research terminology. Although model-agnostic explanation methods, such as Local interpretable model-agnostic explanations, can be applied to explain machine learning classifiers trained on embedding models, these produce results with questionable correspondence to the model. We introduce a new feature importance method, Self-Model Entities Rated (SMER), for logistic regression-based classification models trained on word embeddings. We show that SMER has theoretically perfect fidelity with the explained model, as the average of logits of SMER scores for individual words (SMER explanation) exactly corresponds to the logit of the prediction of the explained model. Quantitative and qualitative evaluation is performed through five diverse experiments conducted on 50,000 research articles (papers) from the CORD-19 corpus. In conclusion, through an AOPC curve analysis, we experimentally demonstrate that SMER produces better explanations than LIME, SHAP and global tree surrogates.

Coarse-grained models↗

RX-ADS: Interpretable Anomaly Detection Using Adversarial ML for Electric Vehicle CAN Data

Recent year has brought considerable advancements in Electric Vehicles (EVs) and associated infrastructures/ communications. Intrusion Detection Systems (IDS) are widely deployed for anomaly detection in such critical infrastructures. This paper presents an Interpretable Anomaly Detection System (RX-ADS) for intrusion detection in CAN protocol communication in EVs. Contributions include: 1) window based feature extraction method; 2) deep Autoencoder based anomaly detection method; and 3) adversarial machine learning based explanation generation methodology. The presented approach was tested on two benchmark CAN datasets: OTIDS and Car Hacking. The anomaly detection performance of RX-ADS was compared against the state-of-the-art approaches on these datasets: HIDS and GIDS. The RX-ADS approach presented performance comparable to the HIDS approach (OTIDS dataset) and has outperformed HIDS and GIDS approaches (Car Hacking dataset). Further, the proposed approach was able to generate explanations for detected abnormal behaviors arising from various intrusions. Furthermore, these explanations were later validated by information used by domain experts to detect anomalies. Other advantages of RX-ADS include: 1) the method can be trained on unlabeled data; 2) explanations help experts in understanding anomalies and root course analysis, and also help with AI model debugging and diagnostics, ultimately improving user trust in AI systems.

42 ENGINEERING↗

Search for NC Delta Radiative Decay Single Photon Events In MicroBooNE

MicroBooNE, a liquid argon time projection chamber at Fermilab, is investigating the MiniBooNE anomaly, which consists of an excess of low-energy electromagnetic showers in a neutrino beam. Recent results from MicroBooNE have ruled out a 3+1 sterile neutrino explanation of the anomaly, leaving the single photon-like explanation as the most likely possibility. In this poster, we describe results of a new expanded search for the neutral current Delta radiative decay topology, a rare type of event which is by far the largest expected source of neutrino-induced single photons. Using two different reconstruction paradigms simultaneously, we study in detail events both with and without visible hadronic activity; either category of event could explain the MiniBooNE anomaly, but the two could have very different implications for underlying physics. We find that events with visible protons are excluded as an explanation of the MiniBooNE anomaly, but events without visible protons remain a possible explanation requiring further study.

Hagaman, Lee [Nevis Labs, Columbia U.] (ORCID:0000↗

Explaining Missing Data in Graphs: A Constraint-based Approach

Abstract: This paper introduces a constraint-based approach to clarify missing values in graphs. Our method capitalizes on a set S of graph data constraints. An explanation is a sequence of operational enforcement of S towards the recovery of interested yet missing data (e.g., attribute values, edges). We show that constraint-based approach helps us to understand not only why a value is missing, but also how to recover the missing value. We study S-explanation problem, which is to compute the optimal explanations with guarantees on the informativeness and conciseness. We show the problem is in ?P^2 for established graph data constraints such as graph keys and graph association rules. We develop an efficient bidirectional algorithm to compute optimal explanations, without enforcing S on the entire graph. We also show our algorithm can be easily extended to support graph refinement within limited time, and to explain missing answers. Using real-world graphs, we experimentally verify the effectiveness and efficiency of our algorithms.

Data Analytics↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Hiding in plain sight: The prevalence and impact of trions and Fermi polarons in transient absorption spectroscopy experiments of 2D semiconductors

Transient absorption (TA) spectroscopy is one of the most popular experimental methods to measure the excited state lifetimes and charge carrier recombination mechanisms in two dimensional (2D) semiconductors. This fundamental information is essential for designing and optimizing the next generation of ultrathin and lightweight 2D semiconductor-based optoelectronic devices. However, the interpretation of TA spectroscopy data varies across the community. The community lacks a unifying physical explanation for how and why experimental variables such as incident light intensity, sample-substrate interactions, and/or applied bias affect TA spectral data. This Perspective (1) compares the physical chemistry TA literature to nanomaterial physics literature from a historical perspective, (2) reviews multiple physical explanations that the TA community developed to explain spectral features and experimental trends, (3) provides a unifying explanation for how and why trions—and, more generally, Fermi polarons—contribute to TA spectra, and (4) quantifies the extent to which various physical interpretations and data analysis procedures yield different timescales and mechanisms for the same set of experimental results. We highlight the importance of considering trions/Fermi polarons in TA measurements and their implications for advancing our understanding of 2D material properties.

2D materials↗

Evidence for the Late Arrival of Hot Jupiters in Systems with High Host-star Obliquities

It has been shown that hot Jupiters systems with massive, hot stellar primaries exhibit a wide range of stellar obliquities. On the other hand, hot Jupiter systems with low-mass, cool primaries often have stellar obliquities close to zero. Efficient tidal interactions between hot Jupiters and the convective envelopes present in lower-mass main-sequence stars have been a popular explanation for these observations. If this explanation is accurate, then aligned systems should be older than misaligned systems. Likewise, the convective envelope mass of a hot Jupiter's host star should be an effective predictor of its obliquity. We derive homogeneous stellar parameters—including convective envelope masses—for hot Jupiter host stars with high-quality sky-projected obliquity inferences. Using a thin-disk stellar population's Galactic velocity dispersion as a relative age proxy, we find that hot Jupiter host stars with larger-than-median obliquities are older than hot Jupiter host stars with smaller-than-median obliquities. The relative age difference between the two populations is larger for hot Jupiter host stars with smaller-than-median fractional convective envelope masses and is significant at the 3.6σ level. We identify stellar mass, not convective envelope mass, as the best predictor of stellar obliquity in hot Jupiter systems. The best explanation for these observations is that many hot Jupiters in misaligned systems arrived in the close proximity of their host stars long after their parent protoplanetary disks dissipated. The dependence of observed age offset on convective envelope mass suggests that tidal realignment contributes to the population of aligned hot Jupiters orbiting stars with convective envelopes.

79 ASTRONOMY AND ASTROPHYSICS↗

Explainable Machine Learning for Functional Data

Black-box machine learning models are recognized as useful tools for prediction applications, but the algorithmic complexity of some models causes interpretation challenges. Explainability methods have been proposed to provide insight into these models, but there is little research focused on supervised modeling with functional data inputs. We argue that, especially in applications of high consequence, it is important to explicitly model the functional dependence in a black-box analysis to not obscure or misrepresent patterns in explanations. As such, we propose the V ariable importance E xplainable E lastic S hape A nalysis (VEESA) pipeline for training supervised machine learning models with functional inputs. The pipeline is an analysis process that includes the data preprocessing, modeling, and post-hoc explanations. The preprocessing is done using elastic functional principal components analysis, which accounts for vertical and horizontal variability in functional data and, ultimately, allows for explanations in the original data space that identify the important functional variability without bias due to correlated variables. Here, we demonstrate the pipeline on two high-consequence applications: explosives classification for national security and inkjet printer identification in forensic science. The applications exhibit the VEESA pipeline’s ability to provide an understanding of the characteristics of the functional data useful for prediction. Code for implementing the pipeline is available in the veesa R package (and supplemental python code).

Elastic Shape Analysis↗

Experimental and Phenomenological Investigations of the MiniBooNE Anomaly

This thesis covers a range of experimental and theoretical efforts to elucidate the origin of the $4.8\sigma$ MiniBooNE low energy excess (LEE). We begin with the follow-up MicroBooNE experiment, which took data along the BNB from 2016 to 2021. This thesis specifically presents MicroBooNE's search for $\nu_e$ charged-current quasi-elastic (CCQE) interactions consistent with two-body scattering. The two-body CCQE analysis uses a novel reconstruction process, including a number of deep-learning-based algorithms, to isolate a sample of $\nu_e$ CCQE interaction candidates with $75\%$ purity. The analysis rules out an entirely $\nu_e$-based explanation of the MiniBooNE excess at the $2.4\sigma$ confidence level. We next perform a combined fit of MicroBooNE and MiniBooNE data to the popular $3+1$ model; even after the MicroBooNE results, allowed regions in $\Delta m^2$-$\sin^2 2_{\theta_{\mu e}}$ parameter space exist at the $3\sigma$ confidence level. This thesis also demonstrates that the MicroBooNE data are consistent with a $\overline{\nu}_e$-based explanation of the MiniBooNE LEE at the $<2\sigma$ confidence level. Next, we investigate a phenomenological explanation of the MiniBooNE excess combining the $3+1$ model with a dipole-coupled heavy neutral lepton (HNL). It is shown that a 500 MeV HNL can accommodate the energy and angular distributions of the LEE at the $2\sigma$ confidence level while avoiding stringent constraints derived from MINER$\nu$A elastic scattering data. Finally, we discuss the Coherent CAPTAIN-Mills experiment--a 10-ton light-based liquid argon detector at Los Alamos National Laboratory. The background rejection achieved from a novel Cherenkov-based reconstruction algorithm will enable world-leading sensitivity to a number of beyond-the-Standard Model physics scenarios, including dipole-coupled HNLs.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Teaching Freight Mode Choice Models New Tricks Using Interpretable Machine Learning Methods

Understanding and forecasting the intricate freight mode choice behavior under various industry, policy, and technology contexts is essential in freight planning and policymaking. Numerous models have been developed in prior studies to provide insights into freight mode selection, the majority of which use discrete choice models such as multinomial logit (MNL) models. However, logit models often rely on linear specifications of independent variables, despite potential nonlinear relationships in the data. Moreover, there often lacks a heuristic and efficient approach to identify such complex relationships to define the logit model specifications. To fill this gap, we developed an MNL model for freight mode choice using the insights from state-of-the- art machine learning (ML) models. ML models can capture the nonlinear nature of the complex decision-making process, and recent advances in 'explainable AI' have greatly improved their interpretability. The interpretable ML methods help enhance the performance of MNL models and advance knowledge of freight mode choice. Specifically, the influential factors and their relationship with individual modes are identified using SHapley Additive exPlanations (SHAP) to improve the MNL's performance. The workflow is demonstrated in a case study of Austin, Texas, and the SHAP results reveal multiple nonlinear relationships predicted by ML models. Incorporating those relationships into MNL model specifications improves the interpretability and accuracy of the MNL model compared to a conventional MNL model. Findings from this study can be used to guide freight planning and inform policymakers and practitioners on how key factors affect freight decision-making.

ADVANCED PROPULSION SYSTEMS,MATHEMATICS AND COMPUT↗

Explainable AI (XAI)-driven vibration sensing scheme for surface quality monitoring in a smart surface grinding process

Local Interpretable and Model-agnostic Explanation (LIME), an explainable artificial intelligence (XAI) approach is adapted to identify the globally important time-frequency bands for predicting average surface roughness (Ra) in a smart grinding process. The smart grinding setup consisted of a Supertech CNC precision surface grinding machine, instrumented with a Dytran piezoelectric accelerometer attached to the tailstock quill along the tangential direction (Y-axis). For every grinding pass, vibration signatures were captured, and the ground truth surface roughness values were recorded using a Mahr Marsurf M300C portable surface roughness profilometer. The roughness values ranged from 0.06 to 0.14 microns over the complete set of experiments. Time-frequency domain spectrogram frames were extracted for each of the vibration signals collected during the grinding process. Convolutional Neural Networks (CNNs) were modeled to predict the surface roughness based on these spectrogram frames and their image augmentations. The best CNN model was able to predict the roughness values with an overall R2-score of 0.95, training R2-score of 0.99, and testing R2-score of 0.81 with only 80 sets of vibration signals corresponding to 4 experiments with 20 trials each. Although the data size is not large enough to guarantee such performance metrics in real-world scenarios, one can extract statistically consistent explanations underlying the relationships these complex deep learning models capture. Further, the LIME methodology was implemented on the developed surface roughness CNN model to identify the important time-frequency bands (i.e., the superpixels of a spectrogram) influencing the predictions. Based on the identified important regions on the spectrogram frames, the corresponding frequency characteristics were determined that influence the surface roughness predictions. The important frequency range based on LIME results was approximately 11.7 to 19.1 kHz. The power of XAI was demonstrated by cutting down the sampling rate from 160 kHz to 30, 20, 10, and 5 kHz based on the important frequency range and considering Nyquist criteria. Separate CNN models were developed for these ranges by only extracting time-frequency contents below their corresponding Nyquist cut-offs. A proper data acquisition strategy is proposed by comparing the model performances to argue the selection of a sufficient sampling rate to capture the grinding process successfully and robustly.

42 ENGINEERING↗

Reprint of: Explainable AI (XAI)-driven vibration sensing scheme for surface quality monitoring in a smart surface grinding process

Local Interpretable and Model-agnostic Explanation (LIME), an explainable artificial intelligence (XAI) approach is adapted to identify the globally important time-frequency bands for predicting average surface roughness (Ra) in a smart grinding process. The smart grinding setup consisted of a Supertech CNC precision surface grinding machine, instrumented with a Dytran piezoelectric accelerometer attached to the tailstock quill along the tangential direction (Y-axis). For every grinding pass, vibration signatures were captured, and the ground truth surface roughness values were recorded using a Mahr Marsurf M300C portable surface roughness profilometer. The roughness values ranged from 0.06 to 0.14 microns over the complete set of experiments. Time-frequency domain spectrogram frames were extracted for each of the vibration signals collected during the grinding process. Convolutional Neural Networks (CNNs) were modeled to predict the surface roughness based on these spectrogram frames and their image augmentations. The best CNN model was able to predict the roughness values with an overall R2-score of 0.95, training R2-score of 0.99, and testing R2-score of 0.81 with only 80 sets of vibration signals corresponding to 4 experiments with 20 trials each. Although the data size is not large enough to guarantee such performance metrics in real-world scenarios, one can extract statistically consistent explanations underlying the relationships these complex deep learning models capture. Further, the LIME methodology was implemented on the developed surface roughness CNN model to identify the important time-frequency bands (i.e., the superpixels of a spectrogram) influencing the predictions. Based on the identified important regions on the spectrogram frames, the corresponding frequency characteristics were determined that influence the surface roughness predictions. The important frequency range based on LIME results was approximately 11.7 to 19.1 kHz. The power of XAI was demonstrated by cutting down the sampling rate from 160 kHz to 30, 20, 10, and 5 kHz based on the important frequency range and considering Nyquist criteria. Separate CNN models were developed for these ranges by only extracting time-frequency contents below their corresponding Nyquist cut-offs. A proper data acquisition strategy is proposed by comparing the model performances to argue the selection of a sufficient sampling rate to capture the grinding process successfully and robustly.

47 OTHER INSTRUMENTATION↗

Explainable Synthesizability Prediction of Inorganic Crystal Polymorphs Using Large Language Models

Abstract We evaluate the ability of machine learning to predict whether a hypothetical crystal structure can be synthesized and explain those predictions to scientists. Fine‐tuned large language models (LLMs) trained on a human‐readable text description of the target crystal structure perform comparably to previous bespoke convolutional graph neural network methods, but better prediction quality can be achieved by training a positive‐unlabeled learning model on a text‐embedding representation of the structure. An LLM‐based workflow can then be used to generate human‐readable explanations for the types of factors governing synthesizability, extract the underlying physical rules, and assess the veracity of those rules. These explanations can guide chemists in modifying or optimizing non‐synthesizable hypothetical structures to make them more feasible for materials design.

Kim, Seongmin [Department of Chemical and Biologic↗

Explainable Synthesizability Prediction of Inorganic Crystal Polymorphs Using Large Language Models

Abstract We evaluate the ability of machine learning to predict whether a hypothetical crystal structure can be synthesized and explain those predictions to scientists. Fine‐tuned large language models (LLMs) trained on a human‐readable text description of the target crystal structure perform comparably to previous bespoke convolutional graph neural network methods, but better prediction quality can be achieved by training a positive‐unlabeled learning model on a text‐embedding representation of the structure. An LLM‐based workflow can then be used to generate human‐readable explanations for the types of factors governing synthesizability, extract the underlying physical rules, and assess the veracity of those rules. These explanations can guide chemists in modifying or optimizing non‐synthesizable hypothetical structures to make them more feasible for materials design.

Kim, Seongmin [Department of Chemical and Biologic↗

Conformal freeze-in, composite dark photon, and asymmetric reheating

Large classes of dark sector models feature mass scales and couplings very different from the ones we observe in the Standard Model (SM). Moreover, in the freeze-in mechanism, often employed by the dark sector models, it is also required that the dark sector cannot be populated during the reheating process like the SM. This is the so called asymmetric reheating. Such disparities in sizes and scales often call for dynamical explanations. In this paper, we explore a scenario in which slow evolving conformal field theories (CFTs) offer such an explanation. Building on the recent work on conformal freeze-in (COFI), we focus on a coupling between the Standard Model Hypercharge gauge boson and an anti-symmetric tensor operator in the dark CFT. We present a scenario which dynamically realizes the asymmetric reheating and COFI production. With a detailed study of dark matter production, and taking into account limits on the dark matter (DM) self-interaction, warm DM bound, and constraints from the stellar evolution, we demonstrate that the correct relic abundance can be obtained with reasonable choices of parameters. The model predicts the existence of a dark photon as an emergent composite particle, with a small kinetic mixing also determined by the CFT dynamics, which correlates it with the generation of the mass scale of the dark sector. At the same time, COFI production of dark matter is very different from those freeze-in mediated by the dark photon. This is an example of the physics in which a realistic dark sector model can often be much richer and with unexpected features.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗