Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “explainable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Quantum biological insights into CRISPR-Cas9 sgRNA efficiency from explainable-AI driven feature engineering

Abstract CRISPR-Cas9 tools have transformed genetic manipulation capabilities in the laboratory. Empirical rules-of-thumb have been developed for only a narrow range of model organisms, and mechanistic underpinnings for sgRNA efficiency remain poorly understood. This work establishes a novel feature set and new public resource, produced with quantum chemical tensors, for interpreting and predicting sgRNA efficiency. Feature engineering for sgRNA efficiency is performed using an explainable-artificial intelligence model: iterative Random Forest (iRF). By encoding quantitative attributes of position-specific sequences for Escherichia coli sgRNAs, we identify important traits for sgRNA design in bacterial species. Additionally, we show that expanding positional encoding to quantum descriptors of base-pair, dimer, trimer, and tetramer sequences captures intricate interactions in local and neighboring nucleotides of the target DNA. These features highlight variation in CRISPR-Cas9 sgRNA dynamics between E. coli and H. sapiens genomes. These novel encodings of sgRNAs enhance our understanding of the elaborate quantum biological processes involved in CRISPR-Cas9 machinery.

59 BASIC BIOLOGICAL SCIENCES↗

Explaining the extra crystal-field mode in 𝐴 ⁢Ce ⁢𝑋 2 (𝐴 = K, Rb, Na; 𝑋 = O, S, Se, Te)

A growing list of Ce-based magnets have shown an extra and heretofore unexplained crystal electric field (CEF) mode at high energies. We describe a process whereby an optical phonon can produce a split CEF mode well above the phonon energy. We use density functional theory and point-charge model calculations to estimate the phonon distortions and coupling to model this effect in KCeO 2 , showing that it accounts for the extra CEF mode observed. Furthermore, this mechanism is generic and may explain the extra modes observed on a variety of Ce 3+ compounds.

36 MATERIALS SCIENCE↗

Explaining Snowball-in-Hell Phenomena in Heavy-Ion Collisions Using a Novel Thermodynamic Variable

A loosely bound hadronic molecule produced by a relativistic heavy-ion collision has been described as a “snowball in hell” since it emerges from a hadron resonance gas whose temperature is orders of magnitude larger than the binding energy of the molecule. This remarkable phenomenon can be explained in terms of a novel thermodynamic variable called the “contact” that is conjugate to the binding momentum of the molecule. The production rate of the molecule can be expressed in terms of the contact density at the kinetic freeze-out of the hadron resonance gas. It approaches a nonzero limit as the binding energy goes to 0.

Relativistic heavy-ion collisions↗

Explaining Neural Spike Activity for Simulated Bio-plausible Network through Deep Sequence Learning

With significant improvements in large-scale simulations of brain models, there is a growing need to develop tools for rapid analysis and interpreting the simulation results. In this work, we explore the potential of sequential deep learning models to understand and explain the network dynamics among the neurons extracted from a large-scale neural simulation in STACS (Simulation Tool for Asynchronous Cortical Stream). Our method employs a representative neuroscience model that abstracts the cortical dynamics with a reservoir of randomly connected spiking neurons with a low stable spike firing rate throughout the simulation duration. We subsequently analyze the spike dynamics of the simulated spiking neural network through an autoencoder model and an attention-based mechanism.

Kulkarni, Shruti↗

Granal thylakoid structure and function: explaining an enduring mystery of higher plants

Summary In higher plants, photosystems II and I are found in grana stacks and unstacked stroma lamellae, respectively. To connect them, electron carriers negotiate tortuous multi‐media paths and are subject to macromolecular blocking. Why does evolution select an apparently unnecessary, inefficient bipartition? Here we systematically explain this perplexing phenomenon. We propose that grana stacks, acting like bellows in accordions, increase the degree of ultrastructural control on photosynthesis through thylakoid swelling/shrinking induced by osmotic water fluxes. This control coordinates with variations in stomatal conductance and the turgor of guard cells, which act like an accordion's air button. Thylakoid ultrastructural dynamics regulate macromolecular blocking/collision probability, direct diffusional pathlengths, division of function of Cytochrome b 6 f complex between linear and cyclic electron transport, luminal pH via osmotic water fluxes, and the separation of pH dynamics between granal and lamellar lumens in response to environmental variations. With the two functionally asymmetrical photosystems located distantly from each other, the ultrastructural control, nonphotochemical quenching, and carbon‐reaction feedbacks maximally cooperate to balance electron transport with gas exchange, provide homeostasis in fluctuating light environments, and protect photosystems in drought. Grana stacks represent a dry/high irradiance adaptation of photosynthetic machinery to improve fitness in challenging land environments. Our theory unifies many well‐known but seemingly unconnected phenomena of thylakoid structure and function in higher plants.

59 BASIC BIOLOGICAL SCIENCES↗

Explainable and trustworthy artificial intelligence for correctable modeling in chemical sciences

Data science has primarily focused on big data, but for many physics, chemistry, and engineering applications, data are often small, correlated and, thus, low dimensional, and sourced from both computations and experiments with various levels of noise. Typical statistics and machine learning methods do not work for these cases. Expert knowledge is essential, but a systematic framework for incorporating it into physics-based models under uncertainty is lacking. Here, we develop a mathematical and computational framework for probabilistic artificial intelligence (AI)–based predictive modeling combining data, expert knowledge, multiscale models, and information theory through uncertainty quantification and probabilistic graphical models (PGMs). We apply PGMs to chemistry specifically and develop predictive guarantees for PGMs generally. Our proposed framework, combining AI and uncertainty quantification, provides explainable results leading to correctable and, eventually, trustworthy models. The proposed framework is demonstrated on a microkinetic model of the oxygen reduction reaction.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Loss of a satellite could explain Saturn’s obliquity and young rings

The origin of Saturn’s ~26.7° obliquity and ~100-million-year-old rings is unknown. The observed rapid outward migration of Saturn’s largest satellite, Titan, could have raised Saturn’s obliquity through a spin-orbit precession resonance with Neptune. We use Cassini data to refine estimates of Saturn’s moment of inertia, finding that it is just outside the range required for the resonance. We propose that Saturn previously had an additional satellite, which we name Chrysalis, that caused Saturn’s obliquity to increase through the Neptune resonance. Destabilization of Chrysalis’s orbit ~100 million years ago can then explain the proximity of the system to the resonance and the formation of the rings through a grazing encounter with Saturn.

Science & Technology - Other Topics↗

Explainable Neural Architecture Search (XNAS)

Code for the paper Learning Interpretable Models Through Multi-Objective Neural Architecture Search by Zachariah Carmichael, Tim Moon, and Sam Ade Jacobs. Monumental advances in deep learning have led to unprecedented achievements across a multitude of domains. While the performance of deep neural networks is indubitable, the architectural design and interpretability of such models are nontrivial. Research has been introduced to automate the design of neural network architectures through neural architecture search (NAS). Recent progress has made these methods more pragmatic by exploiting distributed computation and novel optimization algorithms. However, there is little work in optimizing architectures for interpretability. To this end, we propose a multiobjective distributed NAS framework that optimizes for both task performance and introspection. We leverage the non-dominated sorting genetic algorithm (NSGA-II) and explainable AI (XAI) techniques to reward architectures that can be better comprehended by humans. The framework is evaluated on several image classification datasets. We demonstrate that jointly optimizing for introspection ability and task error leads to more disentangled architectures that perform within tolerable error.

Carmichael, ZachariahJ↗

Vistransformers Explained

The Vistransformers Explained library is a collection of python notebooks that demonstrate the internal mechanics and uses of visual-transformer (ViT) machine learning models. The code implements, with mild modifications, ViT models that have been made publicly available through publication and GitHub code. The value added by this code is in-depth explanations of the mathematics behind the sub-modules of the ViT models, including original figures. Additionally, the library contains the code necessary to implement and train the ViT models. The library does not include example training data for the models; instead, it would rely on users generating their own datasets. The code is based on the PyTorch python library. It does not include any files other than python scripts, modules, or notebooks.

Callis, Skylar↗

Understanding Predictability of Daily Southeast U.S. Precipitation Using Explainable Machine Learning

Abstract We investigate the predictability of the sign of daily southeastern U.S. (SEUS) precipitation anomalies associated with simultaneous predictors of large-scale climate variability using machine learning models. Models using index-based climate predictors and gridded fields of large-scale circulation as predictors are utilized. Logistic regression (LR) and fully connected neural networks using indices of climate phenomena as predictors produce neither accurate nor reliable predictions, indicating that the indices themselves are not good predictors. Using gridded fields as predictors, an LR and convolutional neural network (CNN) are more accurate than the index-based models. However, only the CNN can produce reliable predictions that can be used to identify forecasts of opportunity. Using explainable machine learning we identify which variables and grid points of the input fields are most relevant for confident and correct predictions in the CNN. Our results show that the local circulation is most important as represented by maximum relevance of 850-hPa geopotential heights and zonal winds to making skillful, high-probability predictions. Corresponding composite anomalies identify connections with El Niño–Southern Oscillation during winter and the Atlantic multidecadal oscillation and North Atlantic subtropical high during summer.

Pegion, Kathy↗

Explaining Globally Inhomogeneous Future Changes in Monsoons Using Simple Moist Energy Diagnostics

Here, this study examines the annual cycle of monsoon precipitation simulated by models from phase 6 of the Coupled Model Intercomparison Project (CMIP6), then uses moist energy diagnostics to explain globally inhomogeneous projected future changes. Rainy season characteristics are quantified using a consistent method across the globe. Model bias is shown to include rainy season onsets tens of days later than observed in some monsoon regions (India, Australia, and North America) and overly large summer precipitation in others (North America, South America, and southern Africa). Projected next-century changes include rainy season lengthening in the two largest Northern Hemisphere monsoon regions (South Asia and central Sahel) and shortening in the two largest Southern Hemisphere regions (South America and southern Africa). Changes in the North American and Australian monsoons are less coherent across models. To understand these changes, relative moist static energy (MSE) is defined as the difference between local and tropical-mean surface air MSE. Future changes in relative MSE in each region correlate well with onset and demise date changes. Furthermore, Southern Hemisphere regions projected to undergo rainy season shortening are spanned by an increasing equator-to-pole MSE gradient, suggesting their rainfall will be increasingly inhibited by fluxes of dry extratropical air; Northern Hemisphere regions with projected lengthening of rainy seasons undergo little change in equator-to-pole MSE gradient. Thus, although model biases raise questions as to the reliability of some projections, these results suggest that globally inhomogeneous future changes in monsoon timing may be understood through simple measures of surface air MSE.

54 ENVIRONMENTAL SCIENCES↗

Human Readiness Levels Explained

The Human Readiness Level scale complements and supplements the existing technology readiness level scale to support comprehensive and systematic evaluation of human system aspects throughout a system’s life cycle. The objective is to ensure humans can use a fielded technology or system as intended to support mission operations safely and effectively. This article defines the nine human readiness levels in the scale, explains their meaning, and illustrates their application using a helmet-mounted display example.

60 APPLIED LIFE SCIENCES↗

Potentially adaptive SARS-CoV-2 mutations discovered with novel spatiotemporal and explainable AI models

Abstract Background A mechanistic understanding of the spread of SARS-CoV-2 and diligent tracking of ongoing mutagenesis are of key importance to plan robust strategies for confining its transmission. Large numbers of available sequences and their dates of transmission provide an unprecedented opportunity to analyze evolutionary adaptation in novel ways. Addition of high-resolution structural information can reveal the functional basis of these processes at the molecular level. Integrated systems biology-directed analyses of these data layers afford valuable insights to build a global understanding of the COVID-19 pandemic. Results Here we identify globally distributed haplotypes from 15,789 SARS-CoV-2 genomes and model their success based on their duration, dispersal, and frequency in the host population. Our models identify mutations that are likely compensatory adaptive changes that allowed for rapid expansion of the virus. Functional predictions from structural analyses indicate that, contrary to previous reports, the Asp 614 Gly mutation in the spike glycoprotein (S) likely reduced transmission and the subsequent Pro 323 Leu mutation in the RNA-dependent RNA polymerase led to the precipitous spread of the virus. Our model also suggests that two mutations in the nsp13 helicase allowed for the adaptation of the virus to the Pacific Northwest of the USA. Finally, our explainable artificial intelligence algorithm identified a mutational hotspot in the sequence of S that also displays a signature of positive selection and may have implications for tissue or cell-specific expression of the virus. Conclusions These results provide valuable insights for the development of drugs and surveillance strategies to combat the current and future pandemics.

59 BASIC BIOLOGICAL SCIENCES↗

Explainable Artificial Intelligence in Endocrinological Medical Research

Artificial Intelligence (AI) has been a part of the medical community for decades in the form of Clinical Decision Support Systems (CDSS) to aide physicians in diagnosis and categorization of patients (1). Recent years have seen a shift from expert-derived models to the integration of machine learning (ML) algorithms to drive the output of these AI systems due to the ability of ML models to more accurately make predictions by exploiting higher dimensional and often complex data. In many cases ML models gain their advantage in accuracy by capture complex and often non-linear relationships between features being used to make the prediction. However, the hype and excitement around these methods are tempered by the clinical utility of often black-box solutions driven by skepticism of results that are difficult for practitioners to not only interpret but explain to their patients (1-3). This skepticism is not unfounded as multiple examples of black-box solutions identifying incidental correlates as the key predictors have highlighted the potential bias in a training set, or reward system, that a ML model may exploit; for example a model discerning wolves from huskies based on snow in the background rather than features of the dogs (1, 4, 5).

Endocrinological, artificial intelligence, diabete↗

Explaining machine-learning models for gamma-ray detection and identification

As more complex predictive models are used for gamma-ray spectral analysis, methods are needed to probe and understand their predictions and behavior. Recent work has begun to bring the latest techniques from the field of Explainable Artificial Intelligence (XAI) into the applications of gamma-ray spectroscopy, including the introduction of gradient-based methods like saliency mapping and Gradient-weighted Class Activation Mapping (Grad-CAM), and black box methods like Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). In addition, new sources of synthetic radiological data are becoming available, and these new data sets present opportunities to train models using more data than ever before. In this work, we use a neural network model trained on synthetic NaI(Tl) urban search data to compare some of these explanation methods and identify modifications that need to be applied to adapt the methods to gamma-ray spectral data. We find that the black box methods LIME and SHAP are especially accurate in their results, and recommend SHAP since it requires little hyperparameter tuning. We also propose and demonstrate a technique for generating counterfactual explanations using orthogonal projections of LIME and SHAP explanations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Data and Scripts Associated with the Manuscript “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids”

This data package is associated with the publication “Water Column Respiration in the Yakima River Basin is Explained by Temperature, Nutrients and Suspended Solids” published in EGU Biogeochemistry (Laan et al. 2025). In this research, water column respiration (ERwc) data, surface water chemistry data, organic matter (OM) chemistry data, and publicly available geospatial data were used in analysis to evaluate the variability in ERwc at 47 sites across the Yakima River basin in Washington, USA. In addition to this readme, this data package also includes a file-level metadata (FLMD) file that describes each file and a data dictionary (DD) that describes all column/row headers and variable definitions. The data package includes the data inputs, and outputs, and R scripts to reproduce all the analyses performed in the manuscript and create manuscript figures. The data package is comprised of three main folders (Code, Data, and Figures). The Code folder is comprised of four scripts and three analysis-specific subfolders that contain the R scripts to perform the analyses described in the publication and create publication figures. The Data folder is comprised of two “.csv” files and four subfolders that contain data input and output files. The Published_Data folder contains a readme that directs the user to download the appropriate files and add to this folder when using scripts. The Figures folder includes figures from the manuscript in “.pdf” and “.png” formats and a folder with intermediate figure files. This data package is associated with a GitHub repository which can be found at https://github.com/river-corridors-sfa/rcsfa-RC2-SPS-ERwc. We acknowledge the Yakama Nation as owners and caretakers of the lands where we collected some of these data. We thank the Confederated Tribes and Bands of the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview.

54 ENVIRONMENTAL SCIENCES↗

Methods for Explainable Artificial Intelligence

We explored ways of quantifying information in a neural network. This can be used to determine the right size of a network or to infer the way in which a network is processing information. The first year and a half was somewhat exploratory while the last half of the project focused on approaches that seemed to show the most promise. The introduce a new way of computing explainable artificial intelligence (XAI) saliency maps that is several orders of magnitude faster than methods with similar fidelity We call it FastCAM. The method works be combining a Class Activation Map (CAM) method such as GradCAM with a forward activation map computed with a statistic we call SMOE Scale. The addition of the forward activation maps to CAM methods seems to always improve their fidelity. At the same time, computational overhead is not increased by very much. While Gradients with SmoothGrad scores better on some fidelity measures, it is overall not as good and requires more than 1500 times to compute. We demonstrate two completed applications of FastCAM on tasks outside of the LDRD at LLNL. The source code for FastCAM is currently being implanted into Captum, the official XAI toolkit for the popular deep learning toolkit PyTorch. The LDRD currently has 13 publications released to the public. 10 of them are journal length.

97 MATHEMATICS AND COMPUTING↗

Development of Explainable, Knowledge-Guided AI Models to Enhance the E3SM Land Model Development and Uncertainty Quantification

Focal Area(s): (2)Predictive modeling using AI techniques and AI-derived model components; use of AI and other tools to design a prediction system comprising of a hierarchy of models. (3) Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge- guided AI. Science Challenge: The Energy Exascale Earth System Model (E3SM) is a fully coupled, state-of-the-science Earth system model that uses code optimized for DOE's advanced computers to address the most critical scientific questions facing our nation and society (Golaz et al., 2019). The E3SM Land model (ELM) is designed to understand how the changes in terrestrial land surfaces will interact with other Earth system components and has been used to understand hydrologic cycles, biogeophysics, and ecosystem dynamics. In spite of great successes, the ELM has several known issues that restrain rapid improvements. For example, the ELM uses equilibrium models to simulate dynamic land-climate interactions and it requires long model spin-up time to identify suitable initial conditions for transient simulations. The ELM lacks built-in uncertainty mechanisms that can improve the robustness of model predictions. The ELM is a holistic, deterministic model system with a rigid design, and in many situations, it is hard to modify the ELM system to incorporate new theory/hypothesis and new data across scales to address emerging science problems (such as predicting the impacts of water cycle extremes). In addition, The ELM is technically optimized for traditional CPU-centric computers and it cannot fully utilize the current and incoming leadership computers for model simulations and uncertainty quantification (UQ). The success of artificial intelligence (AI) has inspired scientists to use AI models to discover intrinsic features from simulation data (Chattopadhyay et al., 2020) and observational data (Reichstein et al., 2019) to gain further process understanding of Earth science problems. However, autonomous AI model training through deep learning usually requires a huge amount of annotated data. To overcome the limitations from the data and computing resources, knowledge-guided AI models are necessary where human-knowledge is ingested in model construction (Banino et al., 2018) and training process (Silver et al., 2016) for efficient learning. Herein, we present a new way that leverages the process understanding from the ELM to guide AI model development for the ELM enhancement and UQ. We hope this study can inspire further Earth and environmental system model developments and transformations.

54 ENVIRONMENTAL SCIENCES↗