Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Explainable deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Methods for Explainable Artificial Intelligence

We explored ways of quantifying information in a neural network. This can be used to determine the right size of a network or to infer the way in which a network is processing information. The first year and a half was somewhat exploratory while the last half of the project focused on approaches that seemed to show the most promise. The introduce a new way of computing explainable artificial intelligence (XAI) saliency maps that is several orders of magnitude faster than methods with similar fidelity We call it FastCAM. The method works be combining a Class Activation Map (CAM) method such as GradCAM with a forward activation map computed with a statistic we call SMOE Scale. The addition of the forward activation maps to CAM methods seems to always improve their fidelity. At the same time, computational overhead is not increased by very much. While Gradients with SmoothGrad scores better on some fidelity measures, it is overall not as good and requires more than 1500 times to compute. We demonstrate two completed applications of FastCAM on tasks outside of the LDRD at LLNL. The source code for FastCAM is currently being implanted into Captum, the official XAI toolkit for the popular deep learning toolkit PyTorch. The LDRD currently has 13 publications released to the public. 10 of them are journal length.

97 MATHEMATICS AND COMPUTING↗

Chromium-doped uranium dioxide fuels: A review

UO 2 doped with parts per million CR 2 O 3 powder is considered a potential near term accident tolerant fuel candidate. Here, the results of decades of industry and academic research into Cr-doped UO 2 are analyzed and their shortcomings are critiqued. Focusing on the incorporation mechanisms of Cr into the fuel matrix, we explore a mechanistic understanding of the characteristic properties of Cr-doped UO 2 , notably, enhanced fission gas retention attributed to enlarged grain sizes following sintering, along with marginal improvements in the thermophysical properties. The findings of recent X-ray Adsorption Near Edge Spectroscopy studies were compared and put into conversation with historic data regarding the incorporation of Cr in UO 2 . On the basis of defect mechanisms, the case is made for the substitutional incorporation of Cr governing the lattice solubility but not the enhanced U diffusivity. Instead, Cr/CR 2 O 3 redox chemistry in a well-defined oxygen potential explains the differences in the U diffusivity and O/M ratio. The primary mechanism of doping enhanced grain growth is found to be liquid assisted sintering due to a CRO (1) eutectic phase at the grain boundaries. The role of inhomogeneities in Cr concentration in UO 2 at various length scales across the materials microstructure is highlighted and connected to promising experimental and modeling work to fill in the gaps in the current understanding of Cr-doped UO 2 . The review considers both the open scientific questions and engineering applications to illustrate the deep connections between the practice and theory in the design of accident tolerant nuclear fuels. In conclusion, the review ends with an outline of future works that combine meticulous irradiation studies and high resolution experiments with next generation modeling and simulations techniques empowered by machine learning advances to accelerate the fabrication and adoption of Cr-doped UO 2 light water reactors.

Cleveland, Mack Wesley [Massachusetts Inst. of Tec↗

Machine learning identifies novel signatures of antifungal drug resistance in Saccharomycotina yeasts

Antifungal drug resistance is a major challenge in fungal infection management. Numerous genomic changes are known to contribute to acquired drug resistance in clinical isolates of specific pathogens, but whether they broadly explain natural resistance across entire lineages is unknown. We leveraged genomic, ecological, and phenotypic trait data from naturally sampled strains from nearly all known species in subphylum Saccharomycotina to examine the evolution of resistance to eight antifungal drugs. The phylogenetic distribution of drug resistance varied by drug; fluconazole resistance was widespread, while 5-fluorocytosine resistance was rare, except in Lipomycetales. A random forest algorithm trained on genomic data predicted drug-resistant yeasts with 54–75% accuracy. Fluconazole resistance was consistently predicted with the highest accuracy (75.2%). Furthermore, fluconazole resistance prediction accuracy was similar between models trained on genome-wide variation in the presence and number of InterPro protein annotations across Saccharomycotina (75.2%) and those trained on amino acid sequence alignment data of Erg11, a protein known to be involved in fluconazole resistance (74.3-74.9%). Interestingly, the top Erg11 residues for predicting fluconazole resistance across Saccharomycotina do not overlap with, are not spatially close to, and are less conserved than those previously linked to resistance in clinical isolates of Candida albicans. In silico deep mutational scanning of the C. albicans Erg11 protein reveals that amino acid variants implicated in clinical cases of resistance are almost universally destabilizing while variants in our most informative residues are energetically more neutral, explaining why the latter are much more common than the former in natural populations. Importantly, previous experimental analyses of C. albicans Erg11 have shown that amino acid variation in our most informative residues, despite having never been directly implicated in clinical cases, can directly contribute to resistance. Our results suggest that studies of natural resistance in yeast species never encountered in the clinic will yield a fuller understanding of antifungal drug resistance.

Harrison, Marie-Claire [Vanderbilt Univ., Nashvill↗

Improving Robustness of Spectrogram Classifiers with Neural Stochastic Differential Equations

Signal analysis and classification is fraught with high levels of noise and perturbation. Computer-vision-based deep learning models applied to spectrograms have proven useful in the field of signal classification and detection; however, these methods aren't designed to handle the low signal-to-noise ratios inherent within non-vision signal processing tasks. While they are powerful, they are currently not the method of choice in the inherently noisy and dynamic critical infrastructure domain, such as smart-grid sensing, anomaly detection, and non-intrusive load monitoring. Currently, these models can be brittle, which makes them susceptible to noisy input. This also means they have sub-optimal stability of explanation outputs. Experts and technicians using these models to make decisions in real world scenarios need assurance that a model is performing as it is supposed to. The classification or prediction outputs it generates should be sound and grounded, not likely to change in the presence of shifting noise landscapes. In this work, we explore the idea of Neural Stochastic Differential Equations (NSDE's) to improve the robustness of models trained to classify time series data and the effect of NSDE's on the explainability of outputs. We then test the effectiveness of these approaches by applying them to a non-intrusive load monitoring (NILM) dataset that consists of simulated harmonic signals injected into a real building.

Brogan, Joel↗

Preprocessing for Unintended Conducted Emissions Classification with ResNet

Characterization of Unintended Conducted Emissions (UCE) from electronic devices is important when diagnosing electromagnetic interference, performing nonintrusive load monitoring (NILM) of power systems, and monitoring electronic device health, among other applications. Prior work has demonstrated that UCE analysis can serve as a diagnostic tool for energy efficiency investigations and detailed load analysis. While explaining the feature selection of deep networks with certainty is often not fully comprehensive, or in other applications, quite lacking, additional tools/methods for further corroboration and confirmation can help further the understanding of the researcher. This is true especially in the subject application of the study in this paper. Often the focus of such efforts is the selected features themselves, and there is not as much understanding gained about the noise in the collected data. If selected feature and noise characteristics are known, it can be used to further shape the design of the deep network or associated preprocessing. This is additionally difficult when the available data are limited, as in the case which the authors investigated in this study. Here, the authors present a novel work (which is a proposed complementary portion of the overall solution to the deep network classification explainability problem for this application) by applying a systematic progression of preprocessing and a deep neural network (ResNet architecture) to classify UCE data obtained via current transformers. By using a methodical application of preprocessing techniques prior to a deep classifier, hypotheses can be produced concerning what features the deep network deems important relative to what it perceives as noise. For instance, it is hypothesized in this particular study as a result of execution of the proposed method and periodic inspection of the classifier output that the UCE spectral features are relatively close to each other or to the interferers, as systematically reducing the beta parameter of the Kaiser window produced progressively better classification performance, but only to a point, as going below the Beta of eight produced decreased classifier performance, as well as the hypothesis that further spectral feature resolution was not as important to the classifier as rejection of the leakage from a spectrally distant interference. This can be very important in unpredictable low-FNR applications, where knowing the difference between features and noise is difficult. As a side-benefit, much was learned regarding the best preprocessing to use with the selected deep network for the UCE collected from these low power consumer devices obtained via current transformers. Baseline rectangular windowed FFT preprocessing provided a 62% classification increase versus using raw samples. After performing a more optimal preprocessing, more than 90% classification accuracy was achieved across 18 low-power consumer devices for scenarios in which the in-band features-to-noise ratio (FNR) was very poor.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Why Dissolving Salt in Water Decreases Its Dielectric Permittivity

The dielectric permittivity of salt water decreases on dissolving more salt. For nearly a century, this phenomenon has been explained by invoking saturation in the dielectric response of the solvent water molecules. Herein, we employ an advanced deep neural network (DNN), built using data from density functional theory, to study the dielectric permittivity of sodium chloride solutions. Notably, the decrease in the dielectric permittivity as a function of concentration, computed using the DNN approach, agrees well with experiments. Detailed analysis of the computations reveals that the dominant effect, caused by the intrusion of ionic hydration shells into the solvent hydrogen-bond network, is the disruption of dipolar correlations among water molecules. Accordingly, the observed decrease in the dielectric permittivity is mostly due to increasing suppression of the collective response of solvent waters.

74 ATOMIC AND MOLECULAR PHYSICS↗

The quenching of galaxies, bulges, and disks since cosmic noon

Here, we present an analysis of the quenching of star formation in galaxies, bulges, and disks throughout the bulk of cosmic history, from z = 2 – 0. We utilise observations from the Sloan Digital Sky Survey and the Mapping Nearby Galaxies at Apache Point Observatory survey at low redshifts. We complement these data with observations from the Cosmic Assembly Near-Infrared Deep Extragalactic Legacy Survey at high redshifts. Additionally, we compare the observations to detailed predictions from the LGalaxies semi-analytic model. To analyse the data, we developed a machine learning approach utilising a Random Forest classifier. We first demonstrate that this technique is extremely effective at extracting causal insight from highly complex and inter-correlated model data, before applying it to various observational surveys. Our primary observational results are as follows: at all redshifts studied in this work, we find bulge mass to be the most predictive parameter of quenching, out of the photometric parameter set (incorporating bulge mass, disk mass, total stellar mass, and B/T structure). Moreover, we also find bulge mass to be the most predictive parameter of quenching in both bulge and disk structures, treated separately. Hence, intrinsic galaxy quenching must be due to a stable mechanism operating over cosmic time, and the same quenching mechanism must be effective in both bulge and disk regions. Despite the success of bulge mass in predicting quenching, we find that central velocity dispersion is even more predictive (when available in spectroscopic data sets). In comparison to the LGalaxies model, we find that all of these observational results may be consistently explained through quenching via preventative ‘radio-mode’ active galactic nucleus feedback. Furthermore, many alternative quenching mechanisms (including virial shocks, supernova feedback, and morphological stabilisation) are found to be inconsistent with our observational results and those from the literature.

79 ASTRONOMY AND ASTROPHYSICS↗

Characterization and identification of HPC applications at leadership computing facility

High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. In order to improve our understanding of HPC applications, user demands and resource usage characteristics, we perform correlative analysis of various logs for different subsystems of a leadership supercomputer. This analysis reveals surprising, sometimes counter-intuitive patterns, which, in some cases, conflicts with existing assumptions, and have important implications for future system designs as well as supercomputer operations. For example, our analysis shows that while the applications spend significant time on MPI, most applications spend very little time on file I/O. Combined analysis of hardware event logs and task failure logs show that the probability of a hardware FATAL event causing task failure is low. Combined analysis of control system logs and file I/O logs reveals that pure POSIX I/O is used more widely than higher level parallel I/O. Based on holistic insights of the application gained through combined and co-analysis of multiple logs from different perspectives and general intuition, we engineer features to "fingerprint" HPC applications. We use t-SNE (a machine learning technique for dimensionality reduction) to validate the explainability of our features and finally train machine learning models to identify HPC applications or group those with similar characteristic. To the best of our knowledge, this is the first work that combines logs on file I/O, computing, and inter-node communication for insightful analysis of HPC applications in production.

Liu, Zhengchun↗

Fraction of broad absorption line quasars in different radio morphologies

ABSTRACT In this study, we investigated the orientation model of Broad Absorption Line (BAL) quasars using a sample of sources that are common in Sloan Digital Sky Survey (SDSS) Data Release (DR)-16 quasar catalogue and Very Large Array (VLA)-Faint Images of the Radio Sky at Twenty Centimeters (FIRST) survey. Using the radio cut-out images from the FIRST survey, we first designed a deep-learning model using convolutional neural networks (CNN) to classify the quasar radio morphologies into the core-only, young jet, single lobe, or triples. These radio morphologies are further sub-classified into core-dominated and lobe-dominated sources. The CNN models can classify the sources with a high precision of >98 ${{\ \rm per\ cent}}$ for all the morphological sub-classes. The average BAL fraction in the resolved core, core-dominated, and lobe-dominated quasars are consistent with the BAL fraction inferred from radio and infrared surveys. We also present the distribution of BAL quasars as a function of quasar orientation by using the radio core-dominance as an orientation indicator. A similar analysis is performed for HiBALs, LoBALs, and FeLoBALs. All the radio morphological sub-classes and BAL sub-classes show an increase in BAL fraction at high orientation angles of the jets with respect to the line of sight. Our analysis suggests that BAL quasars are more likely to be found in viewing angles close to the equatorial plane of the quasar. However, a pure orientation model is inadequate, and a combination of orientation and evolution is probably the best way to explain the complete BAL phenomena.

79 ASTRONOMY AND ASTROPHYSICS↗

Towards Geospatial Knowledge Graph Infused Neuro-Symbolic AI for Remote Sensing Scene Understanding

Deep learning has proven its effectiveness in numerous tasks for remote sensing scene understanding. However there is an increasing interest to explore fusion of domain-specific background information to the deep neural network to further improve its performance. Remote sensing researchers are also working towards developing models that generalize and adapt to multiple applications. Generalization challenges coupled with the scarcity of large corpora of high-quality noise-free labelled data, have together fueled an interest for leveraging background information. Knowledge graphs serve as excellent choice to represent domain-specific information in a structured, standardized and extensible manner. Integrating symbolic knowledge representations in the form of Knowledge Graph Embedding (KGE) to perform neuro-symbolic reasoning is an emerging research direction promising significant impacts. This vision paper seeks to position ideas and provoke early thoughts toward advancing neuro-symbolic artificial intelligence in the context of geospatial challenges. Specifically, it conceptualizes and elaborates on an architecture for infusing geospatial knowledge from knowledge graph in a deep neural network pipeline. As guiding case studies - land-use land-cover classification, object detection and instance segmentation can benefit from infusing spatio-contextual information with remote sensing imagery. The discussion further reflects on and articulates the challenges and explainable AI opportunities anticipated when scaling and maintaining large-scale geospatial knowledge graphs.

Potnis, Abhishek↗

Explainable AI classification for parton density theory

Quantitatively connecting properties of parton distribution functions (PDFs, or parton densities) to the theoretical assumptions made within the QCD analyses which produce them has been a longstanding problem in HEP phenomenology. To confront this challenge, we introduce an ML-based explainability framework, XAI4PDF, to classify PDFs by parton flavor or underlying theoretical model using ResNet-like neural networks (NNs). By leveraging the differentiable nature of ResNet models, this approach deploys guided backpropagation to dissect relevant features of fitted PDFs, identifying x-dependent signatures of PDFs important to the ML model classifications. By applying our framework, we are able to sort PDFs according to the analysis which produced them while constructing quantitative, human-readable maps locating the x regions most affected by the internal theory assumptions going into each analysis. This technique expands the toolkit available to PDF analysis and adjacent particle phenomenology while pointing to promising generalizations.

Artificial Intelligence↗

Explaining System-Level Prognostics with Established Machine Learning Methods

System-level prognostics is crucial for ensuring reliability and enabling predictive maintenance in complex systems with interconnected components. This study presents a framework that integrates data-driven methods to predict the remaining useful life (RUL) of a subsystem under multiple and concurrent faults within a nuclear power plant system with explainable artificial intelligence (XAI). A nuclear power plant (NPP) operation was simulated to model the degradation behavior of NPP components, and four machine learning models—Gradient Boosting Regressor (GBR), Support Vector Regressor (SVR), Fully Connected Neural Network (FCNN), and Long Short-Term Memory (LSTM)—were evaluated for prognostics with a novel system RUL parameter. The LSTM model demonstrated potential superior repeatability, while SHAP (SHapley Additive exPlanations) for explainability provided consistent and trustworthy global explanations. In contrast, LIME (Local Interpretable Model-agnostic Explanations) offered localized interpretability but showed reduced stability for sequential data. Key findings include the interplay between component-level degradation and system-wide performance, with LSTM effectively capturing these dynamics through sequence-level predictions. The XAI techniques enhanced transparency by identifying critical features influencing model predictions and aligning with domain knowledge. Furthermore, this framework has significant implications for improving trust and understanding in predictive maintenance, particularly in safety-critical industries like nuclear energy.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Phase Diagram of a Deep Potential Water Model

Using the Deep Potential methodology, we construct a model that reproduces accurately the potential energy surface of the SCAN approximation of density functional theory for water, from low temperature and pressure to about 2400 K and 50 GPa, excluding the vapor stability region. The computational efficiency of the model makes it possible to predict its phase diagram using molecular dynamics. Satisfactory overall agreement with experimental results is obtained. Here, the fluid phases, molecular and ionic, and all the stable ice polymorphs, ordered and disordered, are predicted correctly, with the exception of ice III and XV that are stable in experiments, but metastable in the model. The evolution of the atomic dynamics upon heating, as ice VII transforms first into ice VII" and then into an ionic fluid, reveals that molecular dissociation and breaking of the ice rules coexist with strong covalent fluctuations, explaining why only partial ionization was inferred in experiments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

The suitability of differentiable, physics-informed machine learning hydrologic models for ungauged regions and climate change impact assessment

As a genre of physics-informed machine learning, differentiable process-based hydrologic models (abbreviated as δ or delta models) with regionalized deep-network-based parameterization pipelines were recently shown to provide daily streamflow prediction performance closely approaching that of state-of-the-art long short-term memory (LSTM) deep networks. Meanwhile, δ models provide a full suite of diagnostic physical variables and guaranteed mass conservation. Here, we ran experiments to test (1) their ability to extrapolate to regions far from streamflow gauges and (2) their ability to make credible predictions of long-term (decadal-scale) change trends. We evaluated the models based on daily hydrograph metrics (Nash–Sutcliffe model efficiency coefficient, etc.) and predicted decadal streamflow trends. For prediction in ungauged basins (PUB; randomly sampled ungauged basins representing spatial interpolation), δ models either approached or surpassed the performance of LSTM in daily hydrograph metrics, depending on the meteorological forcing data used. They presented a comparable trend performance to LSTM for annual mean flow and high flow but worse trends for low flow. For prediction in ungauged regions (PUR; regional holdout test representing spatial extrapolation in a highly data-sparse scenario), δ models surpassed LSTM in daily hydrograph metrics, and their advantages in mean and high flow trends became prominent. In addition, an untrained variable, evapotranspiration, retained good seasonality even for extrapolated cases. The δ models' deep-network-based parameterization pipeline produced parameter fields that maintain remarkably stable spatial patterns even in highly data-scarce scenarios, which explains their robustness. Combined with their interpretability and ability to assimilate multi-source observations, the δ models are strong candidates for regional and global-scale hydrologic simulations and climate change impact assessment.

54 ENVIRONMENTAL SCIENCES↗

Understanding the Physics Representation of Deep Learning Models in Environmental Applications

Deep learning (DL) models have been popular in earth and environmental modeling and analysis, which exhibit huge potential in capturing and reconstructing the non-linearity of relevant environmental processes. They are extensively used as analytical tools or emulators for multiple domains (atmosphere, land surface, ocean, and biogeochemistry). Despite their success, their internal working mechanism remains largely unknown. Such a lack of knowledge hinders the identification of physically consistent models that are fully adaptive to non-stationary climate, as well as the development of physics-informed machine learning such as physics-informed neural network (PINN). To establish preliminary knowledge and framework of such physics representation evaluation, this project focuses on an improved understanding of DL models in the environmental applications. DL models are increasingly applied to environmental modeling and prediction. However, they have been evaluated mostly from a performance perspective, and there is a gap in understanding how they represent the known physics internally. Such knowledge is especially critical when applying DL models under climate change conditions, where new inputs are likely outside the ranges of the training datasets. In this project, we reveal how the known physical processes are represented within DL models from both statistical and mechanistic perspectives. Leveraging the traditional model evaluations that focus more on the accuracies of predictions, we establish a framework that examines both the accuracy and physics representation of DL models. This analysis framework can identify DL models that make the correct predictions based on correct physics, thus enhancing the existing explainable artificial intelligence (explainable-AI) portfolio. It lays a foundation for developing novel metrics to evaluate the emerging DL models in environmental applications. This knowledge also informs the development of physics-informed DL models by revealing the direct connections between the known physical processes and specific model components or structures.

54 ENVIRONMENTAL SCIENCES↗

SigTime: Learning and Visually Explaining Time Series Signatures

Understanding and distinguishing temporal patterns in time series data is essential for scientific discovery and decision-making. For example, in biomedical research, uncovering meaningful patterns in physiological signals can improve diagnosis, risk assessment, and patient outcomes. However, existing methods for time series pattern discovery face major challenges, including high computational complexity, limited interpretability, and difficulty in capturing meaningful temporal structures. Here, to address these gaps, we introduce a novel learning framework that jointly trains two Transformer models using complementary time series representations: shapelet-based representations to capture localized temporal structures and traditional feature engineering to encode statistical properties. The learned shapelets serve as interpretable signatures that differentiate time series across classification labels. Additionally, we develop a visual analytics system—SigTime—with coordinated views to facilitate exploration of time series signatures from multiple perspectives, aiding in useful insights generation. We quantitatively evaluate our learning framework on eight publicly available datasets and one proprietary clinical dataset. Additionally, we demonstrate the effectiveness of our system through two usage scenarios along with the domain experts: one involving public ECG data and the other focused on preterm labor analysis.

97 MATHEMATICS AND COMPUTING↗

Development of Explainable, Knowledge-Guided AI Models to Enhance the E3SM Land Model Development and Uncertainty Quantification

Focal Area(s): (2)Predictive modeling using AI techniques and AI-derived model components; use of AI and other tools to design a prediction system comprising of a hierarchy of models. (3) Insight gleaned from complex data (both observed and simulated) using AI, big data analytics, and other advanced methods, including explainable AI and physics- or knowledge- guided AI. Science Challenge: The Energy Exascale Earth System Model (E3SM) is a fully coupled, state-of-the-science Earth system model that uses code optimized for DOE's advanced computers to address the most critical scientific questions facing our nation and society (Golaz et al., 2019). The E3SM Land model (ELM) is designed to understand how the changes in terrestrial land surfaces will interact with other Earth system components and has been used to understand hydrologic cycles, biogeophysics, and ecosystem dynamics. In spite of great successes, the ELM has several known issues that restrain rapid improvements. For example, the ELM uses equilibrium models to simulate dynamic land-climate interactions and it requires long model spin-up time to identify suitable initial conditions for transient simulations. The ELM lacks built-in uncertainty mechanisms that can improve the robustness of model predictions. The ELM is a holistic, deterministic model system with a rigid design, and in many situations, it is hard to modify the ELM system to incorporate new theory/hypothesis and new data across scales to address emerging science problems (such as predicting the impacts of water cycle extremes). In addition, The ELM is technically optimized for traditional CPU-centric computers and it cannot fully utilize the current and incoming leadership computers for model simulations and uncertainty quantification (UQ). The success of artificial intelligence (AI) has inspired scientists to use AI models to discover intrinsic features from simulation data (Chattopadhyay et al., 2020) and observational data (Reichstein et al., 2019) to gain further process understanding of Earth science problems. However, autonomous AI model training through deep learning usually requires a huge amount of annotated data. To overcome the limitations from the data and computing resources, knowledge-guided AI models are necessary where human-knowledge is ingested in model construction (Banino et al., 2018) and training process (Silver et al., 2016) for efficient learning. Herein, we present a new way that leverages the process understanding from the ELM to guide AI model development for the ELM enhancement and UQ. We hope this study can inspire further Earth and environmental system model developments and transformations.

54 ENVIRONMENTAL SCIENCES↗