Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Explainable deep learning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Exploring data-driven modeling of boundary layer transition

Prediction of laminar-turbulent transition in boundary layer flows is an important component of predicting the aerodynamic performance of a number of aerospace configurations. According to the CFD Vision 2030 [1], transition modeling represents acriticalarea in CFD simulation capability that will remain a pacing item for the foreseeable future. The fact thattransition can take placevia either one of a myriad possible paths adds to the challenges inreliable transition predictions, despite a limited knowledge of the relevant input parameters. In the low disturbance environments typical of flight applications, transition is often initiated by small amplitude disturbances in the form of linear instability waves of the laminar boundary layer. These disturbances amplify linearly at first and eventually undergo a sequence of nonlinear interactions that result in transition to turbulence. Because the nonlinear phase is rather rapid, the amplification of boundary layer instabilities is governed by the linearstability theory over a majority of the distance leading up to the onset of transition. Semi-empirical transition correlations based on the linear stability theory have been successful in explaining the observed trends in transition location within a broad class of flows. However, the application of stability theory is highly non-robust and often requires a significant domain expertise. Recent work at the NASA Langley Research Center has beenaimed at bridging the gap between physics based transition analyses such as those based on linear stability theory and practical applications that require transition prediction by users that may not be well versed in transition physics. The applications of deep learning have been at the center of these efforts. This presentation will focus on the progress achieved thus far, highlighting the applications of neural networks to selectedtransition scenarios across a range of Mach numbers and flow configuration, as well as the lessons learnedand remaining challengeswithrespect to the selection of training data and neural networks architectures, hyperparameter tuning, and the physical insights distilled from the otherwise black-box models.

M. R. Malik↗

Data-driven assessment of magnetic charged particle confinement parameter scaling in magnetized liner inertial fusion experiments on Z

In magneto-inertial fusion, the ratio of the characteristic fuel length perpendicular to the applied magnetic field R to the α-particle Larmor radius $ϱ_α$ is a critical parameter setting the scale of electron thermal-conduction loss and charged burn-product confinement. Here, using a previously developed deep-learning-based Bayesian inference tool, we obtain the magnetic-field fuel-radius product BR ∝ R / $ϱ_α$ from an ensemble of 16 magnetized liner inertial fusion (MagLIF) experiments. Observations of the trends in BR are consistent with relative trade-offs between compression and flux loss as well as the impact of mix from 1D resistive radiation magneto-hydrodynamics simulations in all but two experiments, for which 3D effects are hypothesized to play a significant role. Finally, we explain the relationship between BR and the generalized Lawson parameter χ. Our results indicate the ability to improve performance in MagLIF through careful tuning of experimental inputs, while also highlighting key risks from mix and 3D effects that must be mitigated in scaling MagLIF to higher currents with a next-generation driver.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Signatures of a liquid–liquid transition in an ab initio deep neural network model for water

Significance Water is central across much of the physical and biological sciences and exhibits physical properties that are qualitatively distinct from those of most other liquids. Understanding the microscopic basis of water’s peculiar properties remains an active area of research. One intriguing hypothesis is that liquid water can separate into metastable high- and low-density liquid phases at low temperatures and high pressures, and the existence of this liquid–liquid transition could explain many of water’s anomalous properties. We used state-of-the-art approaches in computational quantum chemistry, statistical mechanics, and machine learning and obtained evidence consistent with a liquid–liquid transition, supporting the argument for the existence of this phenomenon in real water.

36 MATERIALS SCIENCE↗

Discovering a reaction–diffusion model for Alzheimer’s disease by combining PINNs with symbolic regression

Misfolded tau proteins play a critical role in the progression and pathology of Alzheimer's disease. Recent studies suggest that the spatio-temporal pattern of misfolded tau follows a reaction-diffusion type equation. However, the precise mathematical model and parameters that characterize the progression of misfolded protein across the brain remain incompletely understood. Here, we use deep learning and artificial intelligence to discover a mathematical model for the progression of Alzheimer's disease using longitudinal tau positron emission tomography from the Alzheimer's Disease Neuroimaging Initiative database. Specifically, we integrate physics informed neural networks (PINNs) and symbolic regression to discover a reaction-diffusion type partial differential equation for tau protein misfolding and spreading. First, we demonstrate the potential of our model and parameter discovery on synthetic data. Then, we apply our method to discover the best model and parameters to explain tau imaging data from 46 individuals who are likely to develop Alzheimer's disease and 30 healthy controls. Our symbolic regression discovers different misfolding models f(c) for two groups, with a faster misfolding for the Alzheimer's group, f(c) = 0.23c 3 – 1.34c 2 + 1.11c, than for the healthy control group, f(c) = –c 3 + 0.62c 2 + 0.39c. Our results suggest that PINNs, supplemented by symbolic regression, can discover a reaction-diffusion type model to explain misfolded tau protein concentrations in Alzheimer's disease. Furthermore, we expect our study to be the starting point for a more holistic analysis to provide image-based technologies for early diagnosis, and ideally early treatment of neurodegeneration in Alzheimer's disease and possibly other misfolding-protein based neurodegenerative disorders.

60 APPLIED LIFE SCIENCES↗

Using Temporal Deep Learning Models to Estimate Daily Snow Water Equivalent Over the Rocky Mountains

Abstract In this study we construct and compare three different deep learning (DL) models for estimating daily snow water equivalent (SWE) from high‐resolution gridded meteorological fields over the Rocky Mountain region. To train the DL models, Snow Telemetry (SNOTEL) station‐based SWE observations are used as the prediction target. All DL models produce higher median Nash‐Sutcliffe Efficiency (NSE) values than a conceptual SWE model and interpolated gridded data sets, although mean squared errors also tend to be higher. Sensitivity of the SWE prediction to the model's input variables is analyzed using an explainable artificial intelligence (XAI) method, yielding insight into the physical relationships learned by the models. This method reveals the dominant role precipitation and temperature play in snowpack dynamics. In applying our models to estimate SWE throughout the Rocky Mountains, an extrapolation problem arises since the statistical properties of SWE (e.g., annual maximum) and geographical properties of individual grid points (e.g., elevation) differ from the training data. This problem is solved by normalizing the SWE with its historical maximum value to alleviate extrapolation for all tested DL models. Our work shows that the DL models are promising tools for estimating SWE, and sufficiently capture relevant physical relationships to make them useful for spatial and temporal extrapolation of SWE values.

54 ENVIRONMENTAL SCIENCES↗

Appendices for Geothermal Exploration Artificial Intelligence Report

The Geothermal Exploration Artificial Intelligence looks to use machine learning to spot geothermal identifiers from land maps. This is done to remotely detect geothermal sites for the purpose of energy uses. Such uses include enhanced geothermal system (EGS) applications, especially regarding finding locations for viable EGS sites. This submission includes the appendices and reports formerly attached to the Geothermal Exploration Artificial Intelligence Quarterly and Final Reports. The appendices below include methodologies, results, and some data regarding what was used to train the Geothermal Exploration AI. The methodology reports explain how specific anomaly detection modes were selected for use with the Geo Exploration AI. This also includes how the detection mode is useful for finding geothermal sites. Some methodology reports also include small amounts of code. Results from these reports explain the accuracy of methods used for the selected sites (Brady Desert Peak and Salton Sea). Data from these detection modes can be found in some of the reports, such as the Mineral Markers Maps, but most of the raw data is included the DOE Database which includes Brady, Desert Peak, and Salton Sea Geothermal Sites.

15 GEOTHERMAL ENERGY↗

Knowledge-guided learning with curated prior genetic biomarkers for robust model interpretation

Abstract Motivation Knowledge-guided learning offers effective and robust model training strategies in data-scarce settings by incorporating established domain knowledge, thereby enhancing generalization, robustness, and interpretability. By contrast, conventional deep learning approaches rely purely on data-driven learning, which can limit robust model interpretability, particularly in high-dimensional settings with limited size samples. In computational biology, knowledge-guided learning has primarily leveraged network- and structural-based knowledge, leading to biologically interpretable representations and enhanced predictive performance compared to conventional approaches. However, curated biomarkers, one of the most accessible forms of biological knowledge, remain largely unexplored within knowledge-guided paradigms. Results In this study, we propose a model-agnostic training paradigm, Biomarker-driven Explainable Prior-guided Learning (BioExPL), that can be applied to any neural networks that incorporates curated prior knowledge. BioExPL enforces neural networks to reflect curated biomarker priors in their latent representations through a novel knowledge-alignment loss. BioExPL consistently demonstrated significantly improved predictive performance and enhanced model interpretability with minimized computational overhead in simulation studies and intensive experiments on multiple cancer datasets. BioExPL not only integrates prior curated knowledge into the model but also accurately identifies unknown associated signals additionally. BioExPL is model-agnostic and domain-independent, enabling its integration into diverse neural network architectures. Availability and implementation The open-source is publicly available at: https://github.com/datax-lab/BioExPL.

Baek, Beomsu [Department of Computer Science, Univ↗

Yet Another Discriminant Analysis (YADA): A Probabilistic Model for Machine Learning Applications

This paper presents a probabilistic model for various machine learning (ML) applications. While deep learning (DL) has produced state-of-the-art results in many domains, DL models are complex and over-parameterized, which leads to high uncertainty about what the model has learned, as well as its decision process. Further, DL models are not probabilistic, making reasoning about their output challenging. In contrast, the proposed model, referred to as Yet Another Discriminate Analysis(YADA), is less complex than other methods, is based on a mathematically rigorous foundation, and can be utilized for a wide variety of ML tasks including classification, explainability, and uncertainty quantification. YADA is thus competitive in most cases with many state-of-the-art DL models. Ideally, a probabilistic model would represent the full joint probability distribution of its features, but doing so is often computationally expensive and intractable. Hence, many probabilistic models assume that the features are either normally distributed, mutually independent, or both, which can severely limit their performance. YADA is an intermediate model that (1) captures the marginal distributions of each variable and the pairwise correlations between variables and (2) explicitly maps features to the space of multivariate Gaussian variables. Numerous mathematical properties of the YADA model can be derived, thereby improving the theoretic underpinnings of ML. Validation of the model can be statistically verified on new or held-out data using native properties of YADA. However, there are some engineering and practical challenges that we enumerate to make YADA more useful.

97 MATHEMATICS AND COMPUTING↗

Deep learning Hamiltonians from disordered image data in quantum materials

The capabilities of image probe experiments are rapidly expanding, providing new information about quantum materials on unprecedented length- and timescales. Many such materials feature inhomogeneous electronic properties with intricate pattern formation on the observable surface. This rich spatial structure contains information about interactions, dimensionality, and disorder—a spatial encoding of the Hamiltonian driving the pattern formation. Image recognition techniques from machine learning are an excellent tool for interpreting information encoded in the spatial relationships in such images. Here, we develop a deep learning framework for using the rich information available in these spatial correlations in order to discover the underlying Hamiltonian driving the patterns. We first vet the method on a known case, scanning near-field optical microscopy on a thin film of V⁢O 2 . We then apply our trained convolutional neural network architecture to new optical microscope images of a different V⁢O 2 film as it goes through the metal-insulator transition. We find that a two-dimensional Hamiltonian with both interactions and random field disorder is required to explain the intricate, fractal intertwining of metal and insulator domains during the transition. This detailed knowledge about the underlying Hamiltonian paves the way for using the model to control the pattern formation via, e.g., tailored hysteresis protocols. Finally, we also introduce a distribution-based confidence measure on the results of a multilabel classifier, which does not rely on adversarial training. In addition, we propose a machine-learning-based criterion for diagnosing a physical system's proximity to criticality.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Extracting the Galactic Center excess’ source-count distribution with neural nets

The two leading hypotheses for the Galactic Center excess (GCE) in the Fermi data are an unresolved population of faint millisecond pulsars (MSPs) and dark-matter (DM) annihilation. The dichotomy between these explanations is typically reflected by modeling them as two separate emission components. However, point sources (PSs) such as MSPs become statistically degenerate with smooth Poisson emission in the ultrafaint limit (formally where each source is expected to contribute much less than one photon on average), leading to an ambiguity that can render questions such as whether the emission is PS-like or Poissonian in nature ill defined. We present a conceptually new approach that describes the PS and Poisson emission in a unified manner and only afterwards derives constraints on the Poissonian component from the so obtained results. For the implementation of this approach, we leverage deep learning techniques, centered around a neural network-based method for histogram regression that expresses uncertainties in terms of quantiles. We demonstrate that our method is robust against a number of systematics that have plagued previous approaches, in particular DM/PS misattribution. In the Fermi data, we find a faint GCE described by a median source-count distribution (SCD) peaked at a flux of ~ 4 x 10 -11 counts cm -2 s -1 (corresponding to ~ 3-4 expected counts per PS), which would require N ~ $\mathscr{O}$(10 4 ) sources to explain the entire excess (median value N = 29,300 across the sky). Although faint, this SCD allows us to derive the constraint η ρ ≤ 66% for the Poissonian fraction of the GCE flux η ρ at 95% confidence, suggesting that a substantial amount of the GCE flux is due to PSs.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Quantum Neural Networks: Issues, Training, and Applications

Our work in the field aims at explaining the limitations and expressive power of Quantum Machine Learning models, as well as finding feasible training algorithms that could be implemented in near-term Quantum Computers. The promise of Quantum Machine Learning is that by incorporating quantum effects, such as entanglement, into machine learning models researchers can improve model performance and understand more complex datasets. This pledge is particularly pronounced in the design of Quantum neural networks (QNNs), a promising framework for creating quantum algorithms, that promise to outperform classical models by combining the speedups of quantum computation with the widespread successes of deep learning. We show that applying this approach alone to quantum deep learning is problematic given that an excess of entanglement between the hidden and visible layers can destroy the predictive power of our QNN models. We address the barren plateau problem by suggesting the use of a generative, unbounded, nonlinear loss function with simple gradients. The loss function quantifies how much the quantum states generated by the QNNs differ from the data and the goal during training is to minimize it. Finally, we showcase how to use generative training to construct a "classical-quantum" neural network to accurately interpolate between the ground states of a Molecular Hamiltonian, a central question in Quantum Chemistry.

97 MATHEMATICS AND COMPUTING↗

Runway Sign Classifier: A DAL C Certifiable Machine Learning System

In recent years, the remarkable progress of Machine Learning (ML) technologies within the domain of Artificial Intelligence (AI) systems has presented unprecedented opportunities for the aviation industry, paving the way for further advancements in automation, including the potential for single pilot or fully autonomous operation of large commercial airplanes. However, ML technology faces major incompatibilities with existing airborne certification standards, such as ML model traceability and explainability issues or the inadequacy of traditional coverage metrics. Certification of ML-based airborne systems using current standards is problematic due to these challenges. This paper presents a case study of an airborne system utilizing a Deep Neural Network (DNN) for airport sign detection and classification. Building upon our previous work, which demonstrates compliance with Design Assurance Level (DAL) ”D”, we upgrade the system to meet the more stringent requirements of Design Assurance Level ”C”. To achieve DAL C, we employ an established architectural mitigation technique involving two redundant and dissimilar Deep Neural Networks. The application of novel ML-specific data management techniques further enhances this approach. This work is intended to illustrate how the certification challenges of ML-based systems can be addressed for medium criticality airborne applications.

Flight Software↗

Development of a Full-Scale Connected U-Net for Reflectivity Inpainting in Spaceborne Radar Blind Zones

CloudSat’s Cloud Profiling Radar is a valuable tool for remotely monitoring high-latitude snowfall, but its ability to observe hydrometeor activity near the Earth’s surface is limited by a radar blind zone caused by ground clutter contamination. This study presents the development of a deeply supervised U-Net-style convolutional neural network to predict cold season reflectivity profiles within the blind zone at two Arctic locations. The network learns to predict the presence and intensity of near-surface hydrometeors by coupling latent features encoded in blind zone-aloft clouds with additional context from collocated atmospheric state variables (i.e., temperature, specific humidity, and wind speed). Results show that the U-Net predictions outperform traditional linear extrapolation methods, with low mean absolute error, a 38% higher Sørensen–Dice coefficient, and vertical reflectivity distributions 60% closer to observed values. The U-Net is also able to detect the presence of near-surface cloud with a critical success index (CSI) of 72% and cases of shallow cumuliform snowfall and virga with 18% higher CSI values compared to linear methods. An explainability analysis shows that reflectivity information throughout the scene, especially at cloud edges and at the 1.2-km blind zone threshold, along with atmospheric state variables near the tropopause, are the most significant contributors to model skill. This surface-trained generative inpainting technique has the potential to enhance current and future remote sensing precipitation missions by providing a better understanding of the nonlinear relationship between blind zone reflectivity values and the surrounding atmospheric state.

54 ENVIRONMENTAL SCIENCES↗

Virtual Assistant for First Responders Using Natural Language Understanding and Optical Character Recognition

Commercial deep learning capabilities are available for many applications such as computer vision processing and intelligent chat bots. The Google Cloud Platform product Google Dialogflow provides lifelike conversational artificial intelligence (AI) using machine learning (ML) to generate natural conversations between computers and humans. This ML utilizes natural language understanding (NLU) to recognize a user’s intent and extracts key information into a form of entities. We have developed a user-friendly application through understanding the hazardous material database, first aid safety guidelines and observing the process of first responders who access this information in the field. We created the Trusted and Explainable Artificial Intelligence for Saving Lives (TruePAL) virtual assistant using Dialogflow1 and TensorFlow2 paired with EasyOCR.3 The chatbot supports first responders by providing voice interaction which helps limit additional steps such as browsing through multiple categories when searching for information. Using feedback from our field interviews, the voice interface has been developed to enable the first responder to focus on the immediate emergency. With less distractions, the first responder is able to engage the incident more effectively. The partial hands-free TruePAL chatbot assistant improves the accessibility to the correct guidance by an average of 1.9 seconds compared to the widely used application, NIH WISER, which requires full attention to operate. We combined this intelligent chatbot with a separate visual processing capability to produce hazardous signage analysis and generate the proper guidance for first responders. With the evolving functionality of AI tools, the use of virtual assistants in first responder technology will be an advancement, benefiting the safety of both first responders and civilians.

Chow, Edward↗

Methodology for interpretable reinforcement learning model for HVAC energy control

Deep reinforcement learning (DRL) approaches have been used in various application areas to improve efficiency, optimization, or automation. However, very little is known about how the DRL algorithms make decisions and what features affect their performance. Using a case study of a DRL based Heating, Ventilation and Air Conditioning (HVAC) optimization methodology, we demonstrate how we can address these challenges by applying interpretability tools and systematically exploring the model inputs for better understanding the DRL behaviour and decision making process. We developed a methodology for interpretable reinforcement learning and evaluated our approach in real-world house located in Knoxville, TN. Our findings explain the reasoning behind DRL-based optimization decisions under different circumstances which has been discussed and confirmed by the experts in the field.

Kotevska, Olivera↗

Surfactant-Specific AI-Driven Molecular Design: Integrating Generative Models, Predictive Modeling, and Reinforcement Learning for Tailored Surfactant Synthesis

Molecular design is a critical aspect of various scientific and industrial fields, where the properties of molecules hold significant importance. In this study, a 3-fold methodology design is presented that leverages the power of generative artificial intelligence (AI), predictive modeling, and reinforcement learning to create tailored molecules with desired properties. This model synergistically combines deep learning techniques with Self-Referencing Embedded Strings (SELFIES) molecular representation to build a generative model that generates valid molecules and a graphical neural network model that accurately forecasts molecular properties. The Variational Autoencoder (VAE) coupled with reinforcement learning helps refine molecule generation based on targeted attributes. Data from an experimental study involving surfactants were used to test the framework. A validation of the structural integrity of the molecules generated was conducted, and Tanimoto similarities were used to quantify the similarity and diversity between the original and generated molecular structures. Also, saliency maps for the generated surfactants were produced to identify the features explaining the property values. Lastly, molecular dynamics simulations were used to validate the stability of the generated molecules. The results showed that the proposed framework can effectively produce valid molecules within the set property threshold value.

36 MATERIALS SCIENCE↗

Explainable AI (XAI)-driven vibration sensing scheme for surface quality monitoring in a smart surface grinding process

Local Interpretable and Model-agnostic Explanation (LIME), an explainable artificial intelligence (XAI) approach is adapted to identify the globally important time-frequency bands for predicting average surface roughness (Ra) in a smart grinding process. The smart grinding setup consisted of a Supertech CNC precision surface grinding machine, instrumented with a Dytran piezoelectric accelerometer attached to the tailstock quill along the tangential direction (Y-axis). For every grinding pass, vibration signatures were captured, and the ground truth surface roughness values were recorded using a Mahr Marsurf M300C portable surface roughness profilometer. The roughness values ranged from 0.06 to 0.14 microns over the complete set of experiments. Time-frequency domain spectrogram frames were extracted for each of the vibration signals collected during the grinding process. Convolutional Neural Networks (CNNs) were modeled to predict the surface roughness based on these spectrogram frames and their image augmentations. The best CNN model was able to predict the roughness values with an overall R2-score of 0.95, training R2-score of 0.99, and testing R2-score of 0.81 with only 80 sets of vibration signals corresponding to 4 experiments with 20 trials each. Although the data size is not large enough to guarantee such performance metrics in real-world scenarios, one can extract statistically consistent explanations underlying the relationships these complex deep learning models capture. Further, the LIME methodology was implemented on the developed surface roughness CNN model to identify the important time-frequency bands (i.e., the superpixels of a spectrogram) influencing the predictions. Based on the identified important regions on the spectrogram frames, the corresponding frequency characteristics were determined that influence the surface roughness predictions. The important frequency range based on LIME results was approximately 11.7 to 19.1 kHz. The power of XAI was demonstrated by cutting down the sampling rate from 160 kHz to 30, 20, 10, and 5 kHz based on the important frequency range and considering Nyquist criteria. Separate CNN models were developed for these ranges by only extracting time-frequency contents below their corresponding Nyquist cut-offs. A proper data acquisition strategy is proposed by comparing the model performances to argue the selection of a sufficient sampling rate to capture the grinding process successfully and robustly.

42 ENGINEERING↗

Reprint of: Explainable AI (XAI)-driven vibration sensing scheme for surface quality monitoring in a smart surface grinding process

Local Interpretable and Model-agnostic Explanation (LIME), an explainable artificial intelligence (XAI) approach is adapted to identify the globally important time-frequency bands for predicting average surface roughness (Ra) in a smart grinding process. The smart grinding setup consisted of a Supertech CNC precision surface grinding machine, instrumented with a Dytran piezoelectric accelerometer attached to the tailstock quill along the tangential direction (Y-axis). For every grinding pass, vibration signatures were captured, and the ground truth surface roughness values were recorded using a Mahr Marsurf M300C portable surface roughness profilometer. The roughness values ranged from 0.06 to 0.14 microns over the complete set of experiments. Time-frequency domain spectrogram frames were extracted for each of the vibration signals collected during the grinding process. Convolutional Neural Networks (CNNs) were modeled to predict the surface roughness based on these spectrogram frames and their image augmentations. The best CNN model was able to predict the roughness values with an overall R2-score of 0.95, training R2-score of 0.99, and testing R2-score of 0.81 with only 80 sets of vibration signals corresponding to 4 experiments with 20 trials each. Although the data size is not large enough to guarantee such performance metrics in real-world scenarios, one can extract statistically consistent explanations underlying the relationships these complex deep learning models capture. Further, the LIME methodology was implemented on the developed surface roughness CNN model to identify the important time-frequency bands (i.e., the superpixels of a spectrogram) influencing the predictions. Based on the identified important regions on the spectrogram frames, the corresponding frequency characteristics were determined that influence the surface roughness predictions. The important frequency range based on LIME results was approximately 11.7 to 19.1 kHz. The power of XAI was demonstrated by cutting down the sampling rate from 160 kHz to 30, 20, 10, and 5 kHz based on the important frequency range and considering Nyquist criteria. Separate CNN models were developed for these ranges by only extracting time-frequency contents below their corresponding Nyquist cut-offs. A proper data acquisition strategy is proposed by comparing the model performances to argue the selection of a sufficient sampling rate to capture the grinding process successfully and robustly.

47 OTHER INSTRUMENTATION↗