Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNN”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Physics-Informed Neural Network Method for Forward and Backward Advection-Dispersion Equations

Advection-dispersion equations (ADEs) are commonly used to describe transport phenomena in porous media. Even though mature discretization-based numerical methods for ADEs exist, some challenges still remain, especially when it comes to solving advection-dominated forward ADEs and diffusion-dominated backward ADEs. The latter problem usually arises in the source identification context and leads to numerically unstable grid-based solutions that require a form of regularization or should be treated as an inverse problem that is computationally more expensive because it requires solving the forward problem multiple times. In this study, we propose a discretization-free approach based on the physics-informed neural network (PINN) method for solving coupled ADE and Darcy flow equations with space-dependent hydraulic conductivity. In this approach, the hydraulic conductivity, hydraulic head, and concentration fields are approximated with deep neural networks (DNNs). We assume that the conductivity field is given by its values on a grid, and we use these values to train the conductivity DNN. The head and concentration DNNs are trained by minimizing the residuals of the flow equation and ADE and using the initial and boundary conditions as additional constraints. The PINN method is applied to one- and two-dimensional forward ADE problems, where its performance for various P\'{e}clet numbers ($Pe$) is compared with the analytical and numerical solutions. We find that the PINN method is accurate with errors of less than 1\% and outperforms some conventional discretization-based methods for $Pe$ larger that 100. Next, we demonstrate that the PINN method remains accurate for the backward ADEs, with the relative errors in most cases staying under 5\% compared to the reference concentration field. Finally, we show that when available, the concentration measurements can be easily incorporated in the PINN method and significantly improve (by more than 50\% in the considered cases) the accuracy of the PINN solution of the backward ADE.

He, Qizhi↗

A Medium‐Sized Paleo‐Tsunami Reconstruction by a Deep Neural Network Processing Sedimentary Deposits

Abstract Reconstructing the magnitude and recurrence time of tsunamis, one of the most destructive and unpredictable natural hazards impacting coastal communities, is essential. While major tsunamis are the most studied due to their disastrous impact, small/medium tsunamis (SMTs) are much more frequent and can still significantly impact the coast. Therefore, SMTs potentially provide an extensive archive of information preserved in the geological record. Analyzing the deposits of small/medium paleo‐tsunamis (SMPTs) opens a window into when their direct observation was unavailable. However, deposits of SMPTs are often degraded, traditional sediment deposition inversion models might fail. Recent research has shown that Deep Neural Networks (DNN) can effectively reconstruct the flow conditions of major tsunamis from their deposits. We evaluate the effectiveness of this approach in reconstructing the characteristics of a recent medium size tsunami (2006 Java) and of a medium paleo‐tsunami (1929 Grand Banks). We successfully reconstruct the flow characteristics of the 2006 Java event and show that an inversion of comparable quality is possible for the 1929 Grand Banks tsunami, despite greater uncertainties due to the deposit degradation. Our research shows that Machine Learning has the potential to unseal the meaning of data of thousands SMPTs.

Batubo, P.↗

Teaching a neural network to attach and detach electrons from molecules

Abstract Interatomic potentials derived with Machine Learning algorithms such as Deep-Neural Networks (DNNs), achieve the accuracy of high-fidelity quantum mechanical (QM) methods in areas traditionally dominated by empirical force fields and allow performing massive simulations. Most DNN potentials were parametrized for neutral molecules or closed-shell ions due to architectural limitations. In this work, we propose an improved machine learning framework for simulating open-shell anions and cations. We introduce the AIMNet-NSE (Neural Spin Equilibration) architecture, which can predict molecular energies for an arbitrary combination of molecular charge and spin multiplicity with errors of about 2–3 kcal/mol and spin-charges with error errors ~0.01e for small and medium-sized organic molecules, compared to the reference QM simulations. The AIMNet-NSE model allows to fully bypass QM calculations and derive the ionization potential, electron affinity, and conceptual Density Functional Theory quantities like electronegativity, hardness, and condensed Fukui functions. We show that these descriptors, along with learned atomic representations, could be used to model chemical reactivity through an example of regioselectivity in electrophilic aromatic substitution reactions.

36 MATERIALS SCIENCE↗

Deep learning for exploring ultra-thin ferroelectrics with highly improved sensitivity of piezoresponse force microscopy

Hafnium oxide-based ferroelectrics have been extensively studied because of their existing ferroelectricity, even in ultra-thin film form. However, studying the weak response from ultra-thin film requires improved measurement sensitivity. In general, resonance-enhanced piezoresponse force microscopy (PFM) has been used to characterize ferroelectricity by fitting a simple harmonic oscillation model with the resonance spectrum. However, an iterative approach, such as traditional least squares (LS) fitting, is sensitive to noise and can result in the misunderstanding of weak responses. In this study, we developed the deep neural network (DNN) hybrid with deep denoising autoencoder (DDA) and principal component analysis (PCA) to extract resonance information. The DDA/PCA-DNN improves the PFM sensitivity down to 0.3 pm, allowing measurement of weak piezoresponse with low excitation voltage in 10-nm-thick Hf 0.5 Zr 0.5 O 2 thin films. Our hybrid approaches could provide more chances to explore the low piezoresponse of the ultra-thin ferroelectrics and could be applied to other microscopic techniques.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Freely scalable and reconfigurable optical hardware for deep learning

Abstract As deep neural network (DNN) models grow ever-larger, they can achieve higher accuracy and solve more complex problems. This trend has been enabled by an increase in available compute power; however, efforts to continue to scale electronic processors are impeded by the costs of communication, thermal management, power delivery and clocking. To improve scalability, we propose a digital optical neural network (DONN) with intralayer optical interconnects and reconfigurable input values. The path-length-independence of optical energy consumption enables information locality between a transmitter and a large number of arbitrarily arranged receivers, which allows greater flexibility in architecture design to circumvent scaling limitations. In a proof-of-concept experiment, we demonstrate optical multicast in the classification of 500 MNIST images with a 3-layer, fully-connected network. We also analyze the energy consumption of the DONN and find that digital optical data transfer is beneficial over electronics when the spacing of computational units is on the order of $$>10\,\upmu $$ > 10 μ m.

42 ENGINEERING↗

Machine learning surrogates for surface complexation model of uranium sorption to oxides

Abstract The safety assessments of the geological storage of spent nuclear fuel require understanding the underground radionuclide mobility in case of a leakage from multi-barrier canisters. Uranium, the most common radionuclide in non-reprocessed spent nuclear fuels, is immobile in reduced form (U(IV) and highly mobile in an oxidized state (U(VI)). The latter form is considered one of the most dangerous environmental threats in the safety assessments of spent nuclear fuel repositories. The sorption of uranium to mineral surfaces surrounding the repository limits their mobility. We quantify uranium sorption using surface complexation models (SCMs). Unfortunately, numerical SCM solvers often encounter convergence problems due to the complex nature of convoluted equations and correlations between model parameters. This study explored two machine learning surrogates for the 2-pK Triple Layer Model of uranium retention by oxide surfaces if released as U(IV) in the oxidizing conditions: random forest regressor and deep neural networks. Our surrogate models, particularly DNN, accurately reproduce SCM model predictions at a fraction of the computational cost without any convergence issues. The safety assessment of spent fuel repositories, specifically the migration of leaked radioactive waste, will benefit from having ultrafast AI/ML surrogates for the computationally expensive sorption models that can be easily incorporated into larger-scale contaminant migration models. One such model is presented here.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Machine learning for design principles for single atom catalysts towards electrochemical reactions

Machine learning (ML) integrated density functional theory (DFT) calculations have recently been used to accelerate the design and discovery of heterogeneous catalysts such as single atom catalysts (SACs) through the establishment of deep structure–activity relationships. Here, this review provides recent progress in the ML-aided rational design of heterogeneous catalysts with the focus on SACs in terms of structure–activity relationships, feature importance analysis, high-throughput screening, stability, and metal–support interactions for electrochemistry. Support vector machine (SVM), random forest regression (RFR), and deep neural networks (DNN) along with atomic properties are mainly used for the design of SACs. The ML results have shown that the number of electrons in the d orbital, oxide formation enthalpy, ionization energy, Bader charge, d-band center, and enthalpy of vaporization are mainly the most important parameters for the defining of the structure–activity relationships for electrochemistry. However, the black-box nature of ML techniques occasionally makes a physical interpretation of descriptors, such as the Bader charge, d-band center, and enthalpy of vaporization, non-trivial. At the current stage, ML application is limited by the lack of a large and high-quality database. Future prospects for the development of a large database and a generalized ML algorithm for SAC design are discussed to give insights for further studies in this field.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

From lattice QCD to in-medium heavy-quark interactions via deep learning

Bottomonium states are key probes for experimental studies of the quark-gluon plasma (QGP) created in high-energy nuclear collisions. Theoretical models of bottomonium productions in high-energy nuclear collisions rely on the in-medium interactions between the bottom and antibottom quarks, which can be characterized by real (V R (T, r)) and imaginary (V I (T, r)) potentials, as functions of temperature and spatial separation. Recently, the masses and thermal widths of up to 3S and 2P bottomonium states in QGP were calculated using lattice quantum chromodynamics (LQCD). Starting from these LQCD results and through a novel application of deep neural network (DNN), here, we obtain model-independent results for V R (T, r) and V I (T, r). The temperature dependence of V R (T, r) was found to be very mild between T ≈ 0 - 330 MeV. Meanwhile, VI(T, r) shows rapid increase with T and r, which is much larger than the perturbation theory based expectations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Deep-Learning based Reconstruction of the Shower Maximum $X_{\mathrm{max}}$ using the Water-Cherenkov Detectors of the Pierre Auger Observatory

The atmospheric depth of the air shower maximum X max is an observable commonly used for the determination of the nuclear mass composition of ultra-high energy cosmic rays. Direct measurements of X max are performed using observations of the longitudinal shower development with fluorescence telescopes. At the same time, several methods have been proposed for an indirect estimation of X max from the characteristics of the shower particles registered with surface detector arrays. In this paper, we present a deep neural network (DNN) for the estimation of X max . The reconstruction relies on the signals induced by shower particles in the ground based water-Cherenkov detectors of the Pierre Auger Observatory. The network architecture features recurrent long short-term memory layers to process the temporal structure of signals and hexagonal convolutions to exploit the symmetry of the surface detector array. We evaluate the performance of the network using air showers simulated with three different hadronic interaction models. Thereafter, we account for long-term detector effects and calibrate the reconstructed X max using fluorescence measurements. Finally, we show that the event-by-event resolution in the reconstruction of the shower maximum improves with increasing shower energy and reaches less than 25 g/cm 2 at energies above 2 × 10 19 eV.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A detailed study of interpretability of deep neural network based top taggers

Abstract Recent developments in the methods of explainable artificial intelligence (XAI) allow researchers to explore the inner workings of deep neural networks (DNNs), revealing crucial information about input–output relationships and realizing how data connects with machine learning models. In this paper we explore interpretability of DNN models designed to identify jets coming from top quark decay in high energy proton–proton collisions at the Large Hadron Collider. We review a subset of existing top tagger models and explore different quantitative methods to identify which features play the most important roles in identifying the top jets. We also investigate how and why feature importance varies across different XAI metrics, how correlations among features impact their explainability, and how latent space representations encode information as well as correlate with physically meaningful quantities. Our studies uncover some major pitfalls of existing XAI methods and illustrate how they can be overcome to obtain consistent and meaningful interpretation of these models. We additionally illustrate the activity of hidden layers as neural activation pattern diagrams and demonstrate how they can be used to understand how DNNs relay information across the layers and how this understanding can help to make such models significantly simpler by allowing effective model reoptimization and hyperparameter tuning. These studies not only facilitate a methodological approach to interpreting models but also unveil new insights about what these models learn. Incorporating these observations into augmented model design, we propose the particle flow interaction network model and demonstrate how interpretability-inspired model augmentation can improve top tagging performance.

97 MATHEMATICS AND COMPUTING↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

Sparse Deep Neural Network Inference using different Programming Models

Sparse deep neural networks have gained increasing attention recently in achieving speedups on inference with reduced memory footprints. However, the real-world applications are little shown with specialized optimizations, yet a wide variety of DNN tasks remain dense without exploiting the advantages of sparsity in networks. Recent work presented by MIT/IEEE/Amazon GraphChallenge has demonstrated significant speedups and various techniques. Still, we find that there is limited investigation of the impact of various Python and C\slash C++ based programming models to explore new opportunities in general cases. In this work, we provide performance evaluation through different programming models using CuPy, cuSPARSE, and OpenMP to discuss the advantages and disadvantages of our sparse implementations on single-GPU and multiple GPUs of NVIDIA DGX-A100 40GB/80GB platforms.

machine learning, HPC↗

Improving Progressive Retrieval for HPC Scientific Data using Deep Neural Network

As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).

Wang, Jinzhen↗

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Baheri, Betis↗

Deep Learning-Assisted Real-Time Forward Modeling of Electromagnetic Logging in Complex Formations

Higher dimensional (i.e., 2-D and 3-D) modeling is indispensable to correctly evaluate the responses of electromagnetic logging tools in complex formation environments. However, limited by the high computational cost of rigorous modeling such as finite difference method and the finite element method, the real-time applications in the well logging industry primarily rely on the 1-D forward solver, which would result in erroneous formation evaluation for complex scenarios. As a result, aiming at realizing fast modeling for electromagnetic logging tools in complex formations, this paper proposes a general framework assisted by deep neural networks (DNNs). The framework consists of three modules: earth model classification, parameter extraction, and surrogate construction. Separate DNNs are trained and tested for different modules. The accuracy and efficiency of the DNN assisted fast modeling are validated by several experiments. Here, this study finds that the fast modeling assisted by DNNs is able to calculate the tool responses and reconstruct the subsurface formations in real-time.

97 MATHEMATICS AND COMPUTING↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer paral- lelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation.

Wang, Tianqi↗