Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “DNN”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Measurement of a radial flow profile with eddy current flow meters and deep neural networks

Eddy current flow meters (ECFMs) measure flows of conductive fluids. Recent interest in ECFMs has increased due to applications in advanced nuclear reactors. ECFMs are well suited for such applications, as they can provide non-invasive measurements of flow in fluids that are often difficult to measure. Traditionally, ECFMs are operated using an alternating current at a single frequency, limiting ECFMs to measure average fluid velocities, blockages, or voids. Here, we expand the capabilities of ECFMs by measuring the fluid radial velocity profile of liquid mercury. To accomplish this, we made several ECFM sensitivity measurements at a range of frequencies. Different frequencies vary the electromagnetic skin depth of the device. By adjusting frequencies, we probed the fluid velocity at various radial locations and constructed a flow-velocity profile. The relationship between the ECFM measurements and velocity profile is nonlinear and requires solving an inverse problem. Using electromagnetic finite-element simulations to train a deep neural network (DNN), we created a model that provides a stable general relationship between the sensitivity measurements of an ECFM and the fluid velocity profile. Using ECFM measurements of liquid mercury, our DNN model calculates a flow profile that agrees well with computational fluid dynamics (CFD) simulations. This technique has potential to improve flow monitoring for optimization, safe operation of conductive fluid loops, and/or validating complex CFD models.

47 OTHER INSTRUMENTATION↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Neural network emulation of flow in heavy-ion collisions at intermediate energies

Applications of new techniques in machine learning are speeding up progress in research in various fields. In this work, we construct and evaluate a deep neural network (DNN) to be used within a Bayesian statistical framework as a faster and more reliable alternative to the Gaussian process (GP) emulator of an isospin-dependent Boltzmann-Uehling-Uhlenbeck (IBUU) transport model simulator of heavy-ion reactions at intermediate beam energies. We found strong evidence of the DNN being able to emulate the IBUU simulator's prediction on the strengths of protons' directed and elliptical flow very efficiently even with small training datasets and with accuracy about ten times higher than the GP. Here, limitations of our present work and future improvements are also discussed.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

Why Dissolving Salt in Water Decreases Its Dielectric Permittivity

The dielectric permittivity of salt water decreases on dissolving more salt. For nearly a century, this phenomenon has been explained by invoking saturation in the dielectric response of the solvent water molecules. Herein, we employ an advanced deep neural network (DNN), built using data from density functional theory, to study the dielectric permittivity of sodium chloride solutions. Notably, the decrease in the dielectric permittivity as a function of concentration, computed using the DNN approach, agrees well with experiments. Detailed analysis of the computations reveals that the dominant effect, caused by the intrusion of ionic hydration shells into the solvent hydrogen-bond network, is the disruption of dipolar correlations among water molecules. Accordingly, the observed decrease in the dielectric permittivity is mostly due to increasing suppression of the collective response of solvent waters.

74 ATOMIC AND MOLECULAR PHYSICS↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

Data-Driven Affinely Adjustable Robust Volt/VAr Control

Recent years have seen the increasing proliferation of distributed energy resources with intermittent power outputs, posing new challenges to the voltage management in distribution networks. To this end, this paper proposes a data-driven affinely adjustable robust Volt/VAr control (AARVVC) scheme, which modulates the smart inverter’s reactive power in an affine function of its active power, based on the voltage sensitivities with respect to real/reactive power injections. To achieve a fast and accurate estimation of voltage sensitivities, we propose a data-driven method based on deep neural network (DNN), together with a rule-based bus-selection process using the bidirectional search method. Our method only uses the operating statuses of selected buses as inputs to DNN, thus significantly improving the training efficiency and reducing information redundancy. Finally, a distributed consensus-based solution, based on the alternating direction method of multipliers (ADMM), for the AARVVC is applied to decide the inverter’s reactive power adjustment rule with respect to its active power. Only limited information exchange is required between each local agent and the central agent to obtain the slope of the reactive power adjustment rule, and there is no need for the central agent to solve any (sub)optimization problems. Finally, numerical results on the modified IEEE-123 bus system validate the effectiveness and superiority of the proposed data-driven AARVVC method.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Train Like a (Var)Pro: Efficient Training of Neural Networks with Variable Projection

Deep neural networks (DNNs) have achieved state-of-the-art performance across a variety of traditional machine learning tasks, e.g., speech recognition, image classification, and segmentation. The ability of DNNs to efficiently approximate high-dimensional functions has also motivated their use in scientific applications, e.g., to solve partial differential equations and to generate surrogate models. In this paper, we consider the supervised training of DNNs, which arises in many of the above applications. We focus on the central problem of optimizing the weights of the given DNN such that it accurately approximates the relation between observed input and target data. Devising effective solvers for this optimization problem is notoriously challenging due to the large number of weights, nonconvexity, data sparsity, and nontrivial choice of hyperparameters. To solve the optimization problem more efficiently, we propose the use of variable projection (VarPro), a method originally designed for separable nonlinear least-squares problems. Our main contribution is the Gauss--Newton VarPro method (GNvpro) that extends the reach of the VarPro idea to nonquadratic objective functions, most notably cross-entropy loss functions arising in classification. These extensions make GNvpro applicable to all training problems that involve a DNN whose last layer is an affine mapping, which is common in many state-of-the-art architectures. In our four numerical experiments from surrogate modeling, segmentation, and classification, GNvpro solves the optimization problem more efficiently than commonly used stochastic gradient descent (SGD) schemes. Finally, GNvpro finds solutions that generalize well, and in all but one example better than well-tuned SGD methods, to unseen data points.

97 MATHEMATICS AND COMPUTING↗

A deep learning-enhanced framework for multiphysics joint inversion

Joint inversion has drawn considerable attention due to the availability of multiple geophysical data sets, ever-increasing computational resources, the development of advanced algorithms, and its ability to reduce inversion uncertainty. A key issue of joint inversion is to develop effective strategies to link different geophysical data in a unified mathematical framework, in which the information obtained from different models can complement each other. We have developed a deep learning-enhanced joint inversion framework to simultaneously reconstruct different physical models by fusing different types of geophysical data. Traditionally, structure similarity constraints are pursued by joint inversion algorithms using manually crafted formulations (e.g., cross gradient). The constraint is constructed by a deep neural network (DNN) during the learning process. The framework is designed to combine the DNN and a traditional independent inversion workflow and improve the joint inversion result iteratively. The network can be easily extended to incorporate multiphysics without structural changes. Numerical experiments on the joint inversion of 2D DC resistivity data and seismic traveltime are used to validate our method. In addition, this learning-based framework demonstrates excellent generalization abilities when tested on data sets using different geologic structures. It also can handle different sensing configurations and nonconforming discretization.

Geochemistry & Geophysics↗

Three-dimensional cooperative inversion of airborne magnetic and gravity gradient data using deep-learning techniques

Using multiple geophysical methods has become a prevailing approach in numerous geophysical applications to investigate subsurface structures and parameters. These multimethod-based exploration strategies have the potential to greatly diminish uncertainties and ambiguities encountered during geophysical data analysis and interpretation. One of the applications is the cooperative inversion of airborne magnetic and gravity gradient data for the interpretation of data obtained in mineral, oil and gas, and geothermal explorations. In this paper, a unified cooperative inversion framework is designed by combining the standard separate inversions with a deep neural network (DNN), which serves as the link between different types of data. A well-trained DNN takes the separately inverted susceptibility and density models as the inputs and provides improved models that will be used as the initial models of deterministic inversions. A two-round iteration strategy is adopted to guarantee the reasonability of the recovered models and overall efficiency of the inversion. In addition, this deep-learning (DL)-based framework demonstrates excellent generalization abilities when tested on models that are entirely distinct from the training data sets. The framework can easily incorporate multiphysics without necessitating any structural changes to the network. Synthetic experiments validate that our DL-based method outperforms conventional separate inversions and cross-gradient-based joint inversion in view of the accuracy of the recovered models and inversion efficiency. Successful application to field data further verifies the effectiveness of our DL-based method.

Geochemistry & Geophysics↗

Optimization of the deep neural network parameters for generating homogenized fuel assembly data for nodal codes

Homogenized fuel assembly (FA) data is a typical input data for nodal codes. Generating that data, however, could be time-consuming. One of promising ways to mitigate the computational burden of generating macroscopic cross-sections is to use trained artificial neural network (ANN) models for predicting nuclear data. However, there is a challenge to make the model support variable FA geometry. In this work, two most common types of FA were combined in one ANN model. Since there could be multiple ways of converting 2-dimensional FA data into 1-dimensional input vector for ANN, three different approaches of data flattening were evaluated. The input parameters included each fuel pin enrichment, fuel temperature, moderator temperature and boron concentration. The output parameters were 2-group macroscopic cross-sections (XS) and pin power distribution (HFF). A fully connected deep neural network (DNN) model was trained and tested using pre-generated data obtained with lattice physics code STREAM. The results of this study showed no statistically significant difference in the accuracy of XS and HFF generation for all 3 tested input vector orders. This means that fully connected DNN for XS generation demonstrated input sequence invariance. Results of comparing predicted XS data with reference solutions were found sufficiently close considering the reduction of computation time offered by ANN. Mean relative difference (MRD) for all output XS parameters was found below 0.7%, while HFF MRD was found higher compared to XS values, in some cases slightly exceeding 1%, mostly near guide tube locations. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Australian Safeguards and Non-Proliferation Office Visit: Nuclear Nonproliferation and Security Program at LANL [Slides]

The Laboratory’s nuclear nonproliferation and security portfolio is managed by the GS-NNS program office: NNSA’s Defense Nuclear Nonproliferation office (DNN) makes up ~80% of our work. We also support State Dept activities closely aligned with our work for DNN and NASAprograms. Our work is executed across the entire Laboratory, approximately half in the Global Security directorate and a third in the Weapons Program.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Learning Optimal Aerodynamic Designs

This project created a framework for efficient, accurate, and scalable deep neural network representations of design optimization problem solutions. The inputs to these DNN representations are the vector of design requirement parameters, the outputs are the optimal design variables, and the goal is to learn the map from inputs to outputs (i.e., inverse design). The team addressed the problem of the optimal shape design of aerodynamic lifting surfaces—in particular aircraft wings—using a Reynolds-Average Navier Stokes model to govern the CFD-based aerodynamic shape optimization. The inverse design map for such problems is very complex and high-dimensional, involving inputs and outputs on the order of 1000s. To approximate this inverse design map, the team developed algorithms to construct parsimonious DNN architectures, which automatically identify low-dimensional manifolds in which design requirements affect optimal shape parameters, and trained these architectures with multifidelity optimization methods. The resulting methodology accurately and automatically designs optimal aerodynamic lifting surfaces with very high accuracy (99%) at interactive speeds, of the order of milliseconds, resulting in factors of one million or more speedup relative to CFD-based design optimization.

97 MATHEMATICS AND COMPUTING↗

AI-Driven Detector Design for the EIC (Final Technical Report)

We developed an optimization workflow based on DNN-based fast-simulation and reconstruction algorithms. We used these methods to advance the design of calorimeter systems for the Electron-Ion Collider (EIC). This DNN-driven optimization provides a blueprint for integrating gradient-based methods into detector-design workflows. All software pipelines and methods have been released publicly and incorporated into the EIC collaboration’s physics studies, broadening their impact. Three journal articles detailing the methods developed here serve as a reference for the design and optimal use of next generation high-granularity calorimeter systems in nuclear and particle physics.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A New Perspective of Post-Weld Baking Effect on Al-Steel Resistance Spot Weld Properties through Machine Learning and Finite Element Modeling

The root cause of post-weld baking on the mechanical performance of Al-steel dissimilar resistance spot welds (RSWs) has been determined by machine learning (ML) and finite element modeling (FEM) in this study. A deep neural network (DNN) model was constructed to associate the spot weld performance with the joint attributes, stacking materials, and other conditions, using a comprehensive experimental dataset. The DNN model positively identified that the post-weld baking reduces the joint performance, and the extent of degradation depends on the thickness of stacking materials. A three-dimensional finite element (FE) model was then used to investigate the root cause and the mechanism of the baking effect. It revealed that the formation of high thermal stresses during baking, from the mismatch of thermal expansion between steel and Al alloy, causes damage and cracking of the brittle intermetallic compound (IMC) formed at the interface of the weld nugget during welding. This in turn reduces the joint performance by promoting undesirable interfacial fracture when the welds were subjected to externally applied loads. The FEM model further revealed that increase in structural stiffness, because of increase in steel sheet thickness, reduces the thermal stresses at the interface caused by the thermal expansion mismatch and consequently lessens the detrimental effect of post-weld baking on the joint performance.

36 MATERIALS SCIENCE↗

Densely Connected G-invariant Deep Neural Networks with Signed Permutation Representations

We introduce and investigate, for finite groups G, G-invariant deep neural network (GDNN) architectures with ReLU activation that are densely connected- i.e., include all possible skip connections. In contrast to other G-invariant architectures in the literature, the preactivations of theG-DNNs presented here are able to transform by signed permutation representations (signed perm-reps) of G. Moreover, the individual layers of the G-DNNs are not required to be G-equivariant; instead, the preactivations are constrained to be G-equivariant functions of the network input in a way that couples weights across all layers. The result is a richer family of G-invariant architectures never seen previously. We derive an efficient implementation of G-DNNs after a reparameterization of weights, as well as necessary and sufficient conditions for an architecture to be "admissible"- i.e., nondegenerate and inequivalent to smaller architectures. We include code that allows a user to build a G-DNN interactively layer-by-layer, with the final architecture guaranteed to be admissible. We show that there are far more admissible G-DNN architectures than those accessible with the "concatenated ReLU" activation function from the literature. Finally, we apply G-DNNs to two example problems--(1) multiplication in --1, 1} (with theoretical guarantees) and (2) 3D object classification--finding that the inclusion of signed perm-reps significantly boosts predictive performance compared to baselines with only ordinary (i.e., unsigned) perm-reps.

97 MATHEMATICS AND COMPUTING↗

Factorized visual representations in the primate visual system and deep neural networks

Object classification has been proposed as a principal objective of the primate ventral visual stream and has been used as an optimization target for deep neural network models (DNNs) of the visual system. However, visual brain areas represent many different types of information, and optimizing for classification of object identity alone does not constrain how other information may be encoded in visual representations. Information about different scene parameters may be discarded altogether (‘invariance’), represented in non-interfering subspaces of population activity (‘factorization’) or encoded in an entangled fashion. In this work, we provide evidence that factorization is a normative principle of biological visual representations. In the monkey ventral visual hierarchy, we found that factorization of object pose and background information from object identity increased in higher-level regions and strongly contributed to improving object identity decoding performance. We then conducted a large-scale analysis of factorization of individual scene parameters – lighting, background, camera viewpoint, and object pose – in a diverse library of DNN models of the visual system. Models which best matched neural, fMRI, and behavioral data from both monkeys and humans across 12 datasets tended to be those which factorized scene parameters most strongly. Notably, invariance to these parameters was not as consistently associated with matches to neural and behavioral data, suggesting that maintaining non-class information in factorized activity subspaces is often preferred to dropping it altogether. Thus, we propose that factorization of visual scene information is a widely used strategy in brains and DNN models thereof.

59 BASIC BIOLOGICAL SCIENCES↗