Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Deep neural networks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Wide‐Field Bond Quality Evaluation Using Frequency Domain Thermoreflectance with Deep Neural Network Feature Reconstruction

Heterogeneous integration of microelectronic components provides a pathway to improve circuit/component performance; however, this comes with assembly challenges, in particular due to complex interfaces via subsurface bump bonds. The ability of these bonds to transmit electrical signals and conduct heat to the carrier substrate limits component performance. In this work, hyperspectral frequency‐domain thermoreflectance (FDTR) imaging is demonstrated as a robust technique for evaluating the quality of subsurface indium bump bonds in a surrogate microelectronic sample. By performing microscale FDTR imaging with coarse motion image stitching, thermal phase maps that cover a 4 mm by 4 mm field‐of‐view with subsurface feature sensitivity at depths greater than 50 µm are obtained. The resulting FDTR hyperspectral data contains more than three million pixels and reveal the quality of subsurface microbump arrays. Wide‐field analysis of bonded versus gap regions is enabled by deep neural network feature reconstruction, that after training, rapidly provides an interpretable representation of bond quality. Utility of noisy higher frequency FDTR phase maps, i.e., near the computationally predicted sensing depth limit, results in an average prediction error of 11%. Taken together, FDTR with neural network‐based analysis demonstrates subsurface bond monitoring at length scales relevant for heterogeneously integrated microelectronics.

FDTR↗

Deep Neural Network Assisted Distributed Strain and Temperature Fiber Sensor System for Natural Gas Pipeline Monitoring

Natural gas pipeline integrity monitoring is crucial to detect potential leaks, find structural issues, and prevent environmental damage. This article presents a system of natural gas pipeline monitoring that uses a specialized double Brillouin peak sensing fiber along with the Brillouin optical time domain analysis (BOTDAs) technique. The calibrated sensing fiber coefficients for strain and temperature are 41.8 kHz/ με and 0.9 MHz/°C for peak 1; and 47.2 kHz/ με , and 1.11 MHz/°C for peak 2, respectively. Initially, lab tests were performed by installing a short section of double Brillouin peak fiber (DBPF) on a 1-in steel pipe under pressure up to 1000 per square inch (psi) at elevated temperatures. Simultaneous distributed measurements of temperature and pressure-induced hoop strain were successfully measured. Considering the long processing speed to extract Brillouin frequency shift (BFS), we employ a novel probabilistic deep neural network (PDNN) framework for rapid BFS prediction. Additionally, using the Finite Element Method, the effects of the pipeline pressure on hoop strain were modeled and compared to the experimental hoop strain under the same set of pipeline conditions. Finally, an actual 4-in outer diameter steel natural gas pipeline was used for pilot-scale tests, where hoop strain was measured at various pressure levels. Leaks were simulated to demonstrate accurate pipeline integrity monitoring. At an internal pipe pressure of 1000 psi, hoop strain of approximately 300 με was observed, and the sensitivity was calculated as 0.28 με /psi. The results of this pilot-scale study demonstrated that the system is capable of performing distributed monitoring sufficient to detect pipeline pressure and the presence of leaks to ensure the safe operation of gas pipelines in the field.

03 NATURAL GAS↗

Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions

Models that predict brain responses to stimuli provide one measure of understanding of a sensory system and have many potential applications in science and engineering. Deep artificial neural networks have emerged as the leading such predictive models of the visual system but are less explored in audition. Prior work provided examples of audio-trained neural networks that produced good predictions of auditory cortical fMRI responses and exhibited correspondence between model stages and brain regions, but left it unclear whether these results generalize to other neural network models and, thus, how to further improve models in this domain. We evaluated model-brain correspondence for publicly available audio neural network models along with in-house models trained on 4 different tasks. Most tested models outpredicted standard spectromporal filter-bank models of auditory cortex and exhibited systematic model-brain correspondence: Middle stages best predicted primary auditory cortex, while deep stages best predicted non-primary cortex. However, some state-of-the-art models produced substantially worse brain predictions. Models trained to recognize speech in background noise produced better brain predictions than models trained to recognize speech in quiet, potentially because hearing in noise imposes constraints on biological auditory representations. The training task influenced the prediction quality for specific cortical tuning properties, with best overall predictions resulting from models trained on multiple tasks. The results generally support the promise of deep neural networks as models of audition, though they also indicate that current models do not explain auditory cortical responses in their entirety.

59 BASIC BIOLOGICAL SCIENCES↗

Simultaneous energy and mass calibration of large-radius jets with the ATLAS detector using a deep neural network

The energy and mass measurements of jets are crucial tasks for the Large Hadron Collider experiments. This paper presents a new calibration method to simultaneously calibrate these quantities for large-radius jets measured with the ATLAS detector using a deep neural network (DNN). To address the specificities of the calibration problem, special loss functions and training procedures are employed, and a complex network architecture, which includes feature annotation and residual connection layers, is used. The DNN-based calibration is compared to the standard numerical approach in an extensive series of tests. The DNN approach is found to perform significantly better in almost all of the tests and over most of the relevant kinematic phase space. In particular, it consistently improves the energy and mass resolutions, with a 30% better energy resolution obtained for transverse momenta $p$ T > $500$ GeV.

47 OTHER INSTRUMENTATION↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Improving missing transverse momentum estimation with a deep neural network

At hadron colliders, the net transverse momentum of particles that do not interact with the detector (missing transverse momentum, $^→_𝑝$$^{miss}_{T}$) is a crucial observable in many analyses. In the standard model, $^→_𝑝$$^{miss}_{T}$ originates from neutrinos. Many beyond-the-standard-model particles, such as dark matter candidates, are also expected to leave the experimental apparatus undetected. This paper presents a novel deep neural network based $^→_𝑝$$^{miss}_{T}$ estimator, DeepMET, developed by the CMS Collaboration at the LHC. The DeepMET algorithm produces a weight for each reconstructed particle based on its properties. The estimator is based on the negative vector sum of the weighted transverse momenta of all reconstructed particles in an event. Compared with other estimators currently employed by CMS, DeepMET improves the $^→_𝑝$$^{miss}_{T}$ resolution by 10%–30%, shows improvement for a wide range of final states, is easier to train, and is more resilient against the effects of additional proton-proton interactions accompanying the collision of interest.

artificial neural networks↗

Integration of Ag-CBRAM crossbars and Mott ReLU neurons for efficient implementation of deep neural networks in hardware

In-memory computing with emerging non-volatile memory devices (eNVMs) has shown promising results in accelerating matrix-vector multiplications. However, activation function calculations are still being implemented with general processors or large and complex neuron peripheral circuits. Here, we present the integration of Ag-based conductive bridge random access memory (Ag-CBRAM) crossbar arrays with Mott rectified linear unit (ReLU) activation neurons for scalable, energy and area-efficient hardware (HW) implementation of deep neural networks. We develop Ag-CBRAM devices that can achieve a high ON/OFF ratio and multi-level programmability. Compact and energy-efficient Mott ReLU neuron devices implementing ReLU activation function are directly connected to the columns of Ag-CBRAM crossbars to compute the output from the weighted sum current. We implement convolution filters and activations for VGG-16 using our integrated HW and demonstrate the successful generation of feature maps for CIFAR-10 images in HW. Our approach paves a new way toward building a highly compact and energy-efficient eNVMs-based in-memory computing system.

Mott insulators↗

Theoretical Prediction of Thermal Expansion Anisotropy for Y 2 Si 2 O 7 Environmental Barrier Coatings Using a Deep Neural Network Potential and Comparison to Experiment

Environmental barrier coatings (EBCs) are an enabling technology for silicon carbide (SiC)-based ceramic matrix composites (CMCs) in extreme environments such as gas turbine engines. However, the development of new coating systems is hindered by the large design space and difficulty in predicting the properties for these materials. Density Functional Theory (DFT) has successfully been used to model and predict some thermodynamic and thermo-mechanical properties of high-temperature ceramics for EBCs, although these calculations are challenging due to their high computational costs. In this work, we use machine learning to train a deep neural network potential (DNP) for Y 2 Si 2 O 7 , which is then applied to calculate the thermodynamic and thermo-mechanical properties at near-DFT accuracy much faster and using less computational resources than DFT. We use this DNP to predict the phonon-based thermodynamic properties of Y 2 Si 2 O 7 with good agreement to DFT and experiments. We also utilize the DNP to calculate the anisotropic, lattice direction-dependent coefficients of thermal expansion (CTEs) for Y 2 Si 2 O 7 . Molecular dynamics trajectories using the DNP correctly demonstrate the accurate prediction of the anisotropy of the CTE in good agreement with the diffraction experiments. In the future, this DNP could be applied to accelerate additional property calculations for Y 2 Si 2 O 7 compared to DFT or experiments.

36 MATERIALS SCIENCE↗

A Medium‐Sized Paleo‐Tsunami Reconstruction by a Deep Neural Network Processing Sedimentary Deposits

Abstract Reconstructing the magnitude and recurrence time of tsunamis, one of the most destructive and unpredictable natural hazards impacting coastal communities, is essential. While major tsunamis are the most studied due to their disastrous impact, small/medium tsunamis (SMTs) are much more frequent and can still significantly impact the coast. Therefore, SMTs potentially provide an extensive archive of information preserved in the geological record. Analyzing the deposits of small/medium paleo‐tsunamis (SMPTs) opens a window into when their direct observation was unavailable. However, deposits of SMPTs are often degraded, traditional sediment deposition inversion models might fail. Recent research has shown that Deep Neural Networks (DNN) can effectively reconstruct the flow conditions of major tsunamis from their deposits. We evaluate the effectiveness of this approach in reconstructing the characteristics of a recent medium size tsunami (2006 Java) and of a medium paleo‐tsunami (1929 Grand Banks). We successfully reconstruct the flow characteristics of the 2006 Java event and show that an inversion of comparable quality is possible for the 1929 Grand Banks tsunami, despite greater uncertainties due to the deposit degradation. Our research shows that Machine Learning has the potential to unseal the meaning of data of thousands SMPTs.

Batubo, P.↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

Deep neural network uncertainty quantification for LArTPC reconstruction

We evaluate uncertainty quantification (UQ) methods for deep learning applied to liquid argon time projection chamber (LArTPC) physics analysis tasks. As deep learning applications enter widespread usage among physics data analysis, neural networks with reliable estimates of prediction uncertainty and robust performance against overconfidence and out-of-distribution (OOD) samples are critical for their full deployment in analyzing experimental data. While numerous UQ methods have been tested on simple datasets, performance evaluations for more complex tasks and datasets are scarce. Here we assess the application of selected deep learning UQ methods on the task of particle classification using the PiLArNet monte carlo 3D LArTPC point cloud dataset. We observe that UQ methods not only allow for better rejection of prediction mistakes and OOD detection, but also generally achieve higher overall accuracy across different task settings. We assess the precision of uncertainty quantification using different evaluation metrics, such as distributional separation of prediction entropy across correctly and incorrectly identified samples, receiver operating characteristic curves (ROCs), and expected calibration error from observed empirical accuracy. We conclude that ensembling methods can obtain well calibrated classification probabilities and generally perform better than other existing methods in deep learning UQ literature.

47 OTHER INSTRUMENTATION↗

Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations

Addressing the "Red-AI" trend of rising energy consumption by large-scale neural networks, this study investigates the measured energy consumption of training various fully connected neural network architectures. We introduce the BUTTER-E dataset, an augmentation to the BUTTER Empirical Deep Learning dataset, containing energy consumption and performance data from 41,129 individual experimental runs spanning 30,582 distinct configurations: 13 datasets, 20 sizes (trainable parameters), 8 "shapes", and 14 depths on both CPUs and GPUs using node-level watt-meters. This dataset reveals the complex relationship between dataset size, network structure, and energy use. Our analysis uncovers a surprising, hardware-mediated non-linear relationship between energy efficiency and network design, challenging the assumption that reducing the number of parameters or FLOPs is the best way to achieve greater energy efficiency. We propose a straightforward and effective energy model that accounts for network size, computing, and memory hierarchy. Highlighting the need for cache-considerate algorithm development, we suggest a codesign approach to energy efficient network, algorithm, and hardware design. This work contributes to the fields of sustainable computing and Green AI, offering practical guidance for creating more energy-efficient neural networks and promoting sustainable AI.

97 MATHEMATICS AND COMPUTING↗

Detecting rare neutral atomic-carbon absorbers with a deep neural network

ABSTRACT C i absorbers play an important role as indicators for exploring the presence of cold gas in the interstellar medium of galaxies. However, the current data base of C i absorbers is very limited due to their weak absorption feature and rarity. Here, we report results from a search of C i λλ1560, 1656 absorption lines using Mg ii absorbers as signposts with modified deep learning algorithms, which provides a very quick way to search for weak C i absorber candidates. A total of 107 C i absorbers were detected, which nearly doubles the size of previously known samples. In addition, we found 17 C i absorbers to be associated with 2175 Å dust absorbers (2DAs), i.e. about 16 per cent C i absorbers are associated with 2DAs. Comparing the average dust depletion patterns of C i absorbers with those of damped Lyman α absorbers (DLAs), Mg ii absorbers, Ca ii absorbers, and 2175 Å dust absorbers (2DAs) shows that C i absorbers generally have environments with more dust than DLAs, Mg ii, and Ca ii absorbers, but similar to dust in 2DAs. Similarity between the dust depletion pattern of C i absorbers to that of the warm disc in the Milky Way indicates that C i absorption clouds are possibly associated with disc components in distant galaxies. Therefore, C i absorbers are confirmed to be excellent probes to trace cold gas and dust in the Universe.

Ge, Jian↗

A fine pore-preserved deep neural network for porosity analytics of a high burnup U-10Zr metallic fuel

Abstract U-10 wt.% Zr (U-10Zr) metallic fuel is the leading candidate for next-generation sodium-cooled fast reactors. Porosity is one of the most important factors that impacts the performance of U-10Zr metallic fuel. The pores generated by the fission gas accumulation can lead to changes in thermal conductivity, fuel swelling, Fuel-Cladding Chemical Interaction (FCCI) and Fuel-Cladding Mechanical Interaction (FCMI). Therefore, it is crucial to accurately segment and analyze porosity to understand the U-10Zr fuel system to design future fast reactors. To address the above issues, we introduce a workflow to process and analyze multi-source Scanning Electron Microscope (SEM) image data. Moreover, an encoder-decoder-based, deep fully convolutional network is proposed to segment pores accurately by integrating the residual unit and the densely-connected units. Two SEM 250 × field of view image datasets with different formats are utilized to evaluate the new proposed model’s performance. Sufficient comparison results demonstrate that our method quantitatively outperforms two popular deep fully convolutional networks. Furthermore, we conducted experiments on the third SEM 2500 × field of view image dataset, and the transfer learning results show the potential capability to transfer the knowledge from low-magnification images to high-magnification images. Finally, we use a pre-trained network to predict the pores of SEM images in the whole cross-sectional image and obtain quantitative porosity analysis. Our findings will guide the SEM microscopy data collection efficiently, provide a mechanistic understanding of the U-10Zr fuel system and bridge the gap between advanced characterization to fuel system design.

36 MATERIALS SCIENCE↗

Quantifying leaf symptoms of sorghum charcoal rot in images of field‐grown plants using deep neural networks

Abstract Charcoal rot of sorghum (CRS) is a significant disease affecting sorghum crops, with limited genetic resistance available. The causative agent, Macrophomina phaseolina (Tassi) Goid, is a highly destructive fungal pathogen that targets over 500 plant species globally, including essential staple crops. Utilizing field image data for precise detection and quantification of CRS could greatly assist in the prompt identification and management of affected fields and thereby reduce yield losses. The objective of this work was to implement various machine learning algorithms to evaluate their ability to accurately detect and quantify CRS in red‐green‐blue images of sorghum plants exhibiting symptoms of infection. EfficientNet‐B3 and a fully convolutional network emerged as the top‐performing models for image classification and segmentation tasks, respectively. Among the classification models evaluated, EfficientNet‐B3 demonstrated superior performance, achieving an accuracy of 86.97%, a recall rate of 0.71, and an F1 score of 0.73. Of the segmentation models tested, FCN proved to be the most effective, exhibiting a validation accuracy of 97.76%, a recall rate of 0.68, and an F1 score of 0.66. As the size of the image patches increased, both models’ validation scores increased linearly, and their inference time decreased exponentially. This trend could be attributed to larger patches containing more information, improving model performance, and fewer patches reducing the computational load, thus decreasing inference time. The models, in addition to being immediately useful for breeders and growers of sorghum, advance the domain of automated plant phenotyping and may serve as a foundation for drone‐based or other automated field phenotyping efforts. Additionally, the models presented herein can be accessed through a web‐based application where users can easily analyze their own images.

Gonzalez, Emmanuel M.↗