Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Accelerating charge estimation in molecular dynamics simulations using physics-informed neural networks: corrosion applications

Molecular Dynamics (MD) simulations are used to understand the effects of corrosion on metallic materials in salt brine. Reactive force fields in classical MD enable accurate modeling of bond formation and breakage in the aqueous medium and at the metal-electrolyte interface, while also facilitating dynamic partial charge equilibration. However, MD simulations are computationally intensive and unsuitable for modeling the long time scales characteristic of corrosive phenomena. To address this, we develop reduced-order machine learning models that provide accurate and efficient predictions of charge density in corrosive environments. Specifically, we use Long Short-Term Memory (LSTM) networks to forecast charge density evolution based on atomic environments represented by Smooth Overlap of Atomic Positions (SOAP) descriptors. A physics-informed loss function enforces charge neutrality and electronegativity equivalence. The atomic charges predicted by the deep learning model trained on this work were obtained two orders of magnitude faster than those from molecular dynamics (MD) simulations, with an error of less than 3% compared to the MD-obtained charges, even in extrapolative scenarios, while adhering to physical constraints. This demonstrates the excellent accuracy, computational efficiency, and validity of the developed model. Lastly, even though developed for corrosion, these protocols are formulated in a phenomenon-agnostic manner, allowing application to various variable-charge interatomic potentials and related fields.

Atomistic models↗

High-Speed Ionic Synaptic Memory Based on 2D Titanium Carbide MXene

Synaptic devices with linear high-speed switching can accelerate learning in artificial neural networks (ANNs) embodied in hardware. Conventional resistive memories however suffer from high write noise and asymmetric conductance tuning, preventing parallel programming of ANN arrays. Electrochemical random-access memories (ECRAMs), where resistive switching occurs by ion insertion into a redox-active channel, aim to address these challenges due to their linear switching and low noise. ECRAMs using 2D materials and metal oxides however suffer from slow ion kinetics, whereas organic ECRAMs enable high-speed operation but face challenges toward on-chip integration due to poor temperature stability of polymers. Here, ECRAMs using 2D titanium carbide (Ti 3 C 2 T x ) MXene that combine the high speed of organics and the integration compatibility of inorganic materials in a single high-performance device are demonstrated. These ECRAMs combine the speed, linearity, write noise, switching energy, and endurance metrics essential for parallel acceleration of ANNs, and importantly, they are stable after heat treatment needed for back-end-of-line integration with Si electronics. The high speed and performance of these ECRAMs introduces MXenes, a large family of 2D carbides and nitrides with more than 30 stoichiometric compositions synthesized to date, as promising candidates for devices operating at the nexus of electrochemistry and electronics.

2D materials↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Finally, our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

42 ENGINEERING↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING↗

DD-AADL

Data Driven Anderson Acceleration for Physics Informed Neural Networks to solve high dimensional partial differential equations

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗

A Path Toward Understanding the Performance Capabilities of SmartNIC Devices [Slides]

SmartNICs, accelerator devices integrated with a network, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. Naturally, there is performance overhead associated with offloading tasks from a host CPU to a SmartNIC. As SmartNICs become more versatile and are used for a wider range of computational tasks, it is critical to understand how well the SmartNICs perform those respective tasks in order to determine whether the cost of offloading is worthwhile. In this work, we lay out a path to understand the important question of when and when not to offload to SmartNIC devices by way of a series of microbenchmarks.

97 MATHEMATICS AND COMPUTING↗

Hydrogen and its Vital Role in a Clean Energy Future

Large-scale, low -cost hydrogen production can enable an economically competitive, secure, and environmentally beneficial future energy system across multiple sectors. Furthermore, clean hydrogen can address specific sectors that are hard to decarbonize (e.g., heavy-duty trucking, load-following electricity, iron, steel, and cement) and can help the U.S. meet the net zero carbon goal by 2050. To achieve this goal, tens of millions of metric tons of clean, reliable, and affordable hydrogen will be needed annually1. In 2021, the Hydrogen Energy Earthshot was launched, and its goal is to reduce the cost of clean hydrogen to $1 per $1 kilogram in 1 decade (1 1 1) 2. One very promising pathway for large-scale hydrogen production is water splitting. Water splitting technologies range from commercial technologies such as electrolyzers to approaches that are at a much earlier stage of development, such as photoelectrochemical (PEC) and thermochemical (TCH) processes. All these water splitting pathways offer diverse benefits in energy storage, grid services, and cross-sector emissions reductions while taking advantage of the diverse domestic resources. However, critical materials-, component- and system-level challenges must be addressed to improve efficiency and durability and reduce cost. To address these barriers and move these promising and high impact technologies forward, the HydroGEN Advanced Water Splitting Materials (AWSM) and the H2 from the Next-generation of Electrolyzers of Water (H2NEW) consortia were formed and supported by the Department of Energy (DOE) EERE Hydrogen and Fuel Cell Technologies Office (HFTO). HydroGEN (https://www.energy.gov/eere/h2awsm/) consortium, established in 2016, is an Energy Materials Network (EMN) that aims to accelerate the materials R&D of low technology readiness level (TRL) advanced water splitting (AWS) technologies. The consortium comprises five core national laboratories and focuses on four early-stage AWS pathways: alkaline exchange membrane (AEM) electrolysis, proton conducting solid oxide electrolysis (p-SOEC), photoelectrochemical, and thermochemical water splitting. Liquid alkaline and PEM electrolyzers are already commercial and significant advancements in oxygen conducting solid oxide electrolysis cells (o-SOECs) have been realized. Yet, these systems are still too expensive and not sufficiently durable for wide-scale commercialization. To enable high-volume manufacturing of affordable, durable, efficient electrolyzers, H2NEW (https://h2new.energy.gov/), another multi-lab consortium, was established in 2020. This comprehensive, concerted effort is focused on overcoming barriers related to components and materials integration and scale-up to achieve performance, durability, with an initial focus to achieve $2/kg H2 by 2026.

AEM↗

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

Deep Learning Based Workflow for Accelerated Industrial X-Ray Computed Tomography

X-ray computed tomography (XCT) is an important tool for high-resolution non-destructive characterization of additively-manufactured metal components. XCT reconstructions of metal components may have beam hardening artifacts such as cupping and streaking which makes reliable detection of flaws and defects challenging. Furthermore, traditional workflows based on using analytic reconstruction algorithms require a large number of projections for accurate characterization - leading to longer measurement times and hindering the adoption of XCT for in-line inspections. In this paper, we introduce a new workflow based on the use of two neural networks to obtain high-quality accelerated reconstructions from sparse-view XCT scans of single material metal parts. The first network, implemented using fully-connected layers, helps reduce the impact of BH in the projection data without the need of any calibration or knowledge of the component material. The second network, a convolutional neural network, maps a low-quality analytic 3D reconstruction to a high-quality reconstruction. Using experimental data, we demonstrate that our method robustly generalizes across several alloys, and for a range of sparsity levels without any need for retraining the networks thereby enabling accurate and fast industrial XCT inspections.

Rahman, Obaid↗

Uncertainty quantification for deep learning in particle accelerator applications

With the advent of increased computational resources and improved algorithms, machine learning-based models are being increasingly applied to complex problems in particle accelerators. However, such data-driven models may provide overly confident predictions with unknown errors and uncertainties. For reliable deployment of machine learning models in high-regret and safety-critical systems such as particle accelerators, estimates of prediction uncertainty are needed along with accurate point predictions. In this investigation, we evaluate Bayesian neural networks (BNN) as an approach that can provide accurate predictions along with reliably quantified uncertainties for particle accelerator problems, and compare their performance with bootstrapped ensembles of neural networks. We select three accelerator setups for this evaluation: a storage ring, a photoinjector, and a linac. The problems span different data volumes and dimensionalities (e.g., scalar predictions as well as image outputs). It is found that BNN provide accurate predictions of the mean along with reliable estimates of predictive uncertainty across the test cases. In this vein, BNN may offer an attractive alternative to deterministic deep learning tools to generate accurate predictions with quantified uncertainties in particle accelerator applications.

43 PARTICLE ACCELERATORS↗

A Study on the Impact of Temperature-Dependent Ferroelectric Switching Behavior in 3D Memory Architecture

The flourishing development of neural networks that require exponentially growing amounts of data has presented an elevated demand for memory footprint. To address this, researchers have been exploring hardware accelerators with innovative memory architectures like 3D memory. These 3D memory architectures offer enhanced storage capacity and processing capabilities, at a cost of rising on-chip temperature during operation. Hafnium Zirconium Oxide (HZO) based Ferroelectric Random Access Memory (FeRAM) is a promising nonvolatile memory candidate in neural network hardware accelerators for its outstanding write performance and reliability. However, its implementation in the architecture regarding the temperature-dependent ferroelectric switching behavior has not been well studied. In this work, we study the thermal impacts on polarization switching through experimental devices and simulation results. We conduct the circuit and architecture-level simulations to showcase that one can exploit this temperature rise to reduce FeRAM's write voltage and write energy due to its unique temperature-activated polarization switching mechanisms. As the on-chip temperature increases to 351K (ambient temperature at 300K) due to neural network workloads, the access energy per bit can be reduced by 27.6% when a dynamic write voltage is applied.

36 MATERIALS SCIENCE↗

Accelerating Finite-Temperature Kohn-Sham Density Functional Theory with Deep Neural Networks

We present a numerical modeling workflow based on machine learning (ML) which reproduces the total energies produced by Kohn-Sham density functional theory (DFT) at finite electronic temperature to within chemical accuracy at negligible computational cost. Based on deep neural networks, our workflow yields the local density of states (LDOS) for a given atomic configuration. From the LDOS, spatially-resolved, energy-resolved, and integrated quantities can be calculated, including the DFT total free energy, which serves as the Born-Oppenheimer potential energy surface for the atoms. We demonstrate the efficacy of this approach for both solid and liquid metals and compare results between independent and unified machine-learning models for solid and liquid aluminum. Our machine-learning density functional theory framework opens up the path towards multiscale materials modeling for matter under ambient and extreme conditions at a computational scale and cost that is unattainable with current algorithms.

97 MATHEMATICS AND COMPUTING↗

From IMPEL to Impact: Lessons Learned in Accelerating Innovative Building Technologies

The built environment is a complex ecosystem of social institutions and physical infrastructures. Innovation and entrepreneurship in the building industry are critical levers for market transformation toward equitable climate action. However, climate tech innovation for the built environment is not moving fast enough for global needs, and it lacks fundamental diversity, leading to inequitable outcomes. IMPEL (Incubating Market-propelled Entrepreneurial-mindset at the Labs and Beyond) - a U.S. Department of Energy incubator–addresses these critical issues. Over five years, IMPEL has enabled 250 innovators, including 55% women and diverse founders, to accelerate their buildings and clean energy technologies towards market and climate impact. IMPEL provides access to strategic mentoring and coaching, carbon tools training, testbeds, and powerful public-private pipelines, including industry demonstrations, non-dilutive grants, and venture capital networks. The IMPEL innovation ecosystem has accelerated the pace of innovation and market adoption of building decarbonization technologies. In this paper, leverage the IMPEL stakeholder ecosystem - from innovators to investors and product industry to policymakers - to analyze the critical barriers to decarbonization still encountered in the building industry. We study the IMPEL approach and highlight lessons learned that benefit young businesses pursuing innovative building and building-edge energy technologies to develop new ideas and products. Finally, we propose a ‘market forming’ framework to improve the quality and efficiency of the entrepreneurial ecosystem in the building industry. This framework could scale vetted technologies and the participation of diverse founders to de-risk the climate tech

Singh, Reshma↗

Generic Multi-Layer Perceptron Inference Accelerator on FPGA (vneuron) v1.0

We have designed and implemented a neural network inference compute engine (vneuron) that can be deployed in the fabric of any FPGA without using special hardware accelerator primitive. The "vneuron" is purely written in verilog, and supports scalable neural network structure with fully connected layers and ReLU activation ( Multi-Layer Perceptron architecture) with 16 bits of precision. We have demonstrated it on an Xilinx Artix 7 FPGA for a 16-input, 8-output MLP with 3 layer, 1600 parameters. It takes 40 DSP48E and 40 BRAM18, and takes 131 clock cycles for computing (1048 ns when clocked at 125MHz). We include PyTorch quantization from a given floating point model, and provide behavioral verification simulation in the disclosed software package.

Du, Qiang↗