Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

OpenSNAPI: Toward a Unified API for SmartNICs

The end of Moore’s Law and Dennard Scaling has produced a renaissance in the field of computer architecture. Unable to continue leveraging silicon-level processor improvements to further enhance performance and scalability, system architects have been forced to explore other options. In this new era of heterogeneous architectures and hardware/software codesign, a new class of devices known as “accelerators” has emerged. Independently designed for optimized execution of distinct workloads, these devices have proven critical to the continued advancement of application performance. SmartNICs, accelerator devices integrated with a network controller, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. In this talk, we introduce the motivation underlying acceleration, explore the fundamentals of SmartNICs, and discuss traditional use cases. We also detail our initial efforts to investigate the feasibility and benefits of SmartNICs as general-purpose accelerators. We present the OpenSNAPI project created to define a uniform application programming interface (API) for this emerging class of devices. Finally, we provide a brief tutorial regarding development of SmartNIC-accelerated applications on Los Alamos National Laboratory’s SmartNIC-enabled platforms.

97 MATHEMATICS AND COMPUTING↗

Functional and transcriptional characterization of complex neuronal co-cultures

Brain-on-a-chip systems are designed to simulate brain activity using traditional in vitro cell culture on an engineered platform. It is a noninvasive tool to screen new drugs, evaluate toxicants, and elucidate disease mechanisms. However, successful recapitulation of brain function on these systems is dependent on the complexity of the cell culture. In this study, we increased cellular complexity of traditional (simple) neuronal cultures by co-culturing with astrocytes and oligodendrocyte precursor cells (complex culture). We evaluated and compared neuronal activity (e.g., network formation and maturation), cellular composition in long-term culture, and the transcriptome of the two cultures. Compared to simple cultures, neurons from complex co-cultures exhibited earlier synapse and network development and maturation, which was supported by localized synaptophysin expression, up-regulation of genes involved in mature neuronal processes, and synchronized neural network activity. Also, mature oligodendrocytes and reactive astrocytes were only detected in complex cultures upon transcriptomic analysis of age-matched cultures. Functionally, the GABA antagonist bicuculline had a greater influence on bursting activity in complex versus simple cultures. Collectively, the cellular complexity of brain-on-a-chip systems intrinsically develops cell type-specific phenotypes relevant to the brain while accelerating the maturation of neuronal networks, important features underdeveloped in traditional cultures.

59 BASIC BIOLOGICAL SCIENCES↗

P38 heterogeneous multi-tiled system with support for message queues (MoSAIC) v0.1

The proposed system is written in the hardware description language (HDL) verilog targeting an FPGA board. It is intended as a testbed to explore architecture tradeoffs in multi-tiled heterogeneous architectures. Although we target FPGAs, the system can be implemented as a monolithic SoC or a package comprised of many chiplets that are interconnected in the same package using a NoC. The proposed NoC is lightweight and follows an axi-lite interface. The endpoints of the NoC are a heterogeneous mix of "tiles" as endpoints that are general purpose processors, fixed function accelerators, and programmable accelerators. We assume that the network interfaces for the NoC endpoints are all addressable in a global name-space in that they represent an address range (for memory addresses) or a range of unique identifiers that are associated with each individual tile. This makes the functionality abstract from the standpoint of the NoC design details. Message queues offer a direct inter-processor interface between peer general purpose cores and diverse accelerators that comprise an SoC. Although they share the same NoC infrastructure for inter-tile communication within an SoC or SiP, the hardware message queues bypass the memory hierarchy and thus do not pollute the memory state or invoke the cache coherence mechanism.

Gonzalez, LouisaPatricia↗

Residual-based error correction for neural operator accelerated infinite-dimensional Bayesian inverse problems

We explore using neural operators, or neural network representations of nonlinear maps between function spaces, to accelerate infinite-dimensional Bayesian inverse problems (BIPs) with models governed by nonlinear parametric partial differential equations (PDEs). Neural operators have gained significant attention in recent years for their ability to approximate the parameter-to-solution maps defined by PDEs using as training data solutions of PDEs at a limited number of parameter samples. The computational cost of BIPs can be drastically reduced if the large number of PDE solves required for posterior characterization are replaced with evaluations of trained neural operators. However, reducing error in the resulting BIP solutions via reducing the approximation error of the neural operators in training can be challenging and unreliable. We provide an a priori error bound result that implies certain BIPs can be ill-conditioned to the approximation error of neural operators, thus leading to inaccessible accuracy requirements in training. To reliably deploy neural operators in BIPs, we consider a strategy for enhancing the performance of neural operators: correcting the prediction of a trained neural operator by solving a linear variational problem based on the PDE residual. We show that a trained neural operator with error correction can achieve a quadratic reduction of its approximation error, all while retaining substantial computational speedups of posterior sampling when models are governed by highly nonlinear PDEs. The strategy is applied to two numerical examples of BIPs based on a nonlinear reaction–diffusion problem and deformation of hyperelastic materials. We demonstrate that posterior representations of the two BIPs produced using trained neural operators are greatly and consistently enhanced by error correction.

97 MATHEMATICS AND COMPUTING↗

Accelerating charge estimation in molecular dynamics simulations using physics-informed neural networks: corrosion applications

Molecular Dynamics (MD) simulations are used to understand the effects of corrosion on metallic materials in salt brine. Reactive force fields in classical MD enable accurate modeling of bond formation and breakage in the aqueous medium and at the metal-electrolyte interface, while also facilitating dynamic partial charge equilibration. However, MD simulations are computationally intensive and unsuitable for modeling the long time scales characteristic of corrosive phenomena. To address this, we develop reduced-order machine learning models that provide accurate and efficient predictions of charge density in corrosive environments. Specifically, we use Long Short-Term Memory (LSTM) networks to forecast charge density evolution based on atomic environments represented by Smooth Overlap of Atomic Positions (SOAP) descriptors. A physics-informed loss function enforces charge neutrality and electronegativity equivalence. The atomic charges predicted by the deep learning model trained on this work were obtained two orders of magnitude faster than those from molecular dynamics (MD) simulations, with an error of less than 3% compared to the MD-obtained charges, even in extrapolative scenarios, while adhering to physical constraints. This demonstrates the excellent accuracy, computational efficiency, and validity of the developed model. Lastly, even though developed for corrosion, these protocols are formulated in a phenomenon-agnostic manner, allowing application to various variable-charge interatomic potentials and related fields.

Atomistic models↗

Accelerating massively parallel hemodynamic models of coarctation of the aorta using neural networks

Comorbidities such as anemia or hypertension and physiological factors related to exertion can influence a patient’s hemodynamics and increase the severity of many cardiovascular diseases. Observing and quantifying associations between these factors and hemodynamics can be difficult due to the multitude of co-existing conditions and blood flow parameters in real patient data. Machine learning-driven, physics-based simulations provide a means to understand how potentially correlated conditions may affect a particular patient. Here, we use a combination of machine learning and massively parallel computing to predict the effects of physiological factors on hemodynamics in patients with coarctation of the aorta. We first validated blood flow simulations against in vitro measurements in 3D-printed phantoms representing the patient’s vasculature. We then investigated the effects of varying the degree of stenosis, blood flow rate, and viscosity on two diagnostic metrics – pressure gradient across the stenosis (ΔP) and wall shear stress (WSS) - by performing the largest simulation study to date of coarctation of the aorta (over 70 million compute hours). Using machine learning models trained on data from the simulations and validated on two independent datasets, we developed a framework to identify the minimal training set required to build a predictive model on a per-patient basis. We then used this model to accurately predict ΔP (mean absolute error within 1.18 mmHg) and WSS (mean absolute error within 0.99 Pa) for patients with this disease.

59 BASIC BIOLOGICAL SCIENCES↗

High-Speed Ionic Synaptic Memory Based on 2D Titanium Carbide MXene

Synaptic devices with linear high-speed switching can accelerate learning in artificial neural networks (ANNs) embodied in hardware. Conventional resistive memories however suffer from high write noise and asymmetric conductance tuning, preventing parallel programming of ANN arrays. Electrochemical random-access memories (ECRAMs), where resistive switching occurs by ion insertion into a redox-active channel, aim to address these challenges due to their linear switching and low noise. ECRAMs using 2D materials and metal oxides however suffer from slow ion kinetics, whereas organic ECRAMs enable high-speed operation but face challenges toward on-chip integration due to poor temperature stability of polymers. Here, ECRAMs using 2D titanium carbide (Ti 3 C 2 T x ) MXene that combine the high speed of organics and the integration compatibility of inorganic materials in a single high-performance device are demonstrated. These ECRAMs combine the speed, linearity, write noise, switching energy, and endurance metrics essential for parallel acceleration of ANNs, and importantly, they are stable after heat treatment needed for back-end-of-line integration with Si electronics. The high speed and performance of these ECRAMs introduces MXenes, a large family of 2D carbides and nitrides with more than 30 stoichiometric compositions synthesized to date, as promising candidates for devices operating at the nexus of electrochemistry and electronics.

2D materials↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication.

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Finally, our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

42 ENGINEERING↗

Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The efficiency of these accelerators comes from employing optimized dataflow (i.e., spatial/temporal partitioning of data across the PEs and fine-grained scheduling) strategies to optimize data reuse. The focus of this work is to evaluate these accelerator architectures using a tiled general matrix-matrix multiplication (GEMM) kernel. To do so, we develop a framework that finds optimized mappings (dataflow and tile sizes) for a tiled GEMM for a given spatial accelerator and workload combination, leveraging an analytical cost model for runtime and energy. Our evaluations over five spatial accelerators demonstrate that the tiled GEMM mappings systematically generated by our framework achieve high performance on various GEMM workloads and accelerators.

43 PARTICLE ACCELERATORS↗

Comparative study of machine learning techniques for post-combustion carbon capture systems

Computational analysis of countercurrent flows in packed absorption columns, often used in solvent-based post-combustion carbon capture systems (CCSs), is challenging. Typically, computational fluid dynamics (CFD) approaches are used to simulate the interactions between a solvent, gas, and column's packing geometry while accounting for the thermodynamics, kinetics, heat, and mass transfer effects of the absorption process. These simulations can then be used explain a column's hydrodynamic characteristics and evaluate its CO 2 -capture efficiency. However, these approaches are computationally expensive, making it difficult to evaluate numerous designs and operating conditions to improve efficiency at industrial scales. In this work, we comprehensively explore the application of statistical ML methods, convolutional neural networks (CNNs), and graph neural networks (GNNs) to aid and accelerate the scale-up and design optimization of solvent-based post-combustion CCSs. We apply these methods to CFD datasets of countercurrent flows in absorption columns with structured packings characterized by several geometric parameters. We train models to use these parameters, inlet velocity conditions, and other model-specific representations of the column to estimate key determinants of CO 2 -capture efficiency without having to simulate additional CFD datasets. We also evaluate the impact of different input types on the accuracy and generalizability of each model. We discuss the strengths and limitations of each approach to further elucidate the role of CNNs, GNNs, and other machine learning approaches for CO 2 -capture property prediction and design optimization.

97 MATHEMATICS AND COMPUTING↗

DD-AADL

Data Driven Anderson Acceleration for Physics Informed Neural Networks to solve high dimensional partial differential equations

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗

A Path Toward Understanding the Performance Capabilities of SmartNIC Devices [Slides]

SmartNICs, accelerator devices integrated with a network, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. Naturally, there is performance overhead associated with offloading tasks from a host CPU to a SmartNIC. As SmartNICs become more versatile and are used for a wider range of computational tasks, it is critical to understand how well the SmartNICs perform those respective tasks in order to determine whether the cost of offloading is worthwhile. In this work, we lay out a path to understand the important question of when and when not to offload to SmartNIC devices by way of a series of microbenchmarks.

97 MATHEMATICS AND COMPUTING↗

Hydrogen and its Vital Role in a Clean Energy Future

Large-scale, low -cost hydrogen production can enable an economically competitive, secure, and environmentally beneficial future energy system across multiple sectors. Furthermore, clean hydrogen can address specific sectors that are hard to decarbonize (e.g., heavy-duty trucking, load-following electricity, iron, steel, and cement) and can help the U.S. meet the net zero carbon goal by 2050. To achieve this goal, tens of millions of metric tons of clean, reliable, and affordable hydrogen will be needed annually1. In 2021, the Hydrogen Energy Earthshot was launched, and its goal is to reduce the cost of clean hydrogen to $1 per $1 kilogram in 1 decade (1 1 1) 2. One very promising pathway for large-scale hydrogen production is water splitting. Water splitting technologies range from commercial technologies such as electrolyzers to approaches that are at a much earlier stage of development, such as photoelectrochemical (PEC) and thermochemical (TCH) processes. All these water splitting pathways offer diverse benefits in energy storage, grid services, and cross-sector emissions reductions while taking advantage of the diverse domestic resources. However, critical materials-, component- and system-level challenges must be addressed to improve efficiency and durability and reduce cost. To address these barriers and move these promising and high impact technologies forward, the HydroGEN Advanced Water Splitting Materials (AWSM) and the H2 from the Next-generation of Electrolyzers of Water (H2NEW) consortia were formed and supported by the Department of Energy (DOE) EERE Hydrogen and Fuel Cell Technologies Office (HFTO). HydroGEN (https://www.energy.gov/eere/h2awsm/) consortium, established in 2016, is an Energy Materials Network (EMN) that aims to accelerate the materials R&D of low technology readiness level (TRL) advanced water splitting (AWS) technologies. The consortium comprises five core national laboratories and focuses on four early-stage AWS pathways: alkaline exchange membrane (AEM) electrolysis, proton conducting solid oxide electrolysis (p-SOEC), photoelectrochemical, and thermochemical water splitting. Liquid alkaline and PEM electrolyzers are already commercial and significant advancements in oxygen conducting solid oxide electrolysis cells (o-SOECs) have been realized. Yet, these systems are still too expensive and not sufficiently durable for wide-scale commercialization. To enable high-volume manufacturing of affordable, durable, efficient electrolyzers, H2NEW (https://h2new.energy.gov/), another multi-lab consortium, was established in 2020. This comprehensive, concerted effort is focused on overcoming barriers related to components and materials integration and scale-up to achieve performance, durability, with an initial focus to achieve $2/kg H2 by 2026.

AEM↗

Moving small files in a networked environment

Globally distributed computing infrastructures, such as clouds and supercomputers, are currently used to manage data that is generated with an unprecedented speed from a variety of resources. Coping with this trend, the volume of data exchanged across distant sites increases substantially. To accelerate data transfer, high-speed networks are provided to connect remote sites. Most existing data movement solutions are optimized for moving large files. However, it is still challenging to transfer a large number of small files across networks. This disadvantage not only lowers data transfer performance, but also decreases overall system utilization. Here, we identify that moving small files is mainly constrained by degraded file system throughput, not just network performance as might be suspected. We have built a data transfer pipeline model to analyze the impact of small network I/O and storage I/O on data movement. Extending one of the widely used open source data movement solutions, GridFTP, we demonstrate several appropriate engineering approaches that mitigate the bottleneck and increase data transfer efficiency. We show optimizations that improve data transfer performance more than 5 times. In comparison to existing solutions, our approaches can save a significant amount of system resources for moving lots of small files.

97 MATHEMATICS AND COMPUTING↗

Deep Learning Based Workflow for Accelerated Industrial X-Ray Computed Tomography

X-ray computed tomography (XCT) is an important tool for high-resolution non-destructive characterization of additively-manufactured metal components. XCT reconstructions of metal components may have beam hardening artifacts such as cupping and streaking which makes reliable detection of flaws and defects challenging. Furthermore, traditional workflows based on using analytic reconstruction algorithms require a large number of projections for accurate characterization - leading to longer measurement times and hindering the adoption of XCT for in-line inspections. In this paper, we introduce a new workflow based on the use of two neural networks to obtain high-quality accelerated reconstructions from sparse-view XCT scans of single material metal parts. The first network, implemented using fully-connected layers, helps reduce the impact of BH in the projection data without the need of any calibration or knowledge of the component material. The second network, a convolutional neural network, maps a low-quality analytic 3D reconstruction to a high-quality reconstruction. Using experimental data, we demonstrate that our method robustly generalizes across several alloys, and for a range of sparsity levels without any need for retraining the networks thereby enabling accurate and fast industrial XCT inspections.

Rahman, Obaid↗

Uncertainty quantification for deep learning in particle accelerator applications

With the advent of increased computational resources and improved algorithms, machine learning-based models are being increasingly applied to complex problems in particle accelerators. However, such data-driven models may provide overly confident predictions with unknown errors and uncertainties. For reliable deployment of machine learning models in high-regret and safety-critical systems such as particle accelerators, estimates of prediction uncertainty are needed along with accurate point predictions. In this investigation, we evaluate Bayesian neural networks (BNN) as an approach that can provide accurate predictions along with reliably quantified uncertainties for particle accelerator problems, and compare their performance with bootstrapped ensembles of neural networks. We select three accelerator setups for this evaluation: a storage ring, a photoinjector, and a linac. The problems span different data volumes and dimensionalities (e.g., scalar predictions as well as image outputs). It is found that BNN provide accurate predictions of the mean along with reliable estimates of predictive uncertainty across the test cases. In this vein, BNN may offer an attractive alternative to deterministic deep learning tools to generate accurate predictions with quantified uncertainties in particle accelerator applications.

43 PARTICLE ACCELERATORS↗

A Study on the Impact of Temperature-Dependent Ferroelectric Switching Behavior in 3D Memory Architecture

The flourishing development of neural networks that require exponentially growing amounts of data has presented an elevated demand for memory footprint. To address this, researchers have been exploring hardware accelerators with innovative memory architectures like 3D memory. These 3D memory architectures offer enhanced storage capacity and processing capabilities, at a cost of rising on-chip temperature during operation. Hafnium Zirconium Oxide (HZO) based Ferroelectric Random Access Memory (FeRAM) is a promising nonvolatile memory candidate in neural network hardware accelerators for its outstanding write performance and reliability. However, its implementation in the architecture regarding the temperature-dependent ferroelectric switching behavior has not been well studied. In this work, we study the thermal impacts on polarization switching through experimental devices and simulation results. We conduct the circuit and architecture-level simulations to showcase that one can exploit this temperature rise to reduce FeRAM's write voltage and write energy due to its unique temperature-activated polarization switching mechanisms. As the on-chip temperature increases to 351K (ambient temperature at 300K) due to neural network workloads, the access energy per bit can be reduced by 27.6% when a dynamic write voltage is applied.

36 MATERIALS SCIENCE↗