Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Acceleration of Graph Neural Network-Based Prediction Models in Chemistry via Co-Design Optimization on Intelligence Processing Units

Atomic structure prediction and associated property calculations are the bedrock of chemical physics. Since high-fidelity ab initio modeling techniques for computing the structure and properties can be prohibitively expensive, this motivates the development of machine-learning (ML) models that make these predictions more efficiently. Training graph neural networks over large atomistic databases introduces unique computational challenges such as the need to process millions of small graphs with variable size and support communication patterns that are distinct from learning over large graphs such as social networks. We demonstrate a novel hardware-software co-design approach to scale up the training of atomistic graph neural networks (GNN) for structure and property prediction. First, to eliminate redundant computation and memory associated with alternative padding techniques and to improve throughput via minimizing communication, we formulate the effective coalescing of the batches of variable-size atomistic graphs as the bin packing problem and introduce a hardware-agnostic algorithm to pack these batches. In addition, we propose hardware-specific optimizations including a planner and vectorization for the gather-scatter operations targeted for Graphcore’s Intelligence Processing Unit (IPU), as well as model-specific optimizations such as merged communication collectives and optimized softplus. Putting these all together, we demonstrate the effectiveness of the proposed co-design approach by providing an implementation of a well-established atomistic GNN on the Graphcore IPUs. We evaluate the training performance on multiple atomistic graph databases with varying degrees of graph counts, sizes and sparsity. Here, we demonstrate that such a co-design approach can reduce the training time of atomistic GNNs and can improve the performance by up to 1.5× compared to the baseline implementation of the model on the IPUs. Additionally, we compare our IPU implementation with a Nvidia GPU-based implementation and show that our atomistic GNN implementation on the IPUs can run 1.8× faster on average compared to the execution time on the GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ChemComp: A Compilation Framework for Computing with Chemical Reaction Networks

The acceleration of scientific computation, data analytics, and artificial intelligence is driving a surge in computational requirements. Yet, state-of-the-art high-performance computing systems are approaching physical limitations that impede further significant improvements in energy efficiency. As we move towards post-exascale computing systems, innovative approaches are necessary to overcome this barrier in power consumption. Novel analog and hybrid digital-analog architectures hold promise for enhancing energy efficiency by several orders of magnitude. Biochemical computation stands out among the various solutions being explored due to its potential to enable new classes of devices with immense computational capabilities. These devices can capitalize on the inherent efficacy of biological cells in solving optimization problems and are scalable through increasing reaction system size or vessel capacity, potentially satisfying scientific computing's high-performance requirements. Nonetheless, several theoretical and practical limitations persist, including problem formulation and mapping to chemical reaction networks (CRNs) and implementation of actual CRN devices. In this paper, we propose a framework for biochemical computation using systems chemistry. We present the initial components of our approach: an abstract chemical reaction dialect implemented as a multi-level intermediate representation (MLIR) compiler extension and a pathway to represent mathematical problems with CRNs. To showcase the potential of this approach, we emulate a simplified chemical reservoir device. This work lays the groundwork for leveraging chemistry's computing potential in creating energy-efficient, high-performance computing systems tailored to contemporary computational needs.

artificial intelligence↗

TorchBraid: High-Performance Layer-Parallel Training of Deep Neural Networks with MPI and GPU Acceleration

TorchBraid is a high-performance implementation of layer-parallel training for deep neural networks (DNNs) supporting MPI-based parallelism and GPU acceleration. Layer-parallel training has been developed to overcome the serialization inherent in forward and backward propagation of DNNs that limits utilization of computational resources in the strong scaling limit. To achieve this, TorchBraid integrates the PyTorch neural network framework with the state-of-the-art XBraid time-parallel library. Furthermore, this article presents the use and performance of TorchBraid, in addition to solutions for overcoming the algorithmic challenges inherent in combining automatic differentiation with layer-parallel. Results are presented with and without GPU acceleration for the Tiny ImageNet and MNIST image classification data sets, as well as recurrent neural networks. Overall, TorchBraid enables fast training of DNNs, both in a strong and weak scaling context. In addition to the TorchBraid software, several new advances in applying layer-parallel algorithms are detailed. Integration of layer-parallel with data-parallel algorithms is presented for the first time, showing the computational advantages of the combination. Standard deep learning techniques, like batch-normalization, are developed for layer-parallel training. Finally, a new approach combining layer-parallel with spatial coarsening in order to accelerate training for 3D image classification shows roughly a 10× speedup over serial execution.

Layer-parallel↗

Phase Space Reconstruction from Accelerator Beam Measurements Using Neural Networks and Differentiable Simulations

Characterizing the phase space distribution of particle beams in accelerators is a central part of accelerator understanding and performance optimization. However, conventional reconstruction-based techniques either use simplifying assumptions or require specialized diagnostics to infer high-dimensional (> $2D$) beam properties. In this Letter, we introduce a general-purpose algorithm that combines neural networks with differentiable particle tracking to efficiently reconstruct high-dimensional phase space distributions without using specialized beam diagnostics or beam manipulations. Furthermore, we demonstrate that our algorithm accurately reconstructs detailed 4D phase space distributions with corresponding confidence intervals in both simulation and experiment using a single focusing quadrupole and diagnostic screen. This technique allows for the measurement of multiple correlated phase spaces simultaneously, which will enable simplified 6D phase space distribution reconstructions in the future.

47 OTHER INSTRUMENTATION↗

Laser Wakefield Accelerator modelling with Variational Neural Networks

A machine learning model was created to predict the electron spectrum generated by a GeV-class laser wakefield accelerator. The model was constructed from variational convolutional neural networks, which mapped the results of secondary laser and plasma diagnostics to the generated electron spectrum. An ensemble of trained networks was used to predict the electron spectrum and to provide an estimation of the uncertainty of that prediction. It is anticipated that this approach will be useful for inferring the electron spectrum prior to undergoing any process that can alter or destroy the beam. In addition, the model provides insight into the scaling of electron beam properties due to stochastic fluctuations in the laser energy and plasma electron density.

43 PARTICLE ACCELERATORS↗

Neuromorphic Accelerator for Deep Spiking Neural Networks with NVM Crossbar Arrays

In this paper, we present a scalable digital hardware accelerator based on non-volatile memory arrays capable of realizing deep convolutional spiking neural networks (SNNs). Our design studies are conducted using a compact model for spin-transfer torque random access memory (STT-RAM) devices. Large networks are realized by tiling multiple cores which communicate by transmitting spike packets via an on-chip routing network. Compared to an equivalent SRAM based core design, we show that the STT-RAM based design achieves nearly 15X higher GSOPS (Synaptic Operations per Second) per Watt per mm 2 making it a promising platform for realizing systems with significant area and power limitations.

Kulkarni, Shruti↗

Control Systems Design for STS Accelerator

The Second Target Station (STS) Project will expand the capabilities of the existing Spallation Neutron Source (SNS), with a suite of neutron instruments optimized for long wavelengths. A new accelerator transport line will be built to deliver one out of four SNS pulses to the new target station. The Integrated Control Systems (ICS) will provide remote control, monitoring, OPI, alarms, and archivers for the accelerator systems, such as magnets power supply, vacuum devices, and beam instrumentation. The ICS will upgrade the existing Linac LLRF controls to allow independent operation of the FTS and STS and support different power levels of the FTS and STS proton beam. The ICS accelerator controls are in the phase of preliminary design for the control systems of magnet power supply, vacuum, LLRF, Timing, Machine protection system (MPS), and computing and machine network. The accelerator control systems build upon the existing SNS Machine Control systems, use the SNS standard hardware and EPICS software, and take full advantage of the performance gains delivered by the PPU Project at SNS.

Yan, Jay↗

Differentiable Neural Architecture, Mixed Precision and Accelerator Co-Search

Quantization, effective Neural Network architecture, and efficient accelerator hardware are three important design paradigms to maximize accuracy and efficiency. Mixed Precision Quantization is a process of assigning different precision to different Neural Network layers for optimized inference. Neural Architecture Search (NAS) is a process of automatically designing the neural network for a task and can also be extended to search for the precision of each weight and activation matrix. In this paper, we develop the following three methods: (i) Fast Differentiable Hardware-aware Mixed Precision Quantization Search method to find optimal precision, (ii) Joint Differentiable hardware-aware Architecture and Mixed Precision Quantization Co-search, (iii) Joint Accelerator, Architecture, and Precision triple co-search to find best possibilities in all the three worlds. We demonstrate the effectiveness of our proposed methods targeting Bitfusion accelerator by searching mixed precision models on MobilenetV2. We achieve better accuracy-latency trade-off models than the manually designed and previously proposed search methods.

97 MATHEMATICS AND COMPUTING↗

Neural Networks for Nuclear Reactions in MAESTROeX

We demonstrate the use of neural networks to accelerate the reaction steps in the MAESTROeX stellar hydrodynamics code. A traditional MAESTROeX simulation uses a stiff ODE integrator for the reactions; here, we employ a ResNet architecture and describe details relating to the architecture, training, and validation of our networks. Our customized approach includes options for the form of the loss functions, a demonstration that the use of parallel neural networks leads to increased accuracy, and a description of a perturbational approach in the training step that robustifies the model. We test our approach on millimeter-scale flames using a single-step, 3-isotope network describing the first stages of carbon fusion occurring in Type Ia supernovae. We train the neural networks using simulation data from a standard MAESTROeX simulation, and show that the resulting model can be effectively applied to different flame configurations. This work lays the groundwork for more complex networks, and iterative time-integration strategies that can leverage the efficiency of the neural networks.

79 ASTRONOMY AND ASTROPHYSICS↗

Integrated reactor architecture of conductive network and catalytic nodes to accelerate polysulfide conversion for durable and high-loading Li-S batteries

The development of carbon-based heterogeneous framework host with synergistic catalytic and conductive effects for sulfur cathode is a promising strategy to realize high performance lithium sulfur batteries (LSBs). Here, an integrated reactor architecture with defective carbon nodes (IRA-DC) is designed for serving as high-loading (92.4 wt%) sulfur host. The hierarchical porous IRA-DC consists of untangled conductive carbon nanotube network and Co/N co-doped catalytic nodes with high dispersity. Therein the optimization of electric field distribution and homogenization of adsorption-catalysis sites offer the multi-electron conversion reaction of polysulfides with excellent kinetics and stability. The resultant IRA-DC/S cathode enables a high areal capacity of 8.86 mAh cm -2 under ultra-high sulfur loading (13.1 mg cm -2 ) and lean electrolyte (8 μL mg sulfur -1 ). It also displays a long-term cycling performance (1200 cycles at 1 C) and ultrahigh rate performance up to 20 C (with a capacity of 473.6 mAh g -1 ). In conclusion, this work provides an electrode building strategy by optimizing the environments of heterogeneous electrocatalysis and micro electric field to activate the polysulfide conversion efficiency and utilization of high-loading sulfur in monolithic sulfur-carbon cathodes.

25 ENERGY STORAGE↗

Disentangling Beam Losses in The Fermilab Main Injector Enclosure Using Real-Time Edge AI

The Fermilab Main Injector enclosure houses two accelerators, the Main Injector and Recycler Ring. During normal operation, high intensity proton beams exist simultaneously in both. The two accelerators share the same beam loss monitors (BLM) and monitoring system. Deciphering the origin of any of the 260 BLM readings is often difficult. The (Accelerator) Real-time Edge AI for Distributed Systems project, or READS, has developed an AI/ML model, and implemented it on fast FPGA hardware, that disentangles mixed beam losses and attributes probabilities to each BLM as to which machine(s) the loss originated from in real-time. The model inferences are then streamed to the Fermilab accelerator controls network (ACNET) where they are available for operators and experts alike to aid in tuning the machines.

43 PARTICLE ACCELERATORS↗

GUI Control System for the Mu2e Electrostatic Septum High Voltage at Fermilab

The Mu2e Experiment has stringent beam structure requirements; namely, its proton bunches with a time structure of 1.7 $\mu$s in the Fermilab Delivery Ring. This beam structure will be delivered using the Fermilab 8-GeV Booster, the 8-GeV Recycler Ring, and the Delivery Ring. The 1.7-$\mu$s period of the Delivery Ring will generate the required beam structure by means of a third order resonant extraction system operating on a single circulating bunch. The electrostatic septum (ESS) for this system is particularly challenging, requiring mechanical precision in a ultra high vacuum of 1 x 10$^-8$ Torr to generate 100 kV across 15 mm. This paper describes a graphical user interface that has been developed to automate the conditioning and commissioning process for the electrostatic septa. It is based on an interface to the Fermilab ACNET system using the ACSys Python Data Pool Manager (DPM) Client produced and maintained by Fermilab Accelerator Controls. Network interfacing between data pool managers made by the application and ACNET devices introduce an inherent (approximately 1 s) latency in throughput of the readouts. This delay is utilized to process and graph incoming data events of devices crucial to conditioning of a electrostatic septum (ESS). 'Ramping' and 'Monitoring' modes adjust settings of the power supply based on internal logic to efficaciously increase and maintain the high voltage (HV) in the ESS, easing the voltage setting on incidence of sparking or other possibly damaging events. A timestamped log file is produced as the application runs.

43 PARTICLE ACCELERATORS↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Beam Synchronous for the Rest of Us!

Fermilab’s Tevatron Clock (TCLK) infrastructure has been an integral part of the accelerator control network since the 1980’s. This 10MHz Manchester encoded protocol has enabled flexible, real-time event distribution for thousands of devices connected to the timing network with a high degree of reliability. Forthcoming upgrades to the Fermilab complex (PIP-II, LBNF, ACORN) necessitate higher levels of precision to maintain inter-bunch timing for Instrumentation and Control purposes. This presents as an opportunity to refine the event distribution protocol for tighter synchronization between machines, experiments, and eventually far-site operations. This paper outlines a method by which beam-synchronous events may be distributed through asynchronous serial protocols via integration with local LLRF and global PPS reference signals. This method is ideal for synchrotron machines with aggressive frequency sweeps (such as Fermilab's 38~53MHz Booster) and allows for precision timing to be maintained across machines without specialized hardware.

43 PARTICLE ACCELERATORS↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗

Analog In-Memory Computing for the Synthetic Aperture Radar Polar Format Algorithm

As the utility of synthetic aperture radar (SAR) systems increases in autonomous vehicles, satellites, and other power- and space-constrained edge applications, there is a growing need for processors that can form SAR images at low power. In recent years, analog in-memory compute (AIMC) has shown immense promise for accelerating neural networks and other matrix-vector multiplication (MVM) heavy workloads at the edge. Here, in this work, we examine how the polar format algorithm (PFA), a popular SAR image formation algorithm, can be mapped to these AIMC systems. The PFA maps readily onto analog MVMs because it primarily consists of two linear operations: interpolation of frequency-domain data to a Cartesian grid, followed by a 2-D Fourier transform. This work presents two approaches to map the interpolation operation onto MVMs in analog hardware: a chirp transform and a modified form of sinc interpolation. These mappings introduce algorithmic errors, and their effect on the quality of SAR image formation is examined, both quantitatively and qualitatively. In addition, the impact of errors introduced by the analog hardware is explored to determine which approach is optimal under varying assumptions about the underlying analog memory devices and circuits.

Analog computing↗