Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Accelerating network layouts using graph neural networks

Graph layout algorithms used in network visualization represent the first and the most widely used tool to unveil the inner structure and the behavior of complex networks. Current network visualization software relies on the force-directed layout (FDL) algorithm, whose high computational complexity makes the visualization of large real networks computationally prohibitive and traps large graphs into high energy configurations, resulting in hard-to-interpret “hairball” layouts. Here we use Graph Neural Networks (GNN) to accelerate FDL, showing that deep learning can address both limitations of FDL: it offers a 10 to 100 fold improvement in speed while also yielding layouts which are more informative. We analytically derive the speedup offered by GNN, relating it to the number of outliers in the eigenspectrum of the adjacency matrix, predicting that GNNs are particularly effective for networks with communities and local regularities. Finally, we use GNN to generate a three-dimensional layout of the Internet, and introduce additional measures to assess the layout quality and its interpretability, exploring the algorithm’s ability to separate communities and the link-length distribution. The novel use of deep neural networks can help accelerate other network-based optimization problems as well, with applications from reaction-diffusion systems to epidemics.

97 MATHEMATICS AND COMPUTING↗

Network acceleration techniques

Splintered offloading techniques with receive batch processing are described for network acceleration. Such techniques offload specific functionality to a NIC while maintaining the bulk of the protocol processing in the host operating system ("OS"). The resulting protocol implementation allows the application to bypass the protocol processing of the received data. Such can be accomplished this by moving data from the NIC directly to the application through direct memory access ("DMA") and batch processing the receive headers in the host OS when the host OS is interrupted to perform other work. Batch processing receive headers allows the data path to be separated from the control path. Unlike operating system bypass, however, the operating system still fully manages the network resource and has relevant feedback about traffic and flows. Embodiments of the present disclosure can therefore address the challenges of networks with extreme bandwidth delay products (BWDP).

Crowley, Patricia↗

Testing a Neural Network Accelerator on a High-Altitude Balloon

The cognitive communications project has been working to re d machine learning approaches to support their deployment and sustained use in space environments. It has historically been difficult to implement such techniques on space platforms, however, due to the computational requirements they levy onto general-purpose avionics hardware. While technologies exist to accelerate the computation of aspects of neural networks, such platforms have not historically been deployed in space environments. Given that testing payloads in such environments can be both cost- and time-prohibitive, high-altitude balloons can be used as a way to approximate a space environment at a much lower cost, thus providing a cost-effective way in which to test newer approaches to hardware acceleration for artificial intelligence which may be deployed onto spacecraft more directly. This paper describes a successful test of a commercial off- the-shelf neural network accelerator on a high-altitude balloon. It begins by explaining our selection criteria when evaluating different commercial neural network acceleration techniques: primary considerations include size, weight, and power (SWaP) as well as ease of integration. Next, the paper describes the development and implementation of an experimental flight test platform: flight and ground components are discussed. Afterward, the paper discusses the experimental payload itself: this includes the experimental procedure as well as the specific image and method used for testing. Finally, the paper concludes with an evaluation of both the experimental device tested at altitude as well as the flight test framework itself, identifying how the existing platform can be used to continue tes g commercial off-the-shelf (COTS) solutions for acceleration.

Clark, Gilbert↗

Complex Parsing for In-Network Acceleration of High-Energy Physics Experiments

This paper describes a novel application and evaluation of programmable networking in High-Energy Physics (HEP): a complete parser for the custom packet format used by Fermilab’s DUNE experiment. Notably, this parser is implemented on a Tofino programmable network switch and evaluated on the FABRIC testbed by using network traffic generated by the ICEBERG DUNE prototype. The parsed network traffic consists of Jumbo Ethernet frames that contain digitizations of sensor readings from ICEBERG’s detector.This work is an early investigation into providing in-network processing support for HEP experiments. The paper describes DUNE’s custom packet format, the challenges encountered when implementing a parser for that format, and an exploration of the techniques that are needed to overcome those challenges. We identify performance bottlenecks and discuss directions for future research.

Sagstad, Bjoern [IIT, Chicago] (ORCID:000900033610↗

Thermodynamic Modeling of Complex Solid Solutions in the Lu-H-N System via Graph Neural Network Accelerated Monte Carlo Simulations

Metal hydrides are important across diverse applications, such as hydrogen storage, batteries, gas sensors, nuclear reactions, and high-temperature superconductivity. Previous computational studies of metal hydrides under extreme pressures, e.g., 𝑂⁡(10 2 ) ⁢GPa, usually treat them as stoichiometric compounds without considering interstitial lattice disorder. As pressures become more moderate in the 𝑂⁡(10 0 ) ⁢GPa and below range, hydrogen disorder at interstitial lattice sites becomes prominent, e.g., in the N-doped Lu hydride that was recently claimed superconducting near 1 GPa. Further adding compositional complexity from alloying and/or multielement interstitial occupation makes elucidating pressure- and temperature-dependent observables intractable by first-principles calculations alone. We therefore propose a lattice graph neural-network surrogate modeling approach to predict configuration- and pressure-dependent equation-of-state properties. Their efficiency permits Monte Carlo simulations to calculate Gibbs energies and pressure-dependent phase diagrams, thereby revealing insights into the synthesis conditions required for achieving desired phase equilibria. We demonstrate this concept for the compositionally complex cubic Lu(H,N,Va) 3 system where three constituents (hydrogen, nitrogen and vacancy) have disordered multielement interstitial occupancies and insights into pressure-dependent phase equilibria are critically needed, e.g., N-doping levels can significantly lower dehydrogenation temperatures and provide a new strategy to optimize hydrogen-storage alloys. This work can improve the thermodynamic understanding of the Lu-H-N system and help rational synthesis of N-doped Lu hydrides, but more generally demonstrates an efficient approach to model pressure-dependent thermodynamics of multicomponent solid solutions.

Monte Carlo methods↗

Statistical data analysis of x-ray spectroscopy data enabled by neural network accelerated Bayesian inference

Bayesian inference applied to x-ray spectroscopy data analysis enables uncertainty quantification necessary to rigorously test theoretical models. However, when comparing to data, detailed atomic physics and radiation transfer calculations of x-ray emission from non-uniform plasma conditions are typically too slow to be performed in line with statistical sampling methods, such as Markov Chain Monte Carlo sampling. Furthermore, differences in transition energies and x-ray opacities often make direct comparisons between simulated and measured spectra unreliable. Here, we present a spectral decomposition method that allows for corrections to line positions and bound–bound opacities to best fit experimental data, with the goal of providing quantitative feedback to improve the underlying theoretical models and guide future experiments. In this work, we use a neural network (NN) surrogate model to replace spectral calculations of isobaric hot-spots created in Kr-doped implosions at the National Ignition Facility. The NN was trained on calculations of x-ray spectra using an isobaric hot-spot model post-processed with Cretin, a multi-species atomic kinetics and radiation code. The speedup provided by the NN model to generate x-ray emission spectra enables statistical analysis of parameterized models with sufficient detail to accurately represent the physical system and extract the plasma parameters of interest.

47 OTHER INSTRUMENTATION↗

Neural network accelerator for quantum control

Efficient quantum control is necessary for practical quantum computing implementations with current technologies. Conventional algorithms for determining optimal control parameters are computationally expensive, largely excluding them from use outside of the simulation. Existing hardware solutions structured as lookup tables are imprecise and costly. By designing a machine learning model to approximate the results of traditional tools, a more efficient method can be produced. Such a model can then be synthesized into a hardware accelerator for use in quantum systems. In this study, we demonstrate a machine learning algorithm for predicting optimal pulse parameters. This algorithm is lightweight enough to fit on a low-resource FPGA and perform inference with a latency of 175 ns and pipeline interval of 5 ns with > 0.99 gate fidelity. In the long term, such an accelerator could be used near quantum computing hardware where traditional computers cannot operate, enabling quantum control at a reasonable cost at low latencies without incurring large data bandwidths outside of the cryogenic environment.

43 PARTICLE ACCELERATORS↗

The Interplanetary Overlay Networking Protocol Accelerator

A document describes the Interplanetary Overlay Networking Protocol Accelerator (IONAC) an electronic apparatus, now under development, for relaying data at high rates in spacecraft and interplanetary radio-communication systems utilizing a delay-tolerant networking protocol. The protocol includes provisions for transmission and reception of data in bundles (essentially, messages), transfer of custody of a bundle to a recipient relay station at each step of a relay, and return receipts. Because of limitations on energy resources available for such relays, data rates attainable in a conventional software implementation of the protocol are lower than those needed, at any given reasonable energy-consumption rate. Therefore, a main goal in developing the IONAC is to reduce the energy consumption by an order of magnitude and the data-throughput capability by two orders of magnitude. The IONAC prototype is a field-programmable gate array that serves as a reconfigurable hybrid (hardware/ firmware) system for implementation of the protocol. The prototype can decode 108,000 bundles per second and encode 100,000 bundles per second. It includes a bundle-cache static randomaccess memory that enables maintenance of a throughput of 2.7Gb/s, and an Ethernet convergence layer that supports a duplex throughput of 1Gb/s.

Pang, Jackson↗

Active sampling for neural network potentials: Accelerated simulations of shear-induced deformation in Cu–Ni multilayers

Neural network potentials (NNPs) can greatly accelerate atomistic simulations relative to ab initio methods, allowing one to sample a broader range of structural outcomes and transformation pathways. In this work, we demonstrate an active sampling algorithm that trains an NNP that is able to produce microstructural evolutions with accuracy comparable to those obtained by density functional theory, exemplified during structure optimizations for a model Cu–Ni multilayer system. We then use the NNP, in conjunction with a perturbation scheme, to stochastically sample structural and energetic changes caused by shear-induced deformation, demonstrating the range of possible intermixing and vacancy migration pathways that can be obtained as a result of the speedups provided by the NNP. The code to implement our active learning strategy and NNP-driven stochastic shear simulations is openly available at https://github.com/pnnl/Active-Sampling-for-Atomistic-Potentials .

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Replace Human Intelligence with Fast and Smart Geometric Reasoning and Graph Neural Network to Accelerate Next Gen ModSim Workflows

We present an agent-guided approach to CAD geometry decomposition that automates hex/hybrid meshing with graph neural networks (GNNs) to accelerate next-generation ModSim workflows. Our end-to-end pipeline (i) reduces 3D boundary-representation (B-Rep) models to a 2D chordal axis skeleton (CAT) and then to a 1D bipartite graph of surface and curve nodes, (ii) assigns per node labels as Cubit® WebCut actions, (iii) trains a multi-action GNN under supervised learning, and (iv) predicts five surface-node and three curve-node actions on out-of-distribution test geometries. Each graph node carries geometric, topological, and meshing attributes drawn from the B-Rep “skin” and CAT “skeleton,” with two-way mappings across 3D↔2D↔1D representations to maintain traceability back to 3D CAD. The supervised learning model exhibits stable convergence of the binary cross-entropy loss and achieves 98.7% accuracy on unseen lattice models. To operationalize decision-making, we rank predicted commands by geometric significance and prototyped the agent-guided workflow through the Cubit® Meshing PowerTool GUI. As a stretch goal, we explore reinforcement learning (RL) to reduce or remove label requirements and to learn policies for action sequences that maximize total reward (e.g., size of hex-meshable regions and resulting hex mesh quality). When all-hex meshing is not feasible, the agent assists in producing hybrid meshes—prioritizing hex in critical regions and transitioning to tetrahedral elements (tets) elsewhere—maintaining fidelity while ensuring robustness. The overarching objective is to replace manual, heuristics-based decomposition with data-driven, reproducible automation, cutting meshing turnaround time by orders of magnitude. We anticipate direct impact on simulation workflows through intelligent, scalable decomposition of complex CAD models into hex-meshable subdomains.

97 MATHEMATICS AND COMPUTING↗

Using convolutional neural networks to accelerate three-dimensional coherent synchrotron radiation computations

Calculating the effects of coherent synchrotron radiation (CSR) is one of the most computationally expensive tasks in accelerator physics. Here, we use convolutional neural networks (CNNs), along with a latent conditional diffusion (LCD) model, trained on physics-based simulations to speed up calculations. Specifically, we produce the 3D CSR wakefields generated by electron bunches in circular orbit in the steady-state condition. Two datasets are used for training and testing the models: wakefields generated by three-dimensional Gaussian electron distributions and wakefields from a sum of up to 25 three-dimensional Gaussian distributions. The CNNs are able to accurately produce the 3D wakefields ∼250–1000 times faster than the numerical calculations, while the LCD achieves a gain of a factor of ∼34. We also test the extrapolation and out-of-distribution generalization ability of the models. They generalize well on distributions with larger spreads than what they were trained on but struggle with smaller spreads.

43 PARTICLE ACCELERATORS↗

Reinforced double-threaded slide-ring networks for accelerated hydrogel discovery and 3D printing

Traditionally, slide-ring gels are stretchable but soft as a result of an elasticity-stretchability trade-off. Herein, we introduce a new approach to breaking this trade-off and creating reinforced slide-ring networks with mobile crosslinkers. Our approach involves the construction of a polyethylene glycol double-threaded γ-cyclodextrin-based pro-slide-ring crosslinker that serves as a modular component for 3D printing and copolymerization. The resulting crystalline-domain-reinforced slide-ring hydrogels, or CrysDoS-gels, exhibit both high elasticity and high stretchability. The modular synthesis allows for high-throughput synthesis of CrysDoS-gels, generating a large amount of data for structure-property analysis. Here, by employing data science techniques, such as machine learning and linear regression, not only were we able to identify which chemical components influence the mechanical properties of CrysDoS-gels, but this analysis also aided in the discovery of better-performing CrysDoS-gels. Finally, we demonstrate the potential application of the newly discovered CrysDoS-gels as sensing devices by 3D printing them as stress sensors with high sensitivity and a broad detection range.

3D-printing↗

Neural Network Based Representation of UH-60A Pilot and Hub Accelerations

Neural network relationships between the full-scale, experimental hub accelerations and the corresponding pilot floor vertical vibration are studied. The present physics-based, quantitative effort represents an initial systematic study on the UH-60A Black Hawk hub accelerations. The NASA/Army UH-60A Airloads Program flight test database was used. A 'maneuver-effect-factor (MEF)', derived using the roll-angle and the pitch-rate, was used. Three neural network based representation-cases were considered. The pilot floor vertical vibration was considered in the first case and the hub accelerations were separately considered in the second case. The third case considered both the hub accelerations and the pilot floor vertical vibration. Neither the advance ratio nor the gross weight alone could be used to predict the pilot floor vertical vibration. However, the advance ratio and the gross weight together could be used to predict the pilot floor vertical vibration over the entire flight envelope. The hub accelerations data were modeled and found to be of very acceptable quality. The hub accelerations alone could not be used to predict the pilot floor vertical vibration. Thus, the hub accelerations alone do not drive the pilot floor vertical vibration. However, the hub accelerations, along with either the advance ratio or the gross weight or both, could be used to satisfactorily predict the pilot floor vertical vibration. The hub accelerations are clearly a factor in determining the pilot floor vertical vibration.

Kottapalli, Sesi↗

SYCL for Performance Portability: Application Experience with Coupled Cluster Formalism in Quantum Chemistry on Exascale Systems

The exascale computing has brought unprecedented heterogeneity in node architectures, with systems such as Frontier and Aurora featuring diverse GPU accelerators, network connectivity among others. Ensuring performance portability across these platforms is a key challenge. To address this, we employ the SYCL programming model to develop portable, high-performance quantum chemistry workloads. As a representative application, we focus on the non-iterative Triples component of the coupled-cluster CCSD(T) method, a key driver in quantum chemistry. In this work, we report on our experience deploying SYCL-based implementations using both DPC++ and AdaptiveCPP across two flagship exascale platforms: OLCF Frontier with AMD MI250X GPUs and ALCF Aurora with Intel GPUs. Our results demonstrate that SYCL enables efficient, single-source implementations that scale to thousands of nodes, delivering performance on par with vendor-optimized HIP solutions. We highlight key insights into runtime behavior, kernel portability, and scaling characteristics, showing that SYCL offers a viable path for performance-portable computing.

Bagusetty, Abhishek [Argonne National Laboratory (↗

Superconducting Hyperdimensional Associative Memory Circuit for Scalable Machine Learning

Here we propose a generalized architecture for the first rapid-single-flux-quantum (RSFQ) associative memory circuit. The circuit employs hyperdimensional computing (HDC), a machine learning (ML) paradigm utilizing vectors with dimensionality in the thousands to represent information. HDC designs have small memory footprints, simple computations, and simple training algorithms compared to superconducting neural network accelerators (SNNAs), making them a better option for scalable SFQ machine learning (ML) solutions. The proposed superconducting HDC (SHDC) circuit uses entirely on-chip RSFQ memory which is tightly integrated with logic, operates at 33.3 GHz, is applicable to general ML tasks, and is manufacturable at practically useful scales given current SFQ fabrication limits. Tailored to a language recognition task, SHDC consists of ~ 2-20 M Josephson junctions (JJs) and consumes up to three times less power than an analogous CMOS HDC circuit while achieving 78-84% higher throughput. SHDC is capable of outperforming the state of the art RSFQ SNNA, SuperNPU, by 48-99% for all benchmark NN architectures tested while occupying up to 90% less area and consuming up to nine times less power. To the best of the authors' knowledge, SHDC is currently the only superconducting ML approach feasible at practically useful scales for real-world ML tasks and capable of online learning.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Modeling of UH-60A Hub Accelerations with Neural Networks

Neural network relationships between the full-scale, flight test hub accelerations and the corresponding three N/rev pilot floor vibration components (vertical, lateral, and longitudinal) are studied. The present quantitative effort on the UH-60A Black Hawk hub accelerations considers the lateral and longitudinal vibrations. An earlier study had considered the vertical vibration. The NASA/Army UH-60A Airloads Program flight test database is used. A physics based "maneuver-effect-factor (MEF)", derived using the roll-angle and the pitch-rate, is used. Fundamentally, the lateral vibration data show high vibration levels (up to 0.3 g's) at low airspeeds (for example, during landing flares) and at high airspeeds (for example, during turns). The results show that the advance ratio and the gross weight together can predict the vertical and the longitudinal vibration. However, the advance ratio and the gross weight together cannot predict the lateral vibration. The hub accelerations and the advance ratio can be used to satisfactorily predict the vertical, lateral, and longitudinal vibration. The present study shows that neural network based representations of all three UH-60A pilot floor vibration components (vertical, lateral, and longitudinal) can be obtained using the hub accelerations along with the gross weight and the advance ratio. The hub accelerations are clearly a factor in determining the pilot vibration. The present conclusions potentially allow for the identification of neural network relationships between the experimental hub accelerations obtained from wind tunnel testing and the experimental pilot vibration data obtained from flight testing. A successful establishment of the above neural network based link between the wind tunnel hub accelerations and the flight test vibration data can increase the value of wind tunnel testing.

UH-60A AIRCRAFT↗