Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Neural Networks for Nuclear Reactions in MAESTROeX

We demonstrate the use of neural networks to accelerate the reaction steps in the MAESTROeX stellar hydrodynamics code. A traditional MAESTROeX simulation uses a stiff ODE integrator for the reactions; here, we employ a ResNet architecture and describe details relating to the architecture, training, and validation of our networks. Our customized approach includes options for the form of the loss functions, a demonstration that the use of parallel neural networks leads to increased accuracy, and a description of a perturbational approach in the training step that robustifies the model. We test our approach on millimeter-scale flames using a single-step, 3-isotope network describing the first stages of carbon fusion occurring in Type Ia supernovae. We train the neural networks using simulation data from a standard MAESTROeX simulation, and show that the resulting model can be effectively applied to different flame configurations. This work lays the groundwork for more complex networks, and iterative time-integration strategies that can leverage the efficiency of the neural networks.

79 ASTRONOMY AND ASTROPHYSICS↗

Integrated reactor architecture of conductive network and catalytic nodes to accelerate polysulfide conversion for durable and high-loading Li-S batteries

The development of carbon-based heterogeneous framework host with synergistic catalytic and conductive effects for sulfur cathode is a promising strategy to realize high performance lithium sulfur batteries (LSBs). Here, an integrated reactor architecture with defective carbon nodes (IRA-DC) is designed for serving as high-loading (92.4 wt%) sulfur host. The hierarchical porous IRA-DC consists of untangled conductive carbon nanotube network and Co/N co-doped catalytic nodes with high dispersity. Therein the optimization of electric field distribution and homogenization of adsorption-catalysis sites offer the multi-electron conversion reaction of polysulfides with excellent kinetics and stability. The resultant IRA-DC/S cathode enables a high areal capacity of 8.86 mAh cm -2 under ultra-high sulfur loading (13.1 mg cm -2 ) and lean electrolyte (8 μL mg sulfur -1 ). It also displays a long-term cycling performance (1200 cycles at 1 C) and ultrahigh rate performance up to 20 C (with a capacity of 473.6 mAh g -1 ). In conclusion, this work provides an electrode building strategy by optimizing the environments of heterogeneous electrocatalysis and micro electric field to activate the polysulfide conversion efficiency and utilization of high-loading sulfur in monolithic sulfur-carbon cathodes.

25 ENERGY STORAGE↗

Disentangling Beam Losses in The Fermilab Main Injector Enclosure Using Real-Time Edge AI

The Fermilab Main Injector enclosure houses two accelerators, the Main Injector and Recycler Ring. During normal operation, high intensity proton beams exist simultaneously in both. The two accelerators share the same beam loss monitors (BLM) and monitoring system. Deciphering the origin of any of the 260 BLM readings is often difficult. The (Accelerator) Real-time Edge AI for Distributed Systems project, or READS, has developed an AI/ML model, and implemented it on fast FPGA hardware, that disentangles mixed beam losses and attributes probabilities to each BLM as to which machine(s) the loss originated from in real-time. The model inferences are then streamed to the Fermilab accelerator controls network (ACNET) where they are available for operators and experts alike to aid in tuning the machines.

43 PARTICLE ACCELERATORS↗

GUI Control System for the Mu2e Electrostatic Septum High Voltage at Fermilab

The Mu2e Experiment has stringent beam structure requirements; namely, its proton bunches with a time structure of 1.7 $\mu$s in the Fermilab Delivery Ring. This beam structure will be delivered using the Fermilab 8-GeV Booster, the 8-GeV Recycler Ring, and the Delivery Ring. The 1.7-$\mu$s period of the Delivery Ring will generate the required beam structure by means of a third order resonant extraction system operating on a single circulating bunch. The electrostatic septum (ESS) for this system is particularly challenging, requiring mechanical precision in a ultra high vacuum of 1 x 10$^-8$ Torr to generate 100 kV across 15 mm. This paper describes a graphical user interface that has been developed to automate the conditioning and commissioning process for the electrostatic septa. It is based on an interface to the Fermilab ACNET system using the ACSys Python Data Pool Manager (DPM) Client produced and maintained by Fermilab Accelerator Controls. Network interfacing between data pool managers made by the application and ACNET devices introduce an inherent (approximately 1 s) latency in throughput of the readouts. This delay is utilized to process and graph incoming data events of devices crucial to conditioning of a electrostatic septum (ESS). 'Ramping' and 'Monitoring' modes adjust settings of the power supply based on internal logic to efficaciously increase and maintain the high voltage (HV) in the ESS, easing the voltage setting on incidence of sparking or other possibly damaging events. A timestamped log file is produced as the application runs.

43 PARTICLE ACCELERATORS↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Beam Synchronous for the Rest of Us!

Fermilab’s Tevatron Clock (TCLK) infrastructure has been an integral part of the accelerator control network since the 1980’s. This 10MHz Manchester encoded protocol has enabled flexible, real-time event distribution for thousands of devices connected to the timing network with a high degree of reliability. Forthcoming upgrades to the Fermilab complex (PIP-II, LBNF, ACORN) necessitate higher levels of precision to maintain inter-bunch timing for Instrumentation and Control purposes. This presents as an opportunity to refine the event distribution protocol for tighter synchronization between machines, experiments, and eventually far-site operations. This paper outlines a method by which beam-synchronous events may be distributed through asynchronous serial protocols via integration with local LLRF and global PPS reference signals. This method is ideal for synchrotron machines with aggressive frequency sweeps (such as Fermilab's 38~53MHz Booster) and allows for precision timing to be maintained across machines without specialized hardware.

43 PARTICLE ACCELERATORS↗

Improving Trustworthiness of Data-Driven Power Grid Contingency Analysis With Bayesian Residual Graph Neural Networks

The evolving energy landscape requires novel tools to efficiently perform contingency analysis and reliability assessment of power grids, potentially in real-time. The high computational cost of traditional power flow solvers limits their applicability in practice. Machine learning (ML) surrogates such as deep neural networks (NNs) accelerate power flow solvers computations, enabling high-order contingency analysis and real-time decision-making by learning highly nonlinear functions and integrating grid topology via graph architectures. However, (graph) NNs lack predictive power away from training data and do not provide predictive confidence estimates. Here, we present a Bayesian residual graph NN that integrates knowledge from low-fidelity data via residual training and embeds granular quantification of uncertainties, improving trustworthiness critical for high-consequence decision-making. Applying Bayesian concepts to NNs is challenging due to the high-dimensionality of both the parameter space, complicating derivation of a meaningful prior, and the output space in large grid systems, requiring enhanced techniques to assess the predicted high-dimensional uncertainties. Our contributions include: (1) Deriving a prior for fully connected and graph NNs that leverages low-fidelity data to guide mean predictions and appropriately control prior predictive uncertainty. (2) Integrating this prior within an ensembling with anchoring scheme for efficient approximate posterior inference. (3) Deriving enhanced metrics to assess accuracy of both the mean and uncertainty predictions in high dimensions, appropriately accounting for correlations propagated through graph layers. The resulting Bayesian residual graph NN is tested on a contingency analysis task for 14-bus and 118-bus grids.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions

Neural network wavefunctions optimized using the variational Monte Carlo method have been shown to produce highly accurate results for the electronic structure of atoms and small molecules, but the high cost of optimizing such wavefunctions prevents their application to larger systems. We propose the Subsampled Projected-Increment Natural Gradient Descent (SPRING) optimizer to reduce this bottleneck. SPRING combines ideas from the recently introduced minimum-step stochastic reconfiguration optimizer (MinSR) and the classical randomized Kaczmarz method for solving linear least-squares problems. We demonstrate that SPRING outperforms both MinSR and the popular Kronecker-Factored Approximate Curvature method (KFAC) across a number of small atoms and molecules, given that the learning rates of all methods are optimally tuned. For example, on the oxygen atom, SPRING attains chemical accuracy after forty thousand training iterations, whereas both MinSR and KFAC fail to do so even after one hundred thousand iterations.

97 MATHEMATICS AND COMPUTING↗

Analog In-Memory Computing for the Synthetic Aperture Radar Polar Format Algorithm

As the utility of synthetic aperture radar (SAR) systems increases in autonomous vehicles, satellites, and other power- and space-constrained edge applications, there is a growing need for processors that can form SAR images at low power. In recent years, analog in-memory compute (AIMC) has shown immense promise for accelerating neural networks and other matrix-vector multiplication (MVM) heavy workloads at the edge. Here, in this work, we examine how the polar format algorithm (PFA), a popular SAR image formation algorithm, can be mapped to these AIMC systems. The PFA maps readily onto analog MVMs because it primarily consists of two linear operations: interpolation of frequency-domain data to a Cartesian grid, followed by a 2-D Fourier transform. This work presents two approaches to map the interpolation operation onto MVMs in analog hardware: a chirp transform and a modified form of sinc interpolation. These mappings introduce algorithmic errors, and their effect on the quality of SAR image formation is examined, both quantitatively and qualitatively. In addition, the impact of errors introduced by the analog hardware is explored to determine which approach is optimal under varying assumptions about the underlying analog memory devices and circuits.

Analog computing↗

ESnet SmartNIC v1.0

The ESnet SmartNIC is a collection of Verilog based FPGA design software, as well as drivers to interact with the FPGA. It provides an infrastructure framework for FPGA based hardware acceleration of network packet use cases. Different research and production applications can be easily written for the SmartNIC. Its advantage is that it reduces the development time for new applications by providing a pre-existing shell library for common functions.

Mah, Bruce↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (Mu2e) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (SRS) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 $μ$s, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab's Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, J. R. [Fermilab]↗

Real time heat load calculation software based on EPICS for Fermilab PIP-II CM tests

Fermilab has a project to improve the proton beam energy which is called PIP-II (the 2nd Proton Improvement Plan). There is a superconducting linear accelerator, LINAC, to improve the proton beam power and the LINAC consists of 5 types of cryomodules (CM), 1 HWR CM, 2 SSR1 CM, 4 SSR2 CM, LB650 CM, and HB650 CM. The prototypes of these cryomodules are being tested at Fermilab’s CryoModule Test Facility (CMTF). Heat load measurements are an important part of the prototype CM testing. The CMTF cryogenic control system was developed based on the ACNET (Accelerator Control NETwork) for CM testing for other projects, but the PIP-II cryogenic control system will be implemented using the Experimental Physics and Industrial Control System (EPICS). As part of the prototype CM testing campaign an EPICS based control system has been implemented at CMTF. This EPICS cryogenic control system includes real time heat load calculation software utilizing the Fortran implementation of Hepak. This paper details the real time heat load calculation software developed for the prototype CM testing including the first results from the HB 650 CM.

Yoon, S. [Fermilab]↗

Heat load measurements for the PIP-II pHB650 cryomodule

This study presents a brief overview of the 1st and 2nd phases and an in-depth analysis of the 3rd phase heat load testing performed on the pHB650 (prototype High Beta 650 MHz) cryomodule at PIP2IT (PIP-II Injector Test Facility), with a focus on both the results and the methodological advancements that have improved testing efficiency and accuracy. A key challenge identified in the testing campaign is the higher-than-expected heat loads observed in the first PIP-II (Proton Improvement Plan II) prototype cryomodules (pSSR1 and pHB650) tested at PIP2IT. Elevated heat loads are concerning given the fixed capacity of the PIP-II cryoplant that is currently being installed at Fermilab. However, understanding the sources of these elevated heat loads offers a critical opportunity to implement effective heat load mitigations on upcoming PIP-II cryomodules to stay within the available capacity of the PIP-II cryoplant. The study includes a summary of test results, descriptions of measurement procedures, and key observations on parameters directly and indirectly related to heat load measurements. Direct observations include measured heat loads and the effectiveness of JT heat exchanger under varying conditions, while indirect observation analyze factors such as the temperature distribution on the two-phase pipe and relief piping under varying conditions. Thermal acoustic oscillations (TAO) were identified during testing, which was mitigated by replacing the original G10 stem with a stainless steel stem equipped with wipers for the cryomodule cooldown valve. A major innovation during pHB650 Phase 3 testing was the development of an automated Python script to streamline data acquisition, analysis, and reporting of heat load results. This script automatically retrieved data from ACNET (Accelerator Control Network), performed heat load calculations, and generated detailed reports featuring plots and tables. This advancement significantly reduced manual labor and enhanced the thoroughness of data analysis compared to earlier campaigns. The heat load test reports were promptly uploaded to the electronic logbook shortly after each test, enabling rapid feedback and collaboration between the SRF and cryogenic teams. The heat load measurements included various components: HTTS (high-temperature thermal shield), LTTS (low-temperature thermal shield), 2K isothermal and non-isothermal heat loads. Results were recorded both within the cryomodule and between the bayonet can supply and return. Measurements were conducted under different operating conditions such as "standard", "linac", and "simulated dynamic". Additionally, HTTS and LTTS heat loads were calculated in real time, allowing for the tracking of thermal stability and identification of changes during testing, both in steady-state and transient conditions. The results of this testing campaign not only provide valuable insights into the performance of the pHB650 cryomodule but also highlight best practices and lessons learned that will inform future cryomodule testing at PIP2IT. These include adopting automated tools for data analysis, refining real-time measurement capabilities, and emphasizing detailed pre-test planning. The framework established in this campaign aims to set an improved standard for cryomodule testing and heat load reporting in future cryomodule test campaigns.

Porwisiak, D. [Fermilab; Wroclaw Tech. U.]↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (\acrshort{Mu2e}) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (\acrshort{SRS}) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 microseconds, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab s Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, Jose Rene [Fermilab]↗

Method Accelerates Training Of Some Neural Networks

Three-layer networks trained faster provided two conditions are satisfied: numbers of neurons in layers are such that majority of work done in synaptic connections between input and hidden layers, and number of neurons in input layer at least as great as number of training pairs of input and output vectors. Based on modified version of back-propagation method.

Shelton, Robert O.↗

NASA Tech Briefs, October 2008

Topics covered include: Control Architecture for Robotic Agent Command and Sensing; Algorithm for Wavefront Sensing Using an Extended Scene; CO2 Sensors Based on Nanocrystalline SnO2 Doped with CuO; Improved Airborne System for Sensing Wildfires; VHF Wide-Band, Dual-Polarization Microstrip-Patch Antenna; Onboard Data Processor for Change-Detection Radar Imaging; Using LDPC Code Constraints to Aid Recovery of Symbol Timing; System for Measuring Flexing of a Large Spaceborne Structure; Integrated Formation Optical Communication and Estimation System; Making Superconducting Welds between Superconducting Wires; Method for Thermal Spraying of Coatings Using Resonant-Pulsed Combustion; Coating Reduces Ice Adhesion; Hybrid Multifoil Aerogel Thermal Insulation; SHINE Virtual Machine Model for In-flight Updates of Critical Mission Software; Mars Image Collection Mosaic Builder; Providing Internet Access to High-Resolution Mars Images; Providing Internet Access to High-Resolution Lunar Images; Expressions Module for the Satellite Orbit Analysis Program Virtual Satellite; Small-Body Extensions for the Satellite Orbit Analysis Program (SOAP); Scripting Module for the Satellite Orbit Analysis Program (SOAP); XML-Based SHINE Knowledge Base Interchange Language; Core Technical Capability Laboratory Management System; MRO SOW Daily Script; Tool for Inspecting Alignment of Twinaxial Connectors; An ATP System for Deep-Space Optical Communication; Polar Traverse Rover Instrument; Expert System Control of Plant Growth in an Enclosed Space; Detecting Phycocyanin-Pigmented Microbes in Reflected Light; DMAC and NMP as Electrolyte Additives for Li-Ion Cells; Mass Spectrometer Containing Multiple Fixed Collectors; Waveguide Harmonic Generator for the SIM; Whispering Gallery Mode Resonator with Orthogonally Reconfigurable Filter Function; Stable Calibration of Raman Lidar Water-Vapor Measurements; Bimaterial Thermal Compensators for WGM Resonators; Root Source Analysis/ValuStream[Trade Mark] - A Methodology for Identifying and Managing Risks; Ensemble: an Architecture for Mission-Operations Software; Object Recognition Using Feature-and Color-Based Methods; On-Orbit Multi-Field Wavefront Control with a Kalman Filter; and The Interplanetary Overlay Networking Protocol Accelerator.

Source record↗