Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network acceleration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Analog In-Memory Computing for the Synthetic Aperture Radar Polar Format Algorithm

As the utility of synthetic aperture radar (SAR) systems increases in autonomous vehicles, satellites, and other power- and space-constrained edge applications, there is a growing need for processors that can form SAR images at low power. In recent years, analog in-memory compute (AIMC) has shown immense promise for accelerating neural networks and other matrix-vector multiplication (MVM) heavy workloads at the edge. Here, in this work, we examine how the polar format algorithm (PFA), a popular SAR image formation algorithm, can be mapped to these AIMC systems. The PFA maps readily onto analog MVMs because it primarily consists of two linear operations: interpolation of frequency-domain data to a Cartesian grid, followed by a 2-D Fourier transform. This work presents two approaches to map the interpolation operation onto MVMs in analog hardware: a chirp transform and a modified form of sinc interpolation. These mappings introduce algorithmic errors, and their effect on the quality of SAR image formation is examined, both quantitatively and qualitatively. In addition, the impact of errors introduced by the analog hardware is explored to determine which approach is optimal under varying assumptions about the underlying analog memory devices and circuits.

Analog computing↗

ESnet SmartNIC v1.0

The ESnet SmartNIC is a collection of Verilog based FPGA design software, as well as drivers to interact with the FPGA. It provides an infrastructure framework for FPGA based hardware acceleration of network packet use cases. Different research and production applications can be easily written for the SmartNIC. Its advantage is that it reduces the development time for new applications by providing a pre-existing shell library for common functions.

Mah, Bruce↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (Mu2e) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (SRS) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 $μ$s, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab's Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, J. R. [Fermilab]↗

Real time heat load calculation software based on EPICS for Fermilab PIP-II CM tests

Fermilab has a project to improve the proton beam energy which is called PIP-II (the 2nd Proton Improvement Plan). There is a superconducting linear accelerator, LINAC, to improve the proton beam power and the LINAC consists of 5 types of cryomodules (CM), 1 HWR CM, 2 SSR1 CM, 4 SSR2 CM, LB650 CM, and HB650 CM. The prototypes of these cryomodules are being tested at Fermilab’s CryoModule Test Facility (CMTF). Heat load measurements are an important part of the prototype CM testing. The CMTF cryogenic control system was developed based on the ACNET (Accelerator Control NETwork) for CM testing for other projects, but the PIP-II cryogenic control system will be implemented using the Experimental Physics and Industrial Control System (EPICS). As part of the prototype CM testing campaign an EPICS based control system has been implemented at CMTF. This EPICS cryogenic control system includes real time heat load calculation software utilizing the Fortran implementation of Hepak. This paper details the real time heat load calculation software developed for the prototype CM testing including the first results from the HB 650 CM.

Yoon, S. [Fermilab]↗

Heat load measurements for the PIP-II pHB650 cryomodule

This study presents a brief overview of the 1st and 2nd phases and an in-depth analysis of the 3rd phase heat load testing performed on the pHB650 (prototype High Beta 650 MHz) cryomodule at PIP2IT (PIP-II Injector Test Facility), with a focus on both the results and the methodological advancements that have improved testing efficiency and accuracy. A key challenge identified in the testing campaign is the higher-than-expected heat loads observed in the first PIP-II (Proton Improvement Plan II) prototype cryomodules (pSSR1 and pHB650) tested at PIP2IT. Elevated heat loads are concerning given the fixed capacity of the PIP-II cryoplant that is currently being installed at Fermilab. However, understanding the sources of these elevated heat loads offers a critical opportunity to implement effective heat load mitigations on upcoming PIP-II cryomodules to stay within the available capacity of the PIP-II cryoplant. The study includes a summary of test results, descriptions of measurement procedures, and key observations on parameters directly and indirectly related to heat load measurements. Direct observations include measured heat loads and the effectiveness of JT heat exchanger under varying conditions, while indirect observation analyze factors such as the temperature distribution on the two-phase pipe and relief piping under varying conditions. Thermal acoustic oscillations (TAO) were identified during testing, which was mitigated by replacing the original G10 stem with a stainless steel stem equipped with wipers for the cryomodule cooldown valve. A major innovation during pHB650 Phase 3 testing was the development of an automated Python script to streamline data acquisition, analysis, and reporting of heat load results. This script automatically retrieved data from ACNET (Accelerator Control Network), performed heat load calculations, and generated detailed reports featuring plots and tables. This advancement significantly reduced manual labor and enhanced the thoroughness of data analysis compared to earlier campaigns. The heat load test reports were promptly uploaded to the electronic logbook shortly after each test, enabling rapid feedback and collaboration between the SRF and cryogenic teams. The heat load measurements included various components: HTTS (high-temperature thermal shield), LTTS (low-temperature thermal shield), 2K isothermal and non-isothermal heat loads. Results were recorded both within the cryomodule and between the bayonet can supply and return. Measurements were conducted under different operating conditions such as "standard", "linac", and "simulated dynamic". Additionally, HTTS and LTTS heat loads were calculated in real time, allowing for the tracking of thermal stability and identification of changes during testing, both in steady-state and transient conditions. The results of this testing campaign not only provide valuable insights into the performance of the pHB650 cryomodule but also highlight best practices and lessons learned that will inform future cryomodule testing at PIP2IT. These include adopting automated tools for data analysis, refining real-time measurement capabilities, and emphasizing detailed pre-test planning. The framework established in this campaign aims to set an improved standard for cryomodule testing and heat load reporting in future cryomodule test campaigns.

Porwisiak, D. [Fermilab; Wroclaw Tech. U.]↗

FPGA-Based Spill Regulation System for the Muon Delivery Ring at Fermilab

The Muon to Electron Experiment (\acrshort{Mu2e}) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (\acrshort{SRS}) has been developed to help achieve consistent spill uniformity. The system is based on a custom-designed carrier board featuring an Arria 10 SoC, capable of executing real-time feedback control. The FPGA processes beam pulses of approximately 200 ns every 1.695 microseconds, allowing for continuous monitoring of the extracted spill intensity through fast bunch integration. The system directly controls three quadrupole magnets, which work in conjunction with sextupole magnets to achieve third-order resonant extraction. Furthermore, the board interfaces with Fermilab s Accelerator Control Network (ACNET), enabling operators to modify spill regulation settings in real-time via the control network while providing diagnostic waveforms. These waveforms help operators monitor the process and fine-tune the feedback mechanisms. This paper presents an overview of the board's architecture and its initial progress toward regulating beam extraction. This initial version of the regulation system aims to evaluate baseline performance to inform future system improvements.

Berlioz, Jose Rene [Fermilab]↗

The Global Climate Action Partnership

The Global Climate Action Partnership (GCAP; formerly the Low Emission Development Strategies Global Partnership) is a worldwide network that accelerates ambitious climate goals and advances implementation in Africa, Asia, and Latin America and the Caribbean (LAC). Our strategic approach emphasizes country-driven implementation actions to achieve resilient, just, and inclusive low-emission and net-zero economies. As an incubator for scalable knowledge and solutions, GCAP is leading the way to equitable, climate-resilient, low-emission development.

African Climate Action Partnership↗

Real time heat load calculation software based on EPICS for Fermilab PIP-II CM tests

Fermilab has a project to improve the proton beam energy which is called PIP-II (the 2nd Proton Improvement Plan). There is a superconducting linear accelerator, LINAC, to improve the proton beam power and the LINAC consists of 5 types of cryomodules (CM), 1 HWR CM, 2 SSR1 CM, 4 SSR2 CM, LB650 CM, and HB650 CM. The prototypes of these cryomodules are being tested at Fermilab’s CryoModule Test Facility (CMTF). Heat load measurements are an important part of the prototype CM testing. The CMTF cryogenic control system was developed based on the ACNET (Accelerator Control NETwork) for CM testing for other projects, but the PIP-II cryogenic control system will be implemented using the Experimental Physics and Industrial Control System (EPICS). As part of the prototype CM testing campaign, an EPICS based control system has been implemented at CMTF. This EPICS cryogenic control system includes real-time heat load calculation software utilizing the Fortran implementation of Hepak. This paper details the real time heat load calculation software developed for the prototype CM testing including the first results from the HB650 CM.

Yoon, S. [Fermilab]↗

Field Insights: Strengthening Digital Assurance Through On-Site Network Monitoring

The accelerating deployment of digital energy infrastructure, ranging from inverter-based resources (IBRs), battery energy storage systems (BESS), to advanced grid control platforms, has brought unprecedented visibility, flexibility, and efficiency to the electric grid. However, this digital transformation also introduces new cybersecurity challenges, particularly in the form of supply chain risks and operational blind spots at the grid edge. Over the past year, the Department of Energy’s Office of Cybersecurity, Energy Security, and Emergency Response (CESER), through its Rapid Risk Assessment initiative, along with the Grid Deployment Office (GDO), through its Technical Assistance for Digital Assurance (TADA) initiative, have supported a series of on-site network engagements led by Idaho National Laboratory (INL). These engagements, conducted in partnership with asset owners across the country, have focused on identifying real-world vulnerabilities and misconfigurations in operational environments, many of which are not detectable through remote assessments or traditional compliance audits. The goal of this report is to distill key findings and lessons learned during network hunt engagements from INL’s fiscal year (FY) 2024 - 2025. It is intended to help asset owners—regardless of their participation in the program—better understand the evolving threat landscape and adopt practical measures to secure their digital energy infrastructure.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

FiberFlex: Real-time FPGA-based Intelligent and Distributed Fiber Sensor System for Pedestrian Recognition

In recent years, security monitoring of public places and critical infrastructure has heavily relied on the widespread use of cameras, raising concerns about personal privacy violations. To balance the need for effective security monitoring with the protection of personal privacy, we explore the potential of optical fiber sensors for this application. This article proposes FiberFlex, an intelligent and distributed fiber sensor system. Ultizing Field Programmable Gate Arrays (FPGA) high-level synthesis (HLS) acceleration, FiberFlex offers real-time pedestrian detection by co-designing the entire pipeline of optical signal acquisition, processing, and recognition networks based on the principles of optical fiber sensing. As a promising alternative to traditional camera-based monitoring systems, FiberFlex achieves pedestrian detection by analyzing the vibration patterns caused by pedestrian footsteps, enabling security monitoring while preserving individual privacy. FiberFlex comprises three modules: First , fiber-optic sensing system: A fiber-optic distributed acoustic sensing (DAS) system is built and used to measure the ground vibration waves generated by people walking. Second , algorithms: We first collect the training data by measuring the ground vibration waves, label the data, and use the data to train the neural network models to perform pedestrian recognition. Third , hardware accelerators: We use HLS tools to design hardware modules on FPGA for data collection and pre-processing and integrate them with the downstream neural network accelerators to perform in-line real-time pedestrian detection. The final detection results are sent back from FPGA to the host CPU. We implement our system FiberFlex with the in-house built DAS system and AMD/Xilinx Kintex7 FPGA KC705 board and verify the whole system using the real-world collected data. We conduct recognition tests on five test subjects of varying ages, heights, and weights in a fixed sensing area. Each subject experienced 20 real-time recognition tests using their daily walking habits, and the subjects were given adequate rest between tests. After 100 tests on five test subjects, the overall real-time recognition accuracy exceeded \(88.0\%\) . The whole system uses 55 W of power, 33 W in the optical DAS system and 22 W in the FPGA. Relying on its end-to-end interdisciplinary design, FiberFlex seamlessly combines fiber-optic sensors with FPGA accelerators to enable low-power real-time security monitoring without compromising privacy, making it a valuable addition to the existing security monitoring network. According to FiberFlex, more valuable research can be conducted in the future, such as fall monitoring for the elderly, migration of identification networks between different application scenarios, and improvement of anti-interference performance in more complex environments. In future perception networks, where the “eyes” are not feasible, let’s use fiber optic touch instead.

Distributed↗

Application of Convolutional and Feedforward Neural Networks for Fault Detection in Particle Accelerator Power Systems

High voltage converter modulators (HVCM) provide power to the accelerating cavities of the spallation neutron source (SNS) facility. HVCM experience catastrophic failures, which increase the downtime of the SNS and reduce beam time. The faults may occur due to different reasons including failures of the resonant capacitor, core saturation due to the magnetic flux, insulated-gate bipolar transistor (IGBT) failures, and others. We recently have setup a HVCM test stand to develop and test machine learning models for anomaly detection and fault prognostics. In this work, we propose binary classifiers and autoencoder architectures based on convolutional (CNN) and feedforward neural networks (FNN) to facilitate distinguishing normal from faulty waveforms coming from the HVCM during operation. The results indicate that the CNN binary classifier is the best model among the four showing very stable performance in the training and testing sets with impressive metrics of precision and recall reaching up to 99\% with a very small uncertainty. The FNN classifier shows the least performance with a large uncertainty in its metrics. The performances of the two autoencoders based on CNN and FNN were in between, showing very good performance nonetheless.

Radaideh, Majdi↗

Modernizing to an AI-Ready Control System

The Accelerator Controls Operations Research Network (ACORN) project will modernize the accelerator control system and upgrade power supplies to enable future operations of the Fermilab Accelerator Complex with megawatt proton beams. By focusing on MLOps, ACORN enables full integration of AI into the control system to allow safe and reliable improvements to accelerator beam operations.

43 PARTICLE ACCELERATORS↗

Upgrading Fermilab’s accelerator control system with ACORN

The Fermilab Accelerator Complex is the largest national user facility in the Office of High Energy Physics (DOE/HEP) program and the only national user facility operating at Fermilab. Fermilab serves as the host to the Long Baseline Neutrino Facility/Deep Underground Neutrino Experiment (LBNF/DUNE), the laboratory’s flagship project for neutrino science that is under construction. LBNF/DUNE will be powered by megawatt beams from an upgraded accelerator, the Proton Improvement Plan II (PIP-II) that will replace the laboratory’s aging linear accelerator with a new one based on superconducting radio-frequency cavities. The Accelerator Controls Operations Research Network (ACORN) Project will support LBNF/DUNE and PIP-II by modernizing the accelerator control system. The project is at the conceptual design phase and looking to achieve Critical Decision 1 (CD-1) later this year. The scope and structure of the project will be presented, along with an overview of how that has changed in the past year. Current design and technology choices will be shared. Specific challenges facing the project will be addressed, along with current thinking on solutions.

Roehrig, Christian [Fermilab]↗

Accelerating iterative ptychography with an integrated neural network

Electron ptychography is a powerful and versatile tool for high-resolution and dose-efficient imaging. Iterative reconstruction algorithms are powerful but also computationally expensive due to their relative complexity and the many hyperparameters that must be optimised. Gradient descent-based iterative ptychography is a popular method, but it may converge slowly when reconstructing low spatial frequencies. Here, in this work, we present a method for accelerating a gradient descent-based iterative reconstruction algorithm by training a neural network (NN) that is applied in the reconstruction loop. The NN works in Fourier space and selectively boosts low spatial frequencies, thus enabling faster convergence in a manner similar to accelerated gradient descent algorithms. We discuss the difficulties that arise when incorporating a NN into an iterative reconstruction algorithm and show how they can be overcome with iterative training. We apply our method to simulated and experimental data of gold nanoparticles on amorphous carbon and show that we can significantly speed up ptychographic reconstruction of the nanoparticles.

4DSTEM↗

A High-Current Pulsed Prototype Power Supply

The Accelerator Controls Operations Research Network (ACORN) project aims to modernize the accelerator control system and replace aging power supplies at Fermilab. As part of this effort, outdated RF ferrite bias power supplies will be redesigned. These power supplies are essential for tuning the resonant frequency of RF cavities by delivering programmable current outputs of up to 2500 A and voltages ranging from −10 V to +35 V. They operate at a repetition rate of 15 Hz in the Booster ring, and 1 Hz at the Main Injector ring. The power supplies utilize a bank of transistors in the linear region, connected in parallel with the load, to actively regulate the output current from a 12-pulse SCR bridge. To support this upgrade, a new bias power supply topology was developed as proof of concept. The design utilizes an IGBT Hbridge operating in Pulse Width Modulation (PWM) mode, controlled by a microcontroller. A prototype, constructed using spare components, successfully delivered an output current of 500 A at a repetition rate of 15 Hz during initial testing. The circuit's bandwidth was measured at 480 Hz, highlighting opportunities for further optimization in the controller design to achieve the target bandwidth of 2 kHz.

Bullman, Austin [ORNL]↗

Accelerating Hamiltonian Monte Carlo for Bayesian inference in neural networks and neural operators

Hamiltonian Monte Carlo (HMC) is a powerful and accurate method to sample from the posterior distribution in Bayesian inference. However, HMC techniques are computationally demanding for Bayesian neural networks due to the high dimensionality of the network’s parameter space and the non-convexity of their posterior distributions. Therefore, various approximation techniques, such as variational inference (VI) or stochastic gradient MCMC, are often employed to infer the posterior distribution of the network parameters. Such approximations introduce inaccuracies in the inferred distributions, resulting in unreliable uncertainty estimates. In this work, we propose a hybrid approach that combines inexpensive VI and accurate HMC methods to efficiently and accurately quantify uncertainties in neural networks and neural operators. The proposed approach leverages an initial VI training on the full network. We examine the influence of individual parameters on the prediction uncertainty, which shows that a large proportion of the parameters do not contribute substantially to uncertainty in the network predictions. This information is then used to significantly reduce the dimension of the parameter space, and HMC is performed only for the subset of network parameters that strongly influence prediction uncertainties. This yields a framework for accelerating the full batch HMC for posterior inference in neural networks. We demonstrate the efficiency and accuracy of the proposed framework on deep neural networks and operator networks, showing that inference can be performed for large networks with tens to hundreds of thousands of parameters. Finally, we show that this method can effectively learn surrogates for complex physical systems by modeling the operator that maps from upstream conditions to wall-pressure data on a cone in hypersonic flow.

Bayesian inference↗