Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “accelerate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

The optimal beam-loading in two-bunch nonlinear plasma wakefield accelerators

Abstract Due to the highly nonlinear nature of the beam-loading, it is currently not possible to analytically determine the beam parameters needed in a two-bunch plasma wakefield accelerator for maintaining a low energy spread. Therefore in this paper, by using the Broyden–Fletcher–Goldfarb–Shanno algorithm for the parameter scanning with the code QuickPIC and the polynomial regression together with k -fold cross-validation method, we obtain two fitting formulas for calculating the parameters of tri-Gaussian electron beams when minimizing the energy spread based on the beam-loading effect in a nonlinear plasma wakefield accelerator. One formula allows the optimization of the normalized charge per unit length of a trailing beam to achieve the minimal energy spread, i.e. the optimal beam-loading. The other one directly gives the transformer ratio when the trailing beam achieves the optimal beam-loading. A simple scaling law for charges of drive beams and trailing beams is obtained from the fitting formula, which indicates that the optimal beam-loading is always achieved for a given charge ratio of the two beams when the length and separation of two beams and the plasma density are fixed. The formulas can also help obtain the optimal plasma densities for the maximum accelerated charge and the maximum acceleration efficiency under the optimal beam-loading respectively. These two fitting formulas will significantly enhance the efficiency for designing and optimizing a two-bunch plasma wakefield acceleration stage.

Physics↗

Experimental demonstration of accelerating a beam with a large transverse emittance ratio in the relativistic heavy ion collider for the electron-ion collider

The electron-ion collider (EIC), to be constructed at Brookhaven National Laboratory, will collide polarized high-energy electron beams with hadron beams, achieving luminosities of up to 1.0 × 10 34 cm −2 s −1 in the center-of-mass energy range of 20–140 GeV. To reach such high luminosity, the EIC will employ small, flat beams at the interaction point. According to the design of the EIC hadron storage ring (HSR), hadron beams with a large transverse emittance ratio of 11:1 will be generated at the injection energy using an electron cooling technique and then accelerated to high energies for collisions. Accelerating hadron beams with such a large emittance ratio had never been demonstrated elsewhere—until our recent beam experiment at the relativistic heavy ion collider (RHIC). In this experiment, we successfully generated a large transverse emittance ratio of 13:1 with a gold-ion beam at 31 GeV/nucleon using stochastic cooling. We then accelerated this beam, with a transverse emittance ratio of 11:1, from 31 to 100 GeV/nucleon. Thanks to RHIC’s high-performance orbit, tune, and decoupling feedback systems, the large emittance ratio was well maintained throughout the 5-min-long acceleration process. This experiment fully validated the EIC/HSR design assumptions—namely, that large-emittance-ratio hadron beams can be generated at injection energy and then accelerated to high energies for collisions.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Ion acceleration from microstructured targets irradiated by high-intensity picosecond laser pulses

Structures on the front surface of thin foil targets for laser-driven ion acceleration have been proposed to increase the ion source maximum energy and conversion efficiency. While structures have been shown to significantly boost the proton acceleration from pulses of moderate-energy fluence, their performance on tightly focused and high-energy lasers remains unclear. Here, we report the results of laser-driven three-dimensional (3D)-printed microtube targets, focusing on their efficacy for ion acceleration. Using the high-contrast (~10 12) PHELIX laser (150 J, 10 21 W / cm 2 ), we studied the acceleration of ions from 1-μm-thick foils covered with micropillars or microtubes, which we compared with flat foils. The front-surface structures significantly increased the conversion efficiency from laser to light ions, with up to a factor of 5 higher proton number with respect to a flat target, albeit without an increase of the cutoff energy. An optimum diameter was found for the microtube targets. Our findings in this work are supported by a systematic particle-in-cell modeling investigation of ion acceleration using 2D simulations with various structure dimensions. Simulations reproduce the experimental data with good agreement, including the observation of the optimum tube diameter, and reveal that the laser is shuttered by the plasma filling the tubes, explaining why the ion cutoff energy was not increased in this regime.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Effects of variable deceleration periods on Rayleigh-Taylor instability with acceleration reversals

The dynamics of an interfacial flow that is initially Rayleigh-Taylor unstable but becomes statically stable for some intermediate period due to the reversal of the externally imposed acceleration field is studied. We discuss scenarios that consider both single and double-acceleration reversals. The accel-decel (AD) case consists of a single reversal imposed at an instant after the constant acceleration instability has entered a self-similar regime. The layer of mixed fluid ceases to grow upon acceleration reversal, and the dominant mechanics are due to internal wave oscillations. Variation of mass flux and the Reynolds stress anisotropy is observed due to the action of the internal waves. Here, a second reversal of the AD case that is termed as accel-decel-accel, ADA is then explored; the response of the mixing layer is shown to depend strongly on the duration and the periodicity of the Reynolds stress anisotropy of the mixing layer during the deceleration period. We explore the effect of this variable deceleration period after the second acceleration reversal where the flow once again becomes Rayleigh-Taylor unstable based on metrics that include the integral mixing-layer width, bubble and spike amplitudes, mass flux, Reynolds stress anisotropy tensor, and the molecular mixing parameter.

42 ENGINEERING↗

Dephasingless Laser Wakefield Acceleration

Laser wakefield accelerators (LWFAs) produce significantly high gradients enabling compact accelerators and radiation sources, but face design limitations, such as dephasing, occurring when trapped electrons outrun the accelerating phase of the wakefield. In this report we combine spherical aberration with a novel cylindrically symmetric echelon optic to spatiotemporally structure an ultra-short, high-intensity laser pulse that can overcome dephasing by propagating at any velocity over any distance. The ponderomotive force of the spatiotemporally shaped pulse can drive a wakefield with a phase velocity equal to the speed of light in vacuum, preventing trapped electrons from outrunning the wake. Simulations in the linear regime and scaling laws in the bubble regime illustrate that this dephasingless LWFA can accelerate electrons to high energies in much shorter distances than a traditional LWFA—a single 4.5 m stage can accelerate electrons to TeV energies without the need for guiding structures.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Galactic Potential and Dark Matter Density from Angular Stellar Accelerations

Here, we present an approach to measure the Milky Way (MW) potential using the angular accelerations of stars in aggregate as measured by astrometric surveys like Gaia. Accelerations directly probe the gradient of the MW potential, as opposed to indirect methods using, e.g., stellar velocities. We show that end-of-mission Gaia stellar acceleration data may be used to measure the potential of the MW disk at approximately 3σ significance and, if recent measurements of the solar acceleration are included, the local dark matter density at ~2σ significance. Since the significance of detection scales steeply as t 5/2 for observing time t, future surveys that include angular accelerations in the astrometric solutions may be combined with Gaia to precisely measure the local dark matter density and shape of the density profile.

79 ASTRONOMY AND ASTROPHYSICS↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

Rayleigh–Taylor Instability With Varying Periods of Zero Acceleration

We present our findings from a numerical investigation of the acceleration-driven Rayleigh–Taylor Instability, modulated by varying periods without an applied acceleration field. It is well known from studies on shock-driven Richtmyer–Meshkov instability that mixing without external forcing grows with a scaling exponent as ≈ t 0.20-0.28 When the Rayleigh–Taylor Instability is subjected to varying periods of “zero” acceleration, the structural changes to the mixing layer remain remarkably small. After the acceleration is re-applied, the mixing layer quickly resumes the profile of development it would have had if there had been no intermission. As a result, this behavior contrasts in particular with the strong sensitivity that is found to other variable acceleration profiles examined previously in the literature.

42 ENGINEERING↗

Evaluating Emerging AI/ML Accelerators: IPU, RDU, and and NVIDIA/AMD GPUs

As the size and complexity of AI/ML workloads continue to increase, the demand for specialized hardware accelerators has grown rapidly. To better understand the landscape of commercial AI/ML accelerators, we evaluate and compare several popular platforms, including Graphcore IPU, Sambanova RDU, and GPUs, on a range of AI workloads. Our research aims to help researchers and developers understand the unique features of commercial AI/ML accelerators and provide reference design and performance numbers for research prototypes. By shedding light on the current landscape of commercial AI/ML accelerators, we hope to offer insights into the future development of specialized hardware accelerators for AI/ML.

Peng, Hongwu↗

Electrolyzer Performance Loss from Accelerated Stress Tests and Corresponding Changes to Catalyst Layers and Interfaces

Stress tests are developed for proton exchange membrane electrolyzers that utilize low catalyst loading, elevated potential, and frequent cycling with square- and triangle-waves to accelerate anode catalyst layer degradation during intermittent operation. Kinetics drive performance losses (ohmic/transport secondary) and are accompanied by decreasing exchange current density, decreasing cyclic voltammetric capacitance, and increasing polarization resistance. Decreased kinetics are likely due to a combination of iridium (Ir) migration into electrochemically inaccessible locations in the anode or membrane, Ir particle growth (supported by X-ray scattering), changes in the extent of the Ir oxidation state (supported by X-ray absorption spectroscopy), and anode catalyst layer reordering. Decreasing catalyst/transport layer contact and catalyst/membrane interfacial tearing may add contact resistances and account for increasing ohmic losses. Performance losses for low and moderate catalyst loading, as well as from accelerated and model wind/solar cycling protocols, were likewise dominated by kinetics but vary in severity. Additionally, accelerated cycling (1 cycle per minute) appears to reasonably accelerate relevant loss mechanisms and can be used to project electrolyzer lifetime from anode deterioration. Ongoing accelerated stress test development and studies into performance loss mechanisms will continue to be critical as electrolysis shifts to intermittent power and low-cost applications.

30 DIRECT ENERGY CONVERSION↗

Assessment of the Pan-Arctic Accelerated Rate of Sea Ice Decline in CMIP6 Historical Simulations

Abstract Arctic sea ice loss in response to a warming climate is assessed in 42 models participating in phase 6 of the Coupled Model Intercomparison Project (CMIP6). Sea ice observations show a significant acceleration in the rate of decline commencing near the turn of the twenty-first century. It is our assertion that state-of-the-art climate models should qualitatively reflect this accelerated trend within the limitations of internal variability and observational uncertainty. Our analysis shows that individual CMIP6 simulations of sea ice depict a wide range of model spread on biases and anomaly trends both across models and among their ensemble members. While the CMIP6 multimodel mean captures the observed sea ice area (SIA) decline relatively well, an individual model’s ability to represent the acceleration in sea ice decline remains a challenge. Seventeen (40%) out of 42 CMIP6 models and 37 (13%) out of the total 286 ensemble members reasonably capture the observed trends and acceleration in SIA decline. In addition, a larger ensemble size appears to increase the odds for a model to include at least one ensemble member skillfully representing the accelerated SIA trends. Simulations of sea ice volume (SIV) show much larger spread and uncertainty than SIA; however, due to limited availability of sea ice thickness data, these are not as well constrained by observations. Finally, we find that models with more ocean heat transport simulate larger sea ice declines, which suggests an emergent constraint in CMIP6 ensembles. This relationship points to the need for better understanding and modeling of ice–ocean interactions, especially with respect to frazil ice growth.

54 ENVIRONMENTAL SCIENCES↗

Advancing Conduction-Cooled 650 MHz SRF Technology for Industrial Accelerators at Fermilab's IARC

The National Nuclear Security Administration (NNSA) funds the Illinois Accelerator Research Center (IARC) at Fermilab in developing a high-power, conduction-cooled Superconducting Radio Frequency (SRF) accelerator tailored for industrial applications requiring robust and efficient operation. A 650 MHz, 1.6 MeV, 20 kW SRF accelerator is currently under development, employing a conduction cooling approach to simplify cryogenic requirements and enhance accessibility for industrial use. The accelerator’s control system is implemented on the Blinky Lite platform, selected for its open-source architecture, secure remote access capabilities, and operational flexibility—attributes advantageous for industrial deployment and sustained operation. A dedicated beamline is designed to measure essential beam parameters and test the integrated performance of the accelerator and control systems, thereby validating their operational readiness for intended applications

Ji, Y. [Fermilab] (ORCID:0000000233981752)↗

Consideration of HTS rapid-cycling magnet for staged muon acceleration

The HTS conductor hysteresis dominates magnet cable power loss but is independent of the magnetic field ramping rate. This makes the HTS conductor suitable to power the rapid-cycling accelerator magnet. We present a possible application of the HTS rapid-cycling magnet as outlined in [1,2] for the staged muon acceleration including the front-end Recirculating Linear Accelerator and the followed-up Rapid Cycling Synchrotrons delivering the muon beams to the Muon Collider.[1] H. Piekarz, S. Otten, A. Kario, H. ten Kate, “Rapid-cycling HTS magnet for muon acceleration”, US MC Inaugural Meeting, FERMILAB-POSTER-24-0219-AD, August 7-9, 2024[2] H. Piekarz, B. Claypool, S. Hays, M. Kufer, V. Shiltsev, “Record High Ramping Rates in HTS Based Supercond. Accelerator Magnet”, MT 27, IEEE Trans. on Applied Superccond, 32 (2022) 6, 4100404

Piekarz, Henryk [Fermilab]↗

Accelerated Irradiation and Qualification of Ceramic Nuclear Fuels

Accelerated neutron irradiation testing is an component of accelerated qualification of new nuclear fuels for light water reactor (LWRs), microreactors, and other special purpose reactors. The qualification and licensing of nuclear fuel is a lengthy process that can take 20-25 years to bring a new fuel into service. Accelerated fuel qualification combines both experimental and modeling work to expedite the total qualification time to 5-10 years timeframe. The experimental aspect of this is accelerated irradiation aims to reduce the total time needed for neutron irradiation to achieve targeted burnup, which can take years using conventional irradiation profiles. The data that results from this irradiation testing can then be entered into BISON models to develop robust and reliable performance simulations to ensure safe operation under normal and off normal conditions. This milestone focused on the fabrication of test articles for accelerated irradiation testing at the Advance Test Reactor (ATR).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

Upgrading Fermilab s Accelerator Control System with ACORN

The Fermilab Accelerator Complex is the largest national user facility in the Office of High Energy Physics (DOE/HEP) program and the only national user facility operating at Fermilab. Fermilab serves as the host to the Long Baseline Neutrino Facility/Deep Underground Neutrino Experiment (LBNF/DUNE), the laboratory’s flagship project for neutrino science that is under construction. LBNF/DUNE will be powered by megawatt beams from an upgraded accelerator, the Proton Improvement Plan II (PIP-II) that will replace the laboratory’s aging linear accelerator with a new one based on superconducting radio-frequency cavities. The Accelerator Controls Operations Research Network (ACORN) Project will support LBNF/DUNE and PIP-II by modernizing the accelerator control system. The project is at the conceptual design phase and looking to achieve Critical Decision 1 (CD-1) later this year. The scope and structure of the project will be presented, along with an overview of how that has changed in the past year. Current design and technology choices will be shared. Specific challenges facing the project will be addressed, along with current thinking on solutions.

Roehrig, Christian [Fermilab]↗

Advancing Conduction-Cooled 650 MHZ SRF Technology for Industrial Accelerators at Fermilab S IARC

The National Nuclear Security Administration (NNSA) funds the Illinois Accelerator Research Center (IARC) at Fermilab in developing a high-power, conduction-cooled Superconducting Radio Frequency (SRF) accelerator tailored for industrial applications requiring robust and efficient operation. A 650 MHz, 1.6 MeV, 20 kW SRF accelerator is currently under development, employing a conduction cooling approach to simplify cryogenic requirements and enhance accessibility for industrial use. The accelerator's control system is implemented on the Blinky Lite platform, selected for its open-source architecture, secure remote access capabilities, and operational flexibility attributes advantageous for industrial deployment and sustained operation. A dedicated beamline is designed to measure essential beam parameters and test the integrated performance of the accelerator and control systems, thereby validating their operational readiness for intended applications.

Ji, Yichen [Fermilab]↗

FPGA Acceleration of GCN in Light of the Symmetry of Graph Adjacency Matrix

Graph Convolutional Neural Networks (GCNs) are widely used to process large-scale graph data. Different from deep neural networks (DNNs), GCNs are sparse, irregular, and unstructured, posing unique challenges to hardware acceleration with regular processing elements (PEs). In particular, the adjacency matrix of a GCN is extremely sparse, leading to frequent but irregular memory access, low spatial/temporal data locality and poor data reuse. Furthermore, a realistic graph usually consists of unstructured data (e.g., unbalanced distributions), creating significantly different processing times and imbalanced workload for each node in GCN acceleration. To overcome these challenges, we propose an end-to-end hardware-software co-design to accelerate GCNs on resource-constrained FPGAs with the features including: (1) A custom dataflow that leverages symmetry along the diagonal of the adjacency matrix to accelerate feature aggregation for undirected graphs. We utilize either the upper or the lower triangular matrix of the adjacency matrix to perform aggregation in GCN to improve data reuse. (2) Unified compute cores for both aggregation and transform phases, with full support to the symmetry-based dataflow. These cores can be dynamically reconfigured to the systolic mode for transformation or as individual accumulators for aggregation in GCN processing. (3) Preprocessing of the graph in software to rearrange the edges and features to match the custom dataflow. This step improves the regularity in memory access and data reuse in the aggregation phase. Moreover, we quantize the GCN precision from FP32 to INT8 to reduce the memory footprint without losing the inference accuracy. We implement our accelerator design in Intel Stratix10 MX FPGA board with HBM2, and demonstrate 1.3x-110.5x improvement in end-to-end GCN latency as compared to the state-of the-art FPGA implementations, on the graph datasets of Cora, Pubmed, Citeseer and Reddit.

Nair, Gopikrishnan R.↗