Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Ground Motion Inputs for the Seismic Shake Table Test

Currently, spent nuclear fuel (SNF) is stored in on-site independent spent-fuel storage installations (ISFSIs) at seventythree (73) nuclear power plants (NPPs) in the US. Because a site for geologic repository for permanent disposal of SNF has not been constructed, the SNF will remain in dry storage significantly longer than planned. During this time, the ISFSIs, and potentially consolidated storage facilities, will experience earthquakes of different magnitudes. The dry storage systems are designed and licensed to withstand large seismic loads. When dry storage systems experience seismic loads, there are little data on the response of SNF assemblies contained within them. The Spent Fuel Waste Disposition (SFWD) program is planning to conduct a full-scale seismic shake table test to close the gap related to the seismic loads on the fuel assemblies in dry storage systems. This test will allow for quantifying the strains and accelerations on surrogate fuel assembly hardware and cladding during earthquakes of different magnitudes and frequency content. The main component of the test unit will be the full-scale NUHOMS 32 PTH2 dry storage canister. The canister will be loaded with three surrogate fuel assemblies and twenty-nine dummy assemblies. Two dry storage configurations will be tested – horizontal and vertical above-ground concrete overpacks. These configurations cover 91% of the current dry storage configurations. The major input into the shake table test are the seismic excitations or the earthquake ground motions – acceleration time histories in two horizontal and one vertical direction that will be applied to the shake table surface during the tests. The shake table surface represents the top of the concrete pad on which a dry storage system is placed. The goal of the ground motion task is to develop the ground motions that would be representative of the range of seismotectonic and other conditions that any site in the Western US (WUS) or Central Eastern US (CEUS) might entail. This task is challenging because of the large number of the ISFSI sites, variety of seismotectonic and site conditions, and effects that soil amplification, soil-structure interaction, and pad flexibility may have on the ground motions.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Position Papers for the ASCR Workshop on Reimagining Codesign

On behalf of the Advanced Scientific Computing Research (ASCR) program in the US Department of Energy (DOE) Office of Science, we are organizing a Workshop on Reimagining Codesign (ReCoDe). Codesign is the process of jointly designing interoperating components of a computing system—in particular: applications, algorithms, system software, programming models, and the hardware on which they run. The goal is to maximize the overall performance, efficiency, and other desirable qualities of the system as a whole. Codesign is a standard methodology in the embedded-systems community, where space, power, and cost constraints are commonly pitted against execution speed for a tightly constrained feature set. Over the last decade, the DOE has invested in codesign efforts to foster the development of exascale computing systems for broad classes of scientific and engineering applications. The ReCoDe workshop hopes to explore how scientific applications of interest to the DOE can be accelerated through close interactions with hardware designers and software-stack developers, in which all components adapt to each other’s requirements and constraints. We want to answer the question of what are the key tools and methodologies for accomplishing codesign in today’s computing landscape, and what will be the highest impact targets for meeting DOE’s emerging mission requirements. This workshop aims to bring together DOE, industry, and academia to identify opportunities to build on past codesign successes and identify new areas that are either emerging or that may need reimagining for the future. We want to continue to find opportunities that can be pursued as a joint effort and continue to break down the traditional customer/vendor dichotomy with true partnerships. From this work, DOE will benefit from increased application performance relative to what stock hardware or existing general-purpose roadmaps can provide, and vendors will benefit from expanding their hardware’s capabilities to address needs they might have not otherwise anticipated and thereby create more widespread interest in their products. The workshop will be structured around a set of breakout sessions, with every attendee expected to participate actively in the discussions. Afterward, workshop attendees—from DOE, industry, and academia—will produce a report for ASCR that summarizes the findings made during the workshop.

97 MATHEMATICS AND COMPUTING↗

Machine learning for reducing noise in RF control signals at industrial accelerators

Industrial particle accelerators typically operate in dirtier environments than research accelerators, leading to increased noise in RF and electronic systems. Furthermore, given that industrial accelerators are mass produced, less attention is given to optimizing the performance of individual systems. As a result, industrial accelerators tend to underperform their own hardware capabilities. Improving signal processing for these machines will improve cost and time margins for deployment, helping to meet the growing demand for accelerators for medical sterilization, food irradiation, cancer treatment, and imaging. Our work focuses on using machine learning techniques to reduce noise in RF signals used for pulse-to-pulse feedback in industrial accelerators. Here we review our algorithms and observed results for simulated RF systems, and discuss next steps with the ultimate goal of deployment on industrial systems.

43 PARTICLE ACCELERATORS↗

ORCHA: A performance portability system for extreme heterogeneity

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Lee, Youngjun↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

Using Likwid and Byfl to Benchmark Hardware Performance

This paper outlines a benchmarking study conducted during my internship at LANL, focusing on CPU (Computer Processing Unit) and program performance assessment. The primary goal was to gather memory access data using three methods across five polybench kernels The data gathered would then be used to compare and contrast to one another and calculate operational intensity for performance comparisons. Benchmarking tools like Byfl and Likwid were employed, with Byfl offering hardware-independent data through LLVM compiler communication and Likwid directly interacting with computer hardware. The study considered various benchmarking factors, including optimization levels, Big O notation ((n)), CPU diversity and specific kernel equations. Big O notation was utilized to simplify code complexity, with detailed breakdwons of operations and memory components for each polybench application. Specific O(n) equations enabled nuanced kernel compariosns, facilitating the identification of performance variations. CPU efficiency assessments were conducted using Likwid tests on two CPUs. The central focus on code optimization aimed at achieving higher speeds and reduced memory usage through streamlined code. Future work propsoes creating a roofline model, synthesizing benchmarking data into a comprehensive data graph to assist in optimizing code and improving hardware performance. The potential impact on the laboratory or national mission was underscored, emphasizing the importance of optimizing applications and hardware to conserve resources and accelerate program execution. The specific relevance to LANL’s operations in math-intensive fields such as Nuclear Fission, Space Exploration, and Nanotechnology highlights the necessity of efficient benchmarking for resource conservation and proram speed. Overall, this study contributes to the understanding of CPU and program performance, providing insights for future optimization efforts in a laboratory setting

97 MATHEMATICS AND COMPUTING↗

Final Seismic Shake Table Test Plan

The Spent Fuel Waste Disposition (SFWD) program is planning to conduct a full-scale seismic shake table test on the dry storage systems of spent nuclear fuel (SNF) to close the gap related to seismic loads on fuel assemblies in dry storage systems. This test will allow for quantifying the strains and accelerations on surrogate fuel assembly hardware and cladding during earthquakes of different magnitudes and frequency content. Full-scale testing is needed because a dry storage system is a complex and highly nonlinear system making it hard to predict (model) the responses to seismic excitations. The non-linearity arises from the multiple spatial gaps in the system – between fuel rods and the basket, between the basket and dry storage canister, between the dry storage canister and the storage cask (overpack), and ventilation gaps. The non-linearities pose significant limitations on the value of tests with scaled systems.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗

Multichannel Analysis of Surface Waves Accelerated (MASWAccelerated): Software for efficient surface wave inversion using MPI and GPUs

Multichannel Analysis of Surface Waves (MASW) is a technique frequently used in geotechnical engineering and engineering geophysics to infer 1D layered models of seismic shear wave velocities in the top tens to hundreds of meters of the subsurface. We aim to accelerate MASW calculations by capitalizing on modern computer hardware available in the workstations of most engineers: multiple cores and graphics processing units (GPUs). We propose new parallel and GPU accelerated algorithms for computing 1D MASW inversion, and provide software implementations in C using Message Passing Interface (MPI) and CUDA. These algorithms take advantage of sparsity that arises in the problem, and the work balance between processes considers typical data trends. We compare our methods to an existing open source Matlab MASW tool. Our serial C implementation achieves a 2x speedup over the Matlab software, and we continue to see improvements by parallelizing the problem with MPI. Here we see nearly perfect strong and weak scaling for uniform data, and improve strong scaling for realistic data by repartitioning the problem to process mapping. By utilizing GPUs available on most modern workstations, we observe an additional 1.3x speedup over the serial C implementation on the first use of the method. We typically repeatedly evaluate theoretical dispersion curves as part of an optimization procedure, and on the GPU the kernel can be cached for faster reuse on later runs. We observe a 3.2x speedup on the cached GPU runs compared to the serial C runs. This work is the first open-source parallel or GPU-accelerated software tool for MASW imaging, and should enable geotechnical engineers to fully utilize all computer hardware at their disposal.

58 GEOSCIENCES↗

EXPERIMENTAL STUDIES OF NONLINEAR INTEGRABLE OPTICS

State-of-the-art accelerators at energy and intensity frontiers require increasingly bright and powerful particle beams. In conventional linear lattices, intense beams suffer from collective instabilities, resulting in beam losses and maximum beam intensity limits. This thesis presents experimental studies of a novel lattice design concept, the nonlinear integrable optics (NIO), aimed at enhancing beam stability limits with little to no beam dynamics degradation. Single-particle beam dynamics measurements of two NIO devices, the quasi-integrable octupole system and the fully integrable Danilov-Nagaitsev system, were carried out at the purpose-built Integrable Optics Test Accelerator (IOTA) at Fermilab. Their simulation, hardware design, and commissioning process are presented. Extensive model and analysis algorithm development and benchmarking is described. Electron beam data from two scientific runs is analyzed, yielding frequency and phase space dynamics consisten t with m odels. These results demonstrate viability and advantages of the NIO design, providing the groundwork for proton studies in the strong space-charge regime and future integrable accelerators.

Kuklev, Nikita↗

FPGA-based HPC accelerators: An evaluation on performance and energy efficiency

Hardware specialization is a promising direction for the future of digital computing. Reconfigurable technologies enable hardware specialization with modest non-recurring engineering cost, but their performance and energy efficiency compared to state-of-the-art processor architectures remain an open question. In this article, we use FPGAs to evaluate the benefits of building specialized hardware for numerical kernels found in scientific applications. In order to properly evaluate performance, we not only compare Intel Arria 10 and Xilinx U280 performance against Intel Xeon, Intel Xeon Phi, and NVIDIA V100 GPUs, but we also extend the Empirical Roofline Toolkit (ERT) to FPGAs in order to assess our results in terms of the Roofline model. We show design optimization and tuning techniques for peak FPGA performance at reasonable hardware usage and power consumption. As FPGA peak performance is known to be far less than that of a GPU, we also benchmark the energy efficiency of each platform for the scientific kernels comparing against microbenchmark and technological limits. Results show that while FPGAs struggle to compete in absolute terms with GPUs on memory- and compute-intensive kernels, they require far less power and can deliver nearly the same energy efficiency.

97 MATHEMATICS AND COMPUTING↗

Real-Time GPU-Accelerated OFDR With an Integrated Auxiliary Interferometer

A GPU-accelerated optical frequency domain reflectometry (OFDR) system with an improved integrated auxiliary interferometer is proposed. Unlike conventional approaches that require separate auxiliary interferometers and multiple detection channels, the proposed OFDR system embeds this functionality directly into the signal via an intentional beat component. This enables self-calibration of laser nonlinearity while maintaining a cost-effective hardware configuration. Building on this simplified configuration, the system leverages GPU acceleration with an NVIDIA RTX 4070 Ti to achieve real-time performance, delivering high-throughput signal processing for continuous OFDR interrogation. The signal processing pipeline comprises signal capture, resampling for nonlinearity compensation, and frequency shift computation, all optimized for parallel execution. Hardware benchmarking demonstrates substantial acceleration over CPU implementations, achieving up to a 45× speedup for resampling and frequency shift computations and enabling processing latencies below 30 ms. Thermal response validation is conducted under two complementary scenarios: localized heating using a water bath and cryogenic-temperature conditions using liquid nitrogen. Under localized heating, the system achieves an accuracy of 0.249 °C with a thermal sensitivity of 5.971 GHz/°C, while cryogenic-temperature validation demonstrates a frequency shift response with a sensitivity of 2.383 GHz/°C and an accuracy of 2.04 °C. The high acceleration of the proposed GPU-accelerated OFDR system and its accuracy are achieved by exploiting CUDA-based stride indexing, enabling efficient parallel segmentation and processing of large datasets without additional memory copies. The benchmarking results confirm the robustness, accuracy, and deployability of the proposed OFDR system across a wide temperature range, establishing it as a practical platform for real-time distributed fiber sensing in structurally dynamic environments.

Harb, Salah [Lawrence Berkeley National Laboratory↗

Evaluation of Portable Acceleration Solutions for LArTPC Simulation Using Wire-Cell Toolkit

The Liquid Argon Time Projection Chamber (LArTPC) technology plays an essential role in many current and future neutrino experiments. Accurate and fast simulation is critical to developing efficient analysis algorithms and precise physics model projections. The speed of simulation becomes more important as Deep Learning algorithms are getting more widely used in LArTPC analysis and their training requires a large simulated dataset. Heterogeneous computing is an efficient way to delegate computationally intensive tasks to specialized hardware. However, as the landscape of compute accelerators quickly evolves, it becomes increasingly difficult to manually adapt the code to the latest hardware or software environments. A solution which is portable to multiple hardware architectures without substantially compromising performance would thus be very beneficial, especially for long-term projects such as the LArTPC simulations. In search of a portable, scalable and maintainable software solution for LArTPC simulations, we have started to explore high-level portable programming frameworks that support several hardware backends. In this paper, we present our experience porting the LArTPC simulation code in the Wire-Cell Toolkit to NVIDIA GPUs, first with the CUDA programming model and then with a portable library called Kokkos. Preliminary performance results on NVIDIA V100 GPUs and multi-core CPUs are presented, followed by a discussion of the factors affiecting the performance and plans for future improvements.

Yu, Haiwang↗

Portability: A Necessary Approach for Future Scientific Software

Today's world of scientific software for High Energy Physics (HEP) is powered by x86 code, while the future will be much more reliant on accelerators like GPUs and FPGAs. The portable parallelization strategies (PPS) project of the High Energy Physics Center for Computational Excellence (HEP/CCE) is investigating solutions for portability techniques that will allow the coding of an algorithm once, and the ability to execute it on a variety of hardware products from many vendors, especially including accelerators. We think without these solutions, the scientific success of our experiments and endeavors is in danger, as software development could be expert driven and costly to be able to run on available hardware infrastructure. We think the best solution for the community would be an extension to the C++ standard with a very low entry bar for users, supporting all hardware forms and vendors. We are very far from that ideal though. We argue that in the future, as a community, we need to request and work on portability solutions and strive to reach this ideal.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SRF Cavity Emulator for PIP-II LLRF Lab and Field Testing

There are many stages in the LLRF and RF system development process for any new accelerator that can take advantage of hardware emulation of the high-power RF system and RF cavities. LLRF development, bench testing, control system development and testing of installed systems must happen well before SRF cavities are available for test. The PIP-II Linac has three frequencies of SRF cavities, 162.5 MHz, 325 MHz and 650 MHz and a simple analog emulator design has been chosen that can meet the cavity bandwidth requirements, provide tuning errors to emulate Lorentz force detuning and microphonics for all cavity types. This emulator design utilizes a quartz crystal with a bandwidth of 65Hz at an IF of ~ 4 MHz, providing a Q of ~ 1.3 x 10^7 at 650MHz. This paper will discuss the design and test results of this emulator.

43 PARTICLE ACCELERATORS↗

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

VWC-BERT: Scaling Vulnerability–Weakness–Exploit Mapping on Modern AI Accelerators

Defending cybersystems needs accurate mapping of software and hardware vulnerabilities to generalized descriptions of weaknesses, and weaknesses to exploits. These mappings enable cyber defenders to build plans for effective defense and assessment of potential risks to a cybersystem. With close to 170k vulnerabilities, manual mapping is not a feasible option. However, automated mapping is challenging due to limited training data, computational intractability, and limitations in computational natural language processing. Tools based on breakthroughs in Transformer-based language models have been demonstrated to classify vulnerabilities with high accuracy. We make three key contributions in this paper: (1) We present a new framework, \VWCBERT, that augments the Transformer-based hierarchical multi-class classification framework of Das et al. (\textsc{V2W-BERT}) with the ability to map weaknesses to exploits. (2) We implement \VWCBERT~ on modern AI accelerator platforms using two data parallel techniques for the pre-training phase and demonstrate nearly linear speedups across NVIDIA and Graphcore accelerator platforms. We observe nearly linear speedups for up to 16 V100 and 8 A100 GPUs, and about 3.4$\times$ speedup for A100 relative to V100 GPUs. We also observe excellent speedups on Graphcore, with $5.7\times$ speedup on 128 IPUs relative to 16 IPUs. Enabled by scaling, we also demonstrate higher accuracy using a larger language model, RoBERTa-Large. We show up to 87\% accuracy for strict and up to 98\% accuracy for relaxed classification. (3) We develop a novel parallel link manager for the link prediction phase and demonstrate up to 21$\times$ speedup with 16 V100 GPUs relative to one V100 GPU, and thus reducing the runtime from 2.5 hours to 10 minutes. We believe that generalizability and scalability of \VWCBERT~ will benefit both the theoretical development and practical deployment of novel cyberdefense solutions and vulnerability classification.

Das, Siddhartha Shankar↗