Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Errant Beam Detection Using the AMD Versal ACAP and Vitis AI

The prevalence of ML and AI-powered solutions along with the slowing of Moore's Law has given rise to novel hardware platforms aimed at accelerating ML and AI. While programming these hardware platforms can be difficult, particularly for non-hardware experts, hardware vendors provide high-level tooling in an effort to address this difficulty. The Versal ACAP is an SoC designed by AMD that combines CPU cores, FPGA fabric, and a tiled, vector architecture called an AI engine all on the same socket. In an effort to more easily program this heterogeneous system, AMD has provided the Vitis AI development stack. In this work, we leverage Vitis AI to program a Versal ACAP to perform errant beam detection in the Spallation Neutron Source at Oak Ridge National Laboratory. Our initial work shows that after quantization and compilation of the model for the Versal ACAP, the classification accuracy, as measured by the AUC metric, is over 95% accurate while achieving this accuracy in 46 microseconds on average.

Cabrera, Anthony↗

Towards Automatic and Agile AI/ML Accelerator Design with End-to-End Synthesis

Domain-specific designs offer greater energy efficiency and performance gain than general-purpose processors. For this reason, modern system-on-chips have a significant portion of their silicon area with custom accelerators. However, designing hardware by hand is laborious and time-consuming, given the large design space and the performance, power, and area constraints that are not realized in the software. Moreover, domain-specific algorithms (e.g., machine learning models) are evolving quickly, challenging the accelerator design further. To address these issues, this paper presents SODA Synthesizer, an automated open-source high-level ML framework to Verilog modular compiler targeting AI/ML Application-Specific Integrated Circuits (ASICs) accelerators. SODA tightly couples the Multi- Level Intermediate Representation (MLIR) compiler infrastructure [24] and open-source HLS approaches. Thus, SODA can support various ML frameworks and algorithms and can perform optimizations that combine specialized architecture templates and conventional HLS to generate the hardware modules. In addition, SODA’s closed-loop design space exploration (DSE) engine allows developers to perform end-to-end design space explorations on different metrics and technology nodes.

Zhang, Jeff↗

Composite Overwrapped Pressure Vessels (COPV): Developing Flight Rationale for the Space Shuttle Program

Introducing composite vessels into the Space Shuttle Program represented a significant technical achievement. Each Orbiter vehicle contains 24 (nominally) Kevlar tanks for storage of pressurized helium (for propulsion) and nitrogen (for life support). The use of composite cylinders saved 752 pounds per Orbiter vehicle compared with all-metal tanks. The weight savings is significant considering each Shuttle flight can deliver 54,000 pounds of payload to the International Space Station. In the wake of the Columbia accident and the ensuing Return to Flight activities, the Space Shuttle Program, in 2005, re-examined COPV hardware certification. Incorporating COPV data that had been generated over the last 30 years and recognizing differences between initial Shuttle Program requirements and current operation, a new failure mode was identified, as composite stress rupture was deemed credible. The Orbiter Project undertook a comprehensive investigation to quantify and mitigate this risk. First, the engineering team considered and later deemed as unfeasible the option to replace existing all flight tanks. Second, operational improvements to flight procedures were instituted to reduce the flight risk and the danger to personnel. Third, an Orbiter reliability model was developed to quantify flight risk. Laser profilometry inspection of several flight COPVs identified deep (up to 20 mil) depressions on the tank interior. A comprehensive analysis was performed and it confirmed that these observed depressions were far less than the criterion which was established as necessary to lead to liner buckling. Existing fleet vessels were exonerated from this failure mechanism. Because full validation of the Orbiter Reliability Model was not possible given limited hardware resources, an Accelerated Stress Rupture Test of a flown flight vessel was performed to provide increased confidence. A Bayesian statistical approach was developed to evaluate possible test results with respect to the model credibility and thus flight rationale for continued operation of the Space Shuttle with existing flight hardware. A non-destructive evaluation (NDE) technique utilizing Raman Spectroscopy was developed to directly measure the overwrap residual stress state. Preliminary results provide optimistic results that patterns of fluctuation in fiber elastic strains over the outside vessel surface could be directly correlated with increased fiber stress ratios and thus reduced reliability.

Kezirian, Michael T.↗

Machine Learning on Heterogeneous, Edge, and Quantum Hardware for Particle Physics (ML-HEQUPP)

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for scientific discovery demands real-time inference and decision-making, intelligent data reduction, and efficient processing architectures beyond current capabilities. Crucial to the success of this experimental paradigm are several emerging technologies, such as artificial intelligence and machine learning (AI/ML) and silicon microelectronics, and the advent of quantum algorithms and processing. Their intersection includes areas of research such as low-power and low-latency devices for edge computing, heterogeneous accelerator systems, reconfigurable hardware, novel codesign and synthesis strategies, readout for cryogenic or high-radiation environments, and analog computing. This white paper presents a community-driven vision to identify and prioritize research and development opportunities in hardware-based ML systems and corresponding physics applications, contributing towards a successful transition to the new data frontier of fundamental science.

Gonski, Julia [SLAC]↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

A Length Adaptive Algorithm-Hardware Co-design of Transformer on FPGA Through Sparse Attention and Dynamic Pipelining

Transformers are considered one of the most important deep learning models since 2018, in part because it establishes state-of-the-art (SOTA) records and could potentially replace existing Deep Neural Networks (DNNs). Despite the remarkable triumphs, the prolonged turnaround time of Transformer models is a widely recognized roadblock. The variety of sequence lengths imposes additional computing overhead where inputs need to be zero-padded to the maximum sentence length in the batch to accommodate the parallel computing platforms. This paper targets the field-programmable gate array (FPGA) and proposes a coherent sequence length adaptive algorithm–hardware co-design for Transformer acceleration. Particularly, we develop a hardware-friendly sparse attention operator and a length-aware hardware resource scheduling algorithm. The proposed sparse attention operator brings the complexity of attention-based models down to linear complexity and alleviates the off-chip memory traffic. The proposed length-aware resource hardware scheduling algorithm dynamically allocates the hardware resources to fill up the pipeline slots and eliminates bubbles for NLP tasks. Experiments show that our design has very small accuracy loss and has 80.2 × and 2.6 × speedup compared to CPU and GPU implementation, and 4 × higher energy efficiency than state-of-the-art GPU accelerator optimized via CUBLAS GEMM.

Peng, Hongwu↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Cabana: A Performance Portable Library for Particle-Based Simulations

Particle-based simulations are ubiquitous throughout many fields of computational science and engineering, spanning the atomistic level with molecular dynamics (MD), to mesoscale particle-in-cell (PIC) simulations for solid mechanics, device-scale modeling with PIC methods for plasma physics, and massive N-body cosmology simulations of galaxy structures, with many other methods in between (Hockney & Eastwood, 1989). While these methods use particles to represent significantly different entities with completely different physical models, many low-level details are shared including performant algorithms for short- and/or long-range particle interactions, multi-node particle communication patterns, and other data management tasks such as particle sorting and neighbor list construction. Cabana is a performance portable library for particle-based simulations, developed as part of the Co-Design Center for Particle Applications (CoPA) within the Exascale Computing Project (ECP) (Alexander et al., 2020). The CoPA project and its full development scope, including ECP partner applications, algorithm development, and similar software libraries for quantum MD, is described in (Mniszewski et al., 2021). Cabana uses the Kokkos library for on-node parallelism (Edwards et al., 2014; Trott et al., 2022), enabling simulation on multi-core CPU and GPU architectures, and MPI for GPU-aware, multi-node communication. Cabana provides particle simulation capabilities on almost all current Kokkos backends, including serial execution, OpenMP (including OpenMP-Target for GPUs), CUDA (NVIDIA GPUs), HIP (AMD GPUs), and SYCL (Intel GPUs), providing a clear path for the coming generation of accelerator-based exascale hardware. Cabana builds on Kokkos by providing new particle data structures and particle algorithms resulting in a similar execution policy-based, node-level programming model that is intended to be used in addition to the core Kokkos library within an application. Cabana is designed as an application and physics agnostic, but particle-specific toolkit which can either be used to generate a new application, or to be used as needed in existing applications at various levels of invasiveness including through interfaces that wrap user memory in existing data structures.

97 MATHEMATICS AND COMPUTING↗

Advanced Module Architecture for Reduced Costs, High Durability and Significantly Improved Manufacturability (Final Report)

This project demonstrated a new module architecture (referred to herein as Glass/Glass) using silicone edge seals and an interlayer polymer that reduces manufacturing costs, significantly streamlines manufacturing processes, and reduces cap-ex costs while improving module reliability. A prototype manufacturing process to fabricate this new module architecture was demonstrated with a significantly improved process cycle time for the lamination step from the current industry standard of 13.5 minutes to approximately 30 seconds for each of the individual edge seal and internal polymer application steps. Samples are fabricated, strenuously stressed, and characterized to quantify the cost benefits, hardware capacities and accelerated stress performance of the new architecture. These results are being compared to a standard, laminated baseline architecture using a Glass/EVA/Glass or Glass/Thermoplastic/Glass design that traditionally has high moisture vapor transmission rates (MVTR).

14 SOLAR ENERGY↗

Microgravity Environment Characterization Program

The Microgravity Science and Applications Division (MSAD), a division within NASA's Office of Life and Microgravity Science and Applications, sponsors a broad range of space-based research in biotechnology, combustion science, fluid physics, fundamental physics, and materials science. To better understand and exploit the orbital environment, MSAD has developed methods and hardware to characterize accelerations on microgravity experiment carriers. MSAD supports research to verify analytically derived acceleration requirements for experiments and provides vibration isolation for sensitive experiments. The Microgravity Measurement and Analysis Project (MMAP), supported by MSAD, incorporates four projects: the Space Acceleration Measurement System (SAMS), the Orbital Acceleration Research Experiment (OARE), the SAMS for International Space Station (SAMS-II), and the Principal Investigator Microgravity Services (PIMS). SAMS was developed to record microgravity accelerations and the OARE was developed to record very low-frequency microgravity accelerations on-board the NASA Orbiters. The SAMS is also used for cooperative investigations on the Russian Mir space station. The SAMS-II is being developed for the same function on-board the International Space Station (ISS). PIMS utilizes microgravity acceleration data to develop a description of each microgravity mission's acceleration environment and to support microgravity investigators in interpreting possible effects of the acceleration environment on their experiments. These elements of the MSAD program will be used to define acceleration requirements for future Orbiter and ISS payloads. This paper describes the MMAP and summarizes the products and services available to principal investigators and other users. This paper also presents some microgravity acceleration characterization results from the last six years of Orbiter microgravity missions.

DeLombard, Richard↗

Parametric Mechanism Design Through Numerical Optimization and Physics Simulation

Design-Build-Test approaches for developing spaceflight hardware are prohibitively time and cost intensive and often lead to suboptimal mechanism designs. Approaches that couple machine learning and high-fidelity physics simulation could eliminate the need for hardware prototyping and dramatically accelerate the engineering design cycle, ultimately reducing cost. This work presents a modular NASA-developed toolchain to optimize hardware mechanisms in a virtual environment using numerical optimization and multi-body physics simulation. The toolchain enables multi-objective optimization, generates parametric CAD files that can be further post-processed by an end user, and can be expanded to optimize full systems and non-mechanical parameters such as feedback control variables. We demonstrate the toolchain through an independently verifiable design problem that optimizes wheel radius to achieve a desired linear velocity in a rigid-body physics environment when the wheel rotates at a constant angular speed, and then post-process the parametric CAD file of the optimal design generated by the tool before ultimately manufacturing it via 3D printing. We end with a discussion of how the toolchain can incorporate other analysis tools, including finite element analysis, computational fluid dynamics, and granular media simulations.

Optimization↗

An Optimization-Based Toolchain for Parametric Mechanism Design

Design-Build-Test approaches for developing spaceflight hardware are prohibitively time and cost intensive and often lead to suboptimal mechanism designs. Approaches that couple machine learning and high-fidelity physics simulation could eliminate the need for hardware prototyping and dramatically accelerate the engineering design cycle, ultimately reducing cost. This work presents a modular NASA-developed toolchain to optimize hardware mechanisms in a virtual environment using numerical optimization and multi-body physics simulation. The toolchain enables multi-objective optimization, generates parametric CAD files that can be further post-processed by an end user, and can be expanded to optimize full systems and non-mechanical parameters such as feedback control variables. We demonstrate the toolchain through an independently verifiable design problem that optimizes wheel radius to achieve a desired linear velocity in a rigid-body physics environment when the wheel rotates at a constant angular speed, and then post-process the parametric CAD file of the optimal design generated by the tool before ultimately manufacturing it via 3D printing. We end with a discussion of how the toolchain can incorporate other analysis tools, including finite element analysis, computational fluid dynamics, and granular media simulations.

Optimization↗

Optimization-Based Parametric Design via High-Fidelity Simulation: Overview + Examples

Design-Build-Test approaches for developing spaceflight hardware are prohibitively time and cost intensive and often lead to suboptimal mechanism designs. Approaches that couple machine learning and high-fidelity physics simulation could eliminate the need for hardware prototyping and dramatically accelerate the engineering design cycle, ultimately reducing cost. This talk presents a modular NASA-developed toolchain to optimize hardware mechanisms in a virtual environment using numerical optimization and multi-body physics simulation and includes example applications related to rigid wheel design for autonomous rovers and computational fluid dynamics.

optimization↗

Heavy ion beam physics at Facility for Rare Isotope Beams

The Facility for Rare Isotope Beams (FRIB) will be the world's premier rare-isotope beam facility. Experiments with the majority (~80%) of the isotope predicted to exist will become available. The FRIB facility is based on a superconducting (SC) heavy ion linac with output energy above 200 MeV/u for any ions at beam power of 400 kW. FRIB includes a target facility for in-flight production of rare isotopes. A three-stage fragment separator will be used to prepare fast rare isotope beams with high-purity for nuclear physics experiments. The installation work of the accelerator and experimental systems is approaching completion and multi-stage beam commissioning activities started in summer 2017 with expected project completion in early 2022. In conclusion, the commencement of operation for users' experiments is planned immediately following the project completion.

43 PARTICLE ACCELERATORS↗

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]↗

X-37 Storable Propulsion System Design and Operations

In a response to NASA's X-37 TA-10 Cycle-1 contract, Boeing assessed nitrogen tetroxide (N2O4) and monomethyl hydrazine (MMH) Storable Propellant Propulsion Systems to select a low risk X-37 propulsion development approach. Space Shuttle lessons learned, planetary spacecraft, and Boeing Satellite HS-601 systems were reviewed to arrive at a low risk and reliable storable propulsion system. This paper describes the requirements, trade studies, design solutions, flight and ground operational issues which drove X-37 toward the selection of a storable propulsion system. The design of storable propulsion systems offers the leveraging of hardware experience that can accelerate progress toward critical design. It also involves the experience gained from launching systems using MMH and N2O4 propellants. Leveraging of previously flight-qualified hardware may offer economic benefits and may reduce risk in cost and schedule. This paper summarizes recommendations based on experience gained from Space Shuttle and similar propulsion systems utilizing MMH and N2O4 propellants. System design insights gained from flying storable propulsion are presented and addressed in the context of the design approach of the X-37 propulsion system.

Rodriguez, Henry↗

X-37 Storable Propulsion System Design and Operations

In a response to NASA's X-37 TA-10 Cycle-1 contract, Boeing assessed nitrogen tetroxide (N2O4) and monomethyl hydrazine (MMH) Storable Propellant Propulsion Systems to select a low risk X-37 propulsion development approach. Space Shuttle lessons learned, planetary spacecraft, and Boeing Satellite HS-601 systems were reviewed to arrive at a low risk and reliable storable propulsion system. This paper describes the requirements, trade studies, design solutions, flight and ground operational issues which drove X-37 toward the selection of a storable propulsion system. The design of storable propulsion systems offers the leveraging of hardware experience that can accelerate progress toward critical design. It also involves the experience gained from launching systems using MMH and N2O4 propellants. Leveraging of previously flight-qualified hardware may offer economic benefits and may reduce risk in cost and schedule. This paper summarizes recommendations based on experience gained from Space Shuttle and similar propulsion systems utilizing MMH and N2O4 propellants. System design insights gained from flying storable propulsion are presented and addressed in the context of the design approach of the X-37 propulsion system.

Rodriguez, Henry↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗