Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Heavy ion beam physics at Facility for Rare Isotope Beams

The Facility for Rare Isotope Beams (FRIB) will be the world's premier rare-isotope beam facility. Experiments with the majority (~80%) of the isotope predicted to exist will become available. The FRIB facility is based on a superconducting (SC) heavy ion linac with output energy above 200 MeV/u for any ions at beam power of 400 kW. FRIB includes a target facility for in-flight production of rare isotopes. A three-stage fragment separator will be used to prepare fast rare isotope beams with high-purity for nuclear physics experiments. The installation work of the accelerator and experimental systems is approaching completion and multi-stage beam commissioning activities started in summer 2017 with expected project completion in early 2022. In conclusion, the commencement of operation for users' experiments is planned immediately following the project completion.

43 PARTICLE ACCELERATORS↗

AI-Powered Knowledge Graphs for Neuromorphic and Energy-Efficient Computing

The surge in scientific literature obscures breakthroughs and hinders the discovery of new research paths. We propose an artificial intelligence (AI) powered framework using large language models (LLMs) and knowledge graphs (KGs) to automate parts of scientific discovery, focusing on energy-efficient AI circuits. Our hybrid approach combines LLMs, structured data, and ontology-based reasoning to construct a comprehensive knowledge graph that integrates insights across computational neuroscience, spiking neuron models, learning rules, architectural motifs, and neuromorphic device technologies. This multi-domain representation enables the generation of hypotheses that connect biological function with implementable, energy-efficient hardware architectures. Using KG embeddings and graph neural networks, the framework generates hypotheses for novel circuits, validates them through optimization on exascale HPC systems, and with tools like SuperNeuro and Fugu, the most promising designs will be prototyped in hardware. This open-source system aims to accelerate discoveries and bridging neuroscience with hardware innovation, drive collaboration, and unlock new opportunities in low-power AI computing.

Gautam, Ashish [ORNL]↗

Shifting Between Compute and Memory Bounds: A Compression-Enabled Roofline Model

In the evolving landscape of high-performance computing, especially to fight the end of Moore’s Law and Dennard’s Scaling, the ability to shift between compute-bound and memory-bound states is critical for enhancing adaptability and flexibility to diverse system and domain-specific architectures. Such capability is vital for optimizing performance across distinguished hardware configurations, such as accelerators, memory hierarchies, and cache systems. Despite that ad hoc optimization techniques, such as compressed/approximate computation, have been enabled for compute-/data-intensive computing for improved performance in distinct hardware settings, there lacks an understanding of 1) the rational behind performance improvement; 2) capability of different optimizations; 3) what optimization to respond to specific computational and memory demands. This work proposes a compression-enabled roofline model to facilitate this adaptability with data compression techniques to balance and transform between computational and memory demands. This model enables applications to adjust in response to the specific strengths and limitations of the underlying hardware and system to optimize resource utilization. The effectiveness of this approach is demonstrated with matrix multiplication kernels on different input sizes, with turning on/off various compression techniques, including 1) low-precision floating point; 2) sparse matrix formulation; and 3) compressed arrays with ZFP. By reducing memory transfer volumes and cache misses and increasing data locality and computational intensity through compression, the specific roofline model can transform between compute and memory bounds to align more efficiently with system capabilities. This advancement not only improves overall performance but also maximizes adaptability in diverse computing environments.

Naraparaju, Ramasoumya [University of Washington]↗

Tournament-Based Pretraining to Accelerate Federated Learning

Advances in hardware, proliferation of compute at the edge, and data creation at unprecedented scales have made federated learning (FL) necessary for the next leap forward in pervasive machine learning. For privacy and network reasons, large volumes of data remain stranded on endpoints located in geographically austere (or at least austere network-wise) locations. However, challenges exist to the effective use of these data. To solve the system and functional level challenges, we present an three novel variants of a serverless federated learning framework. We also present tournament-based pretraining, which we demonstrate significantly improves model performance in some experiments. Overall, these extensions to FL and our novel training method enable greater focus on science rather than ML development.

Baughman, Matt↗

Enabling real-time adaptation of machine learning models at x-ray Free Electron Laser facilities with high-speed training optimized computational hardware

The emergence of novel computational hardware is enabling a new paradigm for rapid machine learning model training. For the Department of Energy’s major research facilities, this developing technology will enable a highly adaptive approach to experimental sciences. In this manuscript we present the per-epoch and end-to-end training times for an example of a streaming diagnostic that is planned for the upcoming high-repetition rate x-ray Free Electron Laser, the Linac Coherent Light Source-II. We explore the parameter space of batch size and data parallel training across multiple Graphics Processing Units and Reconfigurable Dataflow Units. We show the landscape of training times with a goal of full model retraining in under 15 min. Although a full from scratch retraining of a model may not be required in all cases, we nevertheless present an example of the application of emerging computational hardware for adapting machine learning models to changing environments in real-time, during streaming data acquisition, at the rates expected for the data fire hoses of accelerator-based user facilities.

97 MATHEMATICS AND COMPUTING↗

From Atoms to Wheels: The Role of Multi-Scale Modeling in the Future of Transportation Electrification

Traditionally, prototype hardware is built for validation testing to ensure battery systems design changes meet vehicle-level requirements, which is expensive both in cost and time. Virtual engineering (VE) of battery systems for electric vehicle (EV) propulsion offers a reduced-cost alternative to the traditional development process and uses multi-scale modeling to virtually probe the impact of design changes in a particular part on the overall performance of the system. This allows for rapid iteration over multiple design spaces, without committing to build hardware. This perspective article discusses current trends in VE for EV applications and proposes improvements to accelerate EV adoption.

Garrick, Taylor R. (ORCID:0000000322518129)↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

Embedded EPICS server for PowerPMAC motion controllers

An embedded server layer of Experimental Physics and Industrial Control System (EPICS) for PowerPMAC motion controllers has been developed and deployed at two undulator beamlines of the National Institute of General Medical Sciences and the National Cancer Institute (GM/CA) Structural Biology Facility at the Advanced Photon Source (APS). This compact, open source solution makes the power and versatility of PowerPMAC motion controls directly accessible to distributed EPICS clients. At GM/CA the system controls about 200 servo and stepper motors — both encoded and unencoded — and multiple digital and analog I/O accessories. The server stack comprises two sublayers: a lower-level driver and database that communicates directly with PowerPMAC, and a facility-specific soft sublayer built on top. The paper describes installing EPICS on PowerPMAC, the implementation of both layers and client examples, including on-the-fly scanning.

EPICS↗

Caspian: A Neuromorphic Development Platform

Current neuromorphic systems often may be difficult to use and costly to deploy. There exists a need for a simple yet flexible neuromorphic development platform which can allow researchers to quickly prototype ideas and applications. Caspian offers a high-level API along with a fast spiking simulator to enable the rapid development of neuromorphic solutions. It further offers an FPGA architecture that allows for simplified deployment -- particularly in SWaP (size, weight, and power) constrained environments. Leveraging both software and hardware, Caspian aims to accelerate development and deployment while enabling new researchers to quickly become productive with a spiking neural network system.

Mitchell, Parker↗

Shaker Table Test Plan

Currently, spent nuclear fuel (SNF) is stored in onsite independent spent fuel storage facilities (ISFSIs), which is a dry storage facility, at 55 nuclear power plant sites. The majority of SNF in dry storage is in welded metal canisters (2,917 canisters at the end of 2019). The canisters are loaded for storage in storage overpacks (vertical casks or horizontal storage modules) and placed on outdoor concrete pads. Because the SNF will be stored at ISFSIs for an extended period of time, there is growing concern with regards to the behavior of the SNF within these dry storage systems during earthquakes. To address these concerns, the SFWST program is considering conducting an earthquake shaker table test. The goal of this test is to determine the strains and accelerations on fuel assembly hardware and cladding during earthquakes of different magnitudes to better quantify the potential damage an earthquake could inflict on spent nuclear fuel rods. The seismic integrity of the storage system has been addressed in the past by the US Nuclear Regulatory Commission and is not the focus of this potential test. Instead the DOE would benefit from knowing the condition of the fuel cladding from storage, transportation, to disposal so that it can ascertain repository performance for the fuel and packaging in its final state. A seismic event is part of the possible loading events that the fuel could experience in its lifetime. This report proposes several earthquake shaker table tests with different degrees of complexity. Alternative 1 was defined in the FY20 work scope. Alternatives 2 and 3 were recently developed to take advantage of the NUHOMS 32PTH dry storage canister that may be available in FY21 for this test at a minimum cost to the project. The selection of the alternative(s) will depend on the available budget and the SFWST program priorities for the near future.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Develop a Fast Analysis Solver for Welding Sequence Optimization

During the shipbuilding manufacturing process, materials are exposed to significant stresses, as induced both thermally and mechanically, that alter the intended design and significantly affect the production schedule, labor hours (fitting, welding, rework, etc.), and material structural performance. The type and magnitude of deformation of a given structure depends on many factors such as the material, thickness and quality of components, the process heat input, preheat and inter-pass temperatures, type and size of welds, welding sequence and direction, location, sequence, and degree of fixturing. Numerical simulations using finite element analysis (FEA) have long been used to analyze welding-induced structural distortion. For large assemblies, transient thermal elastic-plastic analysis (TEPA) can take days or weeks to run, and optimization of welding sequence is not feasible. Simplified analysis methods were developed to reduce computational time. However, it is challenging to use these techniques to fully optimize welding sequencing because of their applied simplifications in modeling weld details. A fast analysis solver that could be used by the shipbuilding industry is being developed for optimizing welding sequences by taking full advantage of modern GPU-based HPC hardware and incorporating patented acceleration schemes. The accelerated processing factors are up to 2200 times greater for large, multi-pass welded structures.

Yang, Yu-Ping↗

OpenCGRA: An Open-Source Unified Framework for Modeling,Testing, and Evaluating CGRAs

Coarse-grained reconfigurable arrays (CGRAs),loosely defined as arrays of functional units (e.g, adder, sub-tractor, multiplier, divider, or larger multi-operation units, butsmaller than a general-purpose core) interconnected through aNetwork-on-Chip, provide higher flexibility than domain-specificASIC accelerators while offering increased hardware efficiencywith respect to fine-grained reconfigurable devices, such as FieldProgrammable Gate Arrays (FPGAs). The fast evolving fieldsof machine learning and edge computing, which are seeing acontinuous flow of novel algorithms and larger models, makeCGRAs ideal target architectures to allow domain specializationwithout loosing too much generality. They also generally offerquicker and more effective reconfigurability than FPGAs, po-tentially allowing adaptation during actual algorithm execution,and implement a dataflow programming paradigm that adaptswell to these emerging workloads. Designing and generating aCGRA, however, still requires to define the type and number ofthe specific functional units, implement their interconnect andthe network topology, and perform its simulation and validation,given a variety of workloads of interest.In this paper, we propose OpenCGRA, a Python-based unifiedframework that integrates generation, modeling, testing and eval-uation for CGRAs. OpenCGRA is the first open-source integratedframework able to support the full top-to-bottom design flow forspecializing and implementing CGRAs: modeling at different ab-straction levels (functional level, cycle level, register-transfer level),generation, simulation, testing at different granularities (unit test-ing, integration testing, property-based testing), and characteriza-tion (area, power, and timing). OpenCGRAs will be made availableon GitHub.

CGRA, synthesis↗

AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelerators

Coarse-grained reconfigurable arrays (CGRAs), loosely defined as arrays of functional units interconnected through a network-on-chip (NoC), provide higher flexibility than domain-specific ASIC accelerators while offering increased hardware efficiency with respect to fine-grained reconfigurable devices, such as Field Programmable Gate Arrays (FPGAs). Un-fortunately, designing a CGRA for a specific application domain involves enormous software/hardware engineering effort (e.g., designing the CGRA, map operations onto the CGRA, etc) and requires the exploration on a large design space (e.g., applying appropriate loop transformation on each application, specializing the reconfigurable processing elements of the CGRA, refining the network topology, deciding the size of the data memory, etc). Int his paper, we propose AURORA – a software/hardware co-design framework to automatically synthesize optimal CGRA given a set of applications of interest

Tan, Cheng↗

Accelerating Radiation Computations for Dynamical Models With Targeted Machine Learning and Code Optimization

Abstract Atmospheric radiation is the main driver of weather and climate, yet due to a complicated absorption spectrum, the precise treatment of radiative transfer in numerical weather and climate models is computationally unfeasible. Radiation parameterizations need to maximize computational efficiency as well as accuracy, and for predicting the future climate many greenhouse gases need to be included. In this work, neural networks (NNs) were developed to replace the gas optics computations in a modern radiation scheme (RTE+RRTMGP) by using carefully constructed models and training data. The NNs, implemented in Fortran and utilizing BLAS for batched inference, are faster by a factor of 1–6, depending on the software and hardware platforms. We combined the accelerated gas optics with a refactored radiative transfer solver, resulting in clear‐sky longwave (shortwave) fluxes being 3.5 (1.8) faster to compute on an Intel platform. The accuracy, evaluated with benchmark line‐by‐line computations across a large range of atmospheric conditions, is very similar to the original scheme with errors in heating rates and top‐of‐atmosphere radiative forcings typically below 0.1 K day −1 and 0.5 W m −2 , respectively. These results show that targeted machine learning, code restructuring techniques, and the use of numerical libraries can yield material gains in efficiency while retaining accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗