Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Space Experiments with Particle Accelerators (SEPAC)

Plans for SEPAC, an instrument array to be used on Spacelab 1 to study vehicle charging and neutralization, beam-plasma interaction in space, beam-atmospheric interaction exciting artificial aurora and airglow, and the electromagnetic-field configuration of the magnetosphere, are presented. The hardware, consisting of electron beam accelerator, magnetoplasma arcjet, neutral-gas plume generator, power supply, diagnostic package (photometer, plasma probes, particle analyzers, and plasma-wave package), TV monitor, and control and data-management unit, is described. The individual SEPAC experiments, the typical operational sequence, and the general outline of the SEPAC follow-on mission are discussed. Some of the experiments are to be joint ventures with AEPI (INS 003) and will be monitored by low-light-level TV.

Obayashi, T.↗

Evaluating the Performance of the NASA LaRC CMF Motion Base Safety Devices

This paper describes the initial measured performance results of the previously documented NASA Langley Research Center (LaRC) Cockpit Motion Facility (CMF) motion base hardware safety devices. These safety systems are required to prevent excessive accelerations that could injure personnel and damage simulator cockpits or the motion base structure. Excessive accelerations may be caused by erroneous commands or hardware failures driving an actuator to the end of its travel at high velocity, stepping a servo valve, or instantly reversing servo direction. Such commands may result from single order failures of electrical or hydraulic components within the control system itself, or from aggressive or improper cueing commands from the host simulation computer. The safety systems must mitigate these high acceleration events while minimizing the negative performance impacts. The system accomplishes this by controlling the rate of change of valve signals to limit excessive commanded accelerations. It also aids hydraulic cushion performance by limiting valve command authority as the actuator approaches its end of travel. The design takes advantage of inherent motion base hydraulic characteristics to implement all safety features using hardware only solutions.

Gupton, Lawrence E.↗

C++ Resource Intelligent Compilation for GPU Enabled Applications

We are nearing the limits of Moore's Law with current computing technology. As industries push for more performance from smaller systems, alternate methods of computation such as Graphics Processing Units (GPUs) should be considered. Many of these systems utilize the Compute Unified Device Architecture (CUDA) to give programmers access to individual compute elements of the GPU for general purpose computing tasks. Direct access to the GPU's parallel multi-core architecture enables highly efficient computation and can drastically reduce the time required for complex algorithms or data analysis. Of course not all systems have a CUDA-enabled device to leverage, and so applications must consider optional support for users with these devices. Resource Intelligent Compilation (RIC) addresses this situation by enabling GPU-based acceleration of existing applications without affecting users without GPUs. Resource Intelligent Compilation (RIC) creates C/C++ modules that can be compiled to create a standard CPU version or GPU accelerated version of a program, depending on hardware availability. This is accomplished through a toolbox of programming strategies based on features of the CUDA API. Using this toolbox, existing applications can be modified with ease to support GPU acceleration, and new applications can be generated with just a few simple modifications. All of this culminates in an accelerated application for users with the appropriate hardware, with no performance impact to standard systems. This memorandum presents all the important features involved in supporting and implementing RIC and an example of using RIC to accelerate an existing mathematical model, without removing support for standard users. Through this memorandum, NASA engineers can acquire a set of guidelines to follow for RIC-compliant development, seamlessly accelerating C/C++ applications.

GPU↗

Accelerating Radiation Computations for Dynamical Models With Targeted Machine Learning and Code Optimization

Abstract Atmospheric radiation is the main driver of weather and climate, yet due to a complicated absorption spectrum, the precise treatment of radiative transfer in numerical weather and climate models is computationally unfeasible. Radiation parameterizations need to maximize computational efficiency as well as accuracy, and for predicting the future climate many greenhouse gases need to be included. In this work, neural networks (NNs) were developed to replace the gas optics computations in a modern radiation scheme (RTE+RRTMGP) by using carefully constructed models and training data. The NNs, implemented in Fortran and utilizing BLAS for batched inference, are faster by a factor of 1–6, depending on the software and hardware platforms. We combined the accelerated gas optics with a refactored radiative transfer solver, resulting in clear‐sky longwave (shortwave) fluxes being 3.5 (1.8) faster to compute on an Intel platform. The accuracy, evaluated with benchmark line‐by‐line computations across a large range of atmospheric conditions, is very similar to the original scheme with errors in heating rates and top‐of‐atmosphere radiative forcings typically below 0.1 K day −1 and 0.5 W m −2 , respectively. These results show that targeted machine learning, code restructuring techniques, and the use of numerical libraries can yield material gains in efficiency while retaining accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Ground Motion Inputs for the Seismic Shake Table Test

Currently, spent nuclear fuel (SNF) is stored in on-site independent spent-fuel storage installations (ISFSIs) at seventythree (73) nuclear power plants (NPPs) in the US. Because a site for geologic repository for permanent disposal of SNF has not been constructed, the SNF will remain in dry storage significantly longer than planned. During this time, the ISFSIs, and potentially consolidated storage facilities, will experience earthquakes of different magnitudes. The dry storage systems are designed and licensed to withstand large seismic loads. When dry storage systems experience seismic loads, there are little data on the response of SNF assemblies contained within them. The Spent Fuel Waste Disposition (SFWD) program is planning to conduct a full-scale seismic shake table test to close the gap related to the seismic loads on the fuel assemblies in dry storage systems. This test will allow for quantifying the strains and accelerations on surrogate fuel assembly hardware and cladding during earthquakes of different magnitudes and frequency content. The main component of the test unit will be the full-scale NUHOMS 32 PTH2 dry storage canister. The canister will be loaded with three surrogate fuel assemblies and twenty-nine dummy assemblies. Two dry storage configurations will be tested – horizontal and vertical above-ground concrete overpacks. These configurations cover 91% of the current dry storage configurations. The major input into the shake table test are the seismic excitations or the earthquake ground motions – acceleration time histories in two horizontal and one vertical direction that will be applied to the shake table surface during the tests. The shake table surface represents the top of the concrete pad on which a dry storage system is placed. The goal of the ground motion task is to develop the ground motions that would be representative of the range of seismotectonic and other conditions that any site in the Western US (WUS) or Central Eastern US (CEUS) might entail. This task is challenging because of the large number of the ISFSI sites, variety of seismotectonic and site conditions, and effects that soil amplification, soil-structure interaction, and pad flexibility may have on the ground motions.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

A Modular Framework for Modeling Hardware Elements in Distributed Engine Control Systems

Progress toward the implementation of distributed engine control in an aerospace application may be accelerated through the development of a hardware-in-the-loop (HIL) system for testing new control architectures and hardware outside of a physical test cell environment. One component required in an HIL simulation system is a high-fidelity model of the control platform: sensors, actuators, and the control law. The control system developed for the Commercial Modular Aero-Propulsion System Simulation 40k (C-MAPSS40k) provides a verifiable baseline for development of a model for simulating a distributed control architecture. This distributed controller model will contain enhanced hardware models, capturing the dynamics of the transducer and the effects of data processing, and a model of the controller network. A multilevel framework is presented that establishes three sets of interfaces in the control platform: communication with the engine (through sensors and actuators), communication between hardware and controller (over a network), and the physical connections within individual pieces of hardware. This introduces modularity at each level of the model, encouraging collaboration in the development and testing of various control schemes or hardware designs. At the hardware level, this modularity is leveraged through the creation of a Simulink(R) library containing blocks for constructing smart transducer models complying with the IEEE 1451 specification. These hardware models were incorporated in a distributed version of the baseline C-MAPSS40k controller and simulations were run to compare the performance of the two models. The overall tracking ability differed only due to quantization effects in the feedback measurements in the distributed controller. Additionally, it was also found that the added complexity of the smart transducer models did not prevent real-time operation of the distributed controller model, a requirement of an HIL system.

numerical simulation↗

A Modular Framework for Modeling Hardware Elements in Distributed Engine Control Systems

Progress toward the implementation of distributed engine control in an aerospace application may be accelerated through the development of a hardware-in-the-loop (HIL) system for testing new control architectures and hardware outside of a physical test cell environment. One component required in an HIL simulation system is a high-fidelity model of the control platform: sensors, actuators, and the control law. The control system developed for the Commercial Modular Aero-Propulsion System Simulation 40k (40,000 pound force thrust) (C-MAPSS40k) provides a verifiable baseline for development of a model for simulating a distributed control architecture. This distributed controller model will contain enhanced hardware models, capturing the dynamics of the transducer and the effects of data processing, and a model of the controller network. A multilevel framework is presented that establishes three sets of interfaces in the control platform: communication with the engine (through sensors and actuators), communication between hardware and controller (over a network), and the physical connections within individual pieces of hardware. This introduces modularity at each level of the model, encouraging collaboration in the development and testing of various control schemes or hardware designs. At the hardware level, this modularity is leveraged through the creation of a Simulink (R) library containing blocks for constructing smart transducer models complying with the IEEE 1451 specification. These hardware models were incorporated in a distributed version of the baseline C-MAPSS40k controller and simulations were run to compare the performance of the two models. The overall tracking ability differed only due to quantization effects in the feedback measurements in the distributed controller. Additionally, it was also found that the added complexity of the smart transducer models did not prevent real-time operation of the distributed controller model, a requirement of an HIL system.

numerical simulation↗

A Modular Framework for Modeling Hardware Elements in Distributed Engine Control Systems

Progress toward the implementation of distributed engine control in an aerospace application may be accelerated through the development of a hardware-in-the-loop (HIL) system for testing new control architectures and hardware outside of a physical test cell environment. One component required in an HIL simulation system is a high-fidelity model of the control platform: sensors, actuators, and the control law. The control system developed for the Commercial Modular Aero-Propulsion System Simulation 40k (C-MAPSS40k) provides a verifiable baseline for development of a model for simulating a distributed control architecture. This distributed controller model will contain enhanced hardware models, capturing the dynamics of the transducer and the effects of data processing, and a model of the controller network. A multilevel framework is presented that establishes three sets of interfaces in the control platform: communication with the engine (through sensors and actuators), communication between hardware and controller (over a network), and the physical connections within individual pieces of hardware. This introduces modularity at each level of the model, encouraging collaboration in the development and testing of various control schemes or hardware designs. At the hardware level, this modularity is leveraged through the creation of a SimulinkR library containing blocks for constructing smart transducer models complying with the IEEE 1451 specification. These hardware models were incorporated in a distributed version of the baseline C-MAPSS40k controller and simulations were run to compare the performance of the two models. The overall tracking ability differed only due to quantization effects in the feedback measurements in the distributed controller. Additionally, it was also found that the added complexity of the smart transducer models did not prevent real-time operation of the distributed controller model, a requirement of an HIL system.

propulsion simulation↗

Position Papers for the ASCR Workshop on Reimagining Codesign

On behalf of the Advanced Scientific Computing Research (ASCR) program in the US Department of Energy (DOE) Office of Science, we are organizing a Workshop on Reimagining Codesign (ReCoDe). Codesign is the process of jointly designing interoperating components of a computing system—in particular: applications, algorithms, system software, programming models, and the hardware on which they run. The goal is to maximize the overall performance, efficiency, and other desirable qualities of the system as a whole. Codesign is a standard methodology in the embedded-systems community, where space, power, and cost constraints are commonly pitted against execution speed for a tightly constrained feature set. Over the last decade, the DOE has invested in codesign efforts to foster the development of exascale computing systems for broad classes of scientific and engineering applications. The ReCoDe workshop hopes to explore how scientific applications of interest to the DOE can be accelerated through close interactions with hardware designers and software-stack developers, in which all components adapt to each other’s requirements and constraints. We want to answer the question of what are the key tools and methodologies for accomplishing codesign in today’s computing landscape, and what will be the highest impact targets for meeting DOE’s emerging mission requirements. This workshop aims to bring together DOE, industry, and academia to identify opportunities to build on past codesign successes and identify new areas that are either emerging or that may need reimagining for the future. We want to continue to find opportunities that can be pursued as a joint effort and continue to break down the traditional customer/vendor dichotomy with true partnerships. From this work, DOE will benefit from increased application performance relative to what stock hardware or existing general-purpose roadmaps can provide, and vendors will benefit from expanding their hardware’s capabilities to address needs they might have not otherwise anticipated and thereby create more widespread interest in their products. The workshop will be structured around a set of breakout sessions, with every attendee expected to participate actively in the discussions. Afterward, workshop attendees—from DOE, industry, and academia—will produce a report for ASCR that summarizes the findings made during the workshop.

97 MATHEMATICS AND COMPUTING↗

Machine learning for reducing noise in RF control signals at industrial accelerators

Industrial particle accelerators typically operate in dirtier environments than research accelerators, leading to increased noise in RF and electronic systems. Furthermore, given that industrial accelerators are mass produced, less attention is given to optimizing the performance of individual systems. As a result, industrial accelerators tend to underperform their own hardware capabilities. Improving signal processing for these machines will improve cost and time margins for deployment, helping to meet the growing demand for accelerators for medical sterilization, food irradiation, cancer treatment, and imaging. Our work focuses on using machine learning techniques to reduce noise in RF signals used for pulse-to-pulse feedback in industrial accelerators. Here we review our algorithms and observed results for simulated RF systems, and discuss next steps with the ultimate goal of deployment on industrial systems.

43 PARTICLE ACCELERATORS↗

ORCHA: A performance portability system for extreme heterogeneity

Heterogeneity is the prevalent trend in the rapidly evolving high-performance computing (HPC) landscape in both hardware and application software. The diversity in hardware platforms, currently comprising various accelerators and a future possibility of specializable chiplets, poses a significant challenge for scientific software developers aiming to harness optimal performance across different computing platforms while maintaining the quality of solutions when their applications are simultaneously growing more complex. Code synthesis and code generation can provide mechanisms to mitigate this challenge. We have developed a divide and conquer approach where different aspects of performance are handled by different stand-alone tools that are interfaced with the application through generated code. This portability system, ORCHA, enables users to configure and orchestrate their computations among available resources on a platform by specifying a high-level recipe, thereby permitting a many-to-many paradigm where each recipe results in a different variant of the application. The core design goal is to let users decide the application’s hardware mapping and orchestration by editing only the high-level recipe—without modifying the maintained source code or binding the application to a particular runtime system. Tools in ORCHA distribution are: CG-Kit for translating the recipe into an execution graph; Milhoja to execute the graph by orchestrating data and task movement among hardware resources; and Macroprocessor that enables users to define their own code-shorthand for higher composability and easier management of code variants. Additionally, the design of ORCHA permits tools to work in a plug-and-play mode where the application can build and run without CG-Kit and Milhoja, and either tool can be swapped out for other tools with similar capabilities by modifying the code generation portion of ORCHA. In this paper, we describe the design of ORCHA and the role that code-generation plays in isolating applications from tools. We demonstrate the breadth of configurations ORCHA enables with a case study in which an application configuration is realized on three distinct hardware mappings—a GPU-centric, a CPU/GPU balanced, and a CPU/GPU concurrent layouts by using different recipes.

Lee, Youngjun↗

Towards Precision-Aware Fault Tolerance Approaches for Mixed-Precision Applications

Graphics Processing Units (GPUs), the dominantly adopted accelerators in HPC systems, are susceptible to transient hardware fault. New generation of GPUs feature mixed-precision architectures such as NVIDIA Tensor Cores to accelerate matrix multiplications. While being widely adapted, how would they behave under transient hardware faults remain unclear. In this study, we conduct a large-scale fault injection experiments on GEMM kernels implemented with different floating-point data types on the V100 and A100 Tensor Cores, and show distinct error resilience characteristics for the GEMMS with different formats. In the future, we plan to explore this space by building precision-aware floating-point fault tolerance techniques for applications such as DNNs that exercise low-precision computations.

Fang, Bo↗

Using Likwid and Byfl to Benchmark Hardware Performance

This paper outlines a benchmarking study conducted during my internship at LANL, focusing on CPU (Computer Processing Unit) and program performance assessment. The primary goal was to gather memory access data using three methods across five polybench kernels The data gathered would then be used to compare and contrast to one another and calculate operational intensity for performance comparisons. Benchmarking tools like Byfl and Likwid were employed, with Byfl offering hardware-independent data through LLVM compiler communication and Likwid directly interacting with computer hardware. The study considered various benchmarking factors, including optimization levels, Big O notation ((n)), CPU diversity and specific kernel equations. Big O notation was utilized to simplify code complexity, with detailed breakdwons of operations and memory components for each polybench application. Specific O(n) equations enabled nuanced kernel compariosns, facilitating the identification of performance variations. CPU efficiency assessments were conducted using Likwid tests on two CPUs. The central focus on code optimization aimed at achieving higher speeds and reduced memory usage through streamlined code. Future work propsoes creating a roofline model, synthesizing benchmarking data into a comprehensive data graph to assist in optimizing code and improving hardware performance. The potential impact on the laboratory or national mission was underscored, emphasizing the importance of optimizing applications and hardware to conserve resources and accelerate program execution. The specific relevance to LANL’s operations in math-intensive fields such as Nuclear Fission, Space Exploration, and Nanotechnology highlights the necessity of efficient benchmarking for resource conservation and proram speed. Overall, this study contributes to the understanding of CPU and program performance, providing insights for future optimization efforts in a laboratory setting

97 MATHEMATICS AND COMPUTING↗

Final Seismic Shake Table Test Plan

The Spent Fuel Waste Disposition (SFWD) program is planning to conduct a full-scale seismic shake table test on the dry storage systems of spent nuclear fuel (SNF) to close the gap related to seismic loads on fuel assemblies in dry storage systems. This test will allow for quantifying the strains and accelerations on surrogate fuel assembly hardware and cladding during earthquakes of different magnitudes and frequency content. Full-scale testing is needed because a dry storage system is a complex and highly nonlinear system making it hard to predict (model) the responses to seismic excitations. The non-linearity arises from the multiple spatial gaps in the system – between fuel rods and the basket, between the basket and dry storage canister, between the dry storage canister and the storage cask (overpack), and ventilation gaps. The non-linearities pose significant limitations on the value of tests with scaled systems.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

Particle Impact Simulation and Ignition Prediction

An experimentally calibrated tool is needed to predict if a system is susceptible to failure by particle impact ignition (PI) based on use conditions, materials, and flow geometry. This tool will accelerate new components, evaluating existing hardware, and help disposition anomalies. Conduct particle impact testing with in-situ diagnostics and complementary simulations on subset of key engineering materials (IN718, M400, 316L, 6061, Ti64, Zr) to develop a proof-of-concept predictive tool for assessing the risk of PI for idealized geometries (spherical particles) in realistic environments. Assess particle/target interactions (coefficient of restitution, ignition, kindling) using instrumented particle impact rigs while systematically varying key parameters (materials, particle size, environment, target configuration). Determine key field variables (temperature, strain, stress) in particle impacts using Multiphysics finite element and hydrocode simulations validated through comparison with experimental measurements and observations. Synthesize experiments and simulations into constitutive models for PI that can be integrated with existing computational fluid dynamics (CFD) and Debris Transport Analysis (DTA) tools in future efforts

particle impact↗

Particle Impact Simulation and Ignition Prediction

An experimentally calibrated tool is needed to predict if a system is susceptible to failure by particle impact ignition (PI) based on use conditions, materials, and flow geometry. This tool will accelerate new components, evaluating existing hardware, and help disposition anomalies. - Conduct particle impact testing with in-situ diagnostics and complementary simulations on subset of key engineering materials (IN718, M400, 316L, 6061, Ti64, Zr) to develop a proof-of-concept predictive tool for assessing the risk of PI for idealized geometries (spherical particles) in realistic environments. - Assess particle/target interactions (coefficient of restitution, ignition, kindling) using instrumented particle impact rigs while systematically varying key parameters (materials, particle size, environment, target configuration). - Determine key field variables (temperature, strain, stress) in particle impacts using Multiphysics finite element and hydrocode simulations validated through comparison with experimental measurements and observations. - Synthesize experiments and simulations into constitutive models for PI that can be integrated with existing computational fluid dynamics (CFD) and Debris Transport Analysis (DTA) tools in future efforts.

Jonathan Tylka↗

Kernel fusion in atomistic spin dynamics simulations on Nvidia GPUs using tensor core

In atomistic spin dynamics simulations, the time cost of constructing the space- and time-displaced pair correlation function in real space increases quadratically as the number of spins N, leading to significant computational effort. The GEMM subroutine can be adopted to accelerate the calculation of the dynamical spin-spin correlation function, but the computational cost of simulating large spin systems (>40000 spins) on CPUs remains expensive. In this work, we perform the simulation on the graphics processing unit (GPU), a hardware solution widely used as an accelerator for scientific computing and deep learning. Here we show that GPUs can accelerate the simulation up to 25-fold compared to multi-core CPUs when using the GEMM subroutine on both. To hide memory latency, we fuse the element-wise operation into the GEMM kernel using CUTLASS that can improve the performance by 26% ~ 33% compared to implementation based on cuBLAS. Furthermore, we perform the on-the-fly calculation in the epilogue of the GEMM subroutine to avoid saving intermediate results on global memory, which makes the large-scale atomistic spin dynamics simulation feasible and affordable.

97 MATHEMATICS AND COMPUTING↗