Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Tournament-Based Pretraining to Accelerate Federated Learning

Advances in hardware, proliferation of compute at the edge, and data creation at unprecedented scales have made federated learning (FL) necessary for the next leap forward in pervasive machine learning. For privacy and network reasons, large volumes of data remain stranded on endpoints located in geographically austere (or at least austere network-wise) locations. However, challenges exist to the effective use of these data. To solve the system and functional level challenges, we present an three novel variants of a serverless federated learning framework. We also present tournament-based pretraining, which we demonstrate significantly improves model performance in some experiments. Overall, these extensions to FL and our novel training method enable greater focus on science rather than ML development.

Baughman, Matt↗

Enabling real-time adaptation of machine learning models at x-ray Free Electron Laser facilities with high-speed training optimized computational hardware

The emergence of novel computational hardware is enabling a new paradigm for rapid machine learning model training. For the Department of Energy’s major research facilities, this developing technology will enable a highly adaptive approach to experimental sciences. In this manuscript we present the per-epoch and end-to-end training times for an example of a streaming diagnostic that is planned for the upcoming high-repetition rate x-ray Free Electron Laser, the Linac Coherent Light Source-II. We explore the parameter space of batch size and data parallel training across multiple Graphics Processing Units and Reconfigurable Dataflow Units. We show the landscape of training times with a goal of full model retraining in under 15 min. Although a full from scratch retraining of a model may not be required in all cases, we nevertheless present an example of the application of emerging computational hardware for adapting machine learning models to changing environments in real-time, during streaming data acquisition, at the rates expected for the data fire hoses of accelerator-based user facilities.

97 MATHEMATICS AND COMPUTING↗

From Atoms to Wheels: The Role of Multi-Scale Modeling in the Future of Transportation Electrification

Traditionally, prototype hardware is built for validation testing to ensure battery systems design changes meet vehicle-level requirements, which is expensive both in cost and time. Virtual engineering (VE) of battery systems for electric vehicle (EV) propulsion offers a reduced-cost alternative to the traditional development process and uses multi-scale modeling to virtually probe the impact of design changes in a particular part on the overall performance of the system. This allows for rapid iteration over multiple design spaces, without committing to build hardware. This perspective article discusses current trends in VE for EV applications and proposes improvements to accelerate EV adoption.

Garrick, Taylor R. (ORCID:0000000322518129)↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

Age Life Evaluation of Space Shuttle Crew Escape System Pyrotechnic Components Loaded with Hexanitrostilbene (HNS)

Determining deterioration characteristics of the Space Shuttle crew escape system pyrotechnic components loaded with hexanitrostilbene would enable us to establish a hardware life-limit for these items, so we could better plan our equipment use and, possibly, extend the useful life of the hardware. We subjected components to accelerated-age environments to determine degradation characteristics and established a hardware life-limit based upon observed and calculated trends. We extracted samples using manufacturing lots currently installed in the Space Shuttle crew escape system and from other NASA programs. Hardware included in the study consisted of various forms and ages of mild detonating fuse, linear shaped charge, and flexible confined detonating cord. The hardware types were segregated into 5 groups. One was subjected to detonation velocity testing for a baseline. Two were first subjected to prolonged 155 F heat exposure, and the other two were first subjected to 255 F, before undergoing detonation velocity testing and/or chromatography analysis. Test results showed no measurable changes in performance to allow a prediction of an end of life given the storage and elevated temperature environments the hardware experiences. Given the lack of a definitive performance trend, coupled with previous tests on post-flight Space Shuttle hardware showing no significant changes in chemical purity or detonation velocity, we recommend a safe increase in the useful life of the hardware to 20 years, from the current maximum limits of 10 and 15 years, depending on the hardware.

Hoffman, William C., III↗

Mixed-precision numerics in scientific applications: survey and perspectives

The explosive demand for artificial intelligence (AI) workloads has led to a significant increase in silicon area dedicated to lower-precision computations on recent high-performance computing hardware designs. However, mixed-precision capabilities, which can achieve performance improvements of up to 8x compared to double-precision in extreme compute-intensive workloads, remain largely untapped in most scientific applications. A growing number of efforts have shown that mixed-precision algorithmic innovations can deliver superior performance without sacrificing accuracy. These developments should prompt computational scientists to seriously consider whether their scientific modeling and simulation applications could benefit from the acceleration offered by new hardware and mixed-precision algorithms. In this survey, we (1) review progress across diverse scientific domains—fluid dynamics, weather and climate, quantum chemistry, and computational genomics—that have begun adopting mixed-precision strategies; (2) examine state-of-the-art algorithmic techniques such as iterative refinement, splitting and emulation schemes, and adaptive precision solvers; (3) assess their implications for accuracy, performance, and resource utilization; and (4) survey the emerging software ecosystem that enables mixed-precision methods at scale. We conclude with perspectives and recommendations on cross-cutting opportunities, domain-specific challenges, and the role of co-design between application scientists, numerical analysts, and computer scientists. Collectively, this survey underscores that mixed-precision numerics can reshape computational science by aligning algorithms with the evolving landscape of hardware capabilities.

Graphics processing units↗

Impacts of floating-point non-associativity on reproducibility for HPC and deep learning applications

Run to run variability in parallel programs caused by floating-point non-associativity has been known to significantly affect reproducibility in iterative algorithms, due to accumulating errors. Non-reproducibility can critically affect the efficiency and effectiveness of correctness testing for stochastic programs. Recently, the sensitivity of deep learning training and inference pipelines to floating-point non-associativity has been found to sometimes be extreme. It can prevent certification for commercial applications, accurate assessment of robustness and sensitivity, and bug detection. New approaches in scientific computing applications have coupled deep learning models with high-performance computing, leading to an aggravation of debugging and testing challenges. Here we perform an investigation of the statistical properties of floating-point non-associativity within modern parallel programming models, and analyze performance and productivity impacts of replacing atomic operations with deterministic alternatives on GPUs. We examine the recently-added deterministic options in PyTorch within the context of GPU deployment for deep learning, uncovering and quantifying the impacts of input parameters triggering run to run variability and reporting on the reliability and completeness of the documentation. Finally, we evaluate the strategy of exploiting automatic determinism that could be provided by deterministic hardware, using the Groq LPUTM accelerator for inference portions of the deep learning pipeline. We demonstrate the benefits that a hardware-based strategy can provide within reproducibility and correctness efforts.

Shanmugavelu, Sanjif↗

Embedded EPICS server for PowerPMAC motion controllers

An embedded server layer of Experimental Physics and Industrial Control System (EPICS) for PowerPMAC motion controllers has been developed and deployed at two undulator beamlines of the National Institute of General Medical Sciences and the National Cancer Institute (GM/CA) Structural Biology Facility at the Advanced Photon Source (APS). This compact, open source solution makes the power and versatility of PowerPMAC motion controls directly accessible to distributed EPICS clients. At GM/CA the system controls about 200 servo and stepper motors — both encoded and unencoded — and multiple digital and analog I/O accessories. The server stack comprises two sublayers: a lower-level driver and database that communicates directly with PowerPMAC, and a facility-specific soft sublayer built on top. The paper describes installing EPICS on PowerPMAC, the implementation of both layers and client examples, including on-the-fly scanning.

EPICS↗

Caspian: A Neuromorphic Development Platform

Current neuromorphic systems often may be difficult to use and costly to deploy. There exists a need for a simple yet flexible neuromorphic development platform which can allow researchers to quickly prototype ideas and applications. Caspian offers a high-level API along with a fast spiking simulator to enable the rapid development of neuromorphic solutions. It further offers an FPGA architecture that allows for simplified deployment -- particularly in SWaP (size, weight, and power) constrained environments. Leveraging both software and hardware, Caspian aims to accelerate development and deployment while enabling new researchers to quickly become productive with a spiking neural network system.

Mitchell, Parker↗

Shaker Table Test Plan

Currently, spent nuclear fuel (SNF) is stored in onsite independent spent fuel storage facilities (ISFSIs), which is a dry storage facility, at 55 nuclear power plant sites. The majority of SNF in dry storage is in welded metal canisters (2,917 canisters at the end of 2019). The canisters are loaded for storage in storage overpacks (vertical casks or horizontal storage modules) and placed on outdoor concrete pads. Because the SNF will be stored at ISFSIs for an extended period of time, there is growing concern with regards to the behavior of the SNF within these dry storage systems during earthquakes. To address these concerns, the SFWST program is considering conducting an earthquake shaker table test. The goal of this test is to determine the strains and accelerations on fuel assembly hardware and cladding during earthquakes of different magnitudes to better quantify the potential damage an earthquake could inflict on spent nuclear fuel rods. The seismic integrity of the storage system has been addressed in the past by the US Nuclear Regulatory Commission and is not the focus of this potential test. Instead the DOE would benefit from knowing the condition of the fuel cladding from storage, transportation, to disposal so that it can ascertain repository performance for the fuel and packaging in its final state. A seismic event is part of the possible loading events that the fuel could experience in its lifetime. This report proposes several earthquake shaker table tests with different degrees of complexity. Alternative 1 was defined in the FY20 work scope. Alternatives 2 and 3 were recently developed to take advantage of the NUHOMS 32PTH dry storage canister that may be available in FY21 for this test at a minimum cost to the project. The selection of the alternative(s) will depend on the available budget and the SFWST program priorities for the near future.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Develop a Fast Analysis Solver for Welding Sequence Optimization

During the shipbuilding manufacturing process, materials are exposed to significant stresses, as induced both thermally and mechanically, that alter the intended design and significantly affect the production schedule, labor hours (fitting, welding, rework, etc.), and material structural performance. The type and magnitude of deformation of a given structure depends on many factors such as the material, thickness and quality of components, the process heat input, preheat and inter-pass temperatures, type and size of welds, welding sequence and direction, location, sequence, and degree of fixturing. Numerical simulations using finite element analysis (FEA) have long been used to analyze welding-induced structural distortion. For large assemblies, transient thermal elastic-plastic analysis (TEPA) can take days or weeks to run, and optimization of welding sequence is not feasible. Simplified analysis methods were developed to reduce computational time. However, it is challenging to use these techniques to fully optimize welding sequencing because of their applied simplifications in modeling weld details. A fast analysis solver that could be used by the shipbuilding industry is being developed for optimizing welding sequences by taking full advantage of modern GPU-based HPC hardware and incorporating patented acceleration schemes. The accelerated processing factors are up to 2200 times greater for large, multi-pass welded structures.

Yang, Yu-Ping↗

Materials science on parabolic aircraft: The FY 1987-1989 KC-135 microgravity test program

This document covers research results from the KC-135 Materials Science Program managed by MSFC for the period FY87 through FY89. It follows the previous NASA Technical Memorandum for FY84-86 published in August 1988. This volume contains over 30 reports grouped into eight subject areas covering acceleration levels, space flight hardware, transport and interfacial studies, thermodynamics, containerless processing, welding, melt/crucible interactions, and directional solidification. The KC-135 materials science experiments during FY87-89 accomplished direct science, preparation for space flight experiments, and justification for new experiments in orbit.

Curreri, Peter A.↗

Development Status of Amine-based, Combined Humidity, CO2, and Trace Contaminant Control System for CEV

Under a NASA-sponsored technology development project, a multi-disciplinary team consisting of industry, academia, and government organizations lead by Hamilton Sundstrand is developing an amine-based humidity and CO2 removal process and prototype equipment for Vision for Space Exploration (VSE) applications. Originally this project sought to research enhanced amine formulations and incorporate a trace contaminant control capability into the sorbent. In October 2005, NASA re-directed the project team to accelerate the delivery of hardware by approximately one year and emphasize deployment on board the Crew Exploration Vehicle (CEV) as the near-term developmental goal. Preliminary performance requirements were defined based on nominal and off-nominal conditions and the design effort was initiated using the baseline amine sorbent, SA9T. As part of the original project effort, basic sorbent development was continued with the University of Connecticut and dynamic equilibrium trace contaminant adsorption characteristics were evaluated by NASA. This paper summarizes the University sorbent research effort, the basic trace contaminant loading characteristics of the SA9T sorbent, design support testing, and the status of the full-scale system hardware design and manufacturing effort.

Smith, Fred↗

Suited Occupant Injury Potential During Dynamic Spacecraft Flight Phases

In support of the Constellation Space Suit Element [CSSE], a new space-suit architecture will be created for support of Launch, Entry, Abort, Microgravity Extra- Vehicular Activity [EVA], and post-landing crew operations, safety and, under emergency conditions, survival. The space suit is unique in comparison to previous launch, entry, and abort [LEA] suit architectures in that it utilizes rigid mobility elements in the scye (i.e., shoulder) and the upper arm regions. The suit architecture also utilizes rigid thigh disconnect elements to create a quick disconnect approximately located above the knee. This feature allows commonality of the lower portion of the suit (from the thigh disconnect down), making the lower legs common across two suit configurations. This suit must interface with the Orion vehicle seat subsystem, which includes seat components, lateral supports, and restraints. Due to the unique configuration of spacesuit mobility elements, combined with the need to provide occupant protection during dynamic vehicle events, risks have been identified with potential injury due to the suit characteristics described above. To address the risk concerns, a test series has been developed in coordination with the Injury Biomechanics Research Laboratory [IBRL] to evaluate the likelihood and consequences of these potential issues. Testing includes use of Anthropomorphic Test Devices [ATDs; vernacularly referred to as "crash test dummies"], Post Mortem Human Subjects [PMHS], and representative seat/suit hardware in combination with high linear acceleration events. The ensuing treatment focuses on test purpose and objectives; test hardware, facility, and setup; and preliminary results.

Dub, Mark O.↗

OpenCGRA: An Open-Source Unified Framework for Modeling,Testing, and Evaluating CGRAs

Coarse-grained reconfigurable arrays (CGRAs),loosely defined as arrays of functional units (e.g, adder, sub-tractor, multiplier, divider, or larger multi-operation units, butsmaller than a general-purpose core) interconnected through aNetwork-on-Chip, provide higher flexibility than domain-specificASIC accelerators while offering increased hardware efficiencywith respect to fine-grained reconfigurable devices, such as FieldProgrammable Gate Arrays (FPGAs). The fast evolving fieldsof machine learning and edge computing, which are seeing acontinuous flow of novel algorithms and larger models, makeCGRAs ideal target architectures to allow domain specializationwithout loosing too much generality. They also generally offerquicker and more effective reconfigurability than FPGAs, po-tentially allowing adaptation during actual algorithm execution,and implement a dataflow programming paradigm that adaptswell to these emerging workloads. Designing and generating aCGRA, however, still requires to define the type and number ofthe specific functional units, implement their interconnect andthe network topology, and perform its simulation and validation,given a variety of workloads of interest.In this paper, we propose OpenCGRA, a Python-based unifiedframework that integrates generation, modeling, testing and eval-uation for CGRAs. OpenCGRA is the first open-source integratedframework able to support the full top-to-bottom design flow forspecializing and implementing CGRAs: modeling at different ab-straction levels (functional level, cycle level, register-transfer level),generation, simulation, testing at different granularities (unit test-ing, integration testing, property-based testing), and characteriza-tion (area, power, and timing). OpenCGRAs will be made availableon GitHub.

CGRA, synthesis↗

AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelerators

Coarse-grained reconfigurable arrays (CGRAs), loosely defined as arrays of functional units interconnected through a network-on-chip (NoC), provide higher flexibility than domain-specific ASIC accelerators while offering increased hardware efficiency with respect to fine-grained reconfigurable devices, such as Field Programmable Gate Arrays (FPGAs). Un-fortunately, designing a CGRA for a specific application domain involves enormous software/hardware engineering effort (e.g., designing the CGRA, map operations onto the CGRA, etc) and requires the exploration on a large design space (e.g., applying appropriate loop transformation on each application, specializing the reconfigurable processing elements of the CGRA, refining the network topology, deciding the size of the data memory, etc). Int his paper, we propose AURORA – a software/hardware co-design framework to automatically synthesize optimal CGRA given a set of applications of interest

Tan, Cheng↗