Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware accelerators”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Physics successfully implements Lagrange multiplier optimization

Optimization is a major part of human effort. While being mathematical, optimization is also built into physics. For example, physics has the Principle of Least Action; the Principle of Minimum Power Dissipation, also called Minimum Entropy Generation; and the Variational Principle. Physics also has Physical Annealing, which, of course, preceded computational Simulated Annealing. Physics has the Adiabatic Principle, which, in its quantum form, is called Quantum Annealing. Thus, physical machines can solve the mathematical problem of optimization, including constraints. Binary constraints can be built into the physical optimization. In that case, the machines are digital in the same sense that a flip–flop is digital. A wide variety of machines have had recent success at optimizing the Ising magnetic energy. We demonstrate in this paper that almost all those machines perform optimization according to the Principle of Minimum Power Dissipation as put forth by Onsager. Further, we show that this optimization is in fact equivalent to Lagrange multiplier optimization for constrained problems. We find that the physical gain coefficients that drive those systems actually play the role of the corresponding Lagrange multipliers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Reconfigurable Neural Network ASIC for Detector Front-End Data Compression at the HL-LHC

Despite advances in the programmable logic capabilities of modern trigger systems, a significant bottleneck remains in the amount of data to be transported from the detector to off-detector logic where trigger decisions are made. We demonstrate that a neural network (NN) autoencoder model can be implemented in a radiation-tolerant application-specific integrated circuit (ASIC) to perform lossy data compression alleviating the data transmission problem while preserving critical information of the detector energy profile. For our application, we consider the high-granularity calorimeter from the Compact Muon Solenoid (CMS) experiment at the CERN Large Hadron Collider. The advantage of the machine learning approach is in the flexibility and configurability of the algorithm. By changing the NN weights, a unique data compression algorithm can be deployed for each sensor in different detector regions and changing detector or collider conditions. To meet area, performance, and power constraints, we perform quantization-aware training to create an optimized NN hardware implementation. The design is achieved through the use of high-level synthesis tools and the hls4ml framework and was processed through synthesis and physical layout flows based on a low-power (LP)-CMOS 65-nm technology node. The flow anticipates 200 Mrad of ionizing radiation to select gates and reports a total area of 3.6 mm 2 and consumes 95 mW of power. The simulated energy consumption per inference is 2.4 nJ. Furthermore, this is the first radiation-tolerant on-detector ASIC implementation of an NN that has been designed for particle physics applications.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Module-OT: A Hardware Security Module for Operational Technology

Increased penetration levels of renewable energy and other types of distributed energy resources (DERs) on the modern electric grid-combined with technological advancements for electric system monitoring and control-introduce new cyberattack vectors and increase the cyberattack surface of energy systems. According to the IEEE Std. 1547-2018, DERs must use Modbus, Distributed Network Protocol 3 (DNP3), or Smart Energy Profile 2.0 (SEP2) as their communication protocol. Previous research identified several vulnerabilities and security breaches in each one of these communication protocols; despite this, existing standards for DERs do not recommend cybersecurity measures. In order to reduce vulnerabilities in power distribution systems, this paper presents a novel open-source hardware security module that improves both information and operational security to better protect data and communications on the distribution grid. The security hardware is called “module for operational technology,” or simply Module-OT, and it has been validated and tested in an emulated distribution system application. Module-OT is integrated within a communication system in the transport layer of the Open Systems Interconnection (OSI) model. It improves system security through encryption, authentication, authorization, certificate management, and user access control. The main advancement of Module-OT is the addition of hardware cryptographic acceleration that improves the overall communication performance in terms of end-to-end latency.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

ExaFEL: extreme-scale real-time data processing for X-ray free electron laser science

ExaFEL is an HPC-capable X-ray Free Electron Laser (XFEL) data analysis software suite for both Serial Femtosecond Crystallography (SFX) and Single Particle Imaging (SPI) developed in collaboration with the Linac Coherent Lightsource (LCLS), Lawrence Berkeley National Laboratory (LBNL) and Los Alamos National Laboratory. ExaFEL supports real-time data analysis via a cross-facility workflow spanning LCLS and HPC centers such as NERSC and OLCF. Our work therefore constitutes initial path-finding for the US Department of Energy's (DOE) Integrated Research Infrastructure (IRI) program. We present the ExaFEL team's 7 years of experience in developing real-time XFEL data analysis software for the DOE's exascale supercomputers. We present our experiences and lessons learned with the Perlmutter and Frontier supercomputers. Furthermore we outline essential data center services (and the implications for institutional policy) required for real-time data analysis. Finally we summarize our software and performance engineering approaches and our experiences with NERSC's Perlmutter and OLCF's Frontier systems. This work is intended to be a practical blueprint for similar efforts in integrating exascale compute resources into other cross-facility workflows.

59 BASIC BIOLOGICAL SCIENCES↗

Thermal Cycling of Mir Cooperative Solar Array (MCSA) Test Panels

The Mir Cooperative Solar Array (MCSA) project was a joint US/Russian effort to build a photovoltaic (PV) solar array and deliver it to the Russian space station Mir. The MCSA is currently being used to increase the electrical power on Mir and provide PV array performance data in support of Phase 1 of the International Space Station (ISS), which will use arrays based on the same solar cells used in the MCSA. The US supplied the photovoltaic power modules (PPMs) and provided technical and programmatic oversight while Russia provided the array support structures and deployment mechanism and built and tested the array. In order to ensure that there would be no problems with the interface between US and Russian hardware, an accelerated thermal life cycle test was performed at NASA Lewis Research Center on two representative samples of the MCSA. Over an eight-month period (August 1994 - March 1995), two 15-cell MCSA solar array 'mini' panel test articles were simultaneously put through 24,000 thermal cycles (+80 C to -100 C), equivalent to four years on-orbit. The test objectives, facility, procedure and results are described in this paper. Post-test inspection and evaluation revealed no significant degradation in the structural integrity of the test articles and no electrical degradation, not including one cell damaged early as an artifact of the test and removed from consideration. The interesting nature of the performance degradation caused by this one cell, which only occurred at elevated temperatures, is discussed. As a result of this test, changes were made to improve some aspects of the solar cell coupon-to-support frame interface on the flight unit. It was concluded from the results that the integration of the US solar cell modules with the Russian support structure would be able to withstand at least 24,000 thermal cycles (4 years on-orbit).

Hoffman, David J.↗

NASA Tech Briefs, August 2012

Topics covered include: Mars Science Laboratory Drill; Ultra-Compact Motor Controller; A Reversible Thermally Driven Pump for Use in a Sub-Kelvin Magnetic Refrigerator; Shape Memory Composite Hybrid Hinge; Binding Causes of Printed Wiring Assemblies with Card-Loks; Coring Sample Acquisition Tool; Joining and Assembly of Bulk Metallic Glass Composites Through Capacitive Discharge; 670-GHz Schottky Diode-Based Subharmonic Mixer with CPW Circuits and 70-GHz IF; Self-Nulling Lock-in Detection Electronics for Capacitance Probe Electrometer; Discontinuous Mode Power Supply; Optimal Dynamic Sub-Threshold Technique for Extreme Low Power Consumption for VLSI; Hardware for Accelerating N-Modular Redundant Systems for High-Reliability Computing; Blocking Filters with Enhanced Throughput for X-Ray Microcalorimetry; High-Thermal-Conductivity Fabrics; Imidazolium-Based Polymeric Materials as Alkaline Anion-Exchange Fuel Cell Membranes; Electrospun Nanofiber Coating of Fiber Materials: A Composite Toughening Approach; Experimental Modeling of Sterilization Effects for Atmospheric Entry Heating on Microorganisms; Saliva Preservative for Diagnostic Purposes; Hands-Free Transcranial Color Doppler Probe; Aerosol and Surface Parameter Retrievals for a Multi-Angle, Multiband Spectrometer LogScope; TraceContract; AIRS Maps from Space Processing Software; POSTMAN: Point of Sail Tacking for Maritime Autonomous Navigation; Space Operations Learning Center; OVERSMART Reporting Tool for Flow Computations Over Large Grid Systems; Large Eddy Simulation (LES) of Particle-Laden Temporal Mixing Layers; Projection of Stabilized Aerial Imagery Onto Digital Elevation Maps for Geo-Rectified and Jitter-Free Viewing; Iterative Transform Phase Diversity: An Image-Based Object and Wavefront Recovery; 3D Drop Size Distribution Extrapolation Algorithm Using a Single Disdrometer; Social Networking Adapted for Distributed Scientific Collaboration; General Methodology for Designing Spacecraft Trajectories; Hemispherical Field-of-View Above-Water Surface Imager for Submarines; and Quantum-Well Infrared Photodetector (QWIP) Focal Plane Assembly.

Source record↗

ModuleOT: A Hardware Security Module for Operational Technology: Preprint

With increasing penetration levels of distributed energy resources (DERs) on the distribution grid, as well as new technological advancements in the cyber space, new cyberattack vectors are being introduced, and the available attack surface is constantly increasing. Despite this increasing risk, the standard IEEE 1547-2018 does not yet recommend cybersecurity measures for DERs. To address this and to better protect data on the distribution grid - from the standpoints information security as well as operational security - ModuleOT has been developed. The module aims to significantly reduce cyberattack vectors by improving data privacy for user applications. This is accomplished by performing the core functions of encryption, authentication, authorization, certificate management, and user access control. The module integrates a custom security application with hardware cryptographic acceleration. The application secures all communications using Transmission Control Protocol over Internet Protocol (TCP/IP). These include the three most commonly used communications protocols for power systems information exchange: Modbus, Distributed Network Protocol 3 (DNP3), and Smart Energy Profile 2.0 (SEP2.0). These three protocols are also supported by IEEE 1547-2018 for all DER devices. This paper tests the data encryption/decryption feature on a physical networking test bed with emulated Modbus devices reporting grid data and presents the results.

24 POWER TRANSMISSION AND DISTRIBUTION↗

The Orbital Acceleration Research Experiment

The hardware and software of NASA's proposed Orbital Acceleration Research Experiment (OARE) are described. The OARE is to provide aerodynamic acceleration measurements along the Orbiter's principal axis in the free-molecular flow-flight regime at orbital attitude and in the transition regime during reentry. Models considering the effects of electromagnetic effects, solar radiation pressure, orbiter mass attraction, gravity gradient, orbital centripetal acceleration, out-of-orbital-plane effects, orbiter angular velocity, structural noise, mass expulsion signal sources, crew motion, and bias on acceleration are examined. The experiment contains an electrostatically balanced cylindrical proofmass accelerometer sensor with three orthogonal sensing axis outputs. The components and functions of the experimental calibration system and signal processor and control subsystem are analyzed. The development of the OARE software is discussed. The experimental equipment will be enclosed in a cover assembly that will be mounted in the Orbiter close to the center of gravity.

Blanchard, R. C.↗

Modernizing Fermilab s Control Hardware

Modernizing the Fermilab accelerator control system is essential to future operations of the laboratory's accelerator complex. The existing control system has evolved over four decades and uses hardware that is no longer available. The Accelerator Controls Operations Research Network (ACORN) Project will modernize the control system and replace end-of-life power supplies to enable future accelerator complex operations with megawatt particle beams. The ACORN project is planning to replace Fermilab’s obsolete CAMAC crate-and-card controls hardware with modern MicroTCA hardware. There are over 2,000 CAMAC cards and over 250 CAMAC crates serving various functions in Fermilab’s control system. We will present an overview of the existing CAMAC hardware and the proposed MicroTCA replacement plan.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

ModuleOT

ModuleOT is an open hardware security platform which provides all features necessary for securing remote energy resources. The platform consists of a physical bump in-the wire device which runs a custom-built application built with Go and Python and leverages AES-NI Instruction set available on modern hardware for cryptographic acceleration. By combining these features, ModuleOT acts as an all-in-one low-cost solution to enable cryptographically secured communications to any critical remote servers or devices. Because the software application has been built using Golang, this source can be easily compiled for different hardware platforms. The module is designed to provide the following core features: (1) encrypted communications across an untrusted network, (2) certificate-based authentication with secure storage, (3) hardware cryptographic acceleration, (4) IP-based whitelisting, (5) local firewall management, and (6) legacy (RS485) device support.

Hasandka, Adarsh↗

Testing a Neural Network Accelerator on a High-Altitude Balloon

The cognitive communications project has been working to re d machine learning approaches to support their deployment and sustained use in space environments. It has historically been difficult to implement such techniques on space platforms, however, due to the computational requirements they levy onto general-purpose avionics hardware. While technologies exist to accelerate the computation of aspects of neural networks, such platforms have not historically been deployed in space environments. Given that testing payloads in such environments can be both cost- and time-prohibitive, high-altitude balloons can be used as a way to approximate a space environment at a much lower cost, thus providing a cost-effective way in which to test newer approaches to hardware acceleration for artificial intelligence which may be deployed onto spacecraft more directly. This paper describes a successful test of a commercial off- the-shelf neural network accelerator on a high-altitude balloon. It begins by explaining our selection criteria when evaluating different commercial neural network acceleration techniques: primary considerations include size, weight, and power (SWaP) as well as ease of integration. Next, the paper describes the development and implementation of an experimental flight test platform: flight and ground components are discussed. Afterward, the paper discusses the experimental payload itself: this includes the experimental procedure as well as the specific image and method used for testing. Finally, the paper concludes with an evaluation of both the experimental device tested at altitude as well as the flight test framework itself, identifying how the existing platform can be used to continue tes g commercial off-the-shelf (COTS) solutions for acceleration.

Clark, Gilbert↗

Neural network methods for radiation detectors and imaging

Recent advances in image data proccesing through deep learning allow for new optimization and performance-enhancement schemes for radiation detectors and imaging hardware. This enables radiation experiments, which includes photon sciences in synchrotron and X-ray free electron lasers as a subclass, through data-endowed artificial intelligence. We give an overview of data generation at photon sources, deep learning-based methods for image processing tasks, and hardware solutions for deep learning acceleration. Most existing deep learning approaches are trained offline, typically using large amounts of computational resources. However, once trained, DNNs can achieve fast inference speeds and can be deployed to edge devices. A new trend is edge computing with less energy consumption (hundreds of watts or less) and real-time analysis potential. While popularly used for edge computing, electronic-based hardware accelerators ranging from general purpose processors such as central processing units (CPUs) to application-specific integrated circuits (ASICs) are constantly reaching performance limits in latency, energy consumption, and other physical constraints. These limits give rise to next-generation analog neuromorhpic hardware platforms, such as optical neural networks (ONNs), for high parallel, low latency, and low energy computing to boost deep learning acceleration (LA-UR-23-32395).

edge computing↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

$\mathrm{SageNet}$: Fast Neural Network Emulation of the Stiff-amplified Gravitational Waves from Inflation

Accurate modeling of the inflationary gravitational waves (GWs) requires time-consuming, iterative numerical integrations of differential equations to take into account their backreaction on the expansion history. To improve computational efficiency while preserving accuracy, we present the Stiff-amplified Gravitational-wave Emulator Network (SageNet), a deep learning framework designed to replace conventional numerical solvers (code available at https://github.com/YifangLuo/SageNet). SageNet employs a long short-term memory architecture to emulate the present-day energy density spectrum of the inflationary GWs with possible stiff amplification, Ω GW (f). Trained on a data set of 25,689 numerically generated solutions, SageNet allows accurate reconstructions of Ω GW (f) and generalizes well to a wide range of cosmological parameters; 90.9% of the test emulations with randomly distributed parameters exhibit errors of under 4%. In addition, SageNet demonstrates its ability to learn and reproduce the artificial, adaptive sampling patterns in numerical calculations, which implement denser sampling of frequencies around changes in spectral indices in Ω GW (f). The dual capability of learning both physical and artificial features of the numerical GW spectra establishes SageNet as a robust alternative to exact numerical methods. Finally, our benchmark tests show that SageNet reduces the computation time from tens of seconds to milliseconds, achieving a speedup of ∼10 4 times over standard CPU-based numerical solvers with the potential for further acceleration on GPU hardware. These capabilities make SageNet a powerful tool for accelerating Bayesian inference procedures for extended cosmological models. In a broad sense, the SageNet framework offers a fast, accurate, and generalizable solution to modeling cosmological observables whose theoretical predictions demand costly differential equation solvers.

Astronomy data modeling↗

GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design

Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art graph learning model. However, it remains notoriously challenging to inference GCNs over large graph datasets, limiting their application to large real-world graphs and hindering the exploration of deeper and more sophisticated GCN graphs. This is because real-world graphs can be extremely large and sparse. Furthermore, the node degree of GCNs tends to follow the power-law distribution and therefore have highly irregular adjacency matrices, resulting in prohibitive inefficiencies in both data processing and movement and thus substantially limiting the achievable GCN acceleration efficiency. To this end, this paper proposes the first GCN algorithm and accelerator Co-Design framework dubbed GCoD which can largely alleviate the aforementioned GCN irregularity and boost GCNs' inference efficiency. Specifically, on the algorithm level, GCoD integrates a divide and conquer GCN training strategy that polarizes the graphs to be either denser or sparser in local neighborhoods without compromising the model accuracy, resulting in graph adjacency matrices that (mostly) have merely two levels of workload and enjoys largely enhanced regularity and thus ease of acceleration. On the hardware level, we further develop a dedicated two-pronged accelerator with a separated engine to process each of the aforementioned workloads, further boosting the overall utilization and acceleration efficiency. Extensive experiments and ablation studies validate that our GCoD consistently outperforms state-of-the-art designs in terms of accelerator efficiency while maintaining or even improving the task accuracy. Additionally, we visualize GCoD trained graph adjacency matrices to better understand its advantages. All codes and pre-trained models will be released upon acceptance.

You, Haoran↗

Low Power Hardware-In-The-Loop Neuromorphic Training Accelerator

The training process for spiking neural networks can be very computationally intensive. Approaches such as evolutionary algorithms may require evaluating thousands or millions of candidate solutions. In this work, we propose using neuromorphic cores implemented on a Xilinx Zynq system on chip to accelerate and improve the energy efficiency of the evaluation step of an evolutionary training approach. We demonstrate this can significantly reduce the required energy to evolve a network with some cases showing greater than 10 times improvement as compared to a CPU-only system.

Mitchell, Parker↗