Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26

A survey of techniques for optimizing transformer inference

Recent years have seen a phenomenal rise in the performance and applications of transformer neural networks. The family of transformer networks, including Bidirectional Encoder Representations from Transformer (BERT), Generative Pretrained Transformer (GPT) and Vision Transformer (ViT), have shown their effectiveness across Natural Language Processing (NLP) and Computer Vision (CV) domains. Transformer-based networks such as ChatGPT have impacted the lives of common men. However, the quest for high predictive performance has led to an exponential increase in transformers' memory and compute footprint. Researchers have proposed techniques to optimize transformer inference at all levels of abstraction. Further, this paper presents a comprehensive survey of techniques for optimizing the inference phase of transformer networks. We survey techniques such as knowledge distillation, pruning, quantization, neural architecture search and lightweight network design at the algorithmic level. We further review hardware-level optimization techniques and the design of novel hardware accelerators for transformers. We summarize the quantitative results on the number of parameters/FLOPs and the accuracy of several models/techniques to showcase the tradeoff exercised by them. We also outline future directions in this rapidly evolving field of research. We believe that this survey will educate both novice and seasoned researchers and also spark a plethora of research efforts in this field.

97 MATHEMATICS AND COMPUTING↗

Quantum algorithms for geologic fracture networks

Abstract Solving large systems of equations is a challenge for modeling natural phenomena, such as simulating subsurface flow. To avoid systems that are intractable on current computers, it is often necessary to neglect information at small scales, an approach known as coarse-graining. For many practical applications, such as flow in porous, homogenous materials, coarse-graining offers a sufficiently-accurate approximation of the solution. Unfortunately, fractured systems cannot be accurately coarse-grained, as critical network topology exists at the smallest scales, including topology that can push the network across a percolation threshold. Therefore, new techniques are necessary to accurately model important fracture systems. Quantum algorithms for solving linear systems offer a theoretically-exponential improvement over their classical counterparts, and in this work we introduce two quantum algorithms for fractured flow. The first algorithm, designed for future quantum computers which operate without error, has enormous potential, but we demonstrate that current hardware is too noisy for adequate performance. The second algorithm, designed to be noise resilient, already performs well for problems of small to medium size (order 10–1000 nodes), which we demonstrate experimentally and explain theoretically. We expect further improvements by leveraging quantum error mitigation and preconditioning.

58 GEOSCIENCES↗

Characterizing quantum circuits with qubit functional configurations

Abstract We develop a systematic framework for characterizing all quantum circuits with qubit functional configurations. The qubit functional configuration is a mathematical structure that can classify the properties and behaviors of quantum circuits collectively. Major benefits of classifying quantum circuits in this way include: 1. All quantum circuits can be classified into corresponding types; 2. Each type characterizes important properties (such as circuit complexity) of the quantum circuits belonging to it; 3. Each type contains a huge collection of possible quantum circuits allowing systematic investigation of their common properties. We demonstrate the theory’s application to analyzing the hardware-efficient ansatzes of variational quantum algorithms. For potential applications, the functional configuration theory may allow systematic understanding and development of quantum algorithms based on their functional configuration types.

97 MATHEMATICS AND COMPUTING↗

Efficient online quantum circuit learning with no upfront training

Optimization is a promising candidate for studying the utility of variational quantum algorithms (VQAs). However, evaluating cost functions using quantum hardware introduces runtime overheads that limit exploration. Surrogate-based methods can reduce calls to a quantum computer, yet existing approaches require hyperparameter pre-training and have been tested only on small problems. Here, we show that surrogate-based methods can enable successful optimization at scale, without pre-training, by using radial basis function interpolation (RBF) to construct an adaptive, hyperparameter-free surrogate. Using the surrogate as an acquisition function drives hardware queries to the vicinity of the true optima. For 16-qubit random 3-regular Max-Cut instances with the Quantum Approximate Optimization Algorithm (QAOA), our method outperforms state-of-the-art approaches, without considering their upfront training costs. Furthermore, we successfully optimize QAOA circuits for 127-qubit random Ising models on an IBM processor using 10 4 −10 5 measurements. Strong empirical performance demonstrates the promise of automated surrogate-based learning for large-scale VQA applications.

97 MATHEMATICS AND COMPUTING↗

Decentralized modular hybrid supervisory control for the formation of unmanned helicopters

Abstract Formation control of Unmanned Aerial Vehicles (UAVs) requires them to tightly cooperate to reach and keep the formation, while avoiding collision. This paper proposes a novel decentralized hybrid supervisory control approach for the formation control of multiple UAVs. This is achieved by developing a symbolic motion planning technique to polarly partition the motion space resulting in a finite state discrete event model for the motion dynamics of each UAV. Then, a modular discrete supervisor is designed for different components of the formation mission including reaching the formation, keeping the formation, and collision avoidance. Further, for the collision avoidance mechanism, a novel top‐down decomposition‐based approach is developed to design local supervisors decentralizedly. It is formally proved that with the proposed top‐down decomposition‐based approach, the (locally) supervised UAVs, as a whole, can cooperatively satisfy the desired (global) collision avoidance specification. The proposed decentralized supervisory control algorithm is also verified through a hardware‐in‐the‐loop simulator for the formation control of unmanned helicopters.

Karimoddini, Ali↗

Sampling on NISQ Devices: "Who’s the Fairest One of All?"

Modern NISQ devices are subject to a variety of biases and sources of noise that degrade the solution quality of computations carried out on these devices. A natural question that arises in the NISQ era, is how fairly do these devices sample ground state solutions. To this end, we run five fair sampling problems (each with at least three ground state solutions) that are based both on quantum annealing and, on the Grover Mixer, -QAOA algorithm for gate-based NISQ hardware. In particular, we use seven IBM Q devices, the Aspen-9 Rigetti device, the IonQ device, and three D-Wave quantum annealers. For each of the fair sampling problems, we measure the ground state probability, the relative fairness of the frequency of each ground state solution with respect to the other ground state solutions, and the aggregate error as given by each hardware provider. Overall, our results show that NISQ devices do not achieve fair sampling yet. Furthermore, we also observe differences in the software stack with a particular focus on compilation techniques that illustrate what work will still need to be done to achieve a seamless integration of frontend (i.e., quantum circuit description) and backend compilation.

Computer Science↗

HD-Bind: Encoding of Molecular Structure with Low Precision, Hyperdimensional Binary Representations

Publicly available collections of drug-like molecules have grown to comprise tens of billions of compounds due to advances in combinatorial chemistry. Traditional methods for identifying "hit" molecules from a large collection of potential drug-like candidates have relied on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have the major drawback that they require exceptional computing capabilities for even relatively small collections of molecules. Hyperdimensional Computing (HDC) is a recently-proposed learning paradigm that represents data with high-dimension binary vectors; this allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas. We consider existing HDC approaches for molecular property classification and introduce two novel encodings of a commonly-used molecular representation, the extended connectivity fingerprint (ECFP). We show that HDC-based inference methods are as much as 91 times more efficient than traditional machine learning methods, and achieve an acceleration of nearly nine orders of magnitude compared to molecular docking. Our results show that HDC accelerated methods retain competitive accuracy on a number of well-studied tasks such as molecular property predictions using the MoleculeNet dataset, and bind/no-bind activity classification using the DUD-E and LIT-PCBA datasets. Our work thus motivates further investigation into molecular representation learning to develop ultraefficient pre-screening tools.

Jones, William↗

Many-gluon tree amplitudes on modern GPUs: A case study for novel event generators

The compute efficiency of Monte-Carlo event generators for the Large Hadron Collider is expected to become a major bottleneck for simulations in the high-luminosity phase. Aiming at the development of a full-fledged generator for modern GPUs, we study the performance of various recursive strategies to compute multi-gluon tree-level amplitudes. We investigate the scaling of the algorithms on both CPU and GPU hardware. Finally, we provide practical recommendations as well as baseline implementations for the development of future simulation programs. The GPU implementations can be found at: https://www.gitlab.com/ebothmann/blockgen-archive.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Codebase release 1.0 for BlockGen

The compute efficiency of Monte-Carlo event generators for the Large Hadron Collider is expected to become a major bottleneck for simulations in the high-luminosity phase. Aiming at the development of a full-fledged generator for modern GPUs, we study the performance of various recursive strategies to compute multi-gluon tree-level amplitudes. We investigate the scaling of the algorithms on both CPU and GPU hardware. Finally, we provide practical recommendations as well as baseline implementations for the development of future simulation programs. The GPU implementations can be found at: https://www.gitlab.com/ebothmann/blockgen-archive.

Bothmann, Enrico↗

Neuromorphic Processing and Sensing for Interception

Interception of a moving and potentially evading target can be a challenging problem, in particular for conditions in which the target may be moving at high speeds and difficult to detect. We have proposed to merge two Sandia LDRD efforts, the SPARR Spiking/Processing Array (neuromorphic event-driven sensing) and the Dragonfly-Inspired Algorithms for Intercept- Trajectory Planning (neural-inspired algorithms for interception) toward a unified system with direct application to national security. Neuromorphic systems demonstrate the most potential for speed and efficiency gains when communication is event-driven and computations are simple but parallelizable. Accordingly, we anticipate fully realizing potential benefits from a neuromorphic interception system if event-driven sensing is combined with processing and acting also implemented on event-driven (spiking) systems. We have successfully translated a neural-inspired interception algorithm to a neural network architecture for evaluation on neuromorphic hardware. Preliminary implementations of the neural network designed for implementation on the Loihi chip are still too immature for conclusive evaluation, but the results of this effort have demonstrated a viable path for a previously developed dragonfly-inspired interception algorithm to be implemented on neuromorphic hardware.

97 MATHEMATICS AND COMPUTING↗

Non-Intrusive Parallel-in-Time Solvers for Partial Differential Equations (Final Report)

Many time-dependent problems and simulations are often modeled using Partial Differential Equations. Traditional modeling approaches that use sequential time-stepping are reaching a bottleneck in optimizing efficiency. The Center of Applied Science and Computing at Lawrence Livermore National Laboratory extensively works on parallelizing these algorithms to leverage the increasing computational power from the growing number of processors in computer hardware. In particular, they aim to design non-intrusive algorithms that can generalize to a variety of problems and sizes without requiring additional information from or modifications on the original problems. Multigrid Reduction in Time (MGRIT) is a parallel-in-time algorithm that is designed to be non-intrusive. This project focuses on increasing the efficiency of MGRIT by approximating the coarse-grid operator using machine learning approaches as a means to find the most non-intrusive, or general, solution.

97 MATHEMATICS AND COMPUTING↗

SpecFIDLER User Manual: Software Version 2.4.0

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SpecFIDLER User Manual (Software V.2.5.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SpecFIDLER User Manual (Software V.2.6.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

SpecFIDLER User Manual (Software V.2.3.0)

The Spectroscopic Field Instrument for Detection of Low Energy Radiation (SpecFIDLER) allows response teams to detect and quantify plutonium contamination on the ground. Notional scenarios include dispersion from a weapon accident, or the launch failure of a space probe containing a radioisotope thermoelectric generator. Unlike other instruments, the thin-window sodium iodide detector is sensitive to the low-energy gamma rays emitted by plutonium isotopes. The system supports both mobile survey as well as stationary sampling. This manual provides information about installing, maintaining, and troubleshooting the SpecFIDLER. The scope of this document includes the physical hardware, software for data acquisition, and algorithms for data analysis. Recent changes to the software and algorithms aim to streamline the operation of the system.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Demonstrating autonomous controls on hardware test beds is a necessity for successful missions to Mars and beyond

NASA and the Department of Defense are planning for a mission to Mars in the 2030s–2040s using nuclear thermal propulsion (NTP). NTP uses a nuclear reactor to heat flowing hydrogen and create thrust. A serious concern for crewed and uncrewed missions to Mars is the loss of reactor control. The reactor startup and initial rocket impulse are initiated in cislunar or near-earth orbital regions; therefore, radio communications between ground control and the NTP engine should occur in real time. However, radio communications can take more than 20 min, depending on planet positions, to reach Mars orbiters from ground control. To address this delay, local autonomous controls are implemented onboard the NTP engine to ensure acceptable operation. However, autonomous controls have not been demonstrated or implemented in research or power reactor contexts because of safety and reliability concerns. To enable autonomous controls development, demonstration, and validation, Oak Ridge National Laboratory has created a nonnuclear hardware-in-the-loop test bed. Sensors throughout the test bed relay system status and hardware response to the user control algorithm, including measurements of temperature, flow, pressure of a loop, control drum position, and drum speed. This paper discusses the development of this facility and user accessibility.

33 ADVANCED PROPULSION SYSTEMS↗

Applicability of APT aided-inertial system to crustal movement monitoring

The APT system, its stage of development, hardware, and operations are described. The algorithms required to perform the real-time functions of navigation and profiling are presented. The results of computer simulations demonstrate the feasibility of APT for its primary mission: topographic mapping with an accuracy of 15 cm in the vertical. Also discussed is the suitability of modifying APT for the purpose of making vertical crustal movement measurements accurate to 2 cm in the vertical, and at least marginal feasibility is indicated.

Soltz, J. A.↗