Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Hybrid Oscillator-Qubit Quantum Processors: Instruction Set Architectures, Abstract Machine Models, and Applications

This tutorial offers a pedagogical guide to hybrid quantum processors that integrate discrete-variable (DV) qubits and continuous-variable (CV) oscillators. Aimed at computer scientists, engineers, and physicists, it provides an overview of the experimental, algorithmic, and architectural aspects of this novel and rapidly developing hardware model. Experimental realizations of this model include superconducting, trapped-ion, and neutral-atom platforms. By combining DV and CV components, hybrid oscillator-qubit processors enable a powerful new paradigm that offers complementary strengths for quantum control, error correction, computation, and simulation. Working toward the goal of a full-stack system connecting applications to CV-DV hardware, we define and formulate abstract machine models and instruction set architectures. These essential abstractions enable codesign of hardware and software, and resource estimation for exploring the potential of current and future hardware for computational and simulation tasks. Using these abstractions, we present both new and existing examples that illustrate the benefits of hybrid CV-DV processors relative to traditional DV-only hardware in computation as well as quantum simulation of physical models. Examples include algorithms for transferring states between DV and CV systems, performing the quantum Fourier transform, and simulation of lattice gauge theories. Relative to qubit-only hardware, the bosonic degrees of freedom natively available in hybrid architectures can substantially reduce the circuit complexity of simulations for physical models containing bosons. A key technique is the extension of quantum signal processing ideas to CV-DV systems. This work is intended to serve as a timely and comprehensive guide to this relatively unexplored yet promising approach to quantum computation and to provide a road map to guide future development.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

SODA: a New Synthesis Infrastructure for Agile Hardware Design of Machine Learning Accelerators

Next generation systems, such as edge devices, will have to provide efficient processing of machine learning (ML) algorithms along several metrics, including energy, performance, area, and latency. However, the quickly evolving field of ML makes it extremely difficult to generate accelerators able to support a wide variety of algorithms. At the same time, designing accelerators in hardware description languages (HDLs) by hand is hard and time consuming, and does not allow quick exploration of the design space. This paper discusses the SODA synthesizer, an automated open source high-level ML framework-to-Verilog compiler targeting ML Application-Specific Integrated Circuits (ASICs) chiplets based on the LLVM infrastructure. The SODA synthesizers will allow implementing optimal designs by combining templated and fully tunable IPs and macros, and fully custom components generated through high-level synthesis. All these components will be provided through an extendable resource library, characterized with both commercial and open source logic design flows. Through a closed loop design space exploration engine, developers will be able to quickly explore their hardware designs along different dimension

Minutoli, Marco↗

Latency considerations for stochastic optimizers in variational quantum algorithms

Variational quantum algorithms, which have risen to prominence in the noisy intermediate-scale quantum setting, require the implementation of a stochastic optimizer on classical hardware. To date, most research has employed algorithms based on the stochastic gradient iteration as the stochastic classical optimizer. In this work we propose instead using stochastic optimization algorithms that yield stochastic processes emulating the dynamics of classical deterministic algorithms. This approach results in methods with theoretically superior worst-case iteration complexities, at the expense of greater per-iteration sample (shot) complexities. We investigate this trade-off both theoretically and empirically and conclude that preferences for a choice of stochastic optimizer should explicitly depend on a function of both latency and shot execution times.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Error and Correction Analysis for the FFA@CEBAF Energy Upgrade

An energy upgrade design for the Continuous Electron Beam Accelerator Facility (CEBAF) is under development, using fixed field alternating gradient (FFA) return arcs to recirculate electron beam up to an additional five times through the accelerating structures at CEBAF. A necessary component of any large accelerator is a beam steering and optical correction system. Small environmental changes and system errors can lower beam quality or even shut down the machine; and in pursuit of the scientific mission of JLab, high quality electron beams must be delivered to the experimental halls on a predictable schedule. Correction in the novel FFA arcs of the current upgrade design is complicated by several factors. These complexities inform the choice of correction algorithm structure and parameter values. A baseline algorithm in addition to diagnostic and correction hardware configuration is presented. The effect of this correction protocol is shown with respect to estimated errors, and several possible extensions of the algorithm are discussed. This work presents an important proof of concept for the FFA@CEBAF design effort, and provides a functional correction strategy which may be simply adjusted and optimized for future design changes.

Coxe, Alex [Old Dominion Univ., Norfolk, VA (Unite↗

PQML: Enabling the Predictive Reproducibility on NISQ Machines for Quantum ML Applications

Quantum computing represents a groundbreaking approach to high-performance computing. In recent years, quantum computers have progressed from single-qubit processors to systems boasting over 400 qubits. The presence of such a large number of qubits offers significant advantages, including enhanced computational speed—a capability beyond classical computing methods. However, the current stage of quantum computing is referred to as the noisy intermediate-scale quantum (NISQ) era. The existence of noise in this era presents challenges in testing quantum computing applications, leading to considerable variance in application results. Furthermore, the diverse noise characteristics observed across different machines exacerbate this issue, complicating the selection of the appropriate machine for application execution. In response to these challenges, we introduce our Predictive Quantum Machine Learning (PQML) tool. This tool is designed to predict outcomes when executing identical quantum machine learning applications—specifically, a critical suite of variational quantum algorithms—across various quantum computers during the NISQ era. This effort relies on data collected over a 12-month period. To the best of our knowledge, this study represents the first attempt to ensure reproducibility across quantum computers for complex circuits. Additionally, we have developed a model capable of forecasting the accuracy of quantum computers for variational quantum algorithms, with a particular emphasis on quantum machine learning as a case study.

Senapati, Priyabrata [Kent State University]↗

State preparation and evolution in quantum computing: a perspective from Hamiltonian moments

Quantum algorithms on the noisy intermediate-scale quantum (NISQ) devices are expected to simulate quan- tum systems that are classically intractable to demonstrate quantum advantages. However, the non-negligible gate error on the NISQ devices impedes the conventional quantum algorithms to be implemented. Practical strategies usually exploit hybrid quantum-classical quantum algorithms to demonstrate potentially useful ap- plications of quantum computing in the NISQ era. Among the numerous hybrid quantum-classical algorithms, recent efforts highlight the development of quantum algorithms based upon quantum computed Hamiltonian moments, ?f|Hˆn|f? (n = 1, 2, · · · ), with respect to quantum state |f?. In this tutorial, we will give a brief review of these quantum algorithms with focuses on the typical ways of computing Hamiltonian moments using quantum hardware and improving the accuracy of the estimated state energies based on the quantum computed moments. Furthermore, we will present a tutorial to show how we can measure and compute the Hamiltonian moments of a four-site Heisenberg model, and compute the energy and magnetization of the model utilizing the imaginary time evolution in the real IBM-Q NISQ hardware environment. Along this line, we will further discuss some practical issues associated with these algorithms. We will conclude this tutorial review by overviewing some possible developments and applications in this direction in the near future.

Aulicino, Joseph C.↗

Leveraging Qubit Loss Detection in Fault-Tolerant Quantum Algorithms

Qubit loss errors constitute a dominant source of noise in many quantum hardware systems, particularly in neutral-atom quantum computers. We develop a theoretical framework to effectively detect and correct loss errors in logical algorithms and leverage such loss information in decoding. Considering general quantum error correction codes and logical circuits, we introduce a delayed-erasure decoder for experimentally motivated error models which leverages information from delayed loss detection to accurately correct loss errors, even when the precise moment of the error is unknown. Using this decoder, we identify strategies for detecting and correcting loss errors based on the logical circuit structure. For deep circuits prior to logical measurement, we explore methods to integrate loss detection into syndrome extraction with minimal overhead, identifying optimal strategies depending on the qubit loss fraction in the noise and hardware capabilities. In contrast, we find that many key algorithmic subroutines involve frequent gate teleportation, shortening the circuit depth before logical measurement and naturally replacing qubits with no additional experimental overhead. We simulate this setting using a toy model algorithm for small-angle synthesis and find a significant performance improvement as the loss fraction increases. These results provide a path forward for advancing large-scale fault-tolerant quantum computation in systems with loss error detection.

atoms↗

Adaptive Circuit Learning for Quantum Metrology

Quantum sensing is an important application of emerging quantum technologies. We explore whether a hybrid system of quantum sensors and quantum circuits can surpass the classical limit of sensing. In particular, we use optimization techniques to search for encoder and decoder circuits that scalably improve sensitivity under given application and noise characteristics. Furthermore, our approach uses a variational algorithm that can learn a quantum sensing circuit based on platform-specific control capacity, noise, and signal distribution. The quantum circuit is composed of an encoder which prepares the optimal sensing state and a decoder which gives an output distribution containing information of the signal. We optimize the full circuit to maximize the Signal-to-Noise Ratio (SNR). Furthermore, this learning algorithm can be run on real hardware scalably by using the "parameter-shift" rule which enables gradient evaluation on noisy quantum circuits, avoiding the exponential cost of quantum system simulation. We demonstrate up to 13.12x SNR improvement over existing fixed protocol (GHZ), and 3.19x Classical Fisher Information (CFI) improvement over the classical limit on 15 qubits using IBM quantum computer. More notably, our algorithm overcomes the decreasing performance of existing entanglement-based protocols with increased system sizes.

42 ENGINEERING↗

Case Study of Using Kokkos and SYCLs Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.

Dufek, Amanda S↗

Progress on Associate-Particle Imaging Algorithms, 2020

The present work describes progress on the development of imaging algorithms that use fast neutron signatures acquired using the associated-particle imaging (API) method. The present work complements ongoing work to develop neutron source and detector hardware to enable field inspection by investigating algorithms that are capable of discriminating between critical materials or extracting three-dimensional geometrical information from single-sided or transmission measurements. The present work is divided into three approaches:(1)Iterative reconstruction of inelastic gamma-ray emissions to perform three-dimensional time-of-flight imaging in a single view in either transmission or backscatter configurations. Iterative reconstruction enables image resolution better than the inherent TOF resolution.(2)Decomposition of registered neutron and x ray radiographs into an assumed material list for each pixel in the image.(3)Material identification using full spectral analysis that includes the emergent neutron and gamma ray energies, times, and angles.Progress for each approach is summarized for fiscal year 2020.

97 MATHEMATICS AND COMPUTING↗

Progress on Associated-Particle Imaging Algorithms, 2022

The present work describes progress on developing imaging algorithms that use fast neutron signatures acquired using the associated-particle imaging (API) method. The present work complements ongoing work to develop neutron source and detector hardware to enable field inspection by investigating algorithms that are capable of discriminating among critical materials or extracting three-dimensional (3D) geometrical information from single-sided or transmission measurements. The present work is divided into three approaches: 1.Iterative reconstruction of inelastic gamma-ray emissions to perform 3D time-of-flight (TOF) imaging in a single view in either transmission or backscatter configurations. Iterative reconstruction enables image resolution better than the inherent TOF resolution. 2.Decomposition of registered neutron and x-ray radiographs into an assumed material list for each pixel in the image. 3.Material identification using full spectral analysis that includes the emergent neutron and gamma ray energies, times, and angles. Progress for each approach is summarized for fiscal year (FY) 2022.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Encoding trade-offs and design toolkits in quantum algorithms for discrete optimization: coloring, routing, scheduling, and other problems

Challenging combinatorial optimization problems are ubiquitous in science and engineering. Several quantum methods for optimization have recently been developed, in different settings including both exact and approximate solvers. Addressing this field of research, this manuscript has three distinct purposes. First, we present an intuitive method for synthesizing and analyzing discrete (i.e., integer-based) optimization problems, wherein the problem and corresponding algorithmic primitives are expressed using a discrete quantum intermediate representation (DQIR) that is encoding-independent. This compact representation often allows for more efficient problem compilation, automated analyses of different encoding choices, easier interpretability, more complex runtime procedures, and richer programmability, as compared to previous approaches, which we demonstrate with a number of examples. Second, we perform numerical studies comparing several qubit encodings; the results exhibit a number of preliminary trends that help guide the choice of encoding for a particular set of hardware and a particular problem and algorithm. Our study includes problems related to graph coloring, the traveling salesperson problem, factory/machine scheduling, financial portfolio rebalancing, and integer linear programming. Third, we design low-depth graph-derived partial mixers (GDPMs) up to 16-level quantum variables, demonstrating that compact (binary) encodings are more amenable to QAOA than previously understood. We expect this toolkit of programming abstractions and low-level building blocks to aid in designing quantum algorithms for discrete combinatorial problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Benchmarking for AI for Science

AI has been instrumental for recent developments in a number of domains of the sciences. With several hundred machine learning (ML) algorithms and models, and numerous AI-specific hardware platforms, a common quest for all scientists working on AI for Science is around the selection of machine learning algorithm(s) to solve their domain-specific scientific problems. A number of different initiatives around AI Benchmarking have been set up and have been useful in understanding the benefits of different ML algorithms for different tasks.However, with the majority of these AI Benchmarking initiatives focusing on the conventional notions of benchmarking, where the focus is purely runtime performance (such as training time or inference time), their suitability for benchmarking different ML algorithms for solving scientific problems has been viewed as a performance problem even though both are hardly the same. To make reasonable, explainable, and justifiable advancements in science using AI, it is critical to focus on the merits of these algorithms in handling different domain science problems. In other words, more emphasis must be given on Benchmarking for AI for Science than AI Benchmarking. The vision of the former is not only to assess the performance of ML algorithms, but also to assess, and understand the benefits and merits of different ML algorithms in handling scientific problems. Benchmarking for AI for Science, instead of pure performance focused AI Benchmarking, has several benefits: (i) it has the potential to offer advances in the sciences, much more rapidly than through pure performance-based AI methods, (ii) it will encourage the community to focus on developing better domain-specific AI techniques, particularly given the provision for being able to benchmark different techniques, and (iii) it will encourage hardware manufacturers to focus on developing science-specific hardware subsystems.

Thiyagalingam, Jeyan↗

A Holistic Algorithmic Approach to Improving Accuracy, Robustness, and Computational Efficiency for Atmospheric Dynamics

Atmospheric weather and climate models must perform simulations very quickly to be useful. Therefore, modelers have traditionally focused on reducing computations as much as possible. However, in our new era of increasingly compute-capable hardware, data movement is now the prohibiting expense. This study examines the computational benefits of a new algorithmic approach to modeling atmospheric dynamics on scales relevant to weather and climate simulation. Rather than minimizing computations, this new approach considers the larger problem more holistically, including spatial accuracy, temporal accuracy, robustness (i.e., oscillations), on-node efficiency, and internode data transfers together at once. Numerical experiments demonstrate how computations can be strategically increased to simultaneously address each of these constraints while reducing data movement to adapt to modern accelerated hardware. The new algorithm can achieve at times up to 80% peak floating point throughput in single precision on the Nvidia Tesla V100 GPU, where the traditional approach is shown to only achieve single-digit floating point efficiency. Further, the new algorithm is twice as fast as a standard Runge--Kutta time integrator, and high-order accuracy with Weighted Essentially Non-Oscillatory (WENO) limiting came at less than 30% additional runtime cost on a GPU, thus increasing the accuracy per degree of freedom.

54 ENVIRONMENTAL SCIENCES↗

DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware

Spiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches.

Computer Science↗

Performance Evaluation of Distributed Energy Resource Management Algorithm in Large Distribution Networks

This paper presents performance evaluation of hierarchical optimization and control for distributed energy resource management system (DERMS) in large distribution networks via an advanced hardware-in-the-loop (HIL) platform. The HIL platform provides realistic testing in a laboratory environment, including the accurate modeling of a full-scale distribution system of 11,000 nodes, the DERMS software controller, and 90 power hardware photovoltaics (PVs) and battery inverters. The applied DERMS algorithm is designed based on a realtime optimal power flow algorithm and implemented with acceleration design that performs fast dispatch of simulated PVs and real physical hardware DER devices every 4 seconds.

DERMS↗

Union: A Unified HW-SW Co-Design Ecosystem in MLIR for Evaluating Tensor Operationson Spatial Accelerators

To meet the extreme compute demands for deep learning across commercial and scientific applications, dataflow accelerators are becoming increasingly popular. While these“domain-specific” accelerators are not fully programmable like CPUs and GPUs, they retain varying levels of flexibility with respect to data orchestration, i.e., dataflow and tiling optimizations to enhance efficiency. There are several challenges when designing new algorithms and mapping approaches to execute the algorithms for a target problem on new hardware. Previous works have addressed these challenges individually. To address this challenge as a whole, in this work, we present an HW-SW co-design ecosystem for spatial accelerators called Union within the popular MLIR compiler infrastructure. Our framework allows exploring different algorithms and their mappings on several accelerator cost models. Union also includes a plug-and-play library of accelerator cost models and mappers which can easily be extended. The algorithms and accelerator cost models are connected via a novel mapping abstraction that captures the map space of spatial accelerators which can be systematically pruned based on constraints from the hardware, workload, and mapper. We demonstrate the value of Union for the community with several case studies which examine offloading different tensor operations (CONV/GEMM/Tensor Contraction) on diverse accelerator architectures using different mapping schemes.

Jeong, Geonhwa↗

State Dependent Optimization with Quantum Circuit Cutting

Quantum circuits can be reduced through optimization to better fit the constraints of quantum hardware. One such method, initial-state dependent optimization (ISDO), reduces gate count by leveraging knowledge of the input quantum states. Surprisingly, we found that ISDO is broadly applicable to the downstream circuits produced by circuit cutting. Circuit cutting also requires measuring upstream qubits and has some flexibility of selection observables to do reconstruction. Therefore, we propose a state-dependent optimization (SDO) framework that incorporates ISDO, our newly proposed measure-state dependent optimization (MSDO), and a biased observable selection strategy. Building on the strengths of the SDO framework and recognizing the scalability challenges of circuit cutting, we propose nonseparate circuit cutting-a more flexible approach that enables optimizing gates without fully separating them. We validate our methods on noisy simulations of QAOA, QFT, and BV circuits. Results show that our approach consistently mitigates noise and improves overall circuit performance, demonstrating its promise for enhancing quantum algorithm execution on near-term hardware.

Li, Xinpeng↗