Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

A Survey: Handling Irregularities in Neural Network Acceleration with FPGAs

In the last decade, Artificial Intelligence (AI) through Deep Neural Networks (DNNs) has penetrated virtually every aspect of science, technology, and business. Many types of DNNs have been and continue to be developed, including Convolutional Neural Networks (CNNs), Recurrent Neural Net- works (RNNs), and Graph Neural Networks (GNNs). The overall problem for all of these Neural Networks (NNs) is that their target applications generally pose stringent constraints on latency and throughput, while also having strict accuracy requirements. There have been many previous efforts in creating hardware to accelerate NNs. The problem designers face is that optimal NN models typically have significant irregularities, making them hardware-unfriendly. In this paper, we first define the problems in NN acceleration by characterizing common irregularities in NN processing into 4 types; then we summarize the existing works that handle the four types of irregularities efficiently using hardware, especially FPGAs; finally, we provide a new vision of next-generation FPGA-based NN acceleration: that the emerging heterogeneity in the next-generation FPGAs is the key to achieving higher performance.

Geng, Tong↗

Solving sparse finite element problems on neuromorphic hardware

The finite element method (FEM) is one of the most important and ubiquitous numerical methods for solving partial differential equations (PDEs) on computers for scientific and engineering discovery. Applying the FEM to larger and more detailed scientific models has driven advances in high-performance computing for decades. Here we demonstrate that scalable spiking neuromorphic hardware can directly implement the FEM by constructing a spiking neural network that solves the large, sparse, linear systems of equations at the core of the FEM. We show that for the Poisson equation, a fundamental PDE in science and engineering, our neural circuit achieves meaningful levels of numerical accuracy and close to ideal scaling on modern, inherently parallel and energy-efficient neuromorphic hardware, specifically Intel’s Loihi 2 neuromorphic platform. We illustrate extensions to irregular mesh geometries in both two and three dimensions as well as other PDEs such as linear elasticity. Our spiking neural network is constructed from a recurrent network model of the brain’s motor cortex and, in contrast to black-box deep artificial neural network-based methods for PDEs, directly translates the well-understood and trusted mathematics of the FEM to a natively spiking neuromorphic algorithm.

Applied mathematics↗

LuGo: An enhanced quantum phase estimation implementation

Quantum Phase Estimation (QPE) is a cardinal algorithm in quantum computing that plays a crucial role in various applications, including cryptography, molecular simulation, and solving systems of linear equations. However, the standard implementation of QPE faces challenges related to time complexity and circuit depth, which limit its practicality for large-scale computations. We introduce LuGo, a novel framework designed to enhance the performance of QPE by reducing circuit duplication, as well as using parallelization techniques to achieve faster generation of the QPE circuit and gate reduction. We validate the effectiveness of our framework by generating quantum linear solver circuits, which require both QPE and inverse QPE, to solve linear systems of equations. LuGo achieves significant improvements in both computational efficiency and hardware requirements without compromising on accuracy. Compared to a standard QPE implementation, LuGo reduces time consumption to generate a circuit that solves a 2 6 × 2 6 system matrix by a factor of 50.68 and over 31× reduction of quantum gates and circuit depth, with no fidelity loss on an ideal quantum simulator. Furthermore, we demonstrated the versatility and scalability of LuGo enabled HHL algorithm by simulating a canonical Hele-Shaw fluid problem using a quantum simulator. With these advantages, LuGo paves the way for more efficient implementations of QPE, enabling broader applications across several quantum computing domains.

Quantum algorithm↗

Efficient human activity recognition with spatio-temporal spiking neural networks

In this study, we explore Human Activity Recognition (HAR), a task that aims to predict individuals' daily activities utilizing time series data obtained from wearable sensors for health-related applications. Although recent research has predominantly employed end-to-end Artificial Neural Networks (ANNs) for feature extraction and classification in HAR, these approaches impose a substantial computational load on wearable devices and exhibit limitations in temporal feature extraction due to their activation functions. To address these challenges, we propose the application of Spiking Neural Networks (SNNs), an architecture inspired by the characteristics of biological neurons, to HAR tasks. SNNs accumulate input activation as presynaptic potential charges and generate a binary spike upon surpassing a predetermined threshold. This unique property facilitates spatio-temporal feature extraction and confers the advantage of low-power computation attributable to binary spikes. We conduct rigorous experiments on three distinct HAR datasets using SNNs, demonstrating that our approach attains competitive or superior performance relative to ANNs, while concurrently reducing energy consumption by up to 94%.

60 APPLIED LIFE SCIENCES↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗

Multi-Objective Hyperparameter Optimization for Spiking Neural Network Neuroevolution

Neuroevolution has had significant success over recent years, but there has been relatively little work applying neuroevolution approaches to spiking neural networks (SNNs). SNNs are a type of neural network that includes temporal processing component, are not easily trained using other methods, and can be deployed into energy-efficient neuromorphic hardware. In this work, we investigate two evolutionary approaches for training SNNs. We explore the impact of the hyperparameters of the evolutionary approaches, including tournament size, population size, and representation type, on the performance of the algorithms. We present a multi-objective Bayesian-based hyperparameter optimization approach to tune the hyperparameters to produce the most accurate and smallest SNNs. We show that the hyperparameters can significantly affect the performance of these algorithms. We also perform sensitivity analysis and demonstrate that every hyperparameter value has the potential to perform well, assuming other hyperparameter values are set correctly.

Parsa, Maryam↗

ExaSGD: 2021 Kernel Thrust Activities

The Kernel Thrust milestone ADSE22-214 covers the development of device-capable optimization algorithms and solvers technologies required by the ExaSGD project’s software stack in order to solve security-constrained alternating current optimal power flow (SC-ACOPF) problems on emerging exascale architectures. To this extent, in FY21 the main objective of the Kernel Thrust was (i) provide robust optimization solver(s) that run efficiently on hardware accelerator devices (i.e., NVIDIA and AMD GPUs) to perform intra-node computations and (ii) provide coarse-grain parallel optimization capabilities that exploit the decomposition opportunities present in the SC-ACOPF challenge problems to provide exascale-capable solvers.

97 MATHEMATICS AND COMPUTING↗

AOI [1] Advanced Manufacturing of Ceramic Anchors with Embedded Sensors for Process and Health Monitoring of Coal Boilers

Researchers at West Virginia University (WVU) developed methods to fabricate and test ceramic anchors with an embedded sensor technology for monitoring the health and processing conditions within pulverized coal (PC) and fluidized-bed combustion (FBC) boiler systems. The technology included the development of advanced manufacturing processes for 2D/3D printing electroceramic (conductive ceramic) sensor designs within the ceramic anchor microstructure during the manufacturing process. This advanced manufacturing process would allow for the precise control of local microstructure and composition in order to engineer layer-by-layer any protective and electrically active materials within the refractory anchor. This 3D printing technology would permit the rapid and controlled design of the refractory microstructure and embedded sensor design throughout the volume of the ceramic anchor. The work also included a method to interconnect the sensors to boiler shell through the anchor clamp, where the sensor signals will be processed by low-power electronics and transmitted wirelessly to a central processing hub. The end-goal of the program was to produce a ceramic anchor sensor system which would be ready for implementation within a coal boiler, and/or other similar refractory liner systems (such as that in the glass and metal manufacturing areas). The project objectives were to: 1) Define the chemical and microstructural stability, in addition to the electrical properties, of oxide and non-oxide ceramic composites to be embedded within the ceramic anchor compositions that may operate up to 1400ºC; 2) Develop and implement the 2D/3D printing technology to pattern and control the microstructure of the ceramic anchor and embedded sensor circuits; 3) Develop an interconnect technology which will permit easy installation of the ceramic anchors and signal collection at the boiler shell; 4) Develop low power analog electronics and wireless communication hardware to efficiently collect the sensor signal at each processing unit and transmit data to a central hub for data analysis; 5) Demonstrate the smart ceramic anchor system for temperature and liner fracture within a high-temperature processing unit, such as a boiler furnace or glass melting furnace floor/wall liner.

20 FOSSIL-FUELED POWER PLANTS↗

Dendritic Computing with Multigate Ferroelectric Field-Effect Transistors

Although inspired by neuronal systems in the brain, artificial neural networks generally employ point-neurons, which offer computational complexity far less than that of their biological counterparts. Neurons have dendritic arbors that connect to different sets of synapses and offer local nonlinear accumulation – playing a pivotal role in processing and learning. Inspired by this, we propose a novel neuron design based on a multigate ferroelectric field-effect transistor that mimics dendrites. It leverages ferroelectric nonlinearity for local computations within dendritic branches while utilizing the transistor action to generate the neuronal output. The branched architecture enables smaller crossbar arrays in hardware integration, improving efficiency. Using an experimentally calibrated device-circuit-algorithm cosimulation framework, we demonstrate that networks incorporating our dendritic neurons achieve superior performance compared to much larger networks without dendrites (∼ 17× fewer trainable weight parameters). These findings suggest that dendritic hardware can significantly improve computational efficiency and learning capacity of neuromorphic systems optimized for edge applications.

36 MATERIALS SCIENCE↗

Energy-efficient, Large-scale Molecular Dynamics Simulations via Hardware- and Algorithm-level Optimization

This work aims to develop a framework for energy-efficient computing that will enable molecular dynamics (MD) simulations of large-scale phenomena with atomic precision and simultaneously remove computational bottlenecks limiting the speed of MD simulations. We seek to implement such an approach through the development of surrogate models for the interatomic force calculation combined with the use of mixed numerical precision formats. For a model system of neutral atoms (only pairwise interactions), significant force calculation efficiency improvements were achieved, without detrimental effects on atomic structures or average energies, using single precision, by developing a surrogate model (deep neural network), and by quantizing this surrogate model. For a model system of charged atoms, the reciprocal-space calculation of electrostatic interactions was identified as the main bottleneck, and the development of a surrogate model should be pursued to achieve an estimated one-order-of-magnitude additional speedup.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Even Higher-Level Synthesis: An Exploration of AI Hardware Accelerators using HLS4ML

With the rise of artificial intelligence, the popularization of deep learning, and a constantly evolving industry, the demand for flexible and efficient tools has never been greater. As algorithms grow more complex, their runtime and energy consumption increase exponentially. Customized hardware accelerators, long used for specific mathematical operations, remain essential for managing modern applications' computational and power demands. Hardware accelerators can speed up complex computations by orders of magnitude, but their manual design and verification processes are often challenging and time-consuming. High-Level Synthesis (HLS) provides a solution by transforming high-level algorithm descriptions, typically written in C++ or SystemC, into synthesizable RTL suitable for hardware implementation. This approach reduces development time for RTL engineers while offering flexibility beyond what traditional handwritten RTL can provide. We extended this capability to the machine-learning domain with the open-source framework hls4ml, which allows neural networks trained in Python frameworks like Tensorflow or PyTorch to be synthesized into efficient hardware representations for the traditional FPGA and ASIC flows. This breakthrough addresses the growing need for reduced design turnaround and easy verification of ML hardware accelerators with low latency and power efficiency constraints. During this tutorial, we will demonstrate how Python complements HLS by simplifying the ML design process, bridging the gap between software and hardware development. Attendees will explore how we translate neural networks modeled in Python into fixed-point C++ models suitable for HLS workflows. We will dive into strategies like Value-Range Analysis and Quantization-Aware Training, which optimize these designs for deployment and evaluate their accuracy, power consumption, and energy efficiency. To exemplify these concepts, experts from Fermilab will share their experiences applying this technology to high-energy physics experiments, where real-time, low-latency processing is critical. Over the years, Fermilab engineers have demonstrated how deep neural networks, optimized for hardware using hls4ml, can meet the stringent requirements of trigger systems at the CERN Large Hadron Collider. These systems rely on rapid decision-making to process immense data volumes while retaining only the most relevant events for further analysis. The application of hls4ml has also been extended to innovative technologies like smart pixel arrays. These smart pixels integrate ML inference capabilities directly into sensor devices, enabling localized data processing at the pixel level. This approach drastically reduces the need to transmit raw data to external processing units, significantly decreasing power consumption and latency. By embedding neural networks within the pixel architecture, the smart pixels can identify and prioritize relevant data in real time, providing a highly efficient solution for edge computing in scenarios such as particle detectors and imaging systems. Fermilab's work highlights the potential of hardware-accelerated ML in scenarios where both speed and power efficiency are mission-critical. Through this tutorial, attendees will gain valuable insights into the challenges and solutions of deploying ML in hardware. Understanding how HLS and hls4ml streamline the development of neural network-based hardware accelerators is fundamental for the industry's future. Participants will learn how these technologies are shaping the future of AI and scientific computing.

Di Guglielmo, Giuseppe [Fermilab]↗

VQE method: a short survey and recent developments

Abstract The variational quantum eigensolver (VQE) is a method that uses a hybrid quantum-classical computational approach to find eigenvalues of a Hamiltonian. VQE has been proposed as an alternative to fully quantum algorithms such as quantum phase estimation (QPE) because fully quantum algorithms require quantum hardware that will not be accessible in the near future. VQE has been successfully applied to solve the electronic Schrödinger equation for a variety of small molecules. However, the scalability of this method is limited by two factors: the complexity of the quantum circuits and the complexity of the classical optimization problem. Both of these factors are affected by the choice of the variational ansatz used to represent the trial wave function. Hence, the construction of an efficient ansatz is an active area of research. Put another way, modern quantum computers are not capable of executing deep quantum circuits produced by using currently available ansatzes for problems that map onto more than several qubits. In this review, we present recent developments in the field of designing efficient ansatzes that fall into two categories—chemistry–inspired and hardware–efficient—that produce quantum circuits that are easier to run on modern hardware. We discuss the shortfalls of ansatzes originally formulated for VQE simulations, how they are addressed in more sophisticated methods, and the potential ways for further improvements.

Fedorov, Dmitry A. (ORCID:0000000316598580)↗

Bosonic field digitization for quantum computers

Quantum simulation of quantum field theory is a flagship application of quantum computers that promises to deliver capabilities beyond classical computing. The realization of quantum advantage will require methods that can accurately predict error scaling as a function of the resolution and parameters of the model and that can be implemented efficiently on quantum hardware. In this paper, we address the representation of lattice bosonic fields in a discretized field amplitude basis, develop methods to predict error scaling, and present efficient qubit implementation strategies. A low-energy subspace of the bosonic Hilbert space, defined by a boson occupation number cutoff, can be represented with exponentially good accuracy by a low-energy subspace of a finite-size Hilbert space. The finite representation construction and the associated errors are directly related to the accuracy of the Nyquist-Shannon sampling and the finite Fourier transforms of the boson number states in the field and the conjugate-field bases. We analyze the relation between the boson mass, the discretization parameters used for wave function sampling, and the finite representation size. Numerical simulations of small size Φ 4 problems demonstrate that the boson mass optimizing the sampling of the ground state wave function is a good approximation to the optimal boson mass yielding the minimum low-energy subspace size. However, we find that accurate sampling of general wave functions does not necessarily result in accurate representation. Finally, we develop methods for validating and adjusting the discretization parameters to achieve more accurate simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

cuTS: Scaling Subgraph Isomorphism on Distributed Multi-GPUSystems Using Trie Based Data Structure

Subgraph isomorphism is a pattern-matching algorithm widely used in many domains such as chem-informatics, bioinformatics, databases, and social network analysis. It is computationally expensive and is a proven NP-hard problem. The massive parallelism offered by the GPU hardware is well suited for solving the subgraph isomorphism. However, current GPU implementations are far from the achievable performance. Moreover, the enormous memory requirement of current approaches limits the problem size that can be handled. This work analyzes the fundamental challenges associated with processing the subgraph isomorphism on GPUs and develops an efficient GPU hardware-aware implementation. We also develop a new GPU-friendly trie-based data structure to drastically reduce the intermediate storage space requirement. Hence, our approach runs larger benchmarks than the competitors. We also develop the first distributed sub-graph isomorphism algorithm for GPUs. Our experimental evaluation section demonstrates the efficacy of our approach by comparing the execution time and number of cases that we can handle against the state-of-the-art GPU implementations.

Xiang, Lizhi↗

AXI4MLIR: User-Driven Automatic Host Code Generation for Custom AXI-Based Accelerators

Tensor algebra operations represent an important class of algorithms used across many applications, including machine learning, scientific computing, and data analytics. As a result, the efficient generation of custom accelerators for tensor operations has received increased attention. Previous efforts have produced automated tools enabling users to prototype and explore optimized accelerators. However, little effort has been focused on the host-accelerator interaction in these tools. Efficient use of hardware accelerators requires knowledge about the accelerator's capabilities (operations, data formats, and opcode support), the host CPU microarchitecture (e.g., memory hierarchy), the host-accelerator interface, and the application's features (which code regions should be mapped onto an accelerator). Manually rewriting the original applications to facilitate improved custom accelerator mapping is an error-prone and time-consuming endeavor. To cope with this, we propose AXI4MLIR, a new framework to automatically generate and optimize the communication between the host CPU and arbitrary accelerators that implement linear algebra algorithms. AXI4MLIR extends the MLIR compiler framework to automatically generate efficient host-accelerator driver code for accelerators with AXI-based interfaces. Our compiler extensions enable automatic driver code generation while carefully considering the host's memory hierarchy and target accelerator features. To demonstrate the flexibility and utility of AXI4MLIR, we test it with diverse use cases that include different types of accelerators, tiling scenarios, and dataflow schemes. We compare our experimental results to manual implementations of host-accelerator driver code and find that our approach can reduce CPU cache references by 56% and deliver up to a 1.65x speedup.

Bohm Agostini, Nicolas↗

Dish-STARS Commercialization (Final Report)

The goal of this project is to aggressively support the near-term commercialization of a new technology platform – based on the integration of solar concentrators and micro- and meso-channel process technology (MMPT) – that was evaluated and identified as a strong candidate for near-term commercialization at EERE’s inaugural Lab-Corps program during early FY2016. Known as STARS, for Solar Thermochemical Advanced Reactor System, or Dish-STARS TM when paired with parabolic dish concentrators, STARS is a promising energy-related technology developed at the Pacific Northwest National Laboratory (PNNL) that efficiently converts solar energy into chemical energy. Combined with economies through hardware mass production, the efficiency of Dish-STARS TM provides a near-term opportunity for the production of renewable electricity, fuels and chemicals. The project supported the cooperative development of Dish-STARSTM by PNNL and industry partners including California Gas Company (SoCalGas) and the startup company, STARS Technology Corporation (STC), which was founded by the PNNL Lab-Corps team that evaluated STARS on behalf of EERE. Under this project, the team advanced the Technology Readiness Level 6 (TRL 6) STARS reaction system, developed under the previous DOE SunShot project, to TRL 7 through on-sun testing in California by PNNL. The advances made in this project enabled STC to accelerate commercial development and initiate work towards a major technology demonstration in California for a hydrogen filling station application.

14 SOLAR ENERGY↗