Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Implementing a neural network interatomic model with performance portability for emerging exascale architectures

The two main thrusts of computational science are increasingly accurate predictions and faster calculations; to this end, the zeitgeist in molecular dynamics (MD) simulations is pursuing machine learned and data driven interatomic models, e.g. neural network potentials, and novel hardware architectures, e.g. GPUs. Current implementations of neural network potentials are orders of magnitude slower than traditional interatomic models and while looming exascale computing offers the ability to run large, accurate simulations with these models, achieving portable performance for MD with new and varied exascale hardware requires rethinking traditional algorithms, using novel data structures, and library solutions. We re-implement a neural network interatomic model in CabanaMD, an MD proxy application, built on libraries developed for performance portability. Our implementation shows significantly improved thread scaling in this complex kernel as compared to a current LAMMPS implementation, across both strong and weak scaling. Our single-source solution enables simulations up to 20 million atoms on a single CPU node and 4 million atoms with improved performance on a single GPU. Furthermore, we also explore parallelism and data layout choices (using flexible data structures called AoSoAs) and their effect on performance, seeing up to ~50% and ~5% improvements in performance on a GPU by choosing the right level of parallelism and data layout respectively.

97 MATHEMATICS AND COMPUTING↗

High-throughput terahertz imaging: progress and challenges

Abstract Many exciting terahertz imaging applications, such as non-destructive evaluation, biomedical diagnosis, and security screening, have been historically limited in practical usage due to the raster-scanning requirement of imaging systems, which impose very low imaging speeds. However, recent advancements in terahertz imaging systems have greatly increased the imaging throughput and brought the promising potential of terahertz radiation from research laboratories closer to real-world applications. Here, we review the development of terahertz imaging technologies from both hardware and computational imaging perspectives. We introduce and compare different types of hardware enabling frequency-domain and time-domain imaging using various thermal, photon, and field image sensor arrays. We discuss how different imaging hardware and computational imaging algorithms provide opportunities for capturing time-of-flight, spectroscopic, phase, and intensity image data at high throughputs. Furthermore, the new prospects and challenges for the development of future high-throughput terahertz imaging systems are briefly introduced.

36 MATERIALS SCIENCE↗

A physical unclonable neutron sensor for nuclear arms control inspections

Abstract Classical sensor security relies on cryptographic algorithms executed on trusted hardware. This approach has significant shortcomings, however. Hardware can be manipulated, including below transistor level, and cryptographic keys are at risk of extraction attacks. A further weakness is that sensor media themselves are assumed to be trusted, and any authentication and encryption is done ex situ and a posteriori. Here we propose and demonstrate a different approach to sensor security that does not rely on classical cryptography and trusted electronics. We designed passive sensor media that inherently produce secure and trustworthy data, and whose honest and non-malicious nature can be easily established. As a proof-of-concept, we manufactured and characterized the properties of non-electronic, physical unclonable, optically complex media sensitive to neutrons for use in a high-security scenario: the inspection of a military facility to confirm the absence or presence of nuclear weapons and fissile materials.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Evolutionary vs imitation learning for neuromorphic control at the edge*

Abstract Neuromorphic computing offers the opportunity to implement extremely low power artificial intelligence at the edge. Control applications, such as autonomous vehicles and robotics, are also of great interest for neuromorphic systems at the edge. It is not clear, however, what the best neuromorphic training approaches are for control applications at the edge. In this work, we implement and compare the performance of evolutionary optimization and imitation learning approaches on an autonomous race car control task using an edge neuromorphic implementation. We show that the evolutionary approaches tend to achieve better performing smaller network sizes that are well-suited to edge deployment, but they also take significantly longer to train. We also describe a workflow to allow for future algorithmic comparisons for neuromorphic hardware on control applications at the edge.

Schuman, Catherine↗

Performance-Aligned LLMs for Generating Fast HPC Code

Optimizing scientific software is a difficult task because codebases are often large and complex, and performance can depend upon several factors including the algorithm, its implementation, and hardware among others. Causes of poor performance can originate from disparate sources and be difficult to diagnose. Recent years have seen a multitude of work that use large language models (LLMs) to assist in software development tasks. However, these tools are trained to model the distribution of code as text, and are not specifically designed to understand performance aspects of code. In this work, we introduce a reinforcement learning based methodology to align the outputs of code LLMs with performance. This allows us to build upon the current code modeling capabilities of LLMs and extend them to generate better performing code. Here, we demonstrate that our fine-tuned model improves the expected speedup of generated code over base models for a set of benchmark tasks from 0.9 to 1.6 for serial code and 1.9 to 4.5 for OpenMP parallel code.

Computer science↗

Fugu v.0.1

SAND2021-15052 O Fugu provides a common software framework for designing and prototyping algorithms for spiking neuromorphic hardware and compiling to multiple hardware platforms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Aimone, James↗

Mixed Precision xsdk Project: Final Report for Year Ended March 31, 2022

In the past year, as part of the Mixed Precision xsdk Project we have worked an the development and analysis of algorithms for mixed precision hardware. In the following we briefly discuss and summarize our contributions. The postdoctoral researchers were Srikara Pranesh (to September 30, 2021) and Mantas Mikaitis (from October 1, 2021).

97 MATHEMATICS AND COMPUTING↗

Citadels Final Report (GMLC 2.2.1: Citadels)

This is the final project report for the Grid Modernization Laboratory Consortium (GMLC) Resilient Distribution System (RDS) Citadels project. The primary goal of this GMLC project was to increase the operational flexibly of power systems by engaging microgrids distributedly, coordinated using consensus algorithms. The primary goal was successfully achieved. The primary goal was divided into three areas: Implement peer-to-peer control between microgrid controllers using the Open Field Message Bus (OpenFMB) approach; Develop and implement consensus algorithms on commercially available hardware that allows a group of microgrids to distributedly implement operational controls; Develop the architectures and controls to enable groups of microgrids to coordinate their operations to support the bulk power system during abnormal events, and end-use loads in the event the bulk power systems fail.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Quantum Computing for Energy-Related Applications

Growing interest in quantum computing and simulations have created opportunities for its deployment to improve processes pertaining to energy production, distribution, and consumption. While quantum computing is considered as a paradigm shift in our basic understanding of physical computation, effective implementation of quantum computing in energy applications also depends on progress and development in the dimensions of both quantum computing hardware and quantum computing algorithms. To fully address the status and future challenges of quantum information science (QIS) applied within the energy sector, in this presentation, we firstly summarize recent advancements on the applications of quantum computing to energy infrastructure and materials, complex energy system processes, advanced manufacturing, and energy system security. Then, we will demonstrate the results of quantum computing both on a simulator and a quantum device accessing from OLCF. Our first example is to use the variational quantum eigensolver (VQE) with a unitary coupled cluster with singles and doubles (UCCSD) ansatz to simulate a series of LixHyq molecules (q=-1, 0, +1). The obtained results showed that the quantum computing VQE-UCCSD is comparable to classical CCSD for small systems like LiH with respect to full configuration interaction (FCI). Targeting on CO2 capture application, our second example is to use VQE to quantify molecular vibrational energies and reaction pathways between CO2 and a simplified amine-based solvent model—NH3 to form H2NCOOH. This research showcases quantum computing applications in the study of CO2 capture reactions.

Duan, Yuhua↗

2022 Small Turbine Certification Awardee: Ryse Energy LLC - Americas

To meet the growing demands of the distributed wind market and comply with current U.S. standards, Ryse Energy LLC-Americas (Ryse Energy, formerly Primus Wind Power) plans to update all six products in its AIR Range family of micro wind turbines. A new, more cost-effective circuit board will be paired with other hardware upgrades, advanced control algorithms, and a Bluetooth function that allows users to more easily program and control the turbine. This Competitiveness Improvement Project (CIP) award will fund certification testing of the new circuit board to make sure it meets American National Standards Institute/American Clean Power Association (ANSI/ACP) and UL Federal Communications Commission (FCC) safety and quality standards. Primus Wind Power developed the prototype for the new circuit board with an earlier round of CIP funding, and Primus has received CIP awards supporting other certification and optimization projects.

CIP↗

TAMM: Tensor algebra for many-body methods

Tensor algebra operations such as contractions in computational chemistry consume a significant fraction of the computing time on large-scale computing platforms. The widespread use of tensor contractions between large multi-dimensional tensors in describing electronic structure theory has motivated the development of multiple tensor algebra frameworks targeting heterogeneous computing platforms. In this paper, we present Tensor Algebra for Many-body Methods (TAMM), a framework for productive and performance-portable development of scalable computational chemistry methods. TAMM decouples the specification of the computation from the execution of these operations on available high-performance computing systems. With this design choice, the scientific application developers (domain scientists) can focus on the algorithmic requirements using the tensor algebra interface provided by TAMM, whereas high-performance computing developers can direct their attention to various optimizations on the underlying constructs, such as efficient data distribution, optimized scheduling algorithms, and efficient use of intra-node resources (e.g., graphics processing units). The modular structure of TAMM allows it to support different hardware architectures and incorporate new algorithmic advances. We describe the TAMM framework and our approach to the sustainable development of scalable ground- and excited-state electronic structure methods. We present case studies highlighting the ease of use, including the performance and productivity gains compared to other frameworks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On-Sensor Data Filtering using Neuromorphic Computing for High Energy Physics Experiments

This work describes the investigation of neuromorphic computing-based spiking neural network (SNN) models used to filter data from sensor electronics in high energy physics experiments conducted at the High Luminosity Large Hadron Collider. We present our approach for developing a compact neuromorphic model that filters out the sensor data based on the particle's transverse momentum with the goal of reducing the amount of data being sent to the downstream electronics. The incoming charge waveforms are converted to streams of binary-valued events, which are then processed by the SNN. We present our insights on the various system design choices - from data encoding to optimal hyperparameters of the training algorithm - for an accurate and compact SNN optimized for hardware deployment. Our results show that an SNN trained with an evolutionary algorithm and an optimized set of hyperparameters obtains a signal efficiency of about 91% with nearly half as many parameters as a deep neural network.

R. Kulkarni, Shruti↗

Sentinel

Network intrusion detection systems (NIDS) are commonplace in network security but they frequently employ algorithms that are computational demanding requiring hardware and software with significant power requirements. Two examples of such resource-intensive algorithms used for network security are regular expression matching and broader signature pattern matching which are commonly used in deep packet inspection (DPI). Network security algorithms that have large power requirements may be a challenge for low-power internet-of-things (IoT) environments, which generally lack the power resources to implement complex security measures like computationally expensive DPI at the edge. Furthermore, IoT environments incorporating 5G standalone networks have network latency constraints beyond just power that make DPI at the edge even more difficult. Programmable logic is ideally suited for machine learning inference for DPI because of its deep instruction level parallelism and single-cycle memory access. Machine learning approaches for DPI have been explored before using the programmable logic of field programmable gate arrays (FPGA) as a potential solution for NIDS approaches that would be power-suitable for IoT. However, those previous programmable logic NIDS approaches utilize either a supervised or unsupervised learning model. Sentinel utilizes the ensemble of these two machine learning approaches known as a semi-supervised approach which has shown promise in NIDS implementations. Sentinel provides a programmable logic implementation of a semi-supervised approach for DPI which operates at much lower power and latency than a GPU implementation with negligible loss of accuracy due to quantization through a logistic regressor.

Anderson, MatthewW [Idaho National Laboratory (INL↗

Resilience and fault tolerance in high-performance computing for numerical weather and climate prediction

Progress in numerical weather and climate prediction accuracy greatly depends on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increased average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys hardware, application-level and algorithm-level resilience approaches of particular relevance to time-critical numerical weather and climate prediction systems. A selection of applicable existing strategies is analysed, featuring interpolation-restart and compressed checkpointing for the numerical schemes, in-memory checkpointing, user-level failure mitigation and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analysed and some recommendations outlined for future developments.

54 ENVIRONMENTAL SCIENCES↗

How to Build a Quantum Supercomputer: Scaling from Hundreds to Millions of Qubits

In the span of four decades, quantum computation has evolved from an intellectual curiosity to a potentially realizable technology. Today, small-scale demonstrations have become possible for quantum algorithmic primitives on hundreds of physical qubits and proof-of-principle error-correction on a single logical qubit. Nevertheless, despite significant progress and excitement, the path toward a full-stack scalable technology is largely unknown. There are significant outstanding quantum hardware, fabrication, software architecture, and algorithmic challenges that are either unresolved or overlooked. These issues could seriously undermine the arrival of utility-scale quantum computers for the foreseeable future. Here, we provide a comprehensive review of these scaling challenges. We show how the road to scaling could be paved by adopting existing semiconductor technology to build much higher-quality qubits, employing system engineering approaches, and performing distributed quantum computation within heterogeneous high-performance computing infrastructures. These opportunities for research and development could unlock certain promising applications, in particular, efficient quantum simulation/learning of quantum data generated by natural or engineered quantum systems. To estimate the true cost of such promises, we provide a detailed resource and sensitivity analysis for classically hard quantum chemistry calculations on surface-code error-corrected quantum computers given current, target, and desired hardware specifications based on superconducting qubits, accounting for a realistic distribution of errors. Furthermore, we argue that, to tackle industry-scale classical optimization and machine learning problems in a cost-effective manner, heterogeneous quantum-probabilistic computing with custom-designed accelerators should be considered as a complementary path toward scalability.

Mohseni, Masoud↗

Zero and Finite Temperature Quantum Simulations Powered by Quantum Magic

We introduce a quantum information theory-inspired method to improve the characterization of many-body Hamiltonians on near-term quantum devices. We design a new class of similarity transformations that, when applied as a preprocessing step, can substantially simplify a Hamiltonian for subsequent analysis on quantum hardware. By design, these transformations can be identified and applied efficiently using purely classical resources. In practice, these transformations allow us to shorten requisite physical circuit-depths, overcoming constraints imposed by imperfect near-term hardware. Importantly, the quality of our transformations is t u n a b l e : we define a 'ladder' of transformations that yields increasingly simple Hamiltonians at the cost of more classical computation. Using quantum chemistry as a benchmark application, we demonstrate that our protocol leads to significant performance improvements for zero and finite temperature free energy calculations on both digital and analog quantum hardware. Specifically, our energy estimates not only outperform traditional Hartree-Fock solutions, but this performance gap also consistently widens as we tune up the quality of our transformations. In short, our quantum information-based approach opens promising new pathways to realizing useful and feasible quantum chemistry algorithms on near-term hardware.

Physics↗

Project Presentation: Hardware-Based Demonstration of Temperature Control Functions for Reactor Systems

Autonomous thermal regulation in nuclear reactors remains crucial for maintaining stable operation and ensuring the integrity of fuel. To alleviate public skepticism of the safety of nuclear reactors, demonstrating control over this key factor is pivotal. Utilizing electric heat pads to simulate the heat released in a reactor core, thermocouples for temperature monitoring, and an Arduino micro programmable logic controller (PLC) with an embedded proportional-integral-derivative (PID) algorithm for control, a hardware-based demonstration of a reactor heating system will validate the efficacy of reactor control over this key parameter. This report covers internship project presentation as well as relevant experience with nuclear and proposal for project upgrade.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Poster: Hardware-Based Demonstration of Temperature Control Functions for Reactor Systems

Autonomous thermal regulation in nuclear reactors remains crucial for maintaining stable operation and ensuring the integrity of fuel. To alleviate public skepticism of the safety of nuclear reactors, demonstrating control over this key factor is pivotal. Utilizing electric heat pads to simulate the heat released in a reactor core, thermocouples for temperature monitoring, and an Arduino micro programmable logic controller (PLC) with an embedded proportional-integral-derivative (PID) algorithm for control, a hardware-based demonstration of a reactor heating system will validate the efficacy of reactor control over this key parameter.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗