Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

AEDAM: Whole Program Adaptive Error Detection and Mitigation (Final Report)

The overall goals of the AEDAM project were to fundamentally transform software transient-error detection through the design of configurable specialized detectors, quantitative characterization of hardware resilience and software vulnerabilities, and composition of the specialized detectors to most efficiently protect the whole program. Within this scope, UT Austin's research contribution related to enabling and studying the composition of detectors and possible different hardware errors, specifically: (1) developed the Hamartia open-source error injection framework that is designed to make composition studies simple, and (2) develop the methodology and demonstrate the potential benefits of error detector composition.

97 MATHEMATICS AND COMPUTING↗

Keeping LAMMPS cutting edge

Since its inception 30 years ago, LAMMPS has grown to be a world-class molecular dynamics code and a cornerstone of computational materials science research. This project aimed to keep LAMMPS at the forefront of molecular dynamics simulations by adapting LAMMPS to the latest developments in machine learning technology and hardware. Initially, the project set out to provide a unified implementation of active learning for efficient training data generation in LAMMPS, but the research trajectory pivoted to address more immediate and impactful opportunities. On the hardware side, recent record-breaking molecular dynamics simulations were developed on the Cerebras wafer-scale AI chip, and this project has developed an interface between LAMMPS and the hardware-specific molecular dynamics code to accelerate and simplify development and user adoption. On the software side, PyTorch’s Ahead-of-Time (AOT) compilation features promised increased performance for state-of-the-art equivariant neural network potentials, and this project laid the groundwork for their adoption in LAMMPS, resulting in a nearly 20x acceleration in extreme cases. Combined with a comprehensive benchmark study of LAMMPS across all current exascale systems, this project has reinforced LAMMPS’s role as a versatile, high-performance tool for current and future materials science applications.

36 MATERIALS SCIENCE↗

Programmable simulations of molecules and materials with reconfigurable quantum processors

Simulations of quantum chemistry and quantum materials are believed to be among the most important applications of quantum information processors. However, realizing practical quantum advantage for such problems is challenging because of the prohibitive computational cost of programming typical problems into quantum hardware. Here we introduce a simulation framework for strongly correlated quantum systems represented by model spin Hamiltonians that uses reconfigurable qubit architectures to simulate real-time dynamics in a programmable way. Our approach also introduces an algorithm for extracting chemically relevant spectral properties via classical co-processing of quantum measurement results. We develop a digital–analogue simulation toolbox for efficient Hamiltonian time evolution using digital Floquet engineering and hardware-optimized multi-qubit operations to accurately realize complex spin–spin interactions. As an example, we propose an implementation based on Rydberg atom arrays. In addition, we show how detailed spectral information can be extracted from the dynamics through snapshot measurements and single-ancilla control, enabling the evaluation of excitation energies and finite-temperature susceptibilities from a single dataset. To illustrate the approach, we show how to use the method to compute key properties of a polynuclear transition-metal catalyst and two-dimensional magnetic materials.

74 ATOMIC AND MOLECULAR PHYSICS↗

DS-TIDE: Harnessing Dynamical Systems for Efficient Time-Independent Differential Equation Solving

Time-Independent Differential Equations (TIDEs) are central to modeling equilibrium behavior across a wide range of scientific and engineering domains, from electrostatics to porous media flow. Conventional numerical solvers offer reliable solutions but incur significant computational costs due to fine-grained discretization and iterative procedures. Machine learning-based approaches address this by replacing iterative solving processes with one-time inference; however, their sophisticated models require extensive training resources that often exceed those of traditional solvers. Consequently, designing a TIDE solver that achieves high accuracy, broad applicability, and exceptional computational efficiency remains a fundamental challenge. In this paper, we propose DS-TIDE, a novel hardware solver that is inspired by, and subsequently leverages, the intrinsic connection between Dynamical Systems (DS) and Differential Equations (DEs) to efficiently and accurately solve TIDEs. DS-TIDE employs a CMOS-compatible DS-based processor, whose physical states evolve under carefully designed DE-driven dynamics and naturally converge to equilibrium -- the solution of the target TIDE -- within ~1µs on a ~1-watt DS-TIDE processor. To enhance expressivity, DS-TIDE incorporates Heterogeneous Dynamics with Temporal Layering (HDTL), which solves TIDEs through a three-stage DS evolution -- conditioning, solving, and decoding -- each governed by specialized dynamics. The entire evolution process is analogous to an infinitely deep neural network temporally unrolled, offering the system the capability of representing complex equations. Furthermore, DS-TIDE is equipped with an on-device DS-DE Auto-Alignment mechanism that dynamically adapts intrinsic hardware dynamics within milliseconds, effectively aligning the system’s dynamics to diverse target DEs. Experimental results across TIDEs from a wide range of scientific and engineering domains demonstrate that DS-TIDE achieves ~10^3× speedup, ~10^5× energy savings, and competitive or superior accuracy compared to state-of-the-art numerical and ML-based solvers.

Liu, Chuan↗

Fast jet tagging with MLP-Mixers on FPGAs

We explore the innovative use of MLP-Mixer models for real-time jet tagging and establish their feasibility on resource-constrained hardware like FPGAs. MLP-Mixers excel in processing sequences of jet constituents, achieving state-of-the-art performance on datasets mimicking Large Hadron Collider conditions. By using advanced optimization techniques such as High-Granularity Quantization and Distributed Arithmetic, we achieve unprecedented efficiency. These models match or surpass the accuracy of previous architectures, reduce hardware resource usage by up to 97%, double the throughput, and half the latency. Additionally, non-permutation-invariant architectures enable smart feature prioritization and efficient FPGA deployment, setting a new benchmark for machine learning in real-time data processing at particle colliders.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Power Layout Design of a GaN HEMTs-Based High-Power High-Efficiency Three-Level ANPC Inverter for 800 V DC Bus System

Multiple commutation paths exist for switching devices in a three-level active neutral point clamped (3L-ANPC) inverter operation based on the selected switching state and current direction. In addition, the capacitive coupling path of the nonswitching device is a key design aspect for enabling high voltage and high current operation of gallium nitride (GaN) switches in 3L-ANPC topology. A comprehensive study of the switching transient events of inner, outer, and clamping devices of 3L-ANPC is presented in this article. The commutation mechanisms for worst-case transient voltage overshoots (TVOs) are identified. A simplified equivalent circuit model is presented to determine the design criteria for the power layout structure's parasitic inductances. A power layout strategy satisfying the design criteria is then proposed using an insulated metal substrate power printed circuit board (PCB) to enable efficient high-power operation. The proposed design minimizes the commutation and capacitive coupling path inductances to 6 nH and 11.5 nH, respectively. This enables the fast switching operation of GaN HEMTs at 800 V dc, 36 A with a low TVO of 31% verified through experimental three-level double pulse test results. Experimental evaluation of a three-phase 3L-ANPC hardware prototype based on the proposed power layout shows 99% efficiency at 800 V, 9.5 kVA and 50 kHz switching frequency. The proposed design achieves a low case-to-ambient thermal resistance of 2.3 °C/W.

42 ENGINEERING↗

Evaluating Thermostats' Deadbands Using HVAC Hardware-In-the-Loop Experiment for Advanced Control Strategies

Smart thermostats have gained significant popularity due to their potential for optimizing energy consumption and enhanced user control while ensuring occupants' comfort. The deadband, also referred to as temperature differential, is defined as the temperature difference between the desired setpoint and upper threshold or lower threshold for the HVAC equipment to turn on. It is a key factor influencing energy efficiency and user satisfaction. This paper presents a comparative analysis of the deadbands of five different smart thermostats, tested with a heat pump, aiming to identify variations in their deadband settings and implications for energy management. The experimental study was conducted using a HVAC hardware-in-theloop (HIL) system that integrates smart thermostats with physical HVAC equipment in a simulated house environment. The study explores the trade-offs between energy efficiency and occupant comfort and highlights how different thermostats participating in demand response event cycle differently based on their deadband settings. The findings offer valuable insights into how selecting the right thermostat or configuring smart thermostat with appropriate deadband settings can be leveraged to enhance demand response capabilities, shift loads effectively and improve operational flexibility in HVAC systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Evaluating Thermostats' Deadbands Using HVAC Hardware-In-the-Loop Experiment for Advanced Control Strategies: Preprint

Smart thermostats have gained significant popularity due to their potential for optimizing energy consumption and enhanced user control while ensuring occupants' comfort. The deadband, also referred to as temperature differential, is defined as the temperature difference between the desired setpoint and upper threshold or lower threshold for the HVAC equipment to turn on. It is a key factor influencing energy efficiency and user satisfaction. This paper presents a comparative analysis of the deadbands of five different smart thermostats, tested with a heat pump, aiming to identify variations in their deadband settings and implications for energy management. The experimental study was conducted using a HVAC hardware-in-theloop (HIL) system that integrates smart thermostats with physical HVAC equipment in a simulated house environment. The study explores the trade-offs between energy efficiency and occupant comfort and highlights how different thermostats participating in demand response event cycle differently based on their deadband settings. The findings offer valuable insights into how selecting the right thermostat or configuring smart thermostat with appropriate deadband settings can be leveraged to enhance demand response capabilities, shift loads effectively and improve operational flexibility in HVAC systems.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Multiscale modeling and cinematic visualization of photosynthetic energy conversion processes from electronic to cell scales

We report conversion of sunlight into chemical energy, namely photosynthesis, is the primary energy source of life on Earth. A visualization depicting this process, based on multiscale computational models from electronic to cell scales, is presented in the form of an excerpt from the fulldome show Birth of Planet Earth. This accessible visual narrative shows a lay audience, including children, how the energy of sunlight is captured, converted, and stored through a chain of proteins to power living cells. The visualization is the result of a multi-year collaboration among biophysicists, visualization scientists, and artists, which, in turn, is based on a decade-long experimental-computational collaboration on structural and functional modeling that produced an atomic detail description of a bacterial bioenergetic organelle, the chromatophore. Software advancements necessitated by this project have led to significant performance and feature advances, including hardware-accelerated cinematic ray tracing and instanced visualizations for efficient cell-scale modeling. The energy conversion steps depicted feature an integration of function from electronic to cell levels, spanning nearly 12 orders of magnitude in time scales. This atomic detail description uniquely enables a modern retelling of one of humanity’s earliest stories—the interplay between light and life.

97 MATHEMATICS AND COMPUTING↗

Perovskite Catalysts for Pure-Water-Fed Anion-Exchange-Membrane Electrolyzer Anodes: Co-design of Electrically Conductive Nanoparticle Cores and Active Surfaces

Anion-exchange-membrane water electrolyzers (AEMWEs) are a possible low-capital-expense, efficient, and scalable hydrogen-production technology with inexpensive hardware, earth-abundant catalysts, and pure water. However, pure-water-fed AEMWEs remain at an early stage of development and suffer from inferior performance compared with proton-exchange-membrane water electrolyzers (PEMWEs). One challenge is to develop effective non-platinum-group-metal (non-PGM) anode catalysts and electrodes in pure-water-fed AEMWEs. We show how LaNiO3-based perovskite oxides can be tuned by cosubstitution on both A- and B-sites to simultaneously maintain high metallic electrical conductivity along with a degree of surface reconstruction to expose a stable Co-based active catalyst. The optimized perovskite, Sr0.1La0.9Co0.5Ni0.5O3, yielded pure-water AEMWEs operating at 1.97 V at 2.0 A cm-2 at 70 °C with a pure-water feed, thus illustrating the utility of the catalyst design principles.

Zhai, Tingting↗

Evaluation of Portable Acceleration Solutions for LArTPC Simulation Using Wire-Cell Toolkit

The Liquid Argon Time Projection Chamber (LArTPC) technology plays an essential role in many current and future neutrino experiments. Accurate and fast simulation is critical to developing efficient analysis algorithms and precise physics model projections. The speed of simulation becomes more important as Deep Learning algorithms are getting more widely used in LArTPC analysis and their training requires a large simulated dataset. Heterogeneous computing is an efficient way to delegate computationally intensive tasks to specialized hardware. However, as the landscape of compute accelerators quickly evolves, it becomes increasingly difficult to manually adapt the code to the latest hardware or software environments. A solution which is portable to multiple hardware architectures without substantially compromising performance would thus be very beneficial, especially for long-term projects such as the LArTPC simulations. In search of a portable, scalable and maintainable software solution for LArTPC simulations, we have started to explore high-level portable programming frameworks that support several hardware backends. In this paper, we present our experience porting the LArTPC simulation code in the Wire-Cell Toolkit to NVIDIA GPUs, first with the CUDA programming model and then with a portable library called Kokkos. Preliminary performance results on NVIDIA V100 GPUs and multi-core CPUs are presented, followed by a discussion of the factors affiecting the performance and plans for future improvements.

Yu, Haiwang↗

Quantum materials for energy-efficient neuromorphic computing: Opportunities and challenges

Neuromorphic computing approaches become increasingly important as we address future needs for efficiently processing massive amounts of data. The unique attributes of quantum materials can help address these needs by enabling new energy-efficient device concepts that implement neuromorphic ideas at the hardware level. In particular, strong correlations give rise to highly non-linear responses, such as conductive phase transitions that can be harnessed for short- and long-term plasticity. Similarly, magnetization dynamics are strongly non-linear and can be utilized for data classification. This Perspective discusses select examples of these approaches and provides an outlook on the current opportunities and challenges for assembling quantum-material-based devices for neuromorphic functionalities into larger emergent complex network systems.

36 MATERIALS SCIENCE↗

Simulating non-native cubic interactions on noisy quantum machines

As a milestone for general-purpose computing machines, we demonstrate that quantum processors can be programed to efficiently simulate dynamics that are not native to the hardware. Moreover, on noisy devices without error correction, we show that simulation results are significantly improved when the quantum program is compiled using modular gates instead of a restricted set of standard gates. We demonstrate the general methodology by solving a cubic interaction problem, which appears in nonlinear optics, gauge theories, as well as plasma and fluid dynamics. To encode the non-native Hamiltonian evolution, we decompose the Hilbert space into a direct sum of invariant subspaces in which the nonlinear problem is mapped to a finite-dimensional Hamiltonian simulation problem. Furthermore, in a three-states example, the resultant unitary evolution is realized by a product of approximately 20 standard gates, using which approximately ten simulation steps can be carried out on state-of-the-art quantum hardware before results are corrupted by decoherence. In comparison, the simulation depth is improved by more than an order of magnitude when the unitary evolution is realized as a single cubic gate, which is compiled directly using optimal control. Alternatively, parametric gates may also be compiled by interpolating control pulses. Modular gates thus obtained provide high-fidelity building blocks for quantum Hamiltonian simulations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Emerging Threats and Technology Investigation: Industrial Internet of Things - Risk and Mitigation for Nuclear Infrastructure

Industries supporting the global nuclear infrastructure striving for cost savings, expansions in efficiency, and convenience are likely to adopt components (e.g., hardware, software) that comprise the Internet of Things (IoT) and Industrial Internet of Things (IIoT). These devices offer potential improvements along with security challenges. Modern conveniences achieved through application of technology have propagated through society in the form of interconnected devices, from doorbells to microwave ovens, commonly referred to as IoT. IoT devices are often Internet-connected devices that are designed to send data back to a cloud-based server, where a smart phone application then presents device status and control options. Home-based IoT applications carry a different set of risks when compared to a business or security environment, where there is also a history of convenience and interconnection. Industrial settings have long relied on specifically designed Supervisory Control and Data Acquisition (SCADA) systems for process control where IIoT devices are intended to inform business decisions and augment traditional processes. A recent National Institute of Standards and Technology (NIST) report provides a distinction between process control and IIoT in that traditional process control is not replaced by IIoT, but rather IIoT devices are intended to enhance industrial processes through additional monitoring of various sensors and application of data analytics models using artificial intelligence (AI) and machine learning (ML) (Fagan, Marron, et al. 2021) (Ross, et al. 2021).

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

PDV Methods and Analysis for Surveillance of Explosive Components

The Weapons Evaluation Test Laboratory (WETL) at Sandia is collaborating with Lawrence Livermore National Laboratory (LLNL) to enhance explosive surveillance testing by integrating Photon Doppler Velocimetry (PDV) data. This project has streamlined the testing environment, reducing hardware costs and training needs while improving data collection efficiency and usability for lab technicians.

Kress, Matthew Kip [Sandia National Laboratories (↗