Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Simulation of adiabatic quantum computing for molecular ground states

Quantum computation promises to provide substantial speedups in many practical applications with a particularly exciting one being the simulation of quantum many-body systems. Adiabatic state preparation (ASP) is one way that quantum computers could recreate and simulate the ground state of a physical system. In this paper, we explore a novel approach for classically simulating the time dynamics of ASP with high accuracy and with only modest computational resources via an adaptive sampling configuration interaction scheme for truncating the Hilbert space to only the most important determinants. We verify that this truncation introduces negligible error and use this new approach to simulate ASP for sets of small molecular systems and Hubbard models. Furthermore, we examine two approaches to speeding up ASP when performed on quantum hardware: (i) using the complete active space configuration interaction (CASCI) wave function instead of the Hartree–Fock initial state and (ii) a nonlinear interpolation between the initial and target Hamiltonians. We find that starting with a CASCI wave function with a limited active space yields substantial speedups for many of the systems examined, while nonlinear interpolation does not. In additional, we observe interesting trends in the minimum gap location (based on the initial state) as well as how state preparation time can depend on certain molecular properties, such as the number of valence electrons. Importantly, we find that the required state preparation times do not show an immediate exponential wall that would preclude an efficient run of ASP on actual hardware.

Kremenetski, Vladimir↗

DISARM: Target Electronic Device Informed Mitigation of Software Runtime Side-Channel Vulnerabilities

Program runtime/timing attacks exploit variations in a program’s execution times to extract sensitive information from the program (e.g. encryption keys, sensitive variable data, intellectual property). State-of-the-art solutions to runtime side-channel attacks attempt to balance the execution time of the sensitive code for different control flow paths to eliminate the timing leakage. However, during the mitigation process, most techniques do not consider the underlying hardware/device on which the target program is supposed to run on. This can lead to over-fixing (unnecessary extra operations), under-fixing (not solving the imbalance properly), and even failures. Here, we propose DISARM, a joint hardware-software methodology (unlike any existing solution) for mitigating runtime side-channel vulnerabilities that utilizes timing values from real embedded devices to generate targeted software fixes. We implement DISARM to support C/C++/Java source codes and validate it across 22 standard benchmarks. DISARM outperforms state-of-the-art solutions such as PENDULUM and DifFuzzaR in terms of execution time overhead, code size overhead, and correctness on five different embedded/edge devices.

Timing/runtime side-channel↗

Alfalfa Virtual Building Service: Software Engineering Best Practices Applied to Runtime Interaction with Building Energy Models

Buildings are active participants in increasingly complex energy systems. Building Energy Modeling (BEM) has a key role to play in planning and de-risking an equitable energy transition, with BEM-backed "virtual buildings" critical path for diverse applications that include workforce training tools, Hardware-in-the-Loop (HIL) experimentation to study equipment performance under a range of conditions, Control-Hardware-in-the-Loop (CHIL) experimentation to de-risk commercial control implementations at equipment through grid orchestration levels, and integration of dynamic load profiles into grid modeling tools for energy system experimentation at the urban scale. Modeling requirements vary across these applications, but many software engineering tasks do not. The Alfalfa Virtual Building Service (AVBS, see https://github.com/NREL/alfalfa/wiki) is an open-source web service that solves these common tasks robustly in one place, providing a foundational platform for power users to bootstrap their own applications. AVBS abstracts the specifics of runtime interaction with OpenStudio, Modelica, and Spawn of EnergyPlus models behind a unified REST API. Additionally, AVBS provides resources for cloud deployment and scaling to 100s of parallel simulations, a growing library of modular Operational Technology (OT) integrations for emulation of real-world interfaces, and scripts to automate the population of communities of virtual buildings from URBANopt, ResStock and ComStock.

building automation↗

Balance of Plant Modeling and Real-Time Hardware-in-the-Loop Integration with the Microreactor Automated Control System

The advent of novel microreactor technology has driven a focused effort to explore safety and efficiency improvements that can be achieved through the use of automated system control. Development of control strategies, especially for initial demonstration, requires an adequate surrogate environment to safely research failure modes and control integration with realistic hardware delay. However, efficiency gains from control strategies are improved when the scope of controller action is expanded to include system-level dynamics such as downstream heat extraction and mass flow. For this reason, a balance-of-plant (BOP) model of a representative microreactor system has been developed using the TRANsient Simulation Framework of Reconfigurable Models library in Modelica. This model captures a reactor and primary NaK coolant loop that represent corresponding system components of the Microreactor Applications Research Validation and EvaLuation (MARVEL) design as well as a secondary coolant loop and heat extraction representative of the Microreactor Agile Non-Nuclear Experimental Test Bed (MAGNET). This model configuration allows for hardware-in-the-loop (HIL) integration with microreactor automated control system (MACS) hardware in real time through a Python-based gRPC client. Real-time simulation of model performance with emulated hardware and communication delay suggests that under independent proportional-integral-derivative control of BOP model drum dynamics and downstream heat extraction, stable power load following is achievable. A slight delay in load following, filtering of high-frequency dynamics, and localized temperature fluctation suggest room for improvement through the development of higher-level control strategies. The simulated coupling of the MAGNET facility lays the groundwork for future digital twin analysis with a coupled MACS-MAGNET HIL demonstration.

McConnell, Jono [ORNL] (ORCID:0000000238984741)↗

Resilient Operation of Networked Community Microgrids with High Solar Penetration

This project, funded by the US Department of Energy’s Solar Energy Technologies Office (SETO), focused on the operation of microgrids as a coordinated network. The primary objective, which was successfully achieved, was to develop both control strategies and hardware solutions to support the resilient and efficient operation of networked microgrids with high solar penetration. The work was structured around the following four main tasks: • Development of distributed and scalable optimization algorithms for AC-coupled networked microgrids. • Design and implementation of a novel DC interconnection hardware to enable precise power exchange between microgrids. • Laboratory operational validation of the developed technologies using 480 V testbeds and commercially available hardware. • Field operational validation of the complete solution in Adjuntas, Puerto Rico, interconnecting two kW-scale, split-phase microgrids of Casa Pueblo’s microgrids. This project addressed multiple technical challenges across the domains of optimization, control, hardware interconnection, and protection. One of its key contributions was delivering tangible, real-world solutions for networking microgrids. In contrast to purely theoretical or simulation-based work, this project included full-scale hardware operational validation both in the lab and in the field. The work conducted as part of this project—in collaboration with the University of Puerto Rico; the University of Tennessee, Knoxville; the University of Central Florida; and Casa Pueblo—has advanced the state of the art in networked microgrids. Key contributions include the development of distributed control strategies, practical solutions for real-world implementation challenges, and the introduction of a novel DC interlink approach for microgrid interconnection. The project featured both laboratory and field validation using commercial off-the-shelf components. The field deployment successfully validated that a group of microgrids can operate in a coordinated manner, enabling precise power flow between systems and mutual support during extreme events. This project resulted in 15 journal publications and 15 conference papers; 5 graduate students and 15 undergraduate students were supported. The codes of distributed optimization and forecasting were made open-source through OSTI.gov for distributed optimization and forecasting. All the publications are available in the ORNL-hosted project landing page. The DC interlink with state-of-charge balancing control was operationally validated in Adjuntas by interconnecting two real-world, 240 V split-phase microgrids. To the best knowledge of the team, this represents the first operational validation of AC microgrids interconnected via DC-interlinks. As a culmination of this project, a follow-on grant was awarded to support the technology transfer of the distributed optimization framework to a commercial microgrid controller, Stellar Edge, developed by the California-based company New Sun Road.

14 SOLAR ENERGY↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗

NeuroCoreX: Brain-Inspired Computing from Code to Circuit

NeuroCoreX is an open-source codebase that enables the implementation of brain-inspired, energy-efficient neuromorphic computing models on FPGA hardware. Designed to support real-time learning, all-to-all neural connectivity, and flexible network architectures, NeuroCoreX offers a hands-on, accessible platform for exploring biologically inspired models of neural computation. It empowers researchers, students, and developers to implement and experiment with adaptive systems—bringing the power of neuromorphic computing to a broader community through a low-cost, scalable, and reconfigurable framework.

Gautam, Ashish [Oak Ridge National Laboratory (ORN↗

Quantum Information for Fusion Energy Sciences (Final Technical Report)

The simulation of plasma dynamics is a critical area of Fusion Energy Sciences (FES) due to it’s usefulness in predicting, controlling, and confining plasmas in the context of potential fusion reactors. The simulation of plasmas is a computationally difficult problem in both classical and quantum physics, motivating investigation into the potential of quantum computers to simulate these systems. This project took several concrete steps towards this goal by developing tools for improving the control, characterization, and calibration of quantum gates on a superconducting quantum computer, developing error suppression and mitigation tools to reduce errors on the quantum computer, and utilizing these advancements to simulate reduced models of plasma dynamics on the quantum computer. In order to efficiently simulate plasma physics, an optimal control method which synthesizes, directly at the pulse level, any quantum gate on qubit and qutrit systems was developed. Using four superconducting transmon quantum processors at Rigetti and LLNL, it was demonstrated that any arbitrary quantum gate on qubits and qutrits could be implemented with high fidelity, leading to a significantly reduced length of a gate sequence. A problem of interest in FES is the nonlinear optical process of laser pulse compression within a plasma. Since quantum physics is linear, simulating nonlinear operations is not naturally feasible on a quantum computer, however it is possible to simulated a quantized version of the nonlinear process. A quantization approach to convert nonlinear wave-wave interaction problems to Hamiltonian simulation problems was developed and demonstrated using two qubits on a Rigetti device. In this experiment, a number of error suppression and mitigation techniques were investigated to determine how best to utilize the finite quantum resources. This study provides an example of how plasma problems may be solved on near-term, noisy quantum computing platforms and identified a promising set of techniques. Building on the insights of these experiments, the investigation turned to linear electron-plasma wave physics. A connection was identified between a local one-dimensional lattice spin model and linear wave phenomena, allowing a plasma physics problem to be efficiently mapped to the quantum computer. In this framework, reflection and transmission of plasma waves at a sharp boundary was studied, as well as the propagation of waves through an inhomogeneous plasma medium. In addition to the suite of error suppression and mitigation techniques developed, this experiment introduced the use of a digital-analog gate scheme designed to efficiently simulate the plasma Hamiltonian. With hardware available at the conclusion of the project, simulation at the scale of 9 qubits and 15 timesteps (60 entangling layers) was achieved.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Real-time Anomaly Detection for Liquid Argon Time Projection Chambers

We present a real-time anomaly detection framework for liquid argon time projection chambers (LArTPCs), targeting applications in particle physics experiments such as the Short Baseline Near Detector (SBND) or the future Deep Underground Neutrino Experiment (DUNE). These experiments employ detectors that generate and stream high-resolution but sparse images of neutrino and other particle interactions. Our approach utilizes anomaly detection with autoencoders, compressed through knowledge distillation (KD), to enable the detection of anomalous signals in the data through efficient inference on resource-constrained hardware. The framework is targeted for deployment on computing platforms equipped with field-programmable gate arrays (FPGAs), GPUs, or CPUs, allowing low-latency selection of relevant activity directly from the raw detector data stream. We demonstrate that our approach is suitable for the detection and localization of anomalously "high-multiplicity" activity, and outline promising applications for LArTPC online data filtering and triggering.

FOS: Physical sciences↗

Performance Potential of Mixed Data Management Modes for Heterogeneous Memory Systems

Many high-performance systems now include different types of memory devices within the same compute platform to meet strict performance and cost constraints. Such heterogeneous memory systems often include an upper-level tier with better performance, but limited capacity, and lower-level tiers with higher capacity, but less bandwidth and longer latencies for reads and writes. To utilize the different memory layers efficiently, current systems rely on hardware-directed, memory -side caching or they provide facilities in the operating system (OS) that allow applications to make their own data-tier assignments. Since these data management options each come with their own set of trade-offs, many systems also include mixed data management configurations that allow applications to employ hardware- and software-directed management simultaneously, but for different portions of their address space. Despite the opportunity to address limitations of stand-alone data management options, such mixed management modes are under-utilized in practice, and have not been evaluated in prior studies of complex memory hardware. In this work, we develop custom program profiling, configurations, and policies to study the potential of mixed data management modes to outperform hardware- or software-based management schemes alone. Our experiments, conducted on an Intel ® Knights Landing platform with high-bandwidth memory, demonstrate that the mixed data management mode achieves the same or better performance than the best stand-alone option for five memory intensive benchmark applications (run separately and in isolation), resulting in an average speedup compared to the best stand-alone policy of over 10 %, on average.

Effler, Chad↗

Energy Program Innovation Cluster for Equity and Health in Grid-Interactive Efficient Buildings (Final Technical Report)

The EPIC Buildings program was created to target Central New York’s (CNY) Energy Program Innovation Cluster, focused on developing energy hardware innovations for next-generation grid-interactive efficient buildings (GEB). The goals of the program were to establish and grow a sustainable regional innovation cluster for GEB in CNY and support the ambitious transition to achieve net-zero emissions by 2050.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Hardware-in-the-loop testing of a hydraulic wave energy power take-off system

This report describes testing conducted related to the development of a “hydrostatic power takeoff” (HPTO) system for a wave energy converter. Tests were conducted with an experimental electric motor rig to provide preliminary results and de-risk future testing. Efficiency mapping tests were conducted as well as hardware-in-the-loop (HIL) testing. The results of the efficiency mapping tests provide good insight into how to systematically perform efficiency mapping tests. The HIL testing indicates good overall performance of the system and provides a stepping stone towards more complete system tests in the future.

16 TIDAL AND WAVE POWER↗

Low Size, Weight, and Power Neuromorphic Computing to Improve Combustion Engine Efficiency

Neuromorphic computing offers one path forward for AI at the edge. However, accessing and effectively utilizing a neuromorphic hardware platform is non-trivial. In this work, we present a complete pipeline for neuromorphic computing at the edge, including a small, inexpensive, low-power, FPGA-based neuromorphic hardware platform, a training algorithm for designing spiking neural networks for neuromorphic hardware, and a software framework for connecting those components. We demonstrate this pipeline on a real-world application, engine control for a spark-ignition internal combustion engine. We illustrate how we connect engine simulations with neuromorphic hardware simulations and training software to produce hardware-compatible spiking neural networks that perform engine control to improve fuel efficiency. We present initial results on the performance of these spiking neural networks and illustrate that they outperform open-loop engine control. We also give size, weight, and power estimates for a deployed solution of this type.

Schuman, Catherine↗

Breaking the mold: Overcoming the time constraints of molecular dynamics on general-purpose hardware

The evolution of molecular dynamics (MD) simulations has been intimately linked to that of computing hardware. For decades following the creation of MD, simulations have improved with computing power along the three principal dimensions of accuracy, atom count (spatial scale), and duration (temporal scale). Since the mid-2000s, computer platforms have, however, failed to provide strong scaling for MD, as scale-out central processing unit (CPU) and graphics processing unit (GPU) platforms that provide substantial increases to spatial scale do not lead to proportional increases in temporal scale. Important scientific problems therefore remained inaccessible to direct simulation, prompting the development of increasingly sophisticated algorithms that present significant complexity, accuracy, and efficiency challenges. While bespoke MD-only hardware solutions have provided a path to longer timescales for specific physical systems, their impact on the broader community has been mitigated by their limited adaptability to new methods and potentials. In this work, we show that a novel computing architecture, the Cerebras wafer scale engine, completely alters the scaling path by delivering unprecedentedly high simulation rates up to 1.144 M steps/s for 200 000 atoms whose interactions are described by an embedded atom method potential. This enables direct simulations of the evolution of materials using general-purpose programmable hardware over millisecond timescales, dramatically increasing the space of direct MD simulations that can be carried out. In this paper, we provide an overview of advances in MD over the last 60 years and present our recent result in the context of historical MD performance trends.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

On-chip parallel processing of quantum frequency comb

Abstract The frequency degree of freedom of optical photons has been recently explored for efficient quantum information processing. Significant reduction in hardware resources and enhancement of quantum functions can be expected by leveraging the large number of frequency modes. Here, we develope an integrated photonic platform for the generation and parallel processing of quantum frequency combs (QFCs). Cavity-enhanced parametric down-conversion with Sagnac configuration is implemented to generate QFCs with identical spectral distributions. On-chip quantum interference of different frequency modes is simultaneously realized with the same photonic circuit. High interference visibility is maintained across all frequency modes with the identical circuit setting. This enables the on-chip reconfiguration of QFCs. By deterministically separating QFCs without spectral filtering, we further demonstrate high-dimensional Hong-Ou-Mandel effect. Our work provides the critical step for the efficient implementation of quantum information processing with integrated photonics using the frequency degree of freedom.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

In-situ TEM EELS analysis of memristive thin films for neuromorphic computing

Neuromorphic computing stands as a promising frontier for advancing AI algorithms and applications like ChatGBT, offering significant energy efficiency gains. This paper delves into the hardware design intricacies of memristive thin films and their elementary switching mechanisms, including anion migration, electron migration, and phase transitions. Through comprehensive analysis of electron energy loss spectroscopy (EELS) data via in-situ transmission electron microscopy (TEM), we will deduce the primary memristive switching mechanisms vital for optimizing thin film fabrication parameters and achieving desired film thickness, conductivity, and memory retention. A single crystal ptype Si substrate was used with TiN as the bottom metal electrode, TiO x as the insulating dielectric layer, and Pt as the top metal electrode. In-situ TEM was able to tell us the thin film didn’t behave like a filamentary or phase transition material. EELS data deduced that electron trapping/detrapping was one of the primary switching mechanisms. By shedding light on these elementary mechanisms, our study aims to catalyze the development of more 2 efficient and effective neuromorphic computing systems to be deployed into mainstream technologies.

97 MATHEMATICS AND COMPUTING↗

NeuroCoreX: An Open-Source FPGA-Based Spiking Neural Network Emulator with On-Chip Learning

Spiking Neural Networks (SNNs) are computational models inspired by the event-driven communication and connectivity patterns of biological neural circuits. They enable high energy efficiency and natural support for diverse architectures ranging from layered networks to small-world and graphstructured topologies. In this work, we introduce NeuroCoreX, an open-source, FPGA-based spiking neural network emulator that provides real-time, on-chip learning and flexible network organization. NeuroCoreX supports both feedforward sensory inputs streamed directly from sensors or PCs via UART and recurrent on-chip connectivity, enabling simultaneous processing and learning from external stimuli and internal network dynamics-capabilities rarely available in existing FPGA SNN platforms. The system implements a Leaky Integrate-and-Fire (LIF) neuron model with current-based synapses and supports pair-based STDP learning on both feedforward and recurrent synapses. A lightweight Python interface enables interactive configuration, live monitoring, weight read-back, and experiment control. Importantly, NeuroCoreX is tightly integrated with the SuperNeuroMAT simulator, allowing SNN models to be transferred seamlessly from software to hardware for hardware-in-the-loop development. By combining real-time plasticity, flexible connectivity, and an open-source VHDL implementation, NeuroCoreX provides an extensible and accessible platform for neuromorphic research, algorithm-hardware co-design, and energy-efficient edge intelligence.

Gautam, Ashish [ORNL]↗

TrioSim: A Lightweight Simulator for Large-Scale DNN Workloads on Multi-GPU Systems

Deep Neural Networks (DNNs) have become increasingly capable of performing tasks ranging from image recognition to content generation. The training and inference of DNNs heavily rely on GPUs, as GPUs' massively parallel architecture delivers extremely high computing capability. With the growing complexity of DNNs and the size of training datasets, training DNNs with a large number of GPUs is becoming a prevalent strategy. Researchers have been exploring how to design software and hardware systems for GPU farms to achieve the best utilization, efficiency, and DNN accuracy during training or inference. However, when designing and deploying such systems, designers usually rely on testing on physical hardware platforms equipped with many GPUs, incurring high costs that are almost prohibitive for system designers to test different configurations and designs, even for highly resourceful companies. While an alternative solution is to test on GPU simulators, they are often too slow for these l

Li, Ying [William & Mary, Williamsburg, VA, USA] (↗