Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware efficiency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

DISH-STARS™ Commercialization (Abstract)

The goal of this project is to aggressively support the near-term commercialization of a new technology platform – based on the integration of solar concentrators and micro- and meso-channel process technology (MMPT) – that was evaluated and identified as a strong candidate for near-term commercialization at EERE’s inaugural Lab-Corps program during early FY2016. Known as STARS, for Solar Thermochemical Advanced Reactor System, or Dish-STARS™ when paired with parabolic dish concentrators, STARS is a promising energy-related technology developed at the Pacific Northwest National Laboratory (PNNL) that efficiently converts solar energy into chemical energy. Combined with economies through hardware mass production, the efficiency of Dish-STARS™ provides a near-term opportunity for the production of renewable electricity, fuels and chemicals. This proposed CRADA project supports the commercialization of Dish-STARS™ in these important ways: The project will support the cooperative development of Dish-STARS™ by the DOE national laboratory and private partners, including the startup company, STARS Corporation, that is being established by the PNNL Lab-Corps team that evaluated STARS on behalf of EERE. The project will provide important transition funding at the time that the previous DOE SunShot project, which has supported Dish-STARS™ development from Technology Readiness Level 3 (TRL 3) to TRL 6, is scheduled to end.

14 SOLAR ENERGY↗

Efficient exascale discretizations: High-order finite element methods

Efficient exploitation of exascale architectures requires rethinking of the numerical algorithms used in many large-scale applications. These architectures favor algorithms that expose ultra fine-grain parallelism and maximize the ratio of floating point operations to energy intensive data movement. One of the few viable approaches to achieve high efficiency in the area of PDE discretizations on unstructured grids is to use matrix-free/partially assembled high-order finite element methods, since these methods can increase the accuracy and/or lower the computational time due to reduced data motion. In this paper we provide an overview of the research and development activities in the Center for Efficient Exascale Discretizations (CEED), a co-design center in the Exascale Computing Project that is focused on the development of next-generation discretization software and algorithms to enable a wide range of finite element applications to run efficiently on future hardware. CEED is a research partnership involving more than 30 computational scientists from two US national labs and five universities, including members of the Nek5000, MFEM, MAGMA and PETSc projects. We discuss the CEED co-design activities based on targeted benchmarks, miniapps and discretization libraries and our work on performance optimizations for large-scale GPU architectures. We also provide a broad overview of research and development activities in areas such as unstructured adaptive mesh refinement algorithms, matrix-free linear solvers, high-order data visualization, and list examples of collaborations with several ECP and external applications.

97 MATHEMATICS AND COMPUTING↗

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar↗

Computational Tools for Additive Manufacture of Tailored Microstructure and Properties

Additive manufacturing has the potential to revolutionize industrial hardware and unlock efficiency gains through the fabrication of geometries and architectures not possible by conventional processing. Currently most additive builds use a single set of process parameters (e.g. laser power and scan speed) which results in a part with a homogenous microstructure that provides a singular performance level. To move beyond this state, Raytheon Technologies Research Center worked to create a set of computational tools to track material evolution through each step of the additive process. Computational fluid dynamics and phase field models for microstructure evolution as a function of processing parameters, and crystal plasticity models fully coupling microstructure and mechanical properties for performance predictions were leveraged to establish a connection between additive parameters and the final microstructure. This framework was utilized to tailor spatially-varying mechanical properties in a part by appropriately controlling the microstructure evolution during the additive process. Specifically, a turbine blade was 3D printed from nickel superalloy IN718 using laser powder bed fusion with coarse grains in the airfoil section which experiences the highest temperatures and is creep limited while finer grains were printed in the root of the blade which experience higher stresses but at lower temperatures and is therefore fatigue limited. The benefit of being able to intentionally insert coarse grains in the high temperature region of the blade was showcased with a microstructure sensitive creep model that indicates longer creep life for coarser grains.

20 FOSSIL-FUELED POWER PLANTS↗

Optimization of WAAM Process to Produce AUSC Components with Increased Service Life

Additive manufacturing has the potential to revolutionize industrial hardware and unlock efficiency gains through the fabrication of geometries and architectures not possible by conventional processing. Wire Arc Additive manufacturing (WAAM) process is a class of directed energy deposition process enabling higher build rate and allowing custom wire feedstock allowing spatial variation of microstructure. The larger spot size and lower speed creates a larger melt pool, reducing residual stress and often time creating directional/columnar microstructure for Ni-superalloy. The use of flexible platform provides freedom in deposition strategy, which can accommodate complex substrates, including feature addition onto existing structures, non-flat layers, and repair methods. However, the certification of the final component needs to match the strength requirement. Hence, the quality of the component produced is stringently monitored to avoid buildup of residual stress, cracks, porosity and to reduce detrimental segregated phases commonly observed during alloy solidification. Simulation plays a huge role in predicting the melt pool dimension and can be used to optimize the process parameter. Similar development is also required to perform physics-based modeling of microstructural development in WAAM that can predict the microstructural features during solidification and can be used to optimize the process more effectively. To move toward this goal, Raytheon Technologies Research Center together with Siemens worked to create a set of computational tools to control the process parameters, enable on-line measurements and acquisition with feedback to the optimized process parameters, and eventually track material evolution through each step of the additive process. Computational fluid dynamics is used for accurate prediction and calibration of the thermal field during WAAM process. and phase field models for microstructure evolution as a function of processing parameters to establish a connection between additive parameters and the final microstructure. Here we report cellular automata (CA) model development to predict the dendritic microstructure evolution with surface and bulk nuclei for single track and multiple layers. The CA model was developed to account for secondary element addition and predict segregation, local melting, and latent heat release as well as prediction of Euler angles from orientation information and validated against experiments. This framework was utilized to tailor spatially-varying composition in a part by appropriately controlling the microstructure evolution during the additive process. Functionally graded Haynes 282 alloy with high Cr content at the surface was tested for oxidation and mechanical properties. An advanced physics-based reaction-diffusion model predicting the simultaneous creation of chromium oxide and alumina is developed and validated to extend life expectancy of the WAAM manufactured high temperature part. A machine-learning data-driven framework establishing the process-structure relationship from a dataset of real microstructure images and corresponding process history data has been developed and implemented in the NX Siemens design system. The digital twin configuration along with the tool path generation enabled prediction of WAAM component buildup time and the techno-economic analysis provided a favorable option for all 4 cases with 15-40% cost reduction.

33 ADVANCED PROPULSION SYSTEMS↗

A Scalable Hardware-and-Human-in-the-Loop Grid-interactive Efficient Building Equipment Performance Dataset

This project developed a publicly available, high-fidelity dataset about the interactions among humans, homes, and heat pumps supporting grid interactive efficient buildings to balance demand on the grid with comfort for occupants. Laboratory measurements and simulations of the hardware capture the second-scale electric power dynamics of heat pumps providing grid services like load shifting and load shedding. Field measurements, behavior tracking, and qualitative surveys of people in their homes over multiple years—including experimentally adjusting the heating and cooling system to provide grid services—to capture the reciprocal effect of human behavior on grid services, and grid services on human comfort. Taken together, these data capture the complete Hardware and Human in the loop system for residential heat pumps, reducing large uncertainties in simulation for design, and models for control of heat pumps, and grid-interactive buildings.

24 POWER TRANSMISSION AND DISTRIBUTION↗

On-Detector Machine Learning for Beam-Induced Background Rejection at a 10 TeV Muon Collider

A 10 TeV Muon Collider is a compelling candidate for a future energy-frontier facility, offering unprecedented opportunities to explore the fundamental laws of particle physics. Muon decays in the collider ring produce intense beam-induced background (BIB) that can overwhelm detector occupancy and exceed readout bandwidth constraints. We investigate the potential of on-detector Machine Learning for BIB rejection in the vertex detector, exploiting pixel cluster shapes to distinguish background from collision products. We study three classes of lightweight neural-network architectures, and evaluate their implementation feasibility using high-level synthesis. Selected architectures achieve 88 to 90% data reduction at 99% signal efficiency, while requiring hardware resources compatible with potential ASIC implementation. These results demonstrate the potential of performing substantial BIB rejection directly in the pixel readout, providing a strategy for meeting the tracker readout requirements at a future Muon Collider.

Abadjiev, Daniel [Chicago U.]↗

Efficient Simulation of Open Quantum Systems on NISQ Trapped‐Ion Hardware

Abstract Simulating open quantum systems, which interact with external environments, presents significant challenges on noisy intermediate‐scale quantum (NISQ) devices due to limited qubit resources and noise. In this study, an efficient framework is proposed for simulating open quantum systems on NISQ hardware by leveraging a time‐perturbative Kraus operator representation of the system's dynamics. This approach avoids the computationally expensive Trotterization method and exploits the Lindblad master equation to represent time evolution in a compact form, particularly for systems satisfying specific commutation relations. The efficiency of this method is demonstrated by simulating quantum channels, such as the continuous‐time Pauli channel and damped harmonic oscillators, on NISQ trapped‐ion hardware, including IonQ Harmony and Quantinuum H1‐1. Additionally, hardware‐agnostic error mitigation techniques are introduced, including Pauli channel fitting and quantum depolarizing channel inversion, to enhance the fidelity of quantum simulations. These results show strong agreement between the simulations on real quantum hardware and exact solutions, highlighting the potential of Kraus‐based methods for scalable and accurate simulation of open quantum systems on NISQ devices. This framework opens pathways for simulating more complex systems under realistic conditions in the near term.

Burdine, Colin [Department of Electrical and Compu↗

Efficient Reinforcement Learning for Real-Time Hardware-Based Energy System Experiments: Preprint

In the context of urgent climate challenges and the pressing need for rapid technology development, Reinforcement Learning (RL) stands as a compelling data-driven method for controlling real-world physical systems. However, RL implementation often entails time-consuming and computationally intensive data collection and training processes, rendering them inefficient for real-time applications that lack non-real-time models. To address these limitations, real-time emulation techniques have emerged as valuable tools for the lab-scale rapid prototyping of intricate energy systems. While emulated systems offer a bridge between simulation and reality, they too face constraints, hindering comprehensive characterization, testing, and development. In this research, we construct a surrogate model using limited data from simulated systems, enabling an efficient and effective training process for a Double Deep Q-Network (DDQN) agent for future deployment. Our approach is illustrated through a hydropower application, demonstrating the practical impact of our approach on climate-related technology development.

deep Q-learning↗

Computational capacity in hydrodynamic real-time hybrid simulation applied to simulate the dynamic response of floating offshore wind turbines

Real-time hybrid simulation (RTHS) mitigates similitude distortions in model-scale tests of floating offshore wind turbines (FOWTs) by coupling physical experiments with numerical models in real time. The coupling requires faster-than-real-time numerical computations to satisfy temporal similitude with the physical experiment, presenting a bottleneck for using more complex numerical models in RTHS. This paper presents a hydrodynamic-RTHS (hydro-RTHS) framework for FOWTs that simulates the hydrodynamics physically and the aerodynamics numerically with sensor feedback from the physical testing. The framework adapts the three-loop hardware architecture to leverage greater computational resources and mitigate strict temporal requirements, enabling more computationally demanding numerical analyses in hydro-RTHS. The three-loop hardware architecture integrates multiple machines, each dedicated to either numerical analysis or RTHS controls, with a rate-transition algorithm to synchronize the tasks executed across the different machine processors. Virtual and physical tests verified and validated the hydro-RTHS framework, respectively. The ”virtual” tests, which approximates the physical domain numerically, verified the RTHS framework with respect to a numerical full-scale complete FOWT model simulated in the open-source software, OpenFAST. The virtual tests were able to maintain comparable control signals while enabling greater computational resources for the numerical calculations. Real-world physical tests demonstrated that the hydro-RTHS framework computes aerodynamic forces similar to the complete OpenFAST model, validating the hydro-RTHS framework using the three-loop hardware architecture. Findings show that the hydro-RTHS framework with the three-loop hardware architecture is computationally efficient, with reserve capacity to simulate more complex problems due to the customized software, hardware, and rate-transition algorithm.

17 WIND ENERGY↗

Differential Power Processing for Ultra-Efficient Data Storage

Here this paper presents the hardware, software, and power codesign of an ultra-efficient data storage server with differential power processing (DPP). DPP can reduce the power conversion stress, improve the efficiency, and enhance the functionality of modular power electronics systems. The power inputs of a large number of hard disk drives (HDDs) were connected in series and supported by a multiport ac-coupled differential power processing (MAC-DPP) converter through a multiwinding transformer. Methods for controlling the multi-input multi-output power flow in the multiwinding transformer while avoiding core saturation were investigated. A ten-port MAC-DPP prototype with 700-W/in 3 power density was built to support a 450-W HDD storage system with ten series-stacked voltage domains. The prototype was tested on a 50-HDD server testbench, and the overall system loss is below 1 W (99.77% system efficiency). The server was able to maintain high-speed reading and writing operation of all 50 HDDs against the worst hot-swapping scenarios. A variety of hardware/software configurations and many cloud storage techniques were tested on the fully functioning server. Experimental results show that the energy efficiency of large-scale information systems (CPU/GPU clusters, memory banks, HDD arrays, etc.) can be greatly improved by software, hardware, and power codesign.

42 ENGINEERING↗

SPARTAN (Scalable Probabilistic Application Reconfigurable Tensor Autonomous Network)

The technical founder of Ludwig Computing Inc has been competitively selected for support by Cyclotron Road, a U.S. Department of Energy (DOE) Advanced Manufacturing Office (AMO) Lab-Embedded Entrepreneurship Program (LEEP) through an approved merit review process. Ludwig Computing Inc, supported by the U.S. Department of Energy's Advanced Manufacturing Office through the Cyclotron Road program, has investigated the advantages of probabilistic computing for real-world compute-intensive applications. This research adds to the understanding of alternative computing paradigms by exploring a unique hardware-software co-design that integrates quantum computing methods with nature-inspired problem-solving techniques. The project's focus on areas such as combinatorial optimization, graph analytics, and machine learning demonstrates the potential for significant advancements in computational efficiency and performance. By harnessing natural randomness to streamline large circuits into fewer devices, Ludwig's approach enables massive parallelism, potentially offering higher throughput, speed, and energy efficiency compared to conventional hardware solutions. This work benefits the public by paving the way for more efficient computing solutions that could address complex real-world problems while potentially reducing energy consumption in data-intensive industries.

97 MATHEMATICS AND COMPUTING↗

OpenACC offloading of the MFC compressible multiphase flow solver on AMD and NVIDIA GPUs

GPUs are the heart of the latest generations of supercomputers. We efficiently accelerate a compressible multiphase flow solver via OpenACC on NVIDIA and AMD Instinct GPUs. Optimization is accomplished by specifying the directive clauses gang vector and collapse. Further speedups of six and ten times are achieved by packing user-defined types into coalesced multidimensional arrays and manual inlining via metaprogramming. Additional optimizations yield seven-times speedup of array packing and thirty-times speedup of select kernels on Frontier. Weak scaling efficiencies of 97% and 95% are observed when scaling to 50% of Summit and 87% of Frontier. Strong scaling efficiencies of 84% and 81% are observed when increasing the device count by a factor of 8 and 16 on V100 and MI250X hardware. The strong scaling efficiency of AMD’s MI250X increases to 92% when increasing the device count by a factor of 16 when GPU-aware MPI is used for communication.

Wilfong, Benjamin↗

A Hardware and Software Co-design Framework for Energy Efficient Neuromorphic Systems

Neuromorphic systems can be realized by a variety of algorithms and architectures. A common understanding is that spiking neuromorphic designs, which encode information into spatio-temporal spiking events, are both a biologically-accurate and efficient way of processing information. However, representing the information through timing relationships induces sophisticated circuit designs in traditional CMOS-based implementations. In recent years, high-capacity resistive memory (RRAM, aka, memristor) has demonstrated great potential in mimicking synaptic behaviors. Several RRAM-based spiking neuromorphic designs exist, most of which focus on rate coding schemes. These designs simplify circuit implementations of neuron models and explore challenges such as unsatisfactory speed, resolution, and performance. As an alternative, we will explore temporal coding spiking neuromorphic systems that encode information as the relative timing of neuron activations (spikes), which have been proven to be more adaptive and energy-efficient. Developing a neuromorphic system for spiking neural network (SNN) inference and online training, however, faces some major technical challenges: (1) It lacks circuit implementation support for temporal-coding SNN to achieve satisfying power efficiency and accuracy; (2) Although existing research works have investigated memristive synapse and neuron designs for spike-timing-dependent plasticity, the non-ideal conditions in implementation, such as device variations and signal degradation, degrade online learning accuracy of large scale systems; and (3) Non-optimized, inter-layer data traffic in SNNs, leads to unnecessary data communication costs. In this project, we plan to address these challenges by a hardware and software co-design framework that incorporates solutions at the circuit, architecture, and algorithm levels. At the circuit-level, we will elaborate on the in-situ SNN processing element designs for supporting both inference and online training modes. Variation-aware schemes will be studied to improve reliability. At the architecture level, we propose a pipelined, asynchronous architecture to retain the timing resolution of spikes. At the algorithm level, we will investigate an innovative SNN training algorithm for enabling activation sparsification and reducing unnecessary data communication costs. This neuromorphic system will provide an effective solution to real-life energy-constrained applications and significantly contribute to the exploration of next-generation high-performance computing systems under the DOE context.

97 MATHEMATICS AND COMPUTING↗

Grid ancillary services using electrolyzer-Based power-to-Gas systems with increasing renewable penetration

Increasing penetrations of renewable-based generation have led to a decrease in the bulk power system inertia and an increase in intermittency and uncertainty in generation. Energy storage is considered to be an important factor to help manage renewable energy generation at greater penetrations. Hydrogen is a viable long-term storage alternative. This paper analyzes and presents use cases for leveraging electrolyzer-based power-to-gas systems for electric grid support. The paper also discusses some grid services that may favor the use of hydrogen-based storage over other forms such as battery energy storage. Real-time controls are developed, implemented and demonstrated using a power-hardware-in-the-loop(PHIL) setup with a 225-kW proton-exchange-membrane electrolyzer stack. These controls demonstrate frequency and voltage support for the grid for different levels of renewable penetration (0%, 25%, and 50%). A comparison of the results shows the changes in respective frequencies and voltages as seen as different buses as a result of support from the electrolyzers and notes the impact on hydrogen production as a result of grid support. Finally, the paper discusses the practical nuances of implementing the tests with physical hardware, such as inverter/electrolyzer efficiency, as well as the related constraints and opportunities.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exponential concentration in quantum kernel methods

Kernel methods in Quantum Machine Learning (QML) have recently gained significant attention as a potential candidate for achieving a quantum advantage in data analysis. Among other attractive properties, when training a kernel-based model one is guaranteed to find the optimal model’s parameters due to the convexity of the training landscape. However, this is based on the assumption that the quantum kernel can be efficiently obtained from quantum hardware. In this work we study the performance of quantum kernel models from the perspective of the resources needed to accurately estimate kernel values. We show that, under certain conditions, values of quantum kernels over different input data can be exponentially concentrated (in the number of qubits) towards some fixed value. Thus on training with a polynomial number of measurements, one ends up with a trivial model where the predictions on unseen inputs are independent of the input data. We identify four sources that can lead to concentration including expressivity of data embedding, global measurements, entanglement and noise. For each source, an associated concentration bound of quantum kernels is analytically derived. Lastly, we show that when dealing with classical data, training a parametrized data embedding with a kernel alignment method is also susceptible to exponential concentration. Our results are verified through numerical simulations for several QML tasks. Altogether, we provide guidelines indicating that certain features should be avoided to ensure the efficient evaluation of quantum kernels and so the performance of quantum kernel methods.

97 MATHEMATICS AND COMPUTING↗