Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hardware generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Impedance Scan of Inverter-Based Resources and Diesel Generator for Stability Analysis

Impedance-based methods are widely used for power system stability analysis with inverter-based resources (IBRs), e.g., assessing dynamic interactions between the power grid and an IBR, control interactions between multiple IBRs, and the sub-synchronous oscillation and damping phenomenon. Since it is difficult to get a numerical model 100% matching with the hardware IBR, using the hardware inverter directly to obtain its output impedance has become a prominent approach nowadays. Therefore, this article presents the impedance scan using hardware IBRs, and also a hardware diesel generator as it still stays with the grid before the grid completely goes to renewable. The devices under test (DuTs) for the impedance scan includes two 3-..phi.., 480 V, 60 Hz commercial grid-forming IBRs (one of 250 kVA and another of 125 kVA rating) in series with ..delta..-Y transformers, one 3-..phi.., 480 V, 60 Hz commercial grid-following IBR (of 125 kVA rating), and a 3-..phi.., 480 V, 60 Hz commercial diesel generator (of 187.5 kVA rating). Using voltage signals perturbed with sub-, inter-, and higher harmonic components, and measuring the current response, the positive-sequence impedances are computed via an offline-based post-analysis. Moreover, best-fit transfer functions are estimated that closely resemble the measured data points of the positive-sequence impedances. Based on the observations from various outcomes of the hardware experiments, this article also provides some fundamental insights on the equivalent positive-sequence impedance of a combination of multiple hardware components by comparing the estimated and the empirically computed impedances. A comparative insight on the damping capability of the DuTs using the positive-sequence impedances of the hardware is also discussed.

current measurement↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗

Operando microscopy for neuromorphic hardware

Microscopy techniques can uncover the physical properties and dynamic behaviours of materials, driving the discovery of emergent phenomena and guiding the design of next-generation computing hardware. As artificial intelligence becomes pervasive, the demand for high-performance materials to support sustainable information technologies is growing. Here, this Review highlights state-of-the-art imaging from electron and X-ray to optical techniques to probe the dynamics of neuromorphic materials, including operando characterization of devices. We examine design principles for neuromorphic materials, along with obstacles that hinder their development. Emphasis is placed on spatially and temporally resolved approaches that capture state changes including phase transitions, ferroic switching and spin-wave propagation that emulate biological components such as neurons, synapses and their connectivity. We discuss challenges in operando characterization and the integration of artificial intelligence-driven analysis for feedback-guided material discovery. Finally, we outline opportunities for real-time imaging of neuromorphic systems, paving the way towards adaptive, brain-inspired hardware.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Impedance Scan of Inverter-Based Resources and Diesel Generator for Stability Analysis: Preprint

Impedance-based methods are widely used for power system stability analysis with inverter-based resources (IBRs), e.g., assessing dynamic interactions between the power grid and an IBR, control interactions between multiple IBRs, and the sub-synchronous oscillation and damping phenomenon. Since it is difficult to get a numerical model 100% matching with the hardware IBR, using the hardware inverter directly to obtain its output impedance has become a prominent approach nowadays. Therefore, this article presents the impedance scan using hardware IBRs, and also a hardware diesel generator as it still stays with the grid before the grid completely goes to renewable. The devices under test (DuTs) for the impedance scan includes two 3-..phi.., 480 V, 60 Hz commercial grid-forming IBRs (one of 250 kVA and another of 125 kVA rating) in series with ..delta..-Y transformers, one 3-..phi.., 480 V, 60 Hz commercial grid-following IBR (of 125 kVA rating), and a 3-..phi.., 480 V, 60 Hz commercial diesel generator (of 187.5 kVA rating). Using voltage signals perturbed with sub-, inter-, and higher harmonic components, and measuring the current response, the positive-sequence impedances are computed via an offline- based post-analysis. Moreover, best-fit transfer functions are estimated that closely resemble the measured data points of the positive-sequence impedances. Based on the observations from various outcomes of the hardware experiments, this article also provides some fundamental insights on the equivalent positive- sequence impedance of a combination of multiple hardware components by comparing the estimated and the empirically computed impedances. A comparative insight on the damping capability of the DuTs using the positive-sequence impedances of the hardware is also discussed.

grid following inverter↗

Understanding Quantum Control Processor Capabilities and Limitations through Circuit Characterization

Continuing the scaling of quantum computers hinges on building classical control hardware pipelines that are scalable, extensible, and provide real time response. The instruction set architecture (ISA) of the control processor provides functional abstractions that map high-level semantics of quantum programming languages to low-level pulse generation by hardware. Here, we provide a methodology to quantitatively assess the effectiveness of the ISA to encode quantum circuits for intermediate-scale quantum devices with O(10 2 ) qubits. The characterization model that we define reflects performance, the ability to meet timing constraint implications, scalability for future quantum chips, and other important considerations making them useful guides for future designs. Using our methodology, we propose scalar (QUASAR) and vector (qV) quantum ISAs as extensions and compare them with other ISAs in metrics such as circuit encoding efficiency, the ability to meet real-time gate cycle requirements of quantum chips, and the ability to scale to more qubits.

97 MATHEMATICS AND COMPUTING↗

Transforming Science Through Software: Improving While Delivering 100×

The U.S. Department of Energy (DOE) Exascale Computing Project (ECP) funded the development of new (and the transformation of important existing) applications, libraries, and tools that realized improvement in performance and capabilities of often 100 times or more on emerging exascale computers. This exceptional gain inspired the title of this special issue: Transforming Science through Software: Improving while delivering 100X. The term 100X refers to advancing capabilities in modeling, simulation, and analysis by a factor of 100 or more using some combination of new algorithms, optimization techniques, software libraries, and programming models, coupled with the next generation of hardware for high-performance computing (HPC). The papers in this issue share experiences with the practice and science of scientific software development, with an emphasis on developing a coherent, portable, and sustainable HPC software ecosystem for next-generation computational science. Finally, we hope to foster expanded community efforts related to the fundamental role of sustainable scientific software ecosystems in advancing the computing sciences.

97 MATHEMATICS AND COMPUTING↗

Quantum Computer-Aided Design: Digital Quantum Simulation of Quantum Processors

With the increasing size of quantum processors, submodules that constitute the processor hardware will become too large to accurately simulate on a classical computer. Therefore, one would soon have to fabricate and test each new design primitive and parameter choice in time-consuming coordination between design, fabrication, and experimental validation. Here we show how one can design and test the performance of next-generation quantum hardware—by using existing quantum computers. Focusing on superconducting transmon processors as a prominent hardware platform, we compute the static and dynamic properties of individual and coupled transmons. We show how the energy spectra of transmons can be obtained by variational hybrid quantum-classical algorithms that are well suited for near-term noisy quantum computers. In addition, single- and two-qubit gate simulations are demonstrated via Suzuki-Trotter decomposition. Our methods pave a promising way towards designing candidate quantum processors when the demands of calculating submodule properties exceed the capabilities of classical computing resources.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Towards Automatic and Agile AI/ML Accelerator Design with End-to-End Synthesis

Domain-specific designs offer greater energy efficiency and performance gain than general-purpose processors. For this reason, modern system-on-chips have a significant portion of their silicon area with custom accelerators. However, designing hardware by hand is laborious and time-consuming, given the large design space and the performance, power, and area constraints that are not realized in the software. Moreover, domain-specific algorithms (e.g., machine learning models) are evolving quickly, challenging the accelerator design further. To address these issues, this paper presents SODA Synthesizer, an automated open-source high-level ML framework to Verilog modular compiler targeting AI/ML Application-Specific Integrated Circuits (ASICs) accelerators. SODA tightly couples the Multi- Level Intermediate Representation (MLIR) compiler infrastructure [24] and open-source HLS approaches. Thus, SODA can support various ML frameworks and algorithms and can perform optimizations that combine specialized architecture templates and conventional HLS to generate the hardware modules. In addition, SODA’s closed-loop design space exploration (DSE) engine allows developers to perform end-to-end design space explorations on different metrics and technology nodes.

Zhang, Jeff↗

Gate-free state preparation for fast variational quantum eigensolver simulations

Abstract The variational quantum eigensolver is currently the flagship algorithm for solving electronic structure problems on near-term quantum computers. The algorithm involves implementing a sequence of parameterized gates on quantum hardware to generate a target quantum state, and then measuring the molecular energy. Due to finite coherence times and gate errors, the number of gates that can be implemented remains limited. In this work, we propose an alternative algorithm where device-level pulse shapes are variationally optimized for the state preparation rather than using an abstract-level quantum circuit. In doing so, the coherence time required for the state preparation is drastically reduced. We numerically demonstrate this by directly optimizing pulse shapes which accurately model the dissociation of H 2 and HeH + , and we compute the ground state energy for LiH with four transmons where we see reductions in state preparation times of roughly three orders of magnitude compared to gate-based strategies.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Bridging Python to Silicon: The SODA Toolchain

Systems performing scientific computing, data analysis, and machine learning tasks have a growing demand for application-specific accelerators that can provide high computational performance while meeting strict size and power requirements. However, the algorithms and applications that need to be accelerated are evolving at a rate that is incompatible with manual design processes based on hardware description languages. Agile hardware design tools based on compiler techniques can help by quickly producing an application-specific integrated circuit (ASIC) accelerator starting from a high-level algorithmic description. Here, we present the software-defined accelerator (SODA) synthesizer, a modular and open-source hardware compiler that provides automated end-to-end synthesis from high-level software frameworks to ASIC implementation, relying on multilevel representations to progressively lower and optimize the input code. Our approach does not require the application developer to write any register-transfer level code, and it is able to reach up to 364 giga floating point operations per second (GFLOPS)/W efficiency (32-bit precision) on typical convolutional neural network operators.

97 MATHEMATICS AND COMPUTING↗

MIND-MAC: Multi-Level In-memory Quasi Non-Destructive MAC Operation in Compact 2T-nC FeRAM for Efficient DNN Accelerator

We present MIND-MAC, a compact 2T-nC FeRAM architecture that performs multi-level, quasi-non-destructive in-memory multiply–accumulate (MAC) for deep neural networks. By exploiting voltage-controlled partial domain switching in MFM capacitors and read-transistor amplification, the cell stores multi-bit weights and gates bit-serial inputs to produce an accumulated current on shared lines. We combine TCAD-extracted parasitics with experimentally calibrated ferroelectric models in SPICE to validate device-/circuit-level behavior, and validate multi-level sensing and QNRO with measurements on a fabricated 2T-3C test vehicle. An analytical system model maps MIND-MAC to a 6-GB main-memory in-memory compute (IMC) architecture and benchmarks VGG13 inference in 61.08 ms at 964.99 mJ. Results indicate high density, reduced rewrite overhead, and energy efficiency, positioning 2T-nC FeRAM as a promising IMC candidate for next-generation AI hardware.

36 MATERIALS SCIENCE↗

Experiences Readying Applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

exascale↗

AutodiDAQt v1.1.0

AutodiDAQt automates and simplifies writing data acquisition software for spectroscopy and microscopy. After defining only how to communicate with instruments and hardware, autodiDAQt generates user interfaces for long running acquisition applications, handlings data collation and retention, and provides remote communication to analysis computers. This reduces the time to get experiments running from months to hours and increases reliability for scientific experiments. AutodiDAQt metaprograms from instrument drivers directly, where possible.

Stansbury, Conrad↗

Emerging applications: Neuromorphic computing and reservoir computing

The emergence of doped hafnium oxide (HfO 2 )-based ferroelectric films has enabled highly scalable and silicon-compatible ferroelectric devices, opening new frontiers in neuromorphic and reservoir computing. Among these, ferroelectric field-effect transistors (FeFETs) are particularly promising due to their analog memory characteristics and unique polarization dynamics. These properties make FeFETs ideal candidates for artificial synapses in neuromorphic architectures, supporting deep neural networks and spiking neural networks based on leaky-integrate-and-fire (LIF) mechanisms. Beyond neuromorphic computing, FeFETs also play a crucial role in physical reservoir computing, leveraging their intrinsic nonlinear and history-dependent behavior for efficient real-time learning. This approach offers significant advantages for time-series processing and edge artificial intelligence (AI) applications, addressing the growing need for energy-efficient computing. As a result, this article explores the principles, key demonstrations, and future potential of FeFET-based neuromorphic and reservoir computing, highlighting their impact on next-generation AI hardware.

36 MATERIALS SCIENCE↗

Accelerator Real-time Edge AI for Distributed Systems (READS) (Proposal)

Over the last decade, Machine Learning (ML) technologies have slowly made their way into the accelerator community. Rapid advances in recent years in deep learning, particularly reinforcement learning for control system applications and the accessibility of deep learning in embedded hardware, have generated renewed interest and spawned a number of applications. The Fermilab Accelerator Complex, shown in Fig. 1, has provided High Energy Physics (HEP) experiments with proton beams for nearly fifty years. The current focus of the laboratory is its world-class experimental program at the intensity frontier. While increasing beam intensity certainly presents its own challenges, preserving beam size while minimizing beam losses – particles lost through interactions with the beam vacuum pipe – turns out to be, in many ways, the main challenge. The accelerator is controlled via a complex system of hundreds of thousands of devices. Enabling fine tuning and real-time optimization of their parameters using ML methods and stepping beyond experience-based reasoning of human operators are key to the success of future intensity upgrades. Our objective will be to integrate ML into accelerator operations and furthermore, provide an accessible framework, which can also be used by a broad range of other accelerator systems with dynamic tuning needs.

43 PARTICLE ACCELERATORS↗

Experiences readying applications for Exascale

The advent of Exascale computing invites an assessment of existing best practices for developing application readiness on the world's largest supercomputers. This work details observations from the last four years in preparing scientific applications to run on the Oak Ridge Leadership Computing Facility's (OLCF) Frontier system. This paper addresses a range of topics in software including programmability, tuning, and portability considerations that are key to moving applications from existing systems to future installations. A set of representative workloads provides case studies for general system and software testing. We evaluate the use of early access systems for development across several generations of hardware. Finally, we discuss how best practices were identified and disseminated to the community through a wide range of activities including user-guides and trainings. We conclude with recommendations for ensuring application readiness on future leadership computing systems.

Gottiparthi, Kalyan↗