Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Hardware generation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

PCA -BASED COMPRESSOR HARDWARE DESIGN GENERATOR IN CHISEL

SF-25-077 This repository includes a hardware generator written in Chisel that generates lossy compression hardware designs based on principal component analysis (PCA), along with a flexible testbench. It also includes a Python tool for evaluating the accuracy loss of integer quantization.

Kazutomo, Yoshi [Argonne National Laboratory (ANL)↗

Enhancing ICARUS and REDTOP Software and Hardware: Event Generator Interface Development and Calorimeter Tile Prototype

ICARUS (Imaging Cosmic And Rare Underground Signals) is a liquid argon time projection chamber (LArTPC) detector that pursues the sterile neutrino, which relies on accurate simulations of neutrino-argon interactions. REDTOP (Rare Eta Decays To Observe new Physics) is a proposed low-energy, high-intensity meson factory designed to explore rare $\eta$/$\eta'$ meson decays and probe physics beyond the Standard Model. As a next-generation experiment, this requires both accurate simulations and innovative detector technologies. This project contributes to both ICARUS, from a simulation perspective, and REDTOP, from both a simulation and detection perspective, through the event generation of lepton-nucleon interactions and the physical enhancement of the calorimeter technology within the REDTOP detector. We developed an interface between ACHILLES (A CHIcago Land Lepton Event Simulator), a theory-driven lepton-level event generator, and GENIE, a robust event generator framework used for neutrino physics. By incorporating the precise theoretical cross-section calculations of ACHILLES into the experimental realism of GENIE, the interface allows for improved accuracy of neutrino-nucleon simulations, which can be adapted for the proton beam specifications of the REDTOP meson factory as well as for the ICARUS experiment. In parallel, we developed an improved prototype for the ADRIANO2 (A Dual Readout Integrally Active Non-segmented Option) dual-readout calorimeter tiles for the REDTOP detector. To improve the efficiency of the lead-glass tiles trapping Cherenkov light for energy reconstruction and particle identification, we optimized the application of a highly reflective coating. Through viscosity and thickness control, masking, and a custom spray technique, we refined the coating process to reduce surface defects and improve light yield. Together, these efforts strengthen the ICARUS neutrino program and REDTOP's capability of detecting rare decay events.

Visser, Erin [Michigan State U.] (ORCID:0009000184↗

Event-driven readout development: testing of the EDWARD65P1 chip with integrated event generators

Building on a prototype readout integrated circuit for segmented silicon sensors with the EDWARD event-driven readout architecture, the front-end in each pixel was replaced by a hardware generator to verify readout performance, ensuring no data loss, consistent priority handling, and speed verification. Here, this generator produces Poisson-distributed readout requests with individually tunable rates per pixel via a digitally controlled oscillator. The resulting EDWARD65P1 test ASIC is a 32×32 pixel matrix with a 100 μm pitch, equipped with digital event generators simulating radiation hits at user-defined rates. Test results for this new design are presented.

47 OTHER INSTRUMENTATION↗

Proxy Applications for Converged Workloads: DMC LDRD Initiative

Modern scientific applications are complicated and require coordination of several components. Proxy application driven software-hardware co-design plays a vital role in driving innovation among the developments of applications, software infrastructure and hardware architecture. Proxy applications are self-contained and simplified codes that are intended to model the performance-critical computations within applications. Applications executing on modern High Performance Computing (HPC) systems are susceptible to network congestion, insufficient memory bandwidth within and across compute nodes, and inadvertent loss of performance due to bugs and unoptimized programming models. Modern numerical simulations and machine learning models play a critical role in studying physical phenomenon under myriad uncertainties. Such applications often exhibit irregular computation and memory accesses at specific regions of the application code, which can contribute to various performance bottlenecks at scale. To mitigate such issues and prepare the next generation hardware for a variety of computation and data movement contingencies, a well-known practice is to consider "proxy" applications as representative motifs for various classes of scientific applications. While there is disagreement in the HPC community on the mechanisms of construction of the proxy applications, there is a strong consensus on their positive impact in co-design. Proxy Applications for Converged Workloads (PACER) is about facilitating software-hardware co-design through proxy applications with the goal of improving the performance of converged science workflows on heterogeneous systems.

97 MATHEMATICS AND COMPUTING↗

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING↗

Extending High-Level Synthesis with AI/ML Methods

Artificial Intelligence (AI) and Machine Learning (ML) methods provide significant opportunities of improving quality of results when performing high-level synthesis (HLS). For example, they can be used to model and predict metrics of the final design (e.g., area, considering aspects such as interconnect overhead for different device technologies), facilitating exploration when searching for the best design trade-offs. They can also enable identifying hidden correlations across the various phases of the synthesis and the various optimizations performed, identifying the most effective pipelines. Finally, in more general terms, bio-inspired heuristic algorithms can improve the design space exploration for the synthesis process in terms of time and quality of the result. This paper discusses opportunities and challenges to augment HLS with AI/ML using as example flow the SODA Synthesizer, an open-source hardware generation toolchain which includes SODA-OPT, a hardware/software partitioning and pre-optimization tool developed with the MLIR framework, and PandA-Bambu, a state-of-the art HLS tool. SODA interfaces with OpenROAD to provide a complete end-to-end toolchain.

artificial intelligence↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

CrossSim Inference Manual v2.0

Neural networks are largely based on matrix computations. During forward inference, the most heavily used compute kernel is the matrix-vector multiplication (MVM): $W \vec{x} $. Inference is a first frontier for the deployment of next-generation hardware for neural network applications, as it is more readily deployed in edge devices, such as mobile devices or embedded processors with size, weight, and power constraints. Inference is also easier to implement in analog systems than training, which has more stringent device requirements. The main processing kernel used during inference is the MVM.

97 MATHEMATICS AND COMPUTING↗

Field Emission Mitigation in CEBAF SRF Cavities Using Deep Learning

The Continuous Electron Beam Accelerator Facility (CEBAF) operates hundreds of superconducting radio frequency (SRF) cavities in its two main linear accelerators. Field emission can occur when the cavities are set to high operating RF gradients and is an ongoing operational challenge. This is especially true in newer, higher gradient SRF cavities. Field emission results in damage to accelerator hardware, generates high levels of neutron and gamma radiation, and has deleterious effects on CEBAF operations. So, field emission reduction is imperative for the reliable, high gradient operation of CEBAF that is required by experimenters. Here we explore the use of deep learning architectures via multilayer perceptron to simultaneously model radiation measurements at multiple detectors in response to arbitrary gradient distributions. These models are trained on collected data and could be used to minimize the radiation production through gradient redistribution. This work builds on previous efforts in developing machine learning (ML) models, and is able to produce similar model performance as our previous ML model without requiring knowledge of the field emission onset for each cavity.

Ahammed, K.↗

Software Defined Architectures for Portability and Performance

The Software Defined Architectures for Portability and Performance (SODAPOP) project developed a co-design framework to partition and map converged applications on specialized heterogeneous architectures. We started from key domain applications that combine scientific simulation with data analytics and machine learning as drivers to integrate our framework. The framework includes high-level compilers that interfaces with high-level programming frameworks, domain-specific optimization passes, and hardware-oriented optimizations. The framework leverages hardware generators to enable specialization and facilitate exploration of custom system designs.

97 MATHEMATICS AND COMPUTING↗

Characterizing and mitigating coherent errors in a trapped ion quantum processor using hidden inverses

Quantum computing testbeds exhibit high-fidelity quantum control over small collections of qubits, enabling performance of precise, repeatable operations followed by measurements. Currently, these noisy intermediate-scale devices can support a sufficient number of sequential operations prior to decoherence such that near term algorithms can be performed with proximate accuracy (like chemical accuracy for quantum chemistry problems). While the results of these algorithms are imperfect, these imperfections can help bootstrap quantum computer testbed development. Demonstrations of these algorithms over the past few years, coupled with the idea that imperfect algorithm performance can be caused by several dominant noise sources in the quantum processor, which can be measured and calibrated during algorithm execution or in post-processing, has led to the use of noise mitigation to improve typical computational results. Conversely, benchmark algorithms coupled with noise mitigation can help diagnose the nature of the noise, whether systematic or purely random. Here, we outline the use of coherent noise mitigation techniques as a characterization tool in trapped-ion testbeds. We perform model-fitting of the noisy data to determine the noise source based on realistic physics focused noise models and demonstrate that systematic noise amplification coupled with error mitigation schemes provides useful data for noise model deduction. Further, in order to connect lower level noise model details with application specific performance of near term algorithms, we experimentally construct the loss landscape of a variational algorithm under various injected noise sources coupled with error mitigation techniques. This type of connection enables application-aware hardware codesign, in which the most important noise sources in specific applications, like quantum chemistry, become foci of improvement in subsequent hardware generations.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Performance Evaluation of a Microgrid System with Grid-Forming and Grid-Following Inverters with Diesel Generators: Insights From Hardware Experiments

This paper presents a comprehensive performance evaluation of a microgrid system integrating grid-forming (GFM) inverters, grid-following (GFL) inverters, and a diesel generator, focusing on their interactions and behavior under various dynamic events. The study is conducted using a pure hardware setup comprising two GFM inverters, one GFL inverter, a diesel generator, load banks, a point of common coupling (PCC), and an emulated main grid. The evaluation specifically examines dynamic scenarios, including voltage jumps, phase jumps, rate of change of frequency (ROCOF), synchronization, and islanding operations, which pose critical challenges to system stability. Among all the tests conducted, the overloading and phase jump tests proved to be the most challenging. The capacity of the DC side is crucial for withstanding overloading and grid disturbance tests; otherwise, GFM inverters frequently trip due to DC undervoltage. Throughout all grid disturbance tests, the diesel generator consistently stands out as the most robust and reliable GFM unit in the system. Overall, insights from these hardware experiments shed light on the response characteristics of different generation types during grid disturbances and identify potential stability concerns in such hybrid microgrid systems.

14 SOLAR ENERGY↗

An MLIR-based Compiler Flow for System-Level Design and Hardware Acceleration

The generation of custom hardware accelerators for applications implemented within high-level productive programming frameworks requires considerable manual effort. To automate this process, we introduce \sodaopt, a compiler tool that extends the MLIR infrastructure. \sodaopt automatically searches, outlines, tiles, and pre-optimizes relevant code regions to generate high-quality accelerators through high-level synthesis. \sodaopt can support any high-level programming framework and domain-specific language that interface with the MLIR infrastructure. By leveraging MLIR, \sodaopt solves compiler optimization problems with specialized abstractions. Backend synthesis tools connect to \sodaopt through progressive intermediate representation lowerings. \sodaopt interfaces to a design space exploration engine to identify the combination of compiler optimization passes and options that provides high-performance generated designs for different backends and targets. We demonstrate the practical applicability of the compilation flow by exploring the automatic generation of accelerators for deep neural networks operators outlined at arbitrary granularity and by combining outlining with tiling on large convolution layers. Experimental results with kernels from the PolyBench benchmark show that \sodaopt high-level optimizations improve execution delays of synthesized accelerators up to 60x. We also show that for the selected kernels, our solution outperforms the current of state-of-the art in more than 70% of the benchmarks and provides better average speedup in 55% of them.

Bohm Agostini, Nicolas↗

True random number generation using the spin crossover in LaCoO 3

While digital computers rely on software-generated pseudo-random number generators, hardware-based true random number generators (TRNGs), which employ the natural physics of the underlying hardware, provide true stochasticity, and power and area efficiency. Research into TRNGs has extensively relied on the unpredictability in phase transitions, but such phase transitions are difficult to control given their often abrupt and narrow parameter ranges (e.g., occurring in a small temperature window). Here we demonstrate a TRNG based on self-oscillations in LaCoO 3 that is electrically biased within its spin crossover regime. The LaCoO 3 TRNG passes all standard tests of true stochasticity and uses only half the number of components compared to prior TRNGs. Assisted by phase field modeling, we show how spin crossovers are fundamentally better in producing true stochasticity compared to traditional phase transitions. As a validation, by probabilistically solving the NP-hard max-cut problem in a memristor crossbar array using our TRNG as a source of the required stochasticity, we demonstrate solution quality exceeding that using software-generated randomness.

97 MATHEMATICS AND COMPUTING↗

Leveraging dendritic complexity for neuromorphic computing

Abstract Beyond-von Neumann computing approaches are necessary to sustain the growth of microelectronics and the increasing appetite for artificial intelligence/machine learning algorithms. Neuromorphic computing is an emerging paradigm that takes inspiration from the brain to provide a path forward to improve the computational efficiency and computational density of next-generation computing architectures. In nature, we observe brains performing complex computations with a much smaller energy footprint than conventional computing approaches. Current neuromorphic systems are focused primarily on scalability, namely, increasing the number of computational units (neurons) and connections between units (synapses). However, for brain-like cognition and efficiency in next-generation computing hardware, we need increased complexity in function, as well as improved connection density for scalability. Here, we present our work that aims to incorporate dendrites for ‘compute-on-wire’ in neuromorphic architectures to increase the computational complexity (e.g. number of programmable parameters, nonlinear dynamics) as well as computational efficiency (energy/compute) of artificial neural networks (ANNs). We do this by showcasing neuromorphic dendrite elements that can be leveraged for various applications. We will present examples of neuroscience-inspired direction-selective circuits and an ANN with active dendrites leveraging shunting inhibition. We also demonstrate the benefits of using dendrites in deep neural networks. To conclude, we discuss how we can utilize emerging hardware devices in these systems and design next-generation neuromorphic architectures with dendrites.

Cardwell, Suma G. (ORCID:0000000226575545)↗

Operating Wind Turbine as Synchronous Generator: Modeling and Power-Hardware-in-the-Loop Demonstration

Grid-forming (GFM) control of Type 3 and Type 4 wind turbine generators (WTGs) has attracted substantial attention in power systems research; however, the limited overcurrent capability of power electronics converters continues to deteriorate the grid strength of the evolving power systems. Synchronous wind, also referred to as a Type 5 WTG, offers a unique GFM solution to address grid integration and grid strength issues by keeping the grid largely synchronous at very high integration levels of renewable generation. A Type 5 WTG interfaces with the electric grid via a synchronous generator driven by a variable speed hydraulic torque converter; hence, the wind rotor operates in variable-speed mode for maximum power generation, and the generator shaft remains synchronous to the grid. This paper develops and tests a high-fidelity model of a Type 5 WTG in a power-hardware-in-the-loop testing environment, and it presents its operation characteristics under different grid contingencies. The power-hardware-in-the-loop demonstration shows that a Type 5 WTG inherently behaves as a GFM unit and can obtain similar performance in terms of power response, wind rotor dynamics, and stability enhancement compared to a Type 3 WTG in GFM control mode. Furthermore, the paper provides further insight into how Type 5 WTGs can support the smooth transition to power systems with high integration levels of inverter-based resources.

17 - WIND ENERGY↗