Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Field programmable gate arrays”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DynPaC: Coarse-Grained, Dynamic, and Partially Reconfigurable Array for Streaming Applications

Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) and higher efficiency than fine-grained reconfigurable devices such as Field Programmable Gate Arrays (FPGAs). However, CGRAs are generally designed to support offloading of a single kernel. While their design, based on communicating functional units, appears to naturally suit streaming applications composed of multiple cooperating kernels, current approaches only statically partition the resources across kernels. However, streaming applications often are data-dependent, leading to variable kernel execution times depending on the input data and impacting the throughput of the entire pipeline if resources are statically allocated. Therefore, in this paper, we discuss the design of DynPaC — a coarse-grained, dynamically, and partially reconfigurable array for data-dependent streaming applications. We discuss the required software and hardware components to manage partial dynamic reconfiguration. We demonstrate that by supporting partial dynamic reconfiguration, we can obtain an average speedup of 1.44X for a representative set of applications w.r.t. static partitioning, with a limited area overhead (6.4% of the entire chip).

Tan, Cheng↗

Gaia: segmented germanium detector for high-energy X-ray fluorescence and spectroscopic imaging

We present Gaia, a monolithic array of 96 high-purity germanium pixel detectors integrated with a custom low-noise application-specific integrated circuit (ASIC) and a field-programmable gate array (FPGA)-based data acquisition system. The sensor operates at ∼100 K using a commercial closed-cycle cryocooler, with the in-vacuum electronics thermally isolated from the cold finger to ensure thermal stability. The system demonstrates an average energy resolution of 711 eV at 122 keV, measured using a 57 Co source, and 253 eV at 5.89 keV, measured with 55 Fe across all channels. The readout architecture incorporates a high-performance FPGA paired with a dual-core ARM processor, forming a complete embedded Linux-based computing platform. Communication between the processor and FPGA is handled via memory-mapped I/O, and data are streamed over high-speed gigabit Ethernet. A full-scale 384-pixel Gaia detector, based on this 96-element module, is currently under fabrication.

36 MATERIALS SCIENCE↗

First application of a digital mirror Langmuir probe for real-time plasma diagnosis

We report for the first time, a digital Mirror Langmuir Probe (MLP) has successfully sampled plasma temperature, ion saturation current, and floating potential together on a single probe tip in real time in a radio-frequency driven helicon linear plasma device. This is accomplished by feedback control of the bias sweep to ensure a good fit to I–V characteristics with a high frequency, high power digital amplifier, and field-programmable gate array controller. Measurements taken by the MLP were validated by a low speed I–V characteristic manually collected during static plasma conditions. Plasma fluctuations, induced by varying the axial magnetic field($\widetilde{f}$ = 10 Hz), were also successfully monitored with the MLP. Further refinement of the digital MLP pushes it toward a turn-key system that minimizes the time to deployment and lessens the learning curve, positioning the digital MLP as a capable diagnostic for the study of low radio-frequency plasma physics. These demonstrations bolster confidence in fielding such digital MLP diagnostics in magnetic confinement experiments with high spatial and adequate temporal resolution, such as edge plasma, scrape-off layer, and divertor probes.

47 OTHER INSTRUMENTATION↗

First application of a digital mirror Langmuir probe for real-time plasma diagnosis

For the first time, a digital Mirror Langmuir probe (MLP) has successfully sampled plasma temperature, ion saturation current, and floating potential together on a single probe tip in real time in a radio-frequency driven helicon linear plasma device. This is accomplished by feedback control of the bias sweep to ensure a good fit to I-V characteristics with a high frequency, high power digital amplifier and field-programmable gate array (FPGA) controller. Measurements taken by the MLP were validated by a low speed I-V characteristic manually collected during static plasma conditions. Plasma fluctuations, induced by varying the axial magnetic field (f̃ = 10 Hz), were also successfully monitored with the MLP. Further refinement of the digital MLP pushes it towards a turn-key system that minimizes the time to deployment and lessens the learning curve, positioning the digital MLP as a capable diagnostic for the study of low radio-frequency plasma physics. These demonstrations bolster confidence in fielding such digital MLP diagnostics in magnetic confinement experiments with high spatial and adequate temporal resolution such as edge plasma, scrape-off layer, and divertor probes.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

AURORA: Automated Refinement of Coarse-Grained Reconfigurable Accelerators

Coarse-grained reconfigurable arrays (CGRAs), loosely defined as arrays of functional units interconnected through a network-on-chip (NoC), provide higher flexibility than domain-specific ASIC accelerators while offering increased hardware efficiency with respect to fine-grained reconfigurable devices, such as Field Programmable Gate Arrays (FPGAs). Un-fortunately, designing a CGRA for a specific application domain involves enormous software/hardware engineering effort (e.g., designing the CGRA, map operations onto the CGRA, etc) and requires the exploration on a large design space (e.g., applying appropriate loop transformation on each application, specializing the reconfigurable processing elements of the CGRA, refining the network topology, deciding the size of the data memory, etc). Int his paper, we propose AURORA – a software/hardware co-design framework to automatically synthesize optimal CGRA given a set of applications of interest

Tan, Cheng↗

Dtc Commercialization Software Package

This code is the complete software and firmware components supporting DTC model radios H2 and BluSDR6. This software package contains all the hardware boot up code/config files(BSP), user space Linux code (Web, Network, MAC (media access control) & drivers), the field programable gate array HDL (hardware description language) code and the build environment to compile and organize these components together to work in the aforementioned radios. Additional details of these components are as follows: • Hardware support components o Board support package and configuration files o uBoot • Linux Components: o The web components include the user interface for setup, configuration, and status components of the system. o Vulture code configures the radio’s IP network, configures radio parameters and runs the MAC layer of the radio. • The Field Programmable Gate Array HDL contains hardware drivers, interface logic to go between the software to the physical layer and the radio hardware as well as the logic for the physical layer of the radio. • Build environment includes compilers and config files that compile and organize all the other components to be able to be run on the radios.

Loera, Jose [Idaho National Laboratory (INL), Idah↗

Real-Time Artificial Intelligence for Particle Reconstruction and Higgs Physics

With the discovery of the Higgs boson at the CERN LHC, the world's highest-energy particle accelerator complex, scientists have acquired an important tool to study the fundamental building blocks of the universe. Precision measurements of Higgs bosons produced with large momentum allow for unique insights into the structure of the interactions of the Higgs boson with other particles that may shed light on physics beyond the standard model. While experimentally challenging, exploring such interactions with novel artificial intelligence (AI) methods can advance our understanding of the Higgs sector, including the Higgs boson's self-interaction. Moreover, the LHC is undergoing a major upgrade to further increase its particle collision rate and thereby operate for an additional decade. The experimental detectors at the upgraded facility must process at least a factor of ten more data at rates of hundreds of terabytes per second all under challenging conditions. New AI techniques are required to reconstruct and select, or trigger on, the most physics-sensitive events in real-time to handle the resulting avalanche of data. The proposed research will achieve the goals of the LHC program at the CMS experiment by developing a sub-microsecond event reconstruction system using real-time AI algorithms that employ field-programmable gate array technologies. By harnessing sophisticated AI techniques, this research focuses on measuring the production of Higgs bosons at large momentum while enhancing particle reconstruction methods in the trigger and beyond. Overall, the proposed research has broader implications for the use of AI in resource-constrained, low-latency embedded applications across all fields of science.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Ferroelectric FET-based context-switching FPGA enabling dynamic reconfiguration for adaptive deep learning machines

Field programmable gate array (FPGA) is widely used in the acceleration of deep learning applications because of its reconfigurability, flexibility, and fast time-to-market. However, conventional FPGA suffers from the trade-off between chip area and reconfiguration latency, making efficient FPGA accelerations that require switching between multiple configurations still elusive. Here, we propose a ferroelectric field-effect transistor (FeFET)–based context-switching FPGA supporting dynamic reconfiguration to break this trade-off, enabling loading of arbitrary configuration without interrupting the active configuration execution. Leveraging the intrinsic structure and nonvolatility of FeFETs, compact FPGA primitives are proposed and experimentally verified. The evaluation results show our design shows a 63.0%/74.7% reduction in a look-up table (LUT)/connection block (CB) area and 82.7%/53.6% reduction in CB/switch box power consumption with a minimal penalty in the critical path delay (9.6%). Besides, our design yields significant time savings by 78.7 and 20.3% on average for context-switching and dynamic reconfiguration applications, respectively.

97 MATHEMATICS AND COMPUTING↗

Real-Time Anomaly Detection for Charge-Based Triggering in LArTPCs

Modern particle detectors, including liquid argon time projection chambers (LArTPCs), collect a vast amount of data, making it impractical to save everything for offline analysis. As a result, these experiments need to employ different down-selection techniques during data acquisition, referred to as triggering. In this talk, I will present a framework that would enable real-time, data-driven triggering for LArTPCs, using anomaly detection algorithms implemented on Field-Programmable Gate Arrays (FPGAs). Drawing on a study that makes use of collected charge data from the MicroBooNE LArTPC Public Dataset, I will discuss the overall performance of such algorithms and potential applications for future neutrino experiments.

43 PARTICLE ACCELERATORS↗

Implementing machine learning methods on QICK hardware for qubit readout & control

Quantum readout and control is a fundamental aspect of quantum computing that requires accurate measurement of qubit states. Errors emerge in all stages, from initialization to readout, and identifying errors in post-processing necessitates resource-intensive statistical analysis. In our work, we use a lightweight fully-connected neural network (NN) to classify states of a transmon system with no prior processing. Our NN accelerator yields higher fidelities (92%) than the classical matched filter method (84%). By exploiting the natural parallelism of NNs and their placement near the source of data on field-programmable gate arrays (FPGAs), we can achieve ultra-low latency on the Quantum Instrumentation Control Kit (QICK). Integrating machine learning methods on QICK opens several pathways for efficient real-time processing of quantum circuits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Distilling particle knowledge for fast reconstruction at high-energy physics experiments

Knowledge distillation is a form of model compression that allows artificial neural networks of different sizes to learn from one another. Its main application is the compactification of large deep neural networks to free up computational resources, in particular on edge devices. In this article, we consider proton-proton collisions at the High-Luminosity Large Hadron Collider (HL-LHC) and demonstrate a successful knowledge transfer from an event-level graph neural network (GNN) to a particle-level small deep neural network (DNN). Our algorithm, DistillNet, is a DNN that is trained to learn about the provenance of particles, as provided by the soft labels that are the GNN outputs, to predict whether or not a particle originates from the primary interaction vertex. The results indicate that for this problem, which is one of the main challenges at the HL-LHC, there is minimal loss during the transfer of knowledge to the small student network, while improving significantly the computational resource needs compared to the teacher. This is demonstrated for the distilled student network on a CPU, as well as for a quantized and pruned student network deployed on a field programmable gate array. Our study proves that knowledge transfer between networks of different complexity can be used for fast artificial intelligence (AI) in high-energy physics that improves the expressiveness of observables over non-AI-based reconstruction algorithms. Such an approach can become essential at the HL-LHC experiments, e.g. to comply with the resource budget of their trigger stages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Evolution of the ATLAS event data model for the HL-LHC

The upcoming high-luminosity run of the CERN Large Hadron Collider (HL-LHC) will yield an unprecedented volume of data. In order to process this data, the ATLAS collaboration is evolving its offline software to be able to use heterogeneous resources such as graphical processing units (GPUs) and field-programmable gate arrays (FPGAs). To reduce conversion overheads, the event data model (EDM) should be compatible with the requirements of these resources. While the ATLAS EDM has long allowed representing data as a structure of arrays, further evolution of the EDM can enable more efficient sharing of data between CPU and GPU resources. Some of this work will be summarized here, including extensions to allow controlling how memory for event data is allocated and the implementation of jagged vectors.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Nanosecond anomaly detection with decision trees and real-time application to exotic Higgs decays

Abstract We present an interpretable implementation of the autoencoding algorithm, used as an anomaly detector, built with a forest of deep decision trees on FPGA, field programmable gate arrays. Scenarios at the Large Hadron Collider at CERN are considered, for which the autoencoder is trained using known physical processes of the Standard Model. The design is then deployed in real-time trigger systems for anomaly detection of unknown physical processes, such as the detection of rare exotic decays of the Higgs boson. The inference is made with a latency value of 30 ns at percent-level resource usage using the Xilinx Virtex UltraScale+ VU9P FPGA. Our method offers anomaly detection at low latency values for edge AI users with resource constraints.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Searching for Clues for a Matter Dominated Universe in Liquid Argon Time Projection Chambers

Liquid Argon Time Projection Chambers (LArTPCs) represent one of the most widely utilized neutrino detection techniques in neutrino experiments, for instance, in the Short Baseline Neutrino (SBN) program and the future large-scale LArTPC: Deep Underground Neutrino Experiment (DUNE). The high-end technique, facilitating excellent spatial and calorimetric reconstruction resolution, also enables testing exotic Beyond Standard Model (BSM) theories, such as baryon number violation (BNV) processes (e.g., proton-decay, neutron-antineutron oscillation). At the same time, Machine Learning (ML) techniques have demonstrated their ubiquitous use in recent decades; ML techniques have also become some of the most powerful tools in high-energy physics (HEP) analyses. Furthermore, the development of algorithms to cater to the needs of problems in HEP (i.e., triggering, reconstruction, improving sensitivity, etc.) has also become an active area of research. By developing a combined approach using Convolutional Neural Network (CNN) and Boosted Decision Tree (BDT) techniques, the sensitivity of neutron-antineutron oscillation in DUNE is evaluated for a projected exposure of 400kton·years. Additionally, to meet the triggering requirement to select such rare events in DUNE, such a search is only supported with highly efficient self-triggering algorithms. An ML-based self-triggering scheme for large-scale LArTPCs, such as DUNE, is also developed with the intention of implementation on field-programmable gate arrays (FPGAs). The ML-based approach for searching for neutron-antineutron oscillation can be demonstrated and validated on the current LArTPC MicroBooNE. The analysis in MicroBooNE represents the first-ever search for neutron-antineutron oscillation in a LArTPC. DUNE's projected 90% C.L. sensitivity to the neutron antineutron oscillation lifetime is 6.45×10³² years, assuming 1.327×10³⁵ neutron·years, equivalent to 10 years of DUNE far detector exposure (400kton·years). For MicroBooNE, assuming 372 seconds of exposure (equivalent to 3.13×10³⁶ neutron·years), the 90% C.L. lifetime sensitivity is found at 3.07×10²⁵ yrs, after accounting for Monte-Carlo statistical uncertainty and systematic uncertainty from detector effects.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Fast convolutional neural networks on FPGAs with hls4ml

We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

wa-hls4ml: A GNN Surrogate Model for hls4ml

Recent advancements in use of machine learning techniques on field-programmable gate arrays (FPGAs) have allowed for implementation of embedded neural networks with extremely low latency. This is invaluable for particle detectors at the Large Hadron Collider, where latency and used area must be strictly bounded. The hls4ml framework is a procedure for converting from trained machine learning model software, to a synthesis result that can be used on an FPGA. However, running the pipeline is a time-consuming procedure, and there is a strong risk of failure. In particular, it is possible that the model is unable to be converted into a synthesis result, or that the resource consumption of the model will exceed the resources of the target FPGA. To aid with this development, we introduce wa-hls4ml, a surrogate model which uses a graph neural network to emulate the structure of the source models. The goal is to estimate the chance of success and resource consumption of an arbitrary model when passed through the hls4ml procedure, without the time consumption of actually running the pipeline.

43 PARTICLE ACCELERATORS↗

Bringing heterogeneity to the CMS software framework

The advent of computing resources with co-processors, for example Graphics Processing Units (GPU) or Field-Programmable Gate Arrays (FPGA), for use cases like the CMS High-Level Trigger (HLT) or data processing at leadership-class supercomputers imposes challenges for the current data processing frameworks. These challenges include developing a model for algorithms to offload their computations on the co-processors as well as keeping the traditional CPU busy doing other work. The CMS data processing framework, CMSSW, implements multithreading using the Intel Threading Building Blocks (TBB) library, that utilizes tasks as concurrent units of work. In this paper we will discuss a generic mechanism to interact effectively with non-CPU resources that has been implemented in CMSSW. In addition, configuring such a heterogeneous system is challenging. In CMSSW an application is configured with a configuration file written in the Python language. The algorithm types are part of the configuration. The challenge therefore is to unify the CPU and co-processor settings while allowing their implementations to be separate. We will explain how we solved these challenges while minimizing the necessary changes to the CMSSW framework. We will also discuss on a concrete example how algorithms would offload work to NVIDIA GPUs using directly the CUDA API.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The continuous readout stream of the MicroBooNE liquid argon time projection chamber for detection of supernova burst neutrinos

The MicroBooNE continuous readout stream is a parallel readout of the MicroBooNE liquid argon time projection chamber (LArTPC) which enables detection of non-beam events such as those from a supernova neutrino burst. The low energies of the supernova neutrinos and the intense cosmic-ray background flux due to the near-surface detector location makes triggering on these events very challenging. Instead, MicroBooNE relies on a delayed trigger generated by SNEWS (the Supernova Early Warning System) for detecting supernova neutrinos. The continuous readout of the LArTPC generates large data volumes, and requires the use of real-time compression algorithms (zero suppression and Huffman compression) implemented in an FPGA (field-programmable gate array) in the readout electronics. In this paper we present the results of the optimization of the data reduction algorithms, and their operational performance. To demonstrate the capability of the continuous stream to detect low-energy electrons, a sample of Michel electrons from stopping cosmic-ray muons is reconstructed and compared to a similar sample from the lossless triggered readout stream.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗