Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Field programmable gate array”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

The data acquisition system of the LZ dark matter detector: FADR

The Data Acquisition System (DAQ) for the LUX-ZEPLIN (LZ) dark matter detector is described. The signals from 745 PMTs, distributed across three subsystems, are sampled with 100-MHz 32-channel digitizers (DDC-32s). A basic waveform analysis is carried out on the on-board Field Programmable Gate Arrays (FPGAs) to extract information about the observed scintillation and electroluminescence signals. This information is used to determine if the digitized waveforms should be preserved for offline analysis. The system is designed around the Kintex-7 FPGA. In addition to digitizing the PMT signals and providing basic event selection in real time, the flexibility provided by the use of FPGAs allows us to monitor the performance of the detector and the DAQ in parallel to normal data acquisition. Furthermore, the hardware and software/firmware of this FPGA-based Architecture for Data acquisition and Realtime monitoring (FADR) are discussed and performance measurements are described.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Performance of the ATLAS Level-1 topological trigger in Run 2

During LHC Run 2 (2015–2018) the ATLAS Level-1 topological trigger allowed efficient data-taking by the ATLAS experiment at luminosities up to 2.1x10 34 cm -2 s -1 , which exceeds the design value by a factor of two. The system was installed in 2016 and operated in 2017 and 2018. It uses Field Programmable Gate Array processors to select interesting events by placing kinematic and angular requirements on electromagnetic clusters, jets, τ-leptons, muons and the missing transverse energy. It allowed to significantly improve the background event rejection and signal event acceptance, in particular for Higgs and B-physics processes.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

The ATLAS Fast TracKer system

The ATLAS Fast TracKer (FTK) was designed to provide full tracking for the ATLAS high-level trigger by using pattern recognition based on Associative Memory (AM) chips and fitting in high-speed field programmable gate arrays. The tracks found by the FTK are based on inputs from all modules of the pixel and silicon microstrip trackers. The as-built FTK system and components are described, as is the online software used to control them while running in the ATLAS data acquisition system. Also described is the simulation of the FTK hardware and the optimization of the AM pattern banks. An optimization for long-lived particles with large impact parameter values is included. A test of the FTK system with the data playback facility that allowed the FTK to be commissioned during the shutdown between Run 2 and Run 3 of the LHC is reported. The resulting tracks from part of the FTK system covering a limited $\eta$-$\phi$ region of the detector are compared with the output from the FTK simulation. It is shown that FTK performance is in good agreement with the simulation.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A prototype scintillator real‐time beam monitor for ultra‐high dose rate radiotherapy

Background: FLASH Radiotherapy (RT) is an emergent cancer RT modality where an entire therapeutic dose is delivered at more than 1000 times higher dose rate than conventional RT. For clinical trials to be conducted safely, a precise and fast beam monitor that can generate out-of-tolerance beam interrupts is required. This paper describes the overall concept and provides results from a prototype ultra-fast, scintillator-based beam monitor for both proton and electron beam FLASH applications. Purpose: A FLASH Beam Scintillator Monitor (FBSM) is being developed that employs a novel proprietary scintillator material. The FBSM has capabilities that conventional RT detector technologies are unable to simultaneously provide: (1) large area coverage; (2) a low mass profile; (3) a linear response over a broad dynamic range; (4) radiation hardness; (5) real-time analysis to provide an IEC-compliant fast beam-interrupt signal based on true two-dimensional beam imaging, radiation dosimetry and excellent spatial resolution. Methods: The FBSM uses a proprietary low mass, less than 0.5 mm water equivalent, non-hygroscopic, radiation tolerant scintillator material (designated HM: hybrid material) that is viewed by high frame rate CMOS cameras. Folded optics using mirrors enable a thin monitor profile of ∼10 cm. A field programmable gate array (FPGA) data acquisition system generates real-time analysis on a time scale appropriate to the FLASH RT beam modality: 100–1000 Hz for pulsed electrons and 10–20 kHz for quasi-continuous scanning proton pencil beams. An ion beam monitor served as the initial development platform for this work and was tested in low energy heavy-ion beams ( 86 Kr +26 and protons). A prototype FBSM was fabricated and then tested in various radiation beams that included FLASH level dose per pulse electron beams, and a hospital RT clinic with electron beams. Results: Results presented in this report include image quality, response linearity, radiation hardness, spatial resolution, and real-time data processing. Furthermore, the HM scintillator was found to be highly radiation damage resistant. It exhibited a small 0.025%/kGy signal decrease from a 216 kGy cumulative dose resulting from continuous exposure for 15 min at a FLASH compatible dose rate of 237 Gy/s. Measurements of the signal amplitude versus beam fluence demonstrate linear response of the FBSM at FLASH compatible dose rates of >40 Gy/s. Comparison with commercial Gafchromic film indicates that the FBSM produces a high resolution 2D beam image and can reproduce a nearly identical beam profile, including primary beam tails. The spatial resolution was measured at 35–40 µm. Tests of the firmware beta version show successful operation at 20 000 Hz frame rate or 50 µs/frame, where the real-time analysis of the beam parameters is achieved in less than 1 µs. Conclusions: The FBSM is designed to provide real-time beam profile monitoring over a large active area without significantly degrading the beam quality. A prototype device has been staged in particle beams at currents of single particles up to FLASH level dose rates, using both continuous ion beams and pulsed electron beams. Using a novel scintillator, beam profiling has been demonstrated for currents extending from single particles to 10 nA currents. Radiation damage is minimal and even under FLASH conditions would require ≥50 kGy of accumulated exposure in a single spot to result in a 1% decrease in signal output. Beam imaging is comparable to radiochromic films, and provides immediate images without hours of processing. Real-time data processing, taking less than 50 µs (combined data transfer and analysis times), has been implemented in firmware for 20 kHz frame rates for continuous proton beams.

2D beam imaging↗

AEcroscopy: A Software–Hardware Framework Empowering Microscopy Toward Automated and Autonomous Experimentation

Microscopy has been pivotal in improving the understanding of structure-function relationships at the nanoscale and is by now ubiquitous in most characterization labs. However, traditional microscopy operations are still limited largely by a human-centric click-and-go paradigm utilizing vendor-provided software, which limits the scope, utility, efficiency, effectiveness, and at times reproducibility of microscopy experiments. Here, in this work, a coupled software–hardware platform is developed that consists of a software package termed AEcroscopy (short for Automated Experiments in Microscopy), along with a field-programmable-gate-array device with LabView-built customized acquisition scripts, which overcome these limitations and provide the necessary abstractions toward full automation of microscopy platforms. The platform works across multiple vendor devices on scanning probe microscopes and electron microscopes. It enables customized scan trajectories, processing functions that can be triggered locally or remotely on processing servers, user-defined excitation waveforms, standardization of data models, and completely seamless operation through simple Python commands to enable a plethora of microscopy experiments to be performed in a reproducible, automated manner. This platform can be readily coupled with existing machine-learning libraries and simulations, to provide automated decision-making and active theory-experiment optimization to turn microscopes from characterization tools to instruments capable of autonomous model refinement and physics discovery.

47 OTHER INSTRUMENTATION↗

An Analysis of FPGA LUT Bias and Entropy for Physical Unclonable Functions

Process variations within Field Programmable Gate Arrays (FPGAs) provide a rich source of entropy and are therefore well-suited for the implementation of Physical Unclonable Functions (PUFs). However, careful considerations must be given to the design of the PUF architecture as a means of avoiding undesirable localized bias effects that adversely impact randomness, an important statistical quality characteristic of a PUF. Here in this paper, we investigate a ring-oscillator (RO) PUF that leverages localized entropy from individual look-up table (LUT) primitives. A novel RO construction is presented that enables the individual paths through the LUT primitive to be measured and isolated at high precision, and an analysis is presented that demonstrates significant levels of localized design bias. The analysis demonstrates that delay-based PUFs that utilize LUTs as a source of entropy should avoid using FPGA primitives that are localized to specific regions of the FPGA, and instead, a more robust PUF architecture can be constructed by distributing path delay components over a wider region of the FPGA fabric. Compact RO PUF architectures that utilize multiple configurations within a small group of LUTs are particularly susceptible to these types of design-level bias effects. The analysis is carried out on data collected from a set of identically designed, hard macro instantiations of the RO implemented on 30 copies of a Zynq 7010 SoC.

97 MATHEMATICS AND COMPUTING↗

Active control of laser beam pointing for the Zettawatt-Equivalent Ultrashort pulse laser System: a proof-of-principle study with 16-inch optics

We present a proof-of-principle study of active beam-pointing control for the Zettawatt-Equivalent Ultrashort pulse laser System (ZEUS) using a piezo-actuated 16-inch mirror. To the best of our knowledge, this is the largest actively controlled mirror reported in a high-power laser system. A simple proportional feedback control was implemented based on a field-programmable gate array, which reduced the standard deviation of beam-pointing fluctuations by 91% to 0.075 μrad in the horizontal direction and by 78% to 0.25 μrad in the vertical direction. We also demonstrated the elimination of long-term pointing jitter caused by temperature drift using the same apparatus.

laser pointing control↗

Accelerators for Classical Molecular Dynamics Simulations of Biomolecules

Atomistic Molecular Dynamics (MD) simulations provide researchers the ability to model biomolecular structures such as proteins and their interactions with drug-like small molecules with greater spatiotemporal resolution than is otherwise possible using experimental methods. MD simulations are notoriously expensive computational endeavors that have traditionally required massive investment in specialized hardware to access biologically relevant spatiotemporal scales. Our goal is to summarize the fundamental algorithms that are employed in the literature to then highlight the challenges that have affected accelerator implementations in practice. We consider three broad categories of accelerators: Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), and Application Specific Integrated Circuits (ASICs). These categories are comparatively studied to facilitate discussion of their relative trade-offs and to gain context for the current state of the art. We conclude by providing insights into the potential of emerging hardware platforms and algorithms for MD.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A low-latency graph computer to identify metastable particles at the Large Hadron Collider for real-time analysis of potential dark matter signatures

Abstract Image recognition is a pervasive task in many information-processing environments. We present a solution to a difficult pattern recognition problem that lies at the heart of experimental particle physics. Future experiments with very high-intensity beams will produce a spray of thousands of particles in each beam-target or beam-beam collision. Recognizing the trajectories of these particles as they traverse layers of electronic sensors is a massive image recognition task that has never been accomplished in real time. We present a real-time processing solution that is implemented in a commercial field-programmable gate array using high-level synthesis. It is an unsupervised learning algorithm that uses techniques of graph computing. A prime application is the low-latency analysis of dark-matter signatures involving metastable charged particles that manifest as disappearing tracks.

47 OTHER INSTRUMENTATION↗

HDBind: encoding of molecular structure with hyperdimensional binary representations

Traditional methods for identifying “hit” molecules from a large collection of potential drug-like candidates rely on biophysical theory to compute approximations to the Gibbs free energy of the binding interaction between the drug and its protein target. These approaches have a significant limitation in that they require exceptional computing capabilities for even relatively small collections of molecules. Increasingly large and complex state-of-the-art deep learning approaches have gained popularity with the promise to improve the productivity of drug design, notorious for its numerous failures. However, as deep learning models increase in their size and complexity, their acceleration at the hardware level becomes more challenging. Hyperdimensional Computing (HDC) has recently gained attention in the computer hardware community due to its algorithmic simplicity relative to deep learning approaches. The HDC learning paradigm, which represents data with high-dimension binary vectors, allows the use of low-precision binary vector arithmetic to create models of the data that can be learned without the need for the gradient-based optimization required in many conventional machine learning and deep learning methods. This algorithmic simplicity allows for acceleration in hardware that has been previously demonstrated in a range of application areas (computer vision, bioinformatics, mass spectrometery, remote sensing, edge devices, etc.). To the best of our knowledge, our work is the first to consider HDC for the task of fast and efficient screening of modern drug-like compound libraries. We also propose the first HDC graph-based encoding methods for molecular data, demonstrating consistent and substantial improvement over previous work. We compare our approaches to alternative approaches on the well-studied MoleculeNet dataset and the recently proposed LIT-PCBA dataset derived from high quality PubChem assays. We demonstrate our methods on multiple target hardware platforms, including Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs), showing at least an order of magnitude improvement in energy efficiency versus even our smallest neural network baseline model with a single hidden layer. Our work thus motivates further investigation into molecular representation learning to develop ultra-efficient pre-screening tools. We make our code publicly available at https://github.com/LLNL/hdbind.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing SRF cavity stability and minimizing detuning with data-driven resonance control based on dynamic mode decomposition

Effective resonance control of superconducting radio frequency (SRF) cavities is critical for large machines like LCLS-II, as failure to achieve proper control can result in increased RF power consumption, higher cryogenic heat loads, and increased costs. To address this challenge, we have developed a machine learning (ML) model based on the dynamic mode decomposition method to represent the forced cavity dynamics. Using this model, we designed a model predictive controller (MPC) and demonstrated through simulation that the MPC can effectively stabilize the amplitude and phase of SRF cavities using only a frequency actuator, even in the presence of multiple mechanical modes. The lightweight and explicit ML model makes the controller suitable for direct implementation on field-programmable gate arrays, unlocking the full potential of SRF linacs like LCLS-II, enabling higher beam power and energy, and also serving as an advanced motion controller for various applications, such as photon beamlines and storage rings.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Low latency optical-based mode tracking with machine learning deployed on FPGAs on a tokamak

Active feedback control in magnetic confinement fusion devices is desirable to mitigate plasma instabilities and enable robust operation. Optical high-speed cameras provide a powerful, non-invasive diagnostic and can be suitable for these applications. Here, in this study, we process high-speed camera data, at rates exceeding 100 kfps, on in situ field-programmable gate array (FPGA) hardware to track magnetohydrodynamic (MHD) mode evolution and generate control signals in real time. Our system utilizes a convolutional neural network (CNN) model, which predicts the n = 1 MHD mode amplitude and phase using camera images with better accuracy than other tested non-deep-learning-based methods. By implementing this model directly within the standard FPGA readout hardware of the high-speed camera diagnostic, our mode tracking system achieves a total trigger-to-output latency of 17.6 μs and a throughput of up to 120 kfps. This study at the High Beta Tokamak-Extended Pulse (HBT-EP) experiment demonstrates an FPGA-based high-speed camera data acquisition and processing system, enabling application in real-time machine-learning-based tokamak diagnostic and control as well as potential applications in other scientific domains.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A hybrid neural architecture: Online attosecond x-ray characterization

The emergence of high-repetition-rate x-ray free-electron lasers (XFELs), such as SLAC’s LCLS-II, serves as our canonical example for autonomous controls that necessitate high-throughput diagnostics paired with streaming computational pipelines capable of single-shot analysis with extremely low latency. We present the deterministic characterization with an integrated parallelizable hybrid resolver architecture, a hybrid machine learning framework designed for fast, accurate analysis of XFEL diagnostics using angular streaking-based sinogram images. This architecture integrates convolutional neural networks and bidirectional long short-term memory models to denoise input, identify x-ray sub-spike features, and extract sub-spike relative delays with sub-30 attosecond temporal resolution. Deployed on low-latency hardware, it achieves over 10 kHz throughput with 168.3 μs inference latency, indicating scalability to 14 kHz with field-programmable gate array integration. By transforming regression tasks into classification problems and leveraging optimized error encoding, we achieve high precision with low-latency performance that is critical for real-time streaming event selection and experimental control feedback signals. This represents a key development in real-time control pipelines for next-generation autonomous science, generally, and high repetition-rate x-ray experiments in particular.

Accelerator Physics (physics.acc-ph)↗

The Octopus processor for the CMS L1 muon trigger for High Luminosity LHC

The upgraded L1 muon trigger system of the CMS experiment in the High Luminosity Large Hadron Collider is based on custom processors featuring large Field Programmable Gate Arrays (FPGAs) connected by large numbers of optical links. These provide the I/O bandwidth and power necessary to process the complex algorithms used during the collection of physics data. The design and performance requirements of these processors creates significant challenges in signal integrity, power delivery, and thermal management. In this paper we describe the Octopus processor, featuring a large Xilinx Virtex Ultrascale+ FPGA and up to 128 links interfaced to optics through high quality twin-ax copper cables. Results on signal integrity at 25 Gb/s and the first demonstration of 50+ Gb/s links with pluggable optics in CMS are also shown, demonstrating bit error rates below 10 –15 at a 95% confidence level. The thermal performance is measured inside an Advanced-TCA crate with acceptable thermal margins up to 200 W of chip power. Future improvements are mentioned, potentially allowing operation at up to 300 W.

Instruments & Instrumentation↗

High performance FPGA embedded system for machine learning based tracking and trigger in sPhenix and EIC

We present a comprehensive end-to-end pipeline to classify triggers versus background events in this paper. This pipeline makes online decisions to select signal data and enables the intelligent trigger system for efficient data collection in the Data Acquisition System (DAQ) of the upcoming sPHENIX and future EIC (Electron-Ion Collider) experiments. Starting from the coordinates of pixel hits that are lightened by passing particles in the detector, the pipeline applies three-stage of event processing (hits clustering, track reconstruction, and trigger detection) and labels all processed events with the binary tag of trigger versus background events. The pipeline consists of deterministic algorithms such as clustering pixels to reduce event size, tracking reconstruction to predict candidate edges, and advanced graph neural network-based models for recognizing the entire jet pattern. In particular, we apply the message-passing graph neural network to predict links between hits and reconstruct tracks and a hierarchical pooling algorithm (DiffPool) to make the graph-level trigger detection. We obtain an impressive performance (≥70% accuracy) for trigger detection with only 3200 neuron weights in the end-to-end pipeline. We deploy the end-to-end pipeline into a field-programmable gate array (FPGA) and accelerate the three stages with speedup factors of 1152, 280, and 21, respectively.

Instruments & Instrumentation↗

100 Gb/s High Throughput Serial Protocol (HTSP) for data acquisition systems with interleaved streaming

Demands on Field-Programmable Gate Array (FPGA) data transport have been increasing over the years as frame sizes and refresh rates increase. As the bandwidths requirements increase the ability to implement data transport protocol layers using "soft" programmable logic becomes harder and start to require harden IP blocks implementation. To reduce the number of physical links and interconnects, it is common for data acquisition systems to require interleaving of streams on the same link (e.g. streaming data and streaming register access). Further, this paper presents a way to leverage existing FPGA harden IP blocks to achieve a robust, low latency 100 Gb/s point-to-point link with minimal programmable logic overhead geared towards the needs of data acquisition systems with interleaved streaming requirements.

47 OTHER INSTRUMENTATION↗

Nanosecond machine learning regression with deep boosted decision trees in FPGA for high energy physics

We present a novel application of the machine learning / artificial intelligence method called boosted decision trees to estimate physical quantities on field programmable gate arrays (FPGA). The software package fwXmachina features a new architecture called parallel decision paths that allows for deep decision trees with arbitrary number of input variables. It also features a new optimization scheme to use different numbers of bits for each input variable, which produces optimal physics results and ultraefficient FPGA resource utilization. Problems in high energy physics of proton collisions at the Large Hadron Collider (LHC) are considered. Estimation of missing transverse momentum (E T miss ) at the first level trigger system at the High Luminosity LHC (HL-LHC) experiments, with a simplified detector modeled by Delphes, is used to benchmark and characterize the firmware performance. The firmware implementation with a maximum depth of up to 10 using eight input variables of 16-bit precision gives a latency value of $\mathcal{O}$(10) ns, independent of the clock speed, and $\mathcal{O}$(0.1)% of the available FPGA resources without using digital signal processors.

Instruments & Instrumentation↗

Investigating resource-efficient neutron/gamma classification ML models targeting eFPGAs

There has been considerable interest and resulting progress in implementing machine learning (ML) models in hardware over the last several years from the particle and nuclear physics communities. A big driver has been the release of the Python package, hls4ml, which has enabled porting models specified and trained using Python ML libraries to register transfer level (RTL) code. So far, the primary end targets have been commercial field-programmable gate arrays (FPGAs) or synthesized custom blocks on application specific integrated circuits (ASICs). However, recent developments in open-source embedded FPGA (eFPGA) frameworks now provide an alternate, more flexible pathway for implementing ML models in hardware. These customized eFPGA fabrics can be integrated as part of an overall chip design. In general, the decision between a fully custom, eFPGA, or commercial FPGA ML implementation will depend on the details of the end-use application. In this work, we explored the parameter space for eFPGA implementations of fully-connected neural network (fcNN) and boosted decision tree (BDT) models using the task of neutron/gamma classification with a specific focus on resource efficiency. We used data collected using an AmBe sealed source incident on Stilbene, which was optically coupled to an OnSemi J-series silicon photomultiplier (SiPM) to generate training and test data for this study. We investigated relevant input features and the effects of bit-resolution and sampling rate as well as trade-offs in hyperparameters for both ML architectures while tracking total resource usage. The performance metric used to track model performance was the calculated neutron efficiency at a gamma leakage of 10 -3 . The results of the study will be used to aid the specification of an eFPGA fabric, which will be integrated as part of a test chip.

47 OTHER INSTRUMENTATION↗