Engineering PapersSearch

DOE OSTI · 2575369

FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging

Abstract

We present the implementation of a specialized version of our previously published unified embedding model, SpeckleNN, for real-time speckle pattern classification in X-ray Single-Particle Imaging (SPI), using the SLAC Neural Network Library (SNL) on an FPGA platform. This hardware realization transitions SpeckleNN from a prototypic model into a practical edge solution, optimized for running inference near the detector in high-throughput X-ray free-electron laser (XFEL) facilities, such as those found at the Linac Coherent Light Source (LCLS). To address the resource constraints inherent in FPGAs, we developed a more specialized version of SpeckleNN. The original model, which was designed for broader classification across multiple biological samples, comprised ~5.6 million parameters. The new implementation, while reducing the parameter count to 64.6K (a 98.8% reduction), focuses on maintaining the model's essential functionality for real-time operation, achieving an accuracy of 90%. Furthermore, we compressed the latent space from 128 to 50 dimensions. This implementation was demonstrated on the KCU1500 FPGA board, utilizing 71% of available DSPs, 75% of LUTs, and 48% of FFs, with an average power consumption of 9.4W according to the Vivado post-implementation report. The FPGA performed inference on a single image with a latency of 45.015 microseconds at a 200 MHz clock rate. In comparison, running the same inference on an NVIDIA A100 GPU resulted in an average power consumption of ~73W and an image processing latency of around 400 microseconds. Our FPGA-accelerated version of SpeckleNN demonstrated significant improvements, achieving an 8.9 × speedup and a 7.8 × reduction in power consumption compared to the GPU implementation. Key advancements include model specialization and dynamic weight loading through SNL, which eliminates the need for time-consuming FPGA design re-synthesis, allowing fast and continuous deployment of models (re)trained online. These innovations enable real-time adaptive classification and efficient vetoing of speckle patterns, making SpeckleNN more suited for deployment in XFEL facilities. This implementation has the potential to significantly accelerate SPI experiments and enhance adaptability to evolving experimental conditions.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Dave, Abhilasha [SLAC National Accelerator Laboratory (SLAC), Menlo Park, CA (United States)], Wang, Cong [SLAC National Accelerator Laboratory (SLAC), Menlo Park, CA (United States)], Russell, James [SLAC National Accelerator Laboratory (SLAC), Menlo Park, CA (United States)], Herbst, Ryan [SLAC National Accelerator Laboratory (SLAC), Menlo Park, CA (United States)], Thayer, Jana [SLAC National Accelerator Laboratory (SLAC), Menlo Park, CA (United States)]. 2025-06-18. FPGA-accelerated SpeckleNN with SNL for real-time X-ray single-particle imaging. https://doi.org/10.3389/fhpcp.2025.1520151

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related reports

High Multiplicity Trigger for long-lived particles in CMS detector

Searches for long-lived particles (LLPs) at the CMS experiment often involve unconventional event topologies that are difficult to efficiently select using standard trigger strategies. To improve sensitivity to such signatures during LHC Run 3 operation, a dedicated High Multiplicity Trigger (HMT) has been developed and deployed in the CMS trigger system. The trigger targets events containing unusually large numbers of hits in the CMS cathode strip chamber (CSC) muon detectors, a characteristic signature of several LLP scenarios involving displaced decays in the muon system. The HMT implementation, trigger logic, rate dependence with pileup, and operational stability are described. Optimized hit multiplicity thresholds are used to maintain acceptable trigger rates under high-luminosity and high-pileup conditions while preserving high efficiency across a broad range of LLP lifetimes and kinematic regimes. The trigger performance is evaluated using both simulated event samples and proton-proton collision data collected during Run 3 of the LHC. The HMT substantially extends the CMS sensitivity to non-standard signatures associated with LLP decays and provides a flexible platform for future searches for physics beyond the Standard Model.

47 OTHER INSTRUMENTATION