DOE OSTI · 3393807
Distributed Resilience in High-Energy Physics Data Acquisition
Abstract
Historical experience in the High-Performance Computing community teaches us that as computing systems grow, the instance of failures goes from rare to a regular occurrence. A survey of the growth in the size and complexity of Data AcQuisition (DAQ) networks in High-Energy Physics (HEP) experiments reveals that these networks are scaling exponentially, trending to a point where automated fault handling should be considered over the current manual practice, especially given the rarity of data such as in DUNE's mission to observe core-collapse supernovae. We propose a general system, DiDAQt, which is designed to provide fault detection and handling in HEP DAQs specifically, through MPI-like primitives that allow it to be added easily to existing systems. We evaluate the scalability and response time of a prototype on the FABRIC national testbed, with results indicating sufficient scalability for current and near-future DAQs as well as practical response times (under 1 microsecond decision time).
Keep this discovery
Explore connections, maps & timelines
Wolosewicz, A. [IIT, Chicago], Shyamkumar, N. [IIT, Chicago], Dart, E. [Unlisted], Ketchum, W. [Fermilab], Kowalkowski, J. [Fermilab], Wang, M. [Fermilab], Sultana, N. [IIT, Chicago]. 2026-07-30. Distributed Resilience in High-Energy Physics Data Acquisition. https://doi.org/10.2172/3393807
Cite the original work for its findings. Save a collection to share your selection of sources.