Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PCIe”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Development of a High Throughput PCIe Card for DAQ System in the ATLAS and DUNE Experiments

In the Run 3 upgrade of ATLAS experiment, the FELIX (Front-End LInk eXchange) system has been prepared as the interface between front-end electronics and common Data Acquisition (DAQ) systems. Based on a PCIe card hosted in commodity server, FELIX's flexibilty makes it has also been adopted by other experiments, such as the Single-Phase ProtoDUNE (Prototype for the Deep Underground Neutrino Experiment), sPHENIX and CBM experiments. The same PCIe based architecture is proposed for use in the ATLAS HL-LHC (High Luminosity Large Hadron Collider) upgrade and the DUNE experiment. To this end, the next generation of FELIX I/O card FLX-801 has been developed. It supports 25+ Gbps high speed fiber optical links and 16-lane Gen4 PCIe interface. There is an on-card DDR4 module to buffer event data for DUNE experiment. This paper reports on the test results for the demonstrator of this next generation card, with which main functions have been successfully evaluated.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Performance Profile of Transformer Fine-Tuning in Multi-GPU Cloud Environments

The study presented here focuses on performance characteristics and trade-offs associated with running machine-learning tasks in multi-GPU environments on both on-site cloud computing resources and commercial cloud services (Azure). Specifically, this study examines these tradeoffs by examining the performance of training and fine-tuning of transformer-based deep-learning (DL) networks on clinical notes and data, a task of critical importance in the medical domain. To this end, we perform DL-related experiments on the widely deployed NVIDIA V100 GPUs and on the newer A100 GPUs connected via NVLink or PCIe. This study analyzes the execution time of major operations to train DL models and investigate popular options to optimize each of them. We examine and present the findings on the impacts that various operations (e.g. data loading into GPUs, training, fine-tuning), optimizations, and system configurations (single vs. multi-GPU, NVLink vs. PCIe) have on the overall training performance.

Begoli, Edmon↗

Enhancements and Deployment of the TDAQ System for the Mu2e Experiment

The Real Time Processing Systems Division at Fermilab has deployed new features to the Off-The-Shelf Data Acquisition framework (otsdaq) for the Mu2e experiment. The Mu2e experiment will search for the coherent neutrino-less conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement of 10,000 times over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable at present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system uses otsdaq as its online Data Acquisition System (DAQ) framework. otsdaq integrates the artdaq and art frameworks for event transfer, filtering, and processing. otsdaq is a web-based DAQ software suite focusing on flexibility and scalability and provides a multi-user interface accessible through a web browser. artdaq handles the entire data stream, which is read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events combined with the data flux coming from a cosmic-ray veto (CRV) system. Detector front-ends are configured through the PCIe bus by customized otsdaq plugins. The otsdaq slow controls infrastructure has been further developed using the experimental physics and industrial control system (EPICS) open-source platform for monitoring, controlling, alarming, and archiving. The detector control system (DCS) for Mu2e has been integrated into otsdaq. The production TDAQ and DCS system has been deployed at the experimental hall and is being debugged and optimized for experiment operations. We report on the feature enhancements and deployment of otsdaq for Mu2e.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Generic High Bandwidth Data Acquisition Card for Physics Experiments

In high energy physics and nuclear physics experiments particularly the ones based on particle accelerator, the data rate from the detector is usually in the order of Terabytes per second. This high throughput data from detector front-end electronics need to be transmitted to the back-end computing farm for high level event selection and building. A Data Acquisition (DAQ) system with features of high-density, scalable, easily upgradeable is crucial to simplify the readout architecture of whole experiment. This paper will introduce the design of a generic high bandwidth PCIe card which can be used as the important input output card in a scalable DAQ system. It can factorize front-end electronics from data handling, and reduce amount of custom hardware in favor of scalable detectorindependent commercial hardware and software. Besides the 48 channels of bidirectional high speed fiber optical links with frontends, it also supports to synchronize with the experiment timing system, and to fanout the clock and trigger information with a fixed latency to the front-end electronics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

TomocuPy – efficient GPU-based tomographic reconstruction with asynchronous data processing

Fast 3D data analysis and steering of a tomographic experiment by changing environmental conditions or acquisition parameters require fast, close to real-time, 3D reconstruction of large data volumes. Here a performance-optimized TomocuPy package is presented as a GPU alternative to the commonly used central processing unit (CPU) based TomoPy package for tomographic reconstruction. TomocuPy utilizes modern hardware capabilities to organize a 3D asynchronous reconstruction involving parallel read/write operations with storage drives, CPU–GPU data transfers, and GPU computations. In the asynchronous reconstruction, all the operations are timely overlapped to almost fully hide all data management time. Since most cameras work with less than 16-bit digital output, the memory usage and processing speed are furthermore optimized by using 16-bit floating-point arithmetic. As a result, 3D reconstruction with TomocuPy became 20–30 times faster than its multi-threaded CPU equivalent. Full reconstruction (including read/write operations and methods initialization) of a 2048 3 tomographic volume takes less than 7 s on a single Nvidia Tesla A100 and PCIe 4.0 NVMe SSD, and scales almost linearly increasing the data size. To simplify operation at synchrotron beamlines, TomocuPy provides an easy-to-use command-line interface. Efficacy of the package was demonstrated during a tomographic experiment on gas-hydrate formation in porous samples, where a steering option was implemented as a lens-changing mechanism for zooming to regions of interest.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Prototype Data Acquisition and Slow Control Systems for the Mu2e Experiment

The Mu2e experiment at the Fermilab Muon Campus will search for the coherent neutrinoless conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement by a factor of 10 000 over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable with direct searches at either present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system exploits otsdaq as its online Data Acquisition System (DAQ) solution. Furthermore, developed at Fermilab, otsdaq integrates both the artdaq DAQ and the art analysis frameworks for event transfer, filtering, and processing. otsdaq is an online DAQ software suite with a focus on flexibility and scalability and provides a multi-user, web-based, interface accessible through a web browser. The read out controllers (ROCs) stream out zero-suppressed data continuously from the detector subsystems to the data transfer controllers (DTCs). The data stream is then read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events which are combined with the data flux coming from a cosmic-ray veto (CRV) system. The detector control system (DCS) has been developed using the experimental physics and industrial control system (EPICS) open source platform for monitoring, controlling, alarming, and archiving. The DCS has been integrated into otsdaq. A prototype of the TDAQ system and the DCS has been built at Fermilab's Feynman Computing Center. In this article, we report on the progress of the integration of this prototype in the online otsdaq software.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU System

The deep learning revolution has been enabled in large part by GPUs, and more recently accelerators, which make it possible to carry out computationally demanding training and inference in acceptable times. As the size of machine learning networks and workloads continues to increase, multi-GPU machines have emerged as an important platform offered on High Performance Computing and cloud data centers. Since these machines are shared among multiple users, it becomes increasingly important to protect applications against potential attacks. In this paper, we explore the vulnerability of Nvidia's DGX multi-GPU machines to covert and side channel attacks. These machines consist of a number of discrete GPUs that are interconnected through a combination of custom interconnect (NVLink) and PCIe connections. We reverse engineer the interconnected cache hierarchy and show that it is possible for an attacker on one GPU to cause contention on the L2 cache of another GPU. We use this observation to first develop a covert channel attack across two GPUs, achieving the best bandwidth of around 4 MB/s. We also develop a prime and probe attack on a remote GPU allowing an attacker to recover the cache access pattern of another workload. This access pattern can be used in any number of side channel attacks: we demonstrate a proof of concept attack that fingerprints the application running on the remote GPU, with high accuracy. We also develop a proof of concept attack to extract hyperparameters of a machine learning workload. Our work establishes for the first time the vulnerability of these machines to microarchitectural attacks and can guide future research to improve their security.

Dutta, Sankha↗

SAMPA Based Streaming Readout Data Acquisition Prototype

We have assembled a small-scale streaming data acquisition system based on the SAMPA front-end ASIC. The 32-channel SAMPA chip was designed for the high-luminosity upgrade of the ALICE Time Projection Chamber (TPC) detector at the CERN Large Hadron Collider. The goals of the prototype system are to determine if the SAMPA chip is appropriate for use in detector systems at Jefferson Lab, and to gain experience with the hardware and software required to deploy streaming data acquisition systems in nuclear physics experiments. The 800 channel system is composed of components used in the ALICE TPC data acquisition upgrade. Five front-end cards (FEC) support five SAMPA chips each. SAMPA data streams on an FEC are concentrated into two high-speed (4.48 Gb/s) serial data streams by a pair Gigabit Transceiver ASICs (GBTx). These ten streams (44.8 Gb/s) are transmitted from the FECs over fibers to a PCIe based readout unit. The FPGA engine on the readout unit compresses data for transmission to a server via 100 Gb ethernet. Components on the FECs are radiation tolerant. High data rates can be handled. The system is by design scalable and thus provides a functional prototype for high-rate streaming readout at Jefferson Lab and the future Electron Ion Collider. We have made fundamental measurements (noise, linearity, time resolution) on the SAMPA ASIC. We have also coupled the readout system to a small Gas Electron Multiplier (GEM) detector and have studied its response to cosmic rays. A beam test of the system is planned.

ABBOTT, David↗

Lossy checkpoint compression in full waveform inversion: a case study with ZFPv0.5.5 and the overthrust model

This paper proposes a new method that combines checkpointing methods with error-controlled lossy compression for large-scale high-performance full-waveform inversion (FWI), an inverse problem commonly used in geophysical exploration. This combination can significantly reduce data movement, allowing a reduction in run time as well as peak memory. In the exascale computing era, frequent data transfer (e.g., memory bandwidth, PCIe bandwidth for GPUs, or network) is the performance bottleneck rather than the peak FLOPS of the processing unit. Like many other adjoint-based optimization problems, FWI is costly in terms of the number of floating-point operations, large memory footprint during backpropagation, and data transfer overheads. Past work for adjoint methods has developed checkpointing methods that reduce the peak memory requirements during backpropagation at the cost of additional floating-point computations. Combining this traditional checkpointing with error-controlled lossy compression, we explore the three-way tradeoff between memory, precision, and time to solution. We investigate how approximation errors introduced by lossy compression of the forward solution impact the objective function gradient and final inverted solution. Empirical results from these numerical experiments indicate that high lossy-compression rates (compression factors ranging up to 100) have a relatively minor impact on convergence rates and the quality of the final solution.

58 GEOSCIENCES↗

Trigger-DAQ and Slow Controls Systems in the Mu2e Experiment

The muon campus program at Fermilab includes the Mu2e experiment that will search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus, improving by four orders of magnitude the search sensitivity reached so far. Mu2e's Trigger and Data Acquisition System (TDAQ) uses otsdaq as its solution. Developed at Fermilab, otsdaq uses the artdaq DAQ framework and art analysis framework, under the-hood, for event transfer, filtering, and processing. otsdaq is an online DAQ software suite with a focus on flexibility and scalability, while providing a multi-user, web-based, interface accessible through the Chrome or Firefox web browser. The detector Read Out Controller (ROC), from the tracker and calorimeter, stream out zero-suppressed data continuously to the Data Transfer Controller (DTC). Data is then read over the PCIe bus to a software filter algorithm that selects events which are finally combined with the data flux that comes froma Cosmic Ray Ve to System (CRV). A Detector Control System (DCS) for monitoring, controlling, alarming, and archiving has been developed using the Experimental Physics and Industrial Control System (EPICS) Open Source Platform. The DCS System has also been itegrated into otsdaq. The installation of the TDAQ and the DCS systems in the Mu2e building is planned for 2021-2022, and a prototype has been built at Fermilab's Feynman Computing Center. We report here on the developments and achievements of the integration of Mu2e's DCS system into the online otsdaq software.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗