Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PCIe”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Development of a High Throughput PCIe Card for DAQ System in the ATLAS and DUNE Experiments

In the Run 3 upgrade of ATLAS experiment, the FELIX (Front-End LInk eXchange) system has been prepared as the interface between front-end electronics and common Data Acquisition (DAQ) systems. Based on a PCIe card hosted in commodity server, FELIX's flexibilty makes it has also been adopted by other experiments, such as the Single-Phase ProtoDUNE (Prototype for the Deep Underground Neutrino Experiment), sPHENIX and CBM experiments. The same PCIe based architecture is proposed for use in the ATLAS HL-LHC (High Luminosity Large Hadron Collider) upgrade and the DUNE experiment. To this end, the next generation of FELIX I/O card FLX-801 has been developed. It supports 25+ Gbps high speed fiber optical links and 16-lane Gen4 PCIe interface. There is an on-card DDR4 module to buffer event data for DUNE experiment. This paper reports on the test results for the demonstrator of this next generation card, with which main functions have been successfully evaluated.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

AMD Radeon e9173 Low Power PCIE GPU Single Event Effects Test Report

The AMD Radeon Embedded e9170 Graphics Processing Unit (GPU), notably the e9173 Peripheral Component Interconnect Express (PCIE) variant, is of interest to Artemis generation programs with requirements for graphics rendering, compute, artificial intelligence (Ai) with a constraints-requiring piece-part procurement and power consumption of less than 50W. In addition to collecting heavy ion data on this device, a secondary purpose of this test campaign was to validate video capture hardware and software workflows used with GPU, microprocessor and system-on-chip device testing. Five (5) test patterns from the NEPP Processor Enclave (NPE) test suite were used with the e9173. The test patterns covered the operating system’s (OS) idle contribution towards the cross section, matrix math using tensorflow-rocm, two artificial intelligence models developed at NASA GSFC, and an industry standard GPU benchmarking application called Mesa GLXGears.

NASA Technical Memorandum (TM) test report for pos↗

EdgeCortix SAKURA-I Machine-Learning, PCIe Accelerator SEE Heavy Ion Test Report

To enable autonomy in space, machine-learning and computer vision applications become invaluable for sensor processing. However, these algorithms are computationally complex and unfeasible for many embedded central processing units (CPUs) and usually require external coprocessors, such as graphics processing units (GPUs) or accelerators specific to the application, including application specific integrated circuits (ASICs). In power-constrained systems, GPUs tend to consume more power than is acceptable (>40W), so lower-power accelerators have shown promise to provide the performance needed under spacecraft constraints. For radiation engineers, developing methodologies that can properly test CPUs, GPUs, and accelerators, and enable comparisons between them remains a necessary complication to solve as the devices become more complex. The methodology in this test aims to be a start in developing a baseline single-event effect (SEE) test for client-device machine learning accelerators. This category of devices do not host their own operating system. This testing campaign is a continuation of a previous 200 MeV proton test performed in January 2024. This report covers two heavy ion tests of the SAKURA-I card: one in April 2024, and one in June 2024. Additional data was needed after the April test due to ion-range issues experienced at higher linear-energy transfers (LETs). These range issues are described in more detail in Section 8. This experiment characterizes SEEs and data error susceptibility of the EdgeCortix SAKURA-I machine-learning accelerator under heavy ions. The device was monitored for single event upsets (SEUs) and single event functional interrupts (SEFIs) at the Lawrence Berkeley National Laboratory’s 88-inch cyclotron. The SAKURA-I board accelerates machine-learning inference applications on a host computer through a PCIex16 connection. For the purposes of devising an end to end automated analysis workflow for this experiment, the YOLO-V5 and SSD300 objection-detection models, and the ResNet-50, EfficientNet, and MobileNetV2 image classification models were used as a representative suite of analytical machine-learning models.

Seth S Roffe↗

Performance Profile of Transformer Fine-Tuning in Multi-GPU Cloud Environments

The study presented here focuses on performance characteristics and trade-offs associated with running machine-learning tasks in multi-GPU environments on both on-site cloud computing resources and commercial cloud services (Azure). Specifically, this study examines these tradeoffs by examining the performance of training and fine-tuning of transformer-based deep-learning (DL) networks on clinical notes and data, a task of critical importance in the medical domain. To this end, we perform DL-related experiments on the widely deployed NVIDIA V100 GPUs and on the newer A100 GPUs connected via NVLink or PCIe. This study analyzes the execution time of major operations to train DL models and investigate popular options to optimize each of them. We examine and present the findings on the impacts that various operations (e.g. data loading into GPUs, training, fine-tuning), optimizations, and system configurations (single vs. multi-GPU, NVLink vs. PCIe) have on the overall training performance.

Begoli, Edmon↗

Enhancements and Deployment of the TDAQ System for the Mu2e Experiment

The Real Time Processing Systems Division at Fermilab has deployed new features to the Off-The-Shelf Data Acquisition framework (otsdaq) for the Mu2e experiment. The Mu2e experiment will search for the coherent neutrino-less conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement of 10,000 times over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable at present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system uses otsdaq as its online Data Acquisition System (DAQ) framework. otsdaq integrates the artdaq and art frameworks for event transfer, filtering, and processing. otsdaq is a web-based DAQ software suite focusing on flexibility and scalability and provides a multi-user interface accessible through a web browser. artdaq handles the entire data stream, which is read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events combined with the data flux coming from a cosmic-ray veto (CRV) system. Detector front-ends are configured through the PCIe bus by customized otsdaq plugins. The otsdaq slow controls infrastructure has been further developed using the experimental physics and industrial control system (EPICS) open-source platform for monitoring, controlling, alarming, and archiving. The detector control system (DCS) for Mu2e has been integrated into otsdaq. The production TDAQ and DCS system has been deployed at the experimental hall and is being debugged and optimized for experiment operations. We report on the feature enhancements and deployment of otsdaq for Mu2e.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Generic High Bandwidth Data Acquisition Card for Physics Experiments

In high energy physics and nuclear physics experiments particularly the ones based on particle accelerator, the data rate from the detector is usually in the order of Terabytes per second. This high throughput data from detector front-end electronics need to be transmitted to the back-end computing farm for high level event selection and building. A Data Acquisition (DAQ) system with features of high-density, scalable, easily upgradeable is crucial to simplify the readout architecture of whole experiment. This paper will introduce the design of a generic high bandwidth PCIe card which can be used as the important input output card in a scalable DAQ system. It can factorize front-end electronics from data handling, and reduce amount of custom hardware in favor of scalable detectorindependent commercial hardware and software. Besides the 48 channels of bidirectional high speed fiber optical links with frontends, it also supports to synchronize with the experiment timing system, and to fanout the clock and trigger information with a fixed latency to the front-end electronics.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Optically connected memory for disaggregated data centers

Recent advances in integrated photonics enable the implementation of reconfigurable, high-bandwidth, and low energy-per-bit interconnects in next-generation data centers. We propose and evaluate an Optically Connected Memory (OCM) architecture that disaggregates the main memory from the computation nodes in data centers. OCM is based on micro-ring resonators (MRRs), and it does not require any modification to the DRAM memory modules. We calculate energy consumption from real photonic devices and integrate them into a system simulator to evaluate performance. Here, our results show that (1) OCM is capable of interconnecting four DDR4 memory channels to a computing node using two fibers with 1.02 pJ energy-per-bit consumption and (2) OCM performs up to 5.5× faster than a disaggregated memory with 40G PCIe NIC connectors to computing nodes.

97 MATHEMATICS AND COMPUTING↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

TomocuPy – efficient GPU-based tomographic reconstruction with asynchronous data processing

Fast 3D data analysis and steering of a tomographic experiment by changing environmental conditions or acquisition parameters require fast, close to real-time, 3D reconstruction of large data volumes. Here a performance-optimized TomocuPy package is presented as a GPU alternative to the commonly used central processing unit (CPU) based TomoPy package for tomographic reconstruction. TomocuPy utilizes modern hardware capabilities to organize a 3D asynchronous reconstruction involving parallel read/write operations with storage drives, CPU–GPU data transfers, and GPU computations. In the asynchronous reconstruction, all the operations are timely overlapped to almost fully hide all data management time. Since most cameras work with less than 16-bit digital output, the memory usage and processing speed are furthermore optimized by using 16-bit floating-point arithmetic. As a result, 3D reconstruction with TomocuPy became 20–30 times faster than its multi-threaded CPU equivalent. Full reconstruction (including read/write operations and methods initialization) of a 2048 3 tomographic volume takes less than 7 s on a single Nvidia Tesla A100 and PCIe 4.0 NVMe SSD, and scales almost linearly increasing the data size. To simplify operation at synchrotron beamlines, TomocuPy provides an easy-to-use command-line interface. Efficacy of the package was demonstrated during a tomographic experiment on gas-hydrate formation in porous samples, where a steering option was implemented as a lens-changing mechanism for zooming to regions of interest.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Prototype Data Acquisition and Slow Control Systems for the Mu2e Experiment

The Mu2e experiment at the Fermilab Muon Campus will search for the coherent neutrinoless conversion of a muon into an electron in the field of an aluminum nucleus with a sensitivity improvement by a factor of 10 000 over existing limits. Such a charged lepton flavor-violating reaction probes new physics at a scale unavailable with direct searches at either present or planned high-energy colliders. The Mu2e Trigger and Data Acquisition (TDAQ) system exploits otsdaq as its online Data Acquisition System (DAQ) solution. Furthermore, developed at Fermilab, otsdaq integrates both the artdaq DAQ and the art analysis frameworks for event transfer, filtering, and processing. otsdaq is an online DAQ software suite with a focus on flexibility and scalability and provides a multi-user, web-based, interface accessible through a web browser. The read out controllers (ROCs) stream out zero-suppressed data continuously from the detector subsystems to the data transfer controllers (DTCs). The data stream is then read over the peripheral component interconnect express (PCIe) bus to a software filter algorithm that selects events which are combined with the data flux coming from a cosmic-ray veto (CRV) system. The detector control system (DCS) has been developed using the experimental physics and industrial control system (EPICS) open source platform for monitoring, controlling, alarming, and archiving. The DCS has been integrated into otsdaq. A prototype of the TDAQ system and the DCS has been built at Fermilab's Feynman Computing Center. In this article, we report on the progress of the integration of this prototype in the online otsdaq software.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU System

The deep learning revolution has been enabled in large part by GPUs, and more recently accelerators, which make it possible to carry out computationally demanding training and inference in acceptable times. As the size of machine learning networks and workloads continues to increase, multi-GPU machines have emerged as an important platform offered on High Performance Computing and cloud data centers. Since these machines are shared among multiple users, it becomes increasingly important to protect applications against potential attacks. In this paper, we explore the vulnerability of Nvidia's DGX multi-GPU machines to covert and side channel attacks. These machines consist of a number of discrete GPUs that are interconnected through a combination of custom interconnect (NVLink) and PCIe connections. We reverse engineer the interconnected cache hierarchy and show that it is possible for an attacker on one GPU to cause contention on the L2 cache of another GPU. We use this observation to first develop a covert channel attack across two GPUs, achieving the best bandwidth of around 4 MB/s. We also develop a prime and probe attack on a remote GPU allowing an attacker to recover the cache access pattern of another workload. This access pattern can be used in any number of side channel attacks: we demonstrate a proof of concept attack that fingerprints the application running on the remote GPU, with high accuracy. We also develop a proof of concept attack to extract hyperparameters of a machine learning workload. Our work establishes for the first time the vulnerability of these machines to microarchitectural attacks and can guide future research to improve their security.

Dutta, Sankha↗

SAMPA Based Streaming Readout Data Acquisition Prototype

We have assembled a small-scale streaming data acquisition system based on the SAMPA front-end ASIC. The 32-channel SAMPA chip was designed for the high-luminosity upgrade of the ALICE Time Projection Chamber (TPC) detector at the CERN Large Hadron Collider. The goals of the prototype system are to determine if the SAMPA chip is appropriate for use in detector systems at Jefferson Lab, and to gain experience with the hardware and software required to deploy streaming data acquisition systems in nuclear physics experiments. The 800 channel system is composed of components used in the ALICE TPC data acquisition upgrade. Five front-end cards (FEC) support five SAMPA chips each. SAMPA data streams on an FEC are concentrated into two high-speed (4.48 Gb/s) serial data streams by a pair Gigabit Transceiver ASICs (GBTx). These ten streams (44.8 Gb/s) are transmitted from the FECs over fibers to a PCIe based readout unit. The FPGA engine on the readout unit compresses data for transmission to a server via 100 Gb ethernet. Components on the FECs are radiation tolerant. High data rates can be handled. The system is by design scalable and thus provides a functional prototype for high-rate streaming readout at Jefferson Lab and the future Electron Ion Collider. We have made fundamental measurements (noise, linearity, time resolution) on the SAMPA ASIC. We have also coupled the readout system to a small Gas Electron Multiplier (GEM) detector and have studied its response to cosmic rays. A beam test of the system is planned.

ABBOTT, David↗

Lossy checkpoint compression in full waveform inversion: a case study with ZFPv0.5.5 and the overthrust model

This paper proposes a new method that combines checkpointing methods with error-controlled lossy compression for large-scale high-performance full-waveform inversion (FWI), an inverse problem commonly used in geophysical exploration. This combination can significantly reduce data movement, allowing a reduction in run time as well as peak memory. In the exascale computing era, frequent data transfer (e.g., memory bandwidth, PCIe bandwidth for GPUs, or network) is the performance bottleneck rather than the peak FLOPS of the processing unit. Like many other adjoint-based optimization problems, FWI is costly in terms of the number of floating-point operations, large memory footprint during backpropagation, and data transfer overheads. Past work for adjoint methods has developed checkpointing methods that reduce the peak memory requirements during backpropagation at the cost of additional floating-point computations. Combining this traditional checkpointing with error-controlled lossy compression, we explore the three-way tradeoff between memory, precision, and time to solution. We investigate how approximation errors introduced by lossy compression of the forward solution impact the objective function gradient and final inverted solution. Empirical results from these numerical experiments indicate that high lossy-compression rates (compression factors ranging up to 100) have a relatively minor impact on convergence rates and the quality of the final solution.

58 GEOSCIENCES↗

SpaceVPX Interoperability Assessment

The existing VMEbus (VersaModular Eurocard bus) International Trade Association (VITA)-78 industry standard, also known as SpaceVPX, is an avionics board- and chassis-level standard derived from the OpenVPX standard as defined in VITA-65. While VITA-65 defines backplane and board-level profiles from COTS vendors to ensure interoperability of products used in developing systems and subsystems, the VITA-78 standard defines SpaceVPX to incorporate fault tolerance features that are required by many spaceflight systems. However, VITA-78 allows so much flexibility that interoperability between modules cannot be assured. This assessment provides guidelines on the use of, and extensions to, the VITA-78 standard to enable avionics interoperability for future NASA missions. The assessment team was comprised of subject matter experts (SMEs) from Goddard Space Flight Center (GSFC), the Jet Propulsion Laboratory (JPL), Johnson Space Center (JSC), and Langley Research Center (LaRC). The team included valuable external consulting support from a SME who was a key participant in the development of the VITA-78 standard. The team had extensive collaboration with the NASA Space Technology Mission Directorate (STMD) High Performance Spaceflight Computing (HPSC) project, specifically in the development of SpaceVPX interconnect findings, observations, and NESC recommendations. To provide an understanding of the breadth of implementations that SpaceVPX must accommodate, multiple NASA use cases were analyzed to assess the requirements for SpaceVPX implementations across a wide range of NASA missions (Appendix C). Applications included crewed missions, science missions, and orbital and surface robotic systems. Product surveys were conducted to assess the level of industry support for SpaceVPX, applications, and the variations in their implementations (Appendix D). In-depth analysis was conducted in the areas of: (a) power management and distribution, (b) form factors and daughtercards, (c) interconnect, and (d) fault tolerance. Leveraging the use cases, product surveys, and SMEs from multiple NASA Centers, these areas were analyzed to determine the range of implementations permitted by the VITA-78 standard and potential interoperability issues. Applicable findings and NESC recommendations were provided for each area. During this assessment, there were multiple opportunities to engage with other agencies to learn about their interest in SpaceVPX, their strategies for implementing SpaceVPX-based systems, and their internal development efforts. These engagements also generated findings and NESC recommendations. Based on this assessment analysis, NESC recommendations were made regarding the feature set and module profiles to support NASA SpaceVPX implementations. This feature set includes restrictions on features in VITA-78, and extensions to the standard. Key recommendations in this area include the use of 10 Gigabit Ethernet and Peripheral Component Interconnect Express (PCIe) as high bandwidth interconnect on the backplane, the retention of SpaceWire interconnect for control functions, and support for 3U (unit) and 6U, form factors for NASA systems. Restrictions were proposed on the usage of user-defined signals to promote interoperability, and specific power managements and distribution schemes for 3U systems. Beyond the technical implementation of SpaceVPX, recommendations were made on areas that warrant further investigation. Primary among these is the recommendation for NASA to collaborate with other space-going agencies and industry to incorporate recommendations into a future ‘dot spec’ of VITA-78. This would ensure wide adoption and availability of the modules that comply with the specification. The assessment includes appendices with candidate module profiles that can be considered as a starting point for this activity, and example systems based on the recommendations. Follow-on studies are recommended for architectures beyond SpaceVPX to address potential enhancements including condensed set of interconnect, software required to implement protocol layers on the interconnect (and other features), alternative power architectures, and system-level testability.

SpaceVPX↗

Trigger-DAQ and Slow Controls Systems in the Mu2e Experiment

The muon campus program at Fermilab includes the Mu2e experiment that will search for a charged-lepton flavor violating processes where a negative muon converts into an electron in the field of an aluminum nucleus, improving by four orders of magnitude the search sensitivity reached so far. Mu2e's Trigger and Data Acquisition System (TDAQ) uses otsdaq as its solution. Developed at Fermilab, otsdaq uses the artdaq DAQ framework and art analysis framework, under the-hood, for event transfer, filtering, and processing. otsdaq is an online DAQ software suite with a focus on flexibility and scalability, while providing a multi-user, web-based, interface accessible through the Chrome or Firefox web browser. The detector Read Out Controller (ROC), from the tracker and calorimeter, stream out zero-suppressed data continuously to the Data Transfer Controller (DTC). Data is then read over the PCIe bus to a software filter algorithm that selects events which are finally combined with the data flux that comes froma Cosmic Ray Ve to System (CRV). A Detector Control System (DCS) for monitoring, controlling, alarming, and archiving has been developed using the Experimental Physics and Industrial Control System (EPICS) Open Source Platform. The DCS System has also been itegrated into otsdaq. The installation of the TDAQ and the DCS systems in the Mu2e building is planned for 2021-2022, and a prototype has been built at Fermilab's Feynman Computing Center. We report here on the developments and achievements of the integration of Mu2e's DCS system into the online otsdaq software.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗