Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “RDMA”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

NCSU RDMA data

NCSU RDMA Data Level b1: QC checks applied to measurements Data Format: CSV Description: See "instrument description Site: Houston, TX; Tracking Aerosol Convection interactions ExpeRiment (HOU) Location: Houston, TX; AMF1 (main site for TRACER) Facility Code: M1 Category: Aerosol Properties Data Type: PI Data Source Instrument/Data: Nano-Scanning Mobility Particle (SMPS); A Radial Differential Mobility Analyzer (RDMA) coupled with 1 Condensation Particle Counter Start Date: 2022-06-01 End Date: 2022-9-26 Contact PI: Markus Petters (mdpetter@ncsu.edu) Funding Source: DOE ASR award US Department of Energy, Office of Science, Biological and Environment Research (grant no. DE-SC 0021074) Instrument description The NCSU RDMA was operated at a sheath-to-sample flow ratio of 5:1.5 L min−1. The RDMA was configured to scan from 5 to 55 nm. It was located into the temperature-controlled trailer adjacent to the NCSU Flux Tower to observe size distributions of the aerosols.The sample line was dried with three silica-gel driers in series, and then neutralized with X-ray neutralizer. Please contact mdpetter@ncsu.edu for further information.

54 ENVIRONMENTAL SCIENCES↗

Optimizing Data Movement for GPU-Based In-Situ Workflow Using GPUDirect RDMA

The extreme-scale computing landscape is increasingly dominated by GPU-accelerated systems. At the same time, in-situ workflows that employ memory-to-memory inter-application data exchanges have emerged as an effective approach for leveraging these extreme-scale systems. In the case of GPUs, GPUDirect RDMA enables third-party devices, such as network interface cards, to access GPU memory directly and has been adopted for intra-application communications across GPUs. In this paper, we present an interoperable framework for GPU-based in-situ workflows that optimizes data movement using GPUDirect RDMA. Specifically, we analyze the characteristics of the possible data movement pathways between GPUs from an in-situ workflow perspective, and design a strategy that maximizes throughput. Furthermore, we implement this approach as an extension of the DataSpaces data staging service, and experimentally evaluate its performance and scalability on a current leadership GPU cluster. The performance results show that the proposed design reduces data-movement time by up to 53% and 40% for the sender and receiver, respectively, and maintains excellent scalability for up to 256 GPUs.

Zhang, Bo↗

Extending OpenSHMEM with Aggregation Support for Improved Message Rate Performance

OpenSHMEM is a highly efficient one-sided communication API that implements the PGAS parallel programming model, and is known for its low latency communication operations that can be mapped efficiently to RDMA capabilities of network interconnects. However, applications that use OpenSHMEM can be sensitive to point-to-point message rates, as many-to-many communication patterns can generate large amounts of small messages which tend to overwhelm network hardware that has predominantly been optimised for bandwidth over message rate. Additionally, many important emerging classes of problems such as data analytics are similarly troublesome for the irregular access patterns they employ. Message aggregation strategies have been proven to significantly enhance network performance, but their implementation often involves complex restructuring of user code, making them unwieldy. This paper shows how to combine the best qualities of message aggregation within the communication model of OpenSHMEM such that applications with small and irregular access patterns can improve network performance while maintaining their algorithmic simplicity. We do this by providing a path to a message aggregation framework called conveyors through a minimally intrusive OpenSHMEM extension introducing aggregation contexts that fit more naturally to the OpenSHMEM atomics, gets, and puts model. We test these extensions using four of the bale 3.0 applications which contain essential many-to-many access patterns to show how they can produce performance improvements of up to 65×.

Welch, Aaron↗

HOSS!

The Hall-D Online Skim System (HOSS) was developed to simultaneously solve two issues for the high intensity GlueX experiment. One was to parallelize the writing of raw data files to disk in order to improve bandwidth. The other was to distribute the raw data across multiple compute nodes in order to produce calibration skims of the data online. The highly configurable system employs RDMA, RAM disks, and zeroMQ driven by Python to simultaneously store and process the full high intensity GlueX data stream.

Lawrence, David↗

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

Energy–Performance Trade-offs in Privacy-Preserving Federated Learning on SmartNIC-Enabled HPC Systems

Federated learning (FL) is increasingly deployed on accelerator-rich high-performance computing (HPC) systems, yet the system-level energy cost of privacy-aware FL remains poorly understood, particularly across heterogeneous networking and server-placement options. We present a measurement-driven study of energy–performance trade-offs for FL on GH200-class nodes across three deployment configurations: CPU-Ethernet, CPU-InfiniBand (RDMA-capable), and a DPU-hosted FL server over InfiniBand using a BlueField-3 SmartNIC/DPU. Using NVIDIA FLARE (NVFLARE), we align node-level power telemetry with per-round timing extracted from NVFLARE logs to quantify time-to-solution (TTS), energy-to-solution (ETS), energy-delay product (EDP), and synchronization behavior for three transformer models (ALBERT, DistilBERT, BERT), trained with and without differential privacy (DP). We find that interconnect choice is the dominant driver of runtime and energy: host-managed InfiniBand consistently reduces communication overhead versus Ethernet, yielding lower TTS/ETS/EDP. In contrast, in our NVFLARE deployment, placing the FL server on the DPU does not consistently match CPU-InfiniBand performance and can be slower—especially for larger models—highlighting that server placement alone is not sufficient to guarantee end-to-end gains. Finally, under our fixed-round protocol, DP increases per-round cost and runtime variance; ETS increases largely in proportion to TTS because average node power remains relatively stable across configurations.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Hod Carrier

SAND2023-06720O Hod Carrier is a simple proof-of-concept library for moving data from host to NVIDIA BlueField device memory using remote direct memory access. The Hod Carrier software also facilitates research activities related to the potential uses of NVIDIA BlueField data processing unit devices. Using well-established technologies (e.g., RDMA over Infiniband), it will move data from the host to the data processing unit’s memory. This software is a middleware utility for moving data without regard to its semantic meaning. Its operation is fundamentally unaffected by storage allocation. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

HipBone: A performance-portable graphics processing unit-accelerated C++ version of the NekBone benchmark

We present hipBone, an open-source performance-portable proxy application for the Nek5000 (and NekRS) computational fluid dynamics applications. HipBone is a fully GPU-accelerated C++ implementation of the original NekBone CPU proxy application with several novel algorithmic and implementation improvements which optimize its performance on modern fine-grain parallel GPU accelerators. Our optimizations include a conversion to store the degrees of freedom of the problem in assembled form in order to reduce the amount of data moved during the main iteration and a portable implementation of the main Poisson operator kernel. We demonstrate near-roofline performance of the operator kernel on three different modern GPU accelerators from two different vendors. We present a novel algorithm for splitting the application of the Poisson operator on GPUs which aggressively hides MPI communication required for both halo exchange and assembly. Our implementation of nearest-neighbor MPI communication then leverages several different routing algorithms and GPU-Direct RDMA capabilities, when available, which improves scalability of the benchmark. We demonstrate the performance of hipBone on three different clusters housed at Oak Ridge National Laboratory, namely, the Summit supercomputer and the Frontier early-access clusters, Spock and Crusher. Our tests demonstrate both portability across different clusters and very good scaling efficiency, especially on large problems.

Computer Science↗

Spectra Swarm Evaluation

Spectra Swarm product is an ethernet to Serial-attached-SCSI(SAS) bridge that allows a hostserver to communicate with SAS tape drives over ethernet using RDMA over ConvergedEthernet (RoCE).

97 MATHEMATICS AND COMPUTING↗

Particle Flux Measurements during TRACER (TRACER PFM) (Field Campaign Report)

The purpose of this campaign was to deploy an instrument payload collocated with a suite of in situ and remote-sensing aerosol instruments at the urban site at the La Porte, Texas Municipal Airport during the U.S. Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Tracking Aerosol Convection Interactions Experiment (TRACER) in Houston, Texas. The payload was designed to characterize the physicochemical properties, mixing state, and size-resolved vertical fluxes of aerosols. Deployed aerosol instruments associated with the flux measurements included three condensation particle counters (CPCs) with different lower cutoff diameters ( D > 2.5 nm, D > 10 nm, and D > 40 nm), a single-particle soot photometer (SP2) measuring refractory black carbon, and a portable optical particle spectrometer (POPS) measuring the optical size distribution (~180 nm < D < ~5000 nm). A 10-m flux tower hosted a sonic anemometer (SONIC) for vertical velocity measurements. In addition, a humidified tandem differential mobility analyzer (HTDMA) was deployed to measure the growth factors and hygroscopicity parameter of 15-50 nm-sized particles in order to constrain the composition of compounds responsible for modal aerosol growth during new particle formation/growth events. Finally, a radial differential mobility analyzer (RDMA) was deployed to extend the size distribution measurement to 5 nm.

54 ENVIRONMENTAL SCIENCES↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

NCSU HTDMA data

NCSU HTDMA data Data Level b1: QC checks applied to measurements Data Format: CSV Description: See "instrument description Site: Houston, TX; Tracking Aerosol Convection interactions ExpeRiment (HOU) Location: Houston, TX; AMF1 (main site for TRACER) Facility Code: M1 Category: Aerosol Properties Data Type: PI Data Source Instrument/Data: Humidified Tandem Differential Mobility Analyzer; 2 Differential Mobility Analyzers (DMA1 and DMA2). Start Date: 2022-06-01 End Date: 2022-09-26 Contact PI: Markus Petters (mdpetter@ncsu.edu) Funding Source: DOE ASR award US Department of Energy, Office of Science, Biological and Environment Research (grant no. DE-SC 0021074) Instrument description The NCSU HTDMA was operated at a sheath-to-sample flow ratio of 5:1 L min−1. The HTDMA was configured to measure hygroscopic growth factors of dry particles with mobility diameters of D = 15, 20, 30, 40, and 50 nm at RH ~ 70%. A complete cycle for all diameters took ~30 minutes. A sample line brought aerosol inside the trailer at 2.5 L min-1, where it was distributed between NCSU RDMA (1.5 L min-1) and NCSU HTDMA (1 L min-1) lines. The sample line was dried with three silica-gel driers in series, and then neutralized with X-ray neutralizer. The sample line entered DMA1 (operated as an electrostatic classifier). Monodisperse particles with certain fractions were humidified with temperature controlled Nafion membrane immersed in water before entering DMA2 (operated in scanning mobility particle sizer). Please contact mdpetter@ncsu.edu for further information.

54 ENVIRONMENTAL SCIENCES↗