Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22

Dataset of low global warming potential refrigerant refrigeration system for fault detection and diagnostics

Abstract HVAC and refrigeration system fault detection and diagnostics (FDD) has attracted extensive studies for decades; however, FDD of supermarket refrigeration systems has not gained significant attention. Supermarkets consume around 50 kWh/ft 2 of electricity annually. The biggest consumer of energy in a supermarket is its refrigeration system, which accounts for 40%–60% of its total electricity usage and is equivalent to about 2%–3% of the total energy consumed by commercial buildings in the United States. Also, the supermarket refrigeration system is one of the biggest consumers of refrigerants. Reducing refrigerant usage or using environmentally friendly alternatives can result in significant climate benefits. A challenge is the lack of publicly available data sets to benchmark the system performance and record the faulted performance. This paper identifies common faults of supermarket refrigeration systems and conducts an experimental study to collect the faulted performance data and analyze these faults. This work provides a foundation for future research on the development of FDD methods and field automated FDD implementation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

A theoretical investigation of the hydrolysis of uranium hexafluoride: the initiation mechanism and vibrational spectroscopy

Depleted uranium hexafluoride (UF 6 ), a stockpiled byproduct of the nuclear fuel cycle, reacts readily with atmospheric humidity, but the mechanism is poorly understood. Here we compare several potential initiation steps at a consistent level of theory, generating underlying structures and vibrational modes using hybrid density functional theory (DFT) and computing relative energies of stationary points with double-hybrid (DH) DFT. A benchmark comparison is performed to assess the quality of DH-DFT data using reference energy differences obtained using a complete-basis-limit coupled-cluster (CC) composite method. The associated large-basis CC computations were enabled by a new general-purpose pseudopotential capability implemented as part of this work. Dispersion-corrected parameter-free DH-DFT methods, namely PBE0-DH-D3(BJ) and PBE-QIDH-D3(BJ), provided mean unsigned errors within chemical accuracy (1 kcal mol -1 ) for a set of barrier heights corresponding to the most energetically favorable initiation steps. The hydrolysis mechanism is found to proceed via intermolecular hydrogen transfer within van der Waals complexes involving UF 6 , UF 5 OH, and UOF 4 , in agreement with previous studies, followed by the formation of a previously unappreciated dihydroxide intermediate, UF 4 (OH) 2 . The dihydroxide is predicted to form under both kinetic and thermodynamic control, and, unlike the alternate pathway leading to the UO 2 F 2 monomer, its reaction energy is exothermic, in agreement with observation. Finally, harmonic and anharmonic vibrational simulations are performed to reinterpret literature infrared spectroscopy in light of this newly identified species.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty quantification for molecular property predictions with graph neural architecture search

Graph Neural Networks (GNNs) have emerged as a prominent class of data-driven methods for molecular property prediction. However, a key limitation of typical GNN models is their inability to quantify uncertainties in the predictions. This capability is crucial for ensuring the trustworthy use and deployment of models in downstream tasks. To that end, we introduce AutoGNNUQ, an automated uncertainty quantification (UQ) approach for molecular property prediction. AutoGNNUQ leverages architecture search to generate an ensemble of high-performing GNNs, enabling the estimation of predictive uncertainties. Our approach employs variance decomposition to separate data (aleatoric) and model (epistemic) uncertainties, providing valuable insights for reducing them. In our computational experiments, we demonstrate that AutoGNNUQ outperforms existing UQ methods in terms of both prediction accuracy and UQ performance on multiple benchmark datasets, and generalizes well to out-of-distribution datasets. Additionally, we utilize t-SNE visualization to explore correlations between molecular features and uncertainty, offering insight for dataset improvement. AutoGNNUQ has broad applicability in domains such as drug discovery and materials science, where accurate uncertainty quantification is crucial for decision-making.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ddcMD: A fully GPU-accelerated molecular dynamics program for the Martini force field

We have implemented the Martini force field within Lawrence Livermore National Laboratory’s molecular dynamics program, ddcMD. The program is extended to a heterogeneous programming model so that it can exploit graphics processing unit (GPU) accelerators. In addition to the Martini force field being ported to the GPU, the entire integration step, including thermostat, barostat, and constraint solver, is ported as well, which speeds up the simulations to 278-fold using one GPU vs one central processing unit (CPU) core. A benchmark study is performed with several test cases, comparing ddcMD and GROMACS Martini simulations. The average performance of ddcMD for a protein–lipid simulation system of 136k particles achieves 1.04 µs/day on one NVIDIA V100 GPU and aggregates 6.19 µs/day on one Summit node with six GPUs. The GPU implementation in ddcMD offloads all computations to the GPU and only requires one CPU core per simulation to manage the inputs and outputs, freeing up remaining CPU resources on the compute node for alternative tasks often required in complex simulation campaigns.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Physical patterning of high-Q superconducting niobium resonators via ion beam etching

The development of superconducting quantum circuits increasingly involves the exploration of chemically distinct materials and complex multilayered structures. Accelerating this trend may benefit from low-damage, materials-agnostic patterning techniques that are compatible with a broad range of materials. Here, in this work, we investigate the utility of low-energy ion beam etching (IBE), a physical patterning technique, as an alternative to reactive ion etching for fabricating low-loss superconducting resonators. We use niobium (Nb) resonators as a test platform, leveraging their well-characterized performance metrics for benchmarking. To address IBE-induced surface redeposition, we introduce an in situ aluminum capping layer combined with targeted post-fabrication chemical treatment. This strategy yields resonators with internal quality factors as high as 6 × 10 5 in the single-photon regime at 50 mK. These results establish low-energy IBE as a promising patterning technique for superconducting devices, with the potential to accelerate development across chemically diverse and multilayered material platforms.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Concept Study of Robotic Camera-Based Foreign Object Detection for EV Wireless Charging

Wireless charging of an electric vehicle (EV) is an emerging charging technology promising convenient, autonomous, and highly efficient EV charging without requiring heavy gauge cables. However, due to the strong electromagnetic field created by this process that surrounds the wireless charger, the presence of foreign objects can detrimentally interact with it, thus affecting wireless power transfer (WPT) performance or leading to harmful and unwanted safety risks. This paper presents the results for a concept study on a robotic camera-based foreign object detection (FOD) system, as a supplement to the industry-existing overlapped FOD coil array method, for EV wireless charging. A Raspberry PI 4 control board and compatible Raspberry PI Camera Module 2 are used to implement camera-based object detection. The FOD program was developed using a state-of-the-art deep learning object detection model with the OpenCV and Pytorch library and is compatible with camera module hardware. A dry-run test with Raspberry PI and a camera module was conducted and the preliminary FOD function was verified. The feasibility assessment is also validated by comparing the performance of five existing state-of-the-art deep learning object detection models for vehicles, animals, persons, and metals subsets, respectively. Satisfactory performance on the benchmark datasets is observed by the tests, but further improvements are needed in future work when detecting small-sized metallic objects. A programable robotic car is also under development as ongoing work for carrying the Raspberry PI and camera module while moving for the maintenance process.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Application Experiences on a GPU-Accelerated Arm-based HPC Testbed

This paper assesses and reports the experience of ten teams working to port, validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems, each one equipped with a server-class Arm CPU from Ampere Computing and two data center GPUs from NVIDIA Corp. The systems are connected together using InfiniBand interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.

Elwasif, Wael↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

AMG2023

The AMG2023 benchmark solves two diffusion problems with a linear solver preconditioned with algebraic multigrid. The code only contains a driver, a Makefile, and a documentation file. It requires an installation of the open source software library hypre that needs to be downloaded elsewhere and is not included here. Its purpose is to benchmark linear solver performance on high performance computers.

Li, Ruipeng↗

Establishing model credibility for process-microstructure-property relationships in additive manufacturing using exascale computing

Additive Manufacturing (AM) of alloys holds significant promise as a disruptive technology in various industries, yet its adoption is often hindered by challenges in achieving consistent part quality. These issues are primarily due to the complex process-microstructure-property (PSP) relationships inherent to AM. Computational models can greatly aid in understanding these relationships, but their widespread impact and adoption has been limited by a lack of validated, open-source, and computationally efficient PSP modeling frameworks and hardware limitations. Here, this study leverages the ExaAM software suite and data from the AMBench-2018 series of laser powder bed fusion (LPBF) benchmark experiments to perform a comprehensive model assessment, including verification, validation, sensitivity analysis, and uncertainty quantification. The RADICAL-EnTK workflow manager was used to perform an ensemble of heat transport, solidification, and mechanical response simulations on the exascale computer Frontier, considering uncertainties in critical model inputs such as laser spot size and nucleation parameters, and consisting of 125 explicit grain structure simulations and 7875 crystal plasticity simulations. For a selected location within the Inconel 625 AMBench-2018 test artifact, sensitivity analysis and uncertainty quantification were performed using the predicted distributions of grain structure and mechanical properties. Qualitative agreement was found between the predicted grain size and texture and the observed AMBench-2018 microstructure, the mean predicted yield stress was within 5% of the experimental measurement mean, and the mean predicted engineering stress at 5% strain was within 10% of the experimental measurement mean. The insights gained from development and validation of the ExaAM PSP modeling framework will help guide future directions for enhancing the credibility and reliability of PSP models in AM, thereby accelerating the adoption of AM technologies in various industries.

Additive manufacturing↗

Deep-learning-aided forward optical coherence tomography endoscope for percutaneous nephrostomy guidance

Percutaneous renal access is the critical initial step in many medical settings. In order to obtain the best surgical outcome with minimum patient morbidity, an improved method for access to the renal calyx is needed. In our study, we built a forward-view optical coherence tomography (OCT) endoscopic system for percutaneous nephrostomy (PCN) guidance. Porcine kidneys were imaged in our experiment to demonstrate the feasibility of the imaging system. Three tissue types of porcine kidneys (renal cortex, medulla, and calyx) can be clearly distinguished due to the morphological and tissue differences from the OCT endoscopic images. To further improve the guidance efficacy and reduce the learning burden of the clinical doctors, a deep-learning-based computer aided diagnosis platform was developed to automatically classify the OCT images by the renal tissue types. Convolutional neural networks (CNN) were developed with labeled OCT images based on the ResNet34, MobileNetv2 and ResNet50 architectures. Nested cross-validation and testing was used to benchmark the classification performance with uncertainty quantification over 10 kidneys, which demonstrated robust performance over substantial biological variability among kidneys. ResNet50-based CNN models achieved an average classification accuracy of 82.6%±3.0%. The classification precisions were 79%±4% for cortex, 85%±6% for medulla, and 91%±5% for calyx and the classification recalls were 68%±11% for cortex, 91%±4% for medulla, and 89%±3% for calyx. Interpretation of the CNN predictions showed the discriminative characteristics in the OCT images of the three renal tissue types. The results validated the technical feasibility of using this novel imaging platform to automatically recognize the images of renal tissue structures ahead of the PCN needle in PCN surgery.

Wang, Chen↗

PickerXL, A Large Deep Learning Model to Measure Arrival Times from Noisy Seismic Signals

Precisely measuring seismic arrival times is a labor-intensive task but is critical for both earthquake monitoring and subsurface imaging. Recently published deep learning models have demonstrated superior performance compared to traditional automatic approaches for picking arrival times. Although existing deep learning models have shown promising results, further advancements are necessary as their performance is not yet satisfactory especially when applied to new regions and station networks. Increasing model size has led to improved performance in other machine learning applications. Here, we aimed to investigate whether enlarging deep learning models can increase performance on accepted benchmarks. We trained three models of varying sizes, small (1X), medium (4X), and large (16X), using globally distributed local and regional earthquake signals and background noise waveforms from a benchmark dataset, Stanford Earthquake Dataset. Our results indicate that the largest model (PickerXL) outperforms both the smaller models and Seisbench implementation of the PhaseNet model, which has the same number of parameters as our small model. The PickerXL model’s enhanced capacity to extract complex patterns from seismograms contributes to its superior arrival picking abilities compared to the smaller model.

Chai, Chengping [Oak Ridge National Laboratory (OR↗

Imperfection resonance crossing in the AGS Booster

Polarized helions are part of the spin physics program for the EIC, allowing collisions of polarized neutrons with polarized electrons. Helion imperfection resonances are 2.4 times closer than protons. Helions cross two intrinsic resonances (|Gγ| = 12 - ν γ and |Gγ| = 6 + ν γ ) and six imperfection resonances |Gγ| = 5, 6, 7, 8, 9, and 10) in the Booster as they are accelerated to extraction at |Gγ| = 10.5. In this same range of γ, protons cross two imperfection resonances (|Gγ| = 3, and 4) and are extracted from the Booster prior to crossing the |Gγ| = 0 + ν γ . Preliminary benchmarking simulations are performed using protons crossing the |Gγ| = 3 and 4 imperfection resonances, results of which are compared to experimental data. The settings used for protons are extrapolated to the helion case to show there is sufficient corrector strength to preserve polarization at each imperfection resonance up to extraction.

43 PARTICLE ACCELERATORS↗

Comparison of Tritium Dose Calculations from MACCS, UFOTRI, and ETMOD

Tritium exhibits unique environmental behavior because of its potential interactions with water and organic substances. Modeling the environmental consequences of tritium releases can be relatively complex and thus an evaluation of MACCS is needed to understand what updates, if any, are needed in MACCS to account for the behavior of tritium. We examine documented tritium releases and previous benchmarking assessments to perform a model intercomparison between MACCS and state-of-practice tritium-specific codes UFOTRI and ETMOD to quantify the difference between MACCS and state of practice models for assessing tritium consequences. Additionally, information to assist an analyst in judging whether a postulated tritium release is likely to lead to significant doses is provided.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multipleefforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of synthesized ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680 000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin G. [Fermilab]↗

IACMI Project 4.7: Pultruded Textile Carbon Fiber for Spar Caps (Final Report)

The primary objective of this project was to demonstrate the potential to significantly reduce the cost of wind turbine blades with carbon fiber reinforced polymer (CFRP) structure. Applicability of textile carbon fibers (TCF) were evaluated for use in pultruded spar cap (SC) elements as a path to cost reduction for utility scale wind turbine blades. In earlier work for the Department of Energy (DOE) Wind Energy Technologies Office (WETO), a collaboration of Sandia National Laboratory (SNL), Oak Ridge National Laboratory (ORNL), and Montana State University has demonstrated potential for pultruded TCF to compete with infused fiberglass and commercially available carbon fiber pultruded sections for spar cap construction. In the design cases evaluated, the TCF sections fared well when compared on cost per unit composite stiffness and cost per unit composite compressive strength for those designs [1]. Both stiffness and compressive strength tend to be key factors in the design of blade composite Spar Cap which carry the bulk of the blade structural loads in bending. Spar Cap design tends to distribute largely symmetric tensile and compressive stresses to opposite sides of the spar structure, but since carbon fiber composite compressive strength is typically 20-50% lower than tensile strength, the compressive loading reaches failure levels well before the tensile loading. Stiffness is critical in containing the large tip deflection in high wind loading situations. However, materials and process development were very limited in the earlier study and the work in this project was expanded to make the comparative information more representative of what will be required in order to make further inroads towards implementation. Similar to that study, this project team confirmed that the primary materials of interest for pultruded spar cap elements should be thermoset (TS) resins reinforced by carbon fibers, utilizing as high a percentage of TCF as practical to benchmark cost and performance against commercial carbon fibers. To make the closest comparison possible and eliminate specific test article size, resin selection, and equipment/operational nuance effects, the team planned to pultrude sections with 100% commercially available carbon fiber (Panex 35 carbon fiber from Zoltek) as well as samples utilizing high fractions of TCF. The resin system chosen was based on formulations recommended by large wind industry supplier Hexion and consisted of Hexion resin RSL-4597, curing agent CCA-138, and internal mold release additive 117, along with common kaolin filler ASP400P from BASF. As commonly deployed in spar cap configurations, the team had a mold built to pultrude a rectangular spar cap element of 100mm width and 3mm thickness. The extremely limited number of samples produced for the earlier study were produced with a “generic” epoxy utilized for a variety of applications by the pultruder contracted to produce test articles for demonstration purposes. More importantly, those samples were produced at a fiber fraction only slightly over 50%. Based on feedback from our industrial advisory team for that project and strongly recommended by this project team, the consensus is that it is highly desirable to obtain fiber fractions of 65-68% for significant penetration in wind blade spar cap. Although this requirement has yet to be exhaustively confirmed in readily available information, this was established as a project goal and informally decided we needed to exceed 60% fiber fraction to gain serious industry consideration. Previous TCF pultrusion trials have been challenged by the lack of robust TCF packages, resulting in non-uniform tension across and between tows, as well as excess labor and waste for removal of interleaved paper. The non-uniform tension and associated intermingling of tows in textile acrylic fiber tows and associated difficulties created from broken filaments in carbon fiber conversion inhibit the ordered packing necessary to enhance fiber fraction elevation. (These “cross-overs” are not considered undesirable for textile applications and there is some sense that they might be advantageous for those applications). In addition to work that is ongoing at the acrylic fiber manufacturers to improve their formats, The Institute of Advanced Composites Manufacturing Innovation (IACMI) Project 6.12 (report PA16-0349-6.12-01) [2] has developed and demonstrated a more robust packaging and creeling approach that at least partially addresses these issues, thus improving control of the TCF feed into the pultrusion unit. It was hoped that these and other improvements currently being implemented would allow us to achieve fiber fractions at least approaching these fiber fraction targets. During this project, sections utilizing 100% commercially available carbon fiber reinforcement were produced as a baseline, as well as sections reinforced with about 94% TCF and the balance being commercially available fiber for comparison. The most important finding was that similar to results reported in the earlier WETO-funded project and results from tests of TCF reported at IACMI meetings, this work demonstrated that sections pultruded with TCF in an epoxy resin frequently utilized in actual spar cap production had stiffness and compressive strengths largely comparable to similar sections pultruded with a commercially available carbon fiber also frequently utilized in the wind industry. Although the amount of that data is limited, some of the tensile strength results were actually closer than would have been expected based on fiber strength results provided by the TCF and commercial fiber producers. The actual test data are reported and discussed in detail in Section 5. The pultruded sections dominated by TCF reinforcement were approximately 8-10% lower in fiber fraction than for the sections produced using commercial fiber alone, making direct comparison difficult. The COVID-19 project has provided significant insight into the current state-of-the-art with various TCF product forms. The data obtained in this project will guide the planned improvements at the precursor level, especially in attaining uniform tensioning and payout to facilitate enhanced fiber fractions and overall processability of the TCF composites. The project team is providing guidance to stakeholders concerning the attributes, needs, and potential demand for TCF in wind blade spar caps. Results achieved in this project are consistent with findings in the related work cited [1] and support this guidance and the high potential for this product type. TCF precursor-producing partners continue to express interest in enhancing their product forms and the team looks forward to working with these improved materials as they become available.

42 ENGINEERING↗

Update on Radiochemical Assessment of High Burnup Commercially Irradiated Fuel

This work documents an effort to collect burnup measurements on a high burnup rod, designated 6XV, and first cycle accident tolerant fuel (ATF) rod, designated 47I, to enable benchmarking of fuel performance codes and neutronics codes. In addition to measurements, Virtual Environment for Reactor Applications (VERA) full-core-depletion analysis was also performed for the rods that were experimentally analyzed to provide an opportunity for code validation. This effort focuses on collecting data from rods irradiated at Byron Generating Station and shipped to the Oak Ridge National Laboratory (ORNL) hot-cells. This data will also anchor non-destructive examination evaluations of burnup of the various fuel rods undergoing postirradiation examination (PIE) at ORNL. Previous PIE of these fuel rods provides some guidance on the burnup trend across the fuel. Axial gamma spectroscopy scans provide a measure of relative changes in burnup across a fuel pin. Mass spectrometry based burnup measurements performed for this work at specific axial locations in the fuel are fully quantitative. By combining the mass spectrometry data with the gamma scans it is possible to more quantitatively evaluate axial variations in burnup across the entire fuel pin [1]. The combined set of burnup evaluations will be made available to other organizations that have an interest in high burnup radiochemistry data for validation of neutronic simulations and source term evaluation such as the Nuclear Regulatory Commission (NRC).

Harp, Jason [Oak Ridge National Laboratory (ORNL),↗

GASNet-EX Memory Kinds: Support for Device Memory in PGAS Programming Models

There is an emerging need for adaptive, lightweight communication in irregular HPC applications at exascale, where GPU accelerators provide the majority of available compute cycles. To address this need, Lawrence Berkeley National Lab is developing a programming system to support distributed-memory HPC application development using the Partitioned Global Address Space (PGAS) model. This work includes two major components: UPC++ and GASNet-EX. UPC++ is a C++ template library providing Remote Memory Access (RMA) and Remote Procedure Call (RPC) communication interfaces. GASNet-EX is a portable, high-performance communication middleware library, used by the implementations of UPC++ and many other PGAS programming models. We describe recent advances in GASNet-EX to efficiently implement zero-copy Remote Memory Access (RMA) communication to and from memory on accelerator devices such as GPUs. We demonstrate performance improvements via benchmark results from UPC++ (on Summit) and the Legion programming system (on DGX-1), both using GASNet-EX for communication.

Hargrove, Paul H↗