Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “portable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

JACC.shared: Leveraging HPC Metaprogramming and Performance Portability for Computations That Use Shared Memory GPUs

In this work, we present JACC.shared, a new feature of Julia for ACCelerators (JACC), which is the performanceportable and metaprogramming model of the just-in-time and LLVM-based Julia language. This new feature allows JACC applications to leverage the high-performance computing (HPC) capabilities of high-bandwidth, on-chip GPU memory. Historically, exploiting high-bandwidth, shared-memory GPUs has not been a priority for high-level programming solutions. JACC.shared covers that gap for the first time, thereby providing a highlevel, portable, and easy-to-use solution for programmers to exploit this memory and supporting all current major accelerator architectures. Well-known HPC and AI workloads, such as multi/hyperspectral imaging and AI convolutions, have been used to evaluate JACC.shared on two exascale GPU architectures hosted by some of the most powerful US Department of Energy supercomputers: Perlmutter (NVIDIA A100) and Frontier (AMD MI250X). The performance evaluation reports speedup of up to 3.5× by adding only one line of code to the base codes, thus providing important accelerators in a simple, portable, and transparent way and elevating the programming productivity and performance-portability capabilities for Julia/JACC HPC, AI, and scientific applications.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Portability for GPU-accelerated molecular docking applications for cloud and HPC: can portable compiler directives provide performance across all platforms?

High-throughput structure-based screening of drug-like molecules has become a common tool in biomedical research. Recently, acceleration with graphics processing units (GPUs) has provided a large performance boost for molecular docking programs. Both cloud and high-performance computing (HPC) resources have been used for large screens with molecular docking programs; while NVIDIA GPUs have dominated cloud and HPC resources, new vendors such as AMD and Intel are now entering the field, creating the problem of software portability across different GPUs. Ideally, software productivity could be maximized with portable programming models that are able to maintain high performance across architectures. While in many cases compiler directives have been used as an easy way to offload parallel regions of a CPU-based program to a GPU accelerator, they may also be an attractive programming model for providing portability across different GPU vendors, in which case the porting process may proceed in the reverse direction: from low-level, architecture-specific code to higher-level directive-based abstractions. MiniMDock is a new mini-application (miniapp) designed to capture the essential computational kernels found in molecular docking calculations, such as are used in phar-maceutical drug discovery efforts, in order to test different solutions for porting across GPU architectures. Here we extend MiniMDock to GPU offloading with OpenMP directives, and compare to performance of kernels using CUDA and HIP on NVIDIA and AMD GPUs, respectively, as well as across different compilers, exploring performance bottlenecks. We document this reverse-porting process, from highly optimized device code to a higher-level version using directives, compare code structure, and describe barriers that were overcome in this effort.

Thavappiragasam, Mathialakan↗

Implementation of a portable diagnostic system for Thomson scattering measurements on an electrothermal arc source

To fulfill the increasing needs of diagnostic support for researchers in plasma technology, a portable diagnostic package (PDP) equipped for both laser Thomson scattering (TS) and optical emission spectroscopy has been designed and constructed at Oak Ridge National Laboratory (ORNL), aiming to measure the temperature and number density of electrons and temperatures of ions in plasma devices. The PDP has been initially implemented on a high density and low temperature electrothermal arc source (ET-arc) at ORNL to test its TS capability. TS from the plasmas in the ET-arc has been obtained using the PDP. The electron temperature and number density were determined from TS spectra. These results were then compared to measurements from previous studies on the ET-arc. The TS diagnostic measured 0.8 ± 0.1, 1.3 ± 0.2, and 0.7 ± 0.1 eV and (4.4 ± 0.5) × 1021, (5.9 ± 0.7) × 1021, and (4.3 ± 0.5) x 1021 m-3, respectively, from three lines of sight that transect the plasma column.

He, Z. (ORCID:0000000183159882)↗

Portable diagnostic package for Thomson scattering and optical emission spectroscopy on Princeton field-reversed configuration 2 (PFRC 2)

An Advanced Research Projects Agency-Energy funded diagnostic system has been deployed to the Princeton field-reversed configuration 2 (PFRC-2) device, located at Princeton Plasma Physics Laboratory. The Portable Diagnostic Package (PDP), designed at Oak Ridge National Laboratory, allows for the measurement of Thomson Scattering (TS) for electron density and temperature and Optical Emission Spectroscopy (OES) for ion temperature, impurity density, and ion velocity. A tunable spectrometer on the PDP with three gratings provides the flexibility to measure low (1 eV) and high (1000 eV) electron temperature ranges from TS. Additionally, using a second spectrometer, the OES diagnostic can survey light emission from various ion excitation levels for wide wavelength ranges. The electron density (<2 × 10 19 m –3 ) of plasmas generated in PFRC-2 has been below the PDP TS discrimination threshold, which has made TS signal detection challenging against a high-background of laser stray light. The laser stray light was iteratively reduced by making modifications to the entrance and exit geometry on PFRC-2. Rayleigh scattering experiments on PFRC have yielded the TS discrimination sensitivity to be >1 × 10 20 m –3 for the PDP. A recently implemented narrow-band notch spectral filter that masks the second harmonic 532 nm Nd:YAG laser wavelength has increased the system’s TS light discrimination sensitivity 65 times compared to the instance when the notch filter was not implemented. The hardware implementation including design changes to the flight tubes and Brewster windows will be discussed, along with results from Rayleigh and rotational Raman scattering sensitivity analyses, which were used to establish a quantitative figure of merit on the system performance. Further, the Raman scattering calibration with the notch filter has improved the PDP electron density threshold to 1 ± 0.5 × 10 18 m –3 .

47 OTHER INSTRUMENTATION↗

A portable and monoenergetic 24 keV neutron source based on 124Sb-9Be photoneutrons and an iron filter

A portable monoenergetic 24 keV neutron source based on the 124Sb-9Be photoneutron reaction and an iron filter has been constructed and characterized. The coincidence of the neutron energy from SbBe and the low interaction cross-section with iron (mean free path up to 29 cm) makes pure iron specially suited to shield against gamma rays from 124Sb decays while letting through the neutrons. To increase the 124Sb activity and thus the neutron flux, a >1 GBq 124Sb source was produced by irradiating a natural Sb metal pellet with a high flux of thermal neutrons in a nuclear reactor. The design of the source shielding structure makes for easy transportation and deployment. A hydrogen gas proportional counter is used to characterize the neutrons emitted by the source and a NaI detector is used for gamma background characterization. At the exit opening of the neutron beam, the characterization determined the neutron flux in the energy range 20–25 keV to be 6.00±0.30 neutrons per cm2 per second and the total gamma flux to be 245±8 gammas per cm2 per second (numbers scaled to 1 GBq activity of the 124Sb source). A liquid scintillator detector is demonstrated to be sensitive to neutrons with incident kinetic energies from 8 to 17 keV, so it can be paired with the source as a backing detector for neutron scattering calibration experiments. This photoneutron source provides a good tool for in-situ low energy nuclear recoil calibration for dark matter experiments and coherent elastic neutrino-nucleus scattering experiments.

Dark Matter detectors (WIMPs, axions, etc.)↗

JACC: Leveraging HPC Meta-Programming and Performance Portability with the Just-in-Time and LLVM-based Julia Language

We present JACC (Julia for Accelerators), the first high-level, and performance-portable model for the just-in-time and LLVM-based Julia language. JACC provides a unified and lightweight front end across different back ends available in Julia, enabling the same Julia code to run efficiently on many HPC CPU and GPU targets. We evaluated the performance of JACC for common HPC kernels as well as for the most computationally demanding kernels used in applications, HPCCG, a supercomputing benchmark test for sparse domains, and HARVEY, a blood flow simulator to assist in the diagnosis and treatment of patients suffering from vascular diseases. We carried out the performance analysis on the most advanced US DOE supercomputers: Aurora, Frontier, and Perlmutter. Overall, we show that JACC has a negligible overhead versus vendor-specific solutions, reporting GPU speedups with no extra cost to programmability.

Valero-Lara, Pedro↗

Portable EDITOR (PEDITOR): A portable image processing system

The PEDITOR image processing system was created to be readily transferable from one type of computer system to another. While nearly identical in function and operation to its predecessor, EDITOR, PEDITOR employs additional techniques which greatly enhance its portability. These cover system structure and processing. In order to confirm the portability of the software system, two different types of computer systems running greatly differing operating systems were used as target machines. A DEC-20 computer running the TOPS-20 operating system and using a Pascal Compiler was utilized for initial code development. The remaining programmers used a Motorola Corporation 68000-based Forward Technology FT-3000 supermicrocomputer running the UNIX-based XENIX operating system and using the Silicon Valley Software Pascal compiler and the XENIX C compiler for their initial code development.

Angelici, G.↗

Portable programming on parallel/networked computers using the Application Portable Parallel Library (APPL)

The Application Portable Parallel Library (APPL) is a subroutine-based library of communication primitives that is callable from applications written in FORTRAN or C. APPL provides a consistent programmer interface to a variety of distributed and shared-memory multiprocessor MIMD machines. The objective of APPL is to minimize the effort required to move parallel applications from one machine to another, or to a network of homogeneous machines. APPL encompasses many of the message-passing primitives that are currently available on commercial multiprocessor systems. This paper describes APPL (version 2.3.1) and its usage, reports the status of the APPL project, and indicates possible directions for the future. Several applications using APPL are discussed, as well as performance and overhead results.

Quealy, Angela↗

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

Lessons Learned From the Construction of a Portable Cleanroom for NASA OSIRIS-REx Mission Deintegration

NASA Johnson Space Center (JSC) Infrastructure and Astromaterials Acquisition & Curation Office completed construction and commissioning of the OSIRIS-REx (OREx) Deintegration portable cleanroom at the Utah Test and Training Range (UTTR). The new portable cleanroom was designed to receive the OREx sample return capsule from the landing point on the range to an ISO7 environment. Scientists used the portable clean-room to deintegrate the sample canister from the sample return capsule. Once separated, the sample canister was put in a container under nitrogen purge for transportation to B31 at the Johnson Space Center for astromaterial sample extraction, preliminary analysis, and long-term curation. The portable cleanroom was built by a subcontractor at their facility and then deconstructed to be transported to the remote location at UTTR. Since construction was completed in a remote location all tools and materials had to be transported from contractor site in Dallas, TX. The cleanroom was constructed within an existing facility, which provided conditioned air, electric power, and protection from the elements. Careful coordination was required between the host facility, cleanroom contractor, mission scientists, and JSC facilities and curation personnel. An existing anteroom at JSC was transported to UTTR and added to the portable cleanroom after there was concern about contamination without one for personnel entry/exit. The scientific study of organics is critical for the mission, so a stringent contamination control plan was implemented for low organics. Given these mission requirements the cleanroom construction materials were carefully selected to not hinder the scientific search for amino acids and the study of organics in the samples. The same cleanroom contractor that built the long-term astromaterial curation cleanroom back at JSC Houston, TX was selected to build the portable cleanroom and instructed to use the same materials. The cleanroom had double doors to open and allow the sample return capsule to fit into the cleanroom on its stand and be transferred to a clean stand already in the cleanroom. The portable cleanroom successfully completed its mission and the sample canister was safely deintegrated and transported to JSC under nitrogen purge.

astromaterials curation↗

Portable Cleanroom for NASA OSIRIS-REx Mission Deintegration

NASA Johnson Space Center (JSC) Infrastructure and Astromaterials Acquisition & Curation Office completed construction and commissioning of the OSIRIS-REx (OREx) Deintegration portable cleanroom at the Utah Test and Training Range (UTTR). The new portable cleanroom was designed to receive the OREx sample return capsule from the landing point on the range to an ISO7 environment. Scientists used the portable clean-room to deintegrate the sample canister from the sample return capsule. Once separated, the sample canister was put in a container under nitrogen purge for transportation to B31 at the Johnson Space Center for astromaterial sample extraction, preliminary analysis, and long-term curation. The portable cleanroom was built by a subcontractor at their facility and then deconstructed to be transported to the remote location at UTTR. Since construction was completed in a remote location all tools and materials had to be transported from contractor site in Dallas, TX. The cleanroom was constructed within an existing facility, which provided conditioned air, electric power, and protection from the elements. Careful coordination was required between the host facility, cleanroom contractor, mission scientists, and JSC facilities and curation personnel. An existing anteroom at JSC was transported to UTTR and added to the portable cleanroom after there was concern about contamination without one for personnel entry/exit. The scientific study of organics is critical for the mission, so a stringent contamination control plan was implemented for low organics. Given these mission requirements the cleanroom construction materials were carefully selected to not hinder the scientific search for amino acids and the study of organics in the samples. The same cleanroom contractor that built the long-term astromaterial curation cleanroom back at JSC Houston, TX was selected to build the portable cleanroom and instructed to use the same materials. The cleanroom had double doors to open and allow the sample return capsule to fit into the cleanroom on its stand and be transferred to a clean stand already in the cleanroom. The portable cleanroom successfully completed its mission and the sample canister was safely deintegrated and transported to JSC under nitrogen purge.

astromaterials curation↗

Portable aircraft controller devices and systems

A portable computerized device for an aircraft control system includes an input system for inputting commands, a device display for displaying information on the computerized device, a processor, a wireless communication module, and a non-transitory computer readable medium comprising computer executable instructions, the computer executable instructions configured to cause the processor to perform a method. The method can include detecting whether the portable computerized device is in a cockpit state such that the portable computerized device is in and/or docked to an aircraft cockpit or if the portable computerized device is in a remote state such that the portable computerized device is not in an aircraft cockpit or is not docked to an aircraft cockpit. If the portable computerized device is determined to be in a remote state, the method includes operating the remote device in a remote mode. If the portable computerized device is determined to be in a cockpit state, the method includes operating the device in a local mode.

Sahay, Prateek↗

CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems

Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.

Fujita, Norihisa↗