Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “portable”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

JACC.shared: Leveraging HPC Metaprogramming and Performance Portability for Computations That Use Shared Memory GPUs

In this work, we present JACC.shared, a new feature of Julia for ACCelerators (JACC), which is the performanceportable and metaprogramming model of the just-in-time and LLVM-based Julia language. This new feature allows JACC applications to leverage the high-performance computing (HPC) capabilities of high-bandwidth, on-chip GPU memory. Historically, exploiting high-bandwidth, shared-memory GPUs has not been a priority for high-level programming solutions. JACC.shared covers that gap for the first time, thereby providing a highlevel, portable, and easy-to-use solution for programmers to exploit this memory and supporting all current major accelerator architectures. Well-known HPC and AI workloads, such as multi/hyperspectral imaging and AI convolutions, have been used to evaluate JACC.shared on two exascale GPU architectures hosted by some of the most powerful US Department of Energy supercomputers: Perlmutter (NVIDIA A100) and Frontier (AMD MI250X). The performance evaluation reports speedup of up to 3.5× by adding only one line of code to the base codes, thus providing important accelerators in a simple, portable, and transparent way and elevating the programming productivity and performance-portability capabilities for Julia/JACC HPC, AI, and scientific applications.

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

Portability for GPU-accelerated molecular docking applications for cloud and HPC: can portable compiler directives provide performance across all platforms?

High-throughput structure-based screening of drug-like molecules has become a common tool in biomedical research. Recently, acceleration with graphics processing units (GPUs) has provided a large performance boost for molecular docking programs. Both cloud and high-performance computing (HPC) resources have been used for large screens with molecular docking programs; while NVIDIA GPUs have dominated cloud and HPC resources, new vendors such as AMD and Intel are now entering the field, creating the problem of software portability across different GPUs. Ideally, software productivity could be maximized with portable programming models that are able to maintain high performance across architectures. While in many cases compiler directives have been used as an easy way to offload parallel regions of a CPU-based program to a GPU accelerator, they may also be an attractive programming model for providing portability across different GPU vendors, in which case the porting process may proceed in the reverse direction: from low-level, architecture-specific code to higher-level directive-based abstractions. MiniMDock is a new mini-application (miniapp) designed to capture the essential computational kernels found in molecular docking calculations, such as are used in phar-maceutical drug discovery efforts, in order to test different solutions for porting across GPU architectures. Here we extend MiniMDock to GPU offloading with OpenMP directives, and compare to performance of kernels using CUDA and HIP on NVIDIA and AMD GPUs, respectively, as well as across different compilers, exploring performance bottlenecks. We document this reverse-porting process, from highly optimized device code to a higher-level version using directives, compare code structure, and describe barriers that were overcome in this effort.

Thavappiragasam, Mathialakan↗

Using a Customized Portable Deepwater Portable Electrofisher to Assess Larval Lamprey Populations in Irrigation Canals

Larval lamprey densities were quantified at two large-scale water diversions facilities located on the Yakima River near the city of Yakima, Washington. Water diversions are very common in the Yakima River Basin and provide water to irrigate a diverse range of valuable cropland. Out-migrating larval Pacific Lamprey are vulnerable to entrainment through fish protection screens during the withdrawal period and desiccation or predation during dewatering at the end of the season. In order to determine the impact to lamprey populations which inhabit these regions a portable deepwater electroshocking system was utilized to determine lamprey densities which could otherwise not be determined due to water depth. Surveys were conducted in the fall during the dewatering period in 2015 and 2017 and larval lamprey densities ranged from 0.8 to 8.8 fish m-2 at the Sunnyside Diversion Facility to 4.2 to 6.5 fish m-2 at the Wapato Diversion Facility. The estimates for the total lamprey numbers at each of the region using a portable deepwater electrofishing platform (PDEP) was very similar to those estimates using backpack electrofishing post-dewatering; the PDEP method was 12- 36% higher where comparisons were available. During 2017, the estimated lamprey entrainment at the Sunnyside Diversion increased by 484% from August to November. Lamprey lengths were also determined and ranged from 50 to 130 mm. Our results indicate that the use of the deepwater shocking system was very safe and effective at determining larval lamprey densities at hard to sample regions which are present near irrigation facilities.

larval lamprey, electrofishing, water diversion↗

Implementation of a portable diagnostic system for Thomson scattering measurements on an electrothermal arc source

To fulfill the increasing needs of diagnostic support for researchers in plasma technology, a portable diagnostic package (PDP) equipped for both laser Thomson scattering (TS) and optical emission spectroscopy has been designed and constructed at Oak Ridge National Laboratory (ORNL), aiming to measure the temperature and number density of electrons and temperatures of ions in plasma devices. The PDP has been initially implemented on a high density and low temperature electrothermal arc source (ET-arc) at ORNL to test its TS capability. TS from the plasmas in the ET-arc has been obtained using the PDP. The electron temperature and number density were determined from TS spectra. These results were then compared to measurements from previous studies on the ET-arc. The TS diagnostic measured 0.8 ± 0.1, 1.3 ± 0.2, and 0.7 ± 0.1 eV and (4.4 ± 0.5) × 1021, (5.9 ± 0.7) × 1021, and (4.3 ± 0.5) x 1021 m-3, respectively, from three lines of sight that transect the plasma column.

He, Z. (ORCID:0000000183159882)↗

Portable diagnostic package for Thomson scattering and optical emission spectroscopy on Princeton field-reversed configuration 2 (PFRC 2)

An Advanced Research Projects Agency-Energy funded diagnostic system has been deployed to the Princeton field-reversed configuration 2 (PFRC-2) device, located at Princeton Plasma Physics Laboratory. The Portable Diagnostic Package (PDP), designed at Oak Ridge National Laboratory, allows for the measurement of Thomson Scattering (TS) for electron density and temperature and Optical Emission Spectroscopy (OES) for ion temperature, impurity density, and ion velocity. A tunable spectrometer on the PDP with three gratings provides the flexibility to measure low (1 eV) and high (1000 eV) electron temperature ranges from TS. Additionally, using a second spectrometer, the OES diagnostic can survey light emission from various ion excitation levels for wide wavelength ranges. The electron density (<2 × 10 19 m –3 ) of plasmas generated in PFRC-2 has been below the PDP TS discrimination threshold, which has made TS signal detection challenging against a high-background of laser stray light. The laser stray light was iteratively reduced by making modifications to the entrance and exit geometry on PFRC-2. Rayleigh scattering experiments on PFRC have yielded the TS discrimination sensitivity to be >1 × 10 20 m –3 for the PDP. A recently implemented narrow-band notch spectral filter that masks the second harmonic 532 nm Nd:YAG laser wavelength has increased the system’s TS light discrimination sensitivity 65 times compared to the instance when the notch filter was not implemented. The hardware implementation including design changes to the flight tubes and Brewster windows will be discussed, along with results from Rayleigh and rotational Raman scattering sensitivity analyses, which were used to establish a quantitative figure of merit on the system performance. Further, the Raman scattering calibration with the notch filter has improved the PDP electron density threshold to 1 ± 0.5 × 10 18 m –3 .

47 OTHER INSTRUMENTATION↗

A portable and monoenergetic 24 keV neutron source based on 124Sb-9Be photoneutrons and an iron filter

A portable monoenergetic 24 keV neutron source based on the 124Sb-9Be photoneutron reaction and an iron filter has been constructed and characterized. The coincidence of the neutron energy from SbBe and the low interaction cross-section with iron (mean free path up to 29 cm) makes pure iron specially suited to shield against gamma rays from 124Sb decays while letting through the neutrons. To increase the 124Sb activity and thus the neutron flux, a >1 GBq 124Sb source was produced by irradiating a natural Sb metal pellet with a high flux of thermal neutrons in a nuclear reactor. The design of the source shielding structure makes for easy transportation and deployment. A hydrogen gas proportional counter is used to characterize the neutrons emitted by the source and a NaI detector is used for gamma background characterization. At the exit opening of the neutron beam, the characterization determined the neutron flux in the energy range 20–25 keV to be 6.00±0.30 neutrons per cm2 per second and the total gamma flux to be 245±8 gammas per cm2 per second (numbers scaled to 1 GBq activity of the 124Sb source). A liquid scintillator detector is demonstrated to be sensitive to neutrons with incident kinetic energies from 8 to 17 keV, so it can be paired with the source as a backing detector for neutron scattering calibration experiments. This photoneutron source provides a good tool for in-situ low energy nuclear recoil calibration for dark matter experiments and coherent elastic neutrino-nucleus scattering experiments.

Dark Matter detectors (WIMPs, axions, etc.)↗

JACC: Leveraging HPC Meta-Programming and Performance Portability with the Just-in-Time and LLVM-based Julia Language

We present JACC (Julia for Accelerators), the first high-level, and performance-portable model for the just-in-time and LLVM-based Julia language. JACC provides a unified and lightweight front end across different back ends available in Julia, enabling the same Julia code to run efficiently on many HPC CPU and GPU targets. We evaluated the performance of JACC for common HPC kernels as well as for the most computationally demanding kernels used in applications, HPCCG, a supercomputing benchmark test for sparse domains, and HARVEY, a blood flow simulator to assist in the diagnosis and treatment of patients suffering from vascular diseases. We carried out the performance analysis on the most advanced US DOE supercomputers: Aurora, Frontier, and Perlmutter. Overall, we show that JACC has a negligible overhead versus vendor-specific solutions, reporting GPU speedups with no extra cost to programmability.

Valero-Lara, Pedro↗

Dual Channel Dual Staging: Hierarchical and Portable Staging for GPU-Based In-Situ Workflow

In-situ workflows have emerged as an attractive approach for addressing data movement challenges at very large scales. Since GPU-based architectures dominate the HPC landscapes, porting these in-situ workflows, and, specifically, the inter-application data exchange, to GPU-based systems can be challenging. Technologies such as GPUDirect RDMA (GDR), which is typically used for I/O in GPU applications as an optimization that circumvents the CPU overhead, can be leveraged to support bulk data exchanges between GPU applications. However, current GDR design often lacks performance portability across HPC clusters built with different hardware configurations. Furthermore, the local CPU may also be effectively used as an auxiliary communication mechanism to offload data exchanges. In this paper, we present a dual channel dual staging approach for efficient, scalable, and performance-portable inter-application data exchange for in-situ workflows. This approach exploits the data access pattern within in-situ workflows along with the inherent execution asynchrony to accelerate data exchanges and, at the same time, improve performance portability. Specifically, the dual channel dual staging method leverages both the local CPU and the remote data staging server to build a hierarchical joint staging area and uses this staging area to transform blocking inter-application bulk data exchanges into best-effort local data movements between GPU and CPU. The dual channel dual staging is implemented as a portability extension of the Dataspaces-GPU staging framework. We present an experimental evaluation of its performance, portability, and scalability using this implementation on three leadership GPU clusters. The evaluation results demonstrate that the dual channel dual staging method saves up to 75% in data-exchange time compared to host-based, GDR, and alternate portable designs, while maintaining scalability (up to 512 GPUs) and performance portability across the three platforms.

Zhang, Bo [University of Utah]↗

Portable aircraft controller devices and systems

A portable computerized device for an aircraft control system includes an input system for inputting commands, a device display for displaying information on the computerized device, a processor, a wireless communication module, and a non-transitory computer readable medium comprising computer executable instructions, the computer executable instructions configured to cause the processor to perform a method. The method can include detecting whether the portable computerized device is in a cockpit state such that the portable computerized device is in and/or docked to an aircraft cockpit or if the portable computerized device is in a remote state such that the portable computerized device is not in an aircraft cockpit or is not docked to an aircraft cockpit. If the portable computerized device is determined to be in a remote state, the method includes operating the remote device in a remote mode. If the portable computerized device is determined to be in a cockpit state, the method includes operating the device in a local mode.

Sahay, Prateek↗

CHARM-SYCL & IRIS: A Tool Chain for Performance Portability on Extremely Heterogeneous Systems

Performance portability is becoming crucial as high-performance computing systems become increasingly heterogeneous. We have many options for CPUs and accelerators (e.g., GPUs) but also for non-Von Neumann architectures such as field-programmable gate arrays. This paper presents the CHARM-SYCL unified programming environment for multiple accelerator types as a performance-portable programming environment. It uses the IRIS library developed at Oak Ridge National Laboratory as the back end accelerator runtime. IRIS has a high-performance scheduler to distribute tasks across accelerators. This design allows us to run an application from the same source on multiple systems with multiple configurations. We provide three types of portability with CHARM-SYCL: Portable Workflow, Compiler and Runtime Portability, and Application and Performance Portability. We implement a Monte Carlo simulation benchmark code on the CHARM-SYCL execution environment and demonstrate that our programming environment can accommodate extremely heterogeneous systems.

Fujita, Norihisa↗

Grating-Based Imaging-Scattering with Portable Neutron Generator

Company: Adelphi Technology, Inc. Title: Grating-Based Imaging-Scattering with Portable Neutron Generator PI: Dr. Jay Theodore Cremer, Jr. Topic: 26a Statement of the problem or situation that is being addressed. Successful plant growth depends upon an efficient and robust root system. The plant root is part of a larger system of water and microbial flows in the soil system. While much effort has been exerted to develop an imaging system for water, microbes, and roots, the problem is challenging, and no widely accepted imaging method currently exists. The optical solutions use a highly modified soil system. X-ray imaging methods are insensitive to the soft tissues in the presence of sand. Thermal neutron imaging has been often tested, but found inadequate, due to limited access and low image resolution. This project will develop a new strategy for neutron imaging of plant/soil systems. The project will allow long duration experiments in greenhouse environments and increase the image information content to the micron scale. General statement of how this problem is being addressed. Portable, rugged thermal and fast neutron sources are being developed where portable means a two-soldier team can carry the source and power unit to survey rough terrain for explosives. In the past decade, microfabrication of X-ray and thermal/cold neutron optics has opened a new imaging strategy. The standard transmission image is now supplemented with simultaneous acquisition of a phase contrast image and an image revealing scattering features. In materials science, the interferometric neutron scattering image has been used to detect early crack formation in stressed additive manufacturing test samples. The detection requires sensitivity to scattering features at the 1-micron scale. By addition of our proposed grating-optic to Adelphi Technology’s radiographic/tomographic imaging system, which is based on portable thermal neutron source, the resulting thermal neutron scatter image of the plant/soil system, will reveal details at 1-micron. Commercial Applications and Other Benefits Neutron interferometry imaging has greater penetration through large metal components compared to industrial X-ray imaging. The low-cost, large area optics developed for greenhouse applications, combined with the robust, portable neutron generator, can then be marketed as a system for inspection of additive manufactured components. In the aerospace industry, all freshly printed components are validated with X-ray CT. Scheduled maintenance again requires X-ray CT as the ability to predict aerospace component lifetime does not yet exist. Large aerospace components are only partially observed with X-ray imaging. Key Words Neutron radiography/tomography, grating interferometry, thermal neutron generator imaging, plant root and soil imaging, rhizosphere imaging, deployable neutron imaging systems Summary for Members of Congress A rugged, portable source of thermal neutrons is adapted for neutron interferometry imaging with the addition of low-cost, 3D printed optics. The first application of our proposed deployable, compact thermal neutron generator imaging system, using a grating optic, is plant root/soil science in greenhouse settings and agricultural laboratories.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Neutron Transmission Imaging with a Portable D-T Neutron Generator

Fast-neutron transmission imaging provides complementary information to x-ray transmission imaging. While fast neutron imaging resolution is generally below x-ray imaging, 14-MeV neutrons have an advantage over portable x-ray systems. Neutrons have higher transmission through high-Z materials due to a more uniform attenuation as a function of material atomic number Z compared to X-rays, and can therefore image low-Z materials inside high-Z materials. As a result, portable neutron transmission imaging has many applications, including inspection of concrete and welds for corrosion in vehicles, bridges, and other infrastructure, measurement of material levels in containers, and inspection of suspicious packages. Fast-neutron imaging is also more practical for field use than thermal-neutron imaging due to the size and shielding requirements typical of thermal-imaging systems compared to the availability of small 14.1 MeV D-T neutron generators. However, there are limitations in portable fast-neutron imaging systems, including limited neutron output, limited light produced by neutron scintillators, and lower resolution due to neutron source spot size and 2-3 mm scintillator thickness. In addition, digital-panel dark-noise is roughly 100x higher than neutron scintillator light, and variations in noise across the panel and in time is comparable to the imaging signal. Here we discuss recent efforts in developing a portable fast-neutron radiography system, including an improved neutron scintillator, mitigation of panel noise, and new commercial portable D-T neutron generators. We also present MCNP efforts to model neutron imaging, including scintillator resolution and the effects of neutron scattering from the object and surrounding materials.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Making Uintah Performance Portable for Department of Energy Exascale Testbeds

To help ease ports to forthcoming Department of Energy (DOE) exascale systems, testbeds have been made available to select users. These testbeds are helpful for preparing codes to run on the same hardware and similar software as in their respective exascale systems. This paper describes how the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system, has been modified to be performance portable across the DOE Crusher, DOE Polaris, and DOE Sunspot testbeds in preparation for portable simulations across the exascale DOE Frontier and DOE Aurora systems. The Crusher, Polaris, and Sunspot testbeds feature the AMD MI250X, NVIDIA A100, and Intel PVC GPUs, respectively. This performance portability has been made possible by extending Uintah’s intermediate portability layer [18] to additionally support the Kokkos::HIP, Kokkos::OpenMPTarget, and Kokkos::SYCL back-ends. This paper also describes notable updates to Uintah’s support for Kokkos, which were required to make this extension possible. Results are shown for a challenging radiative heat transfer calculation, central to the University of Utah’s predictive boiler simulations. These results demonstrate single-source portability across AMD-, NVIDIA-, and Intel-based GPUs using various Kokkos back-ends.

Holmen, John↗

Application of Portable Parallelization Strategies for GPUs on track reconstruction kernels

Utilizing the computational power of GPUs is one of the key ingredients to meet the computing challenges presented to the next generation of High-Energy Physics (HEP) experiments. Unlike CPUs, developing software for GPUs often involves using architecturespecific programming languages promoted by the GPU vendors and hence limits the platform that the code can run on. Various portability solutions have been developed to achieve portable, performant software across different GPU vendors. Given the rapid evolution of these portability solutions, an early adoption of them in simple HEP testbed applications will help us understand the strengths and weaknesses of respective approaches.We apply several portability solutions, including Alpaka, Kokkos, SYCL and std::execution::par, on kernels for track propagation extracted from the mkFit project. We report on the development experience of the same application with different portability solutions, as well as their performance on GPUs, measured as the throughput of the kernels, from different manufacturers such as NVIDIA, AMD and Intel.

Kwok, Martin [Fermilab] (ORCID:0000000286936146)↗