Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “offloading”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

DoCeph: DPU-Offloaded Messaging in Ceph for Reduced Host CPU Utilization

Ceph is a widely used distributed object store, but its messenger layer imposes substantial CPU overhead on the host. To address this limitation, we propose DoCeph, a DPU-offloaded storage architecture for Ceph that disaggregates the system by offloading the communication-intensive messaging component to the DPU while retaining the storage backend on the host. The DPU efficiently manages communication, using lightweight RPC for metadata operations and DMA for data transfer. Moreover, DoCeph introduces a pipelining technique that overlaps data transmission with buffer preparation, mitigating hardware-imposed transfer size limitations. We implemented DoCeph on a Ceph cluster with NVIDIA BlueField-3 DPUs. Evaluation results indicate that DoCeph cuts host CPU usage by up to 92% while sustaining stable throughput and providing larger performance benefits for object writes over 1 MB.

Park, Kuri [Sogang University]↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

Offloading techniques for large deployable space structures

The validation and verification of large deployable space structures are continual challenges which face the integration and test engineer today. Spar Aerospace Limited has worked on various programs in which such structure validation was required and faces similar tasks in the future. This testing is reported and the different offloading and deployment methods which were used, as well as the proposed methods which will be used on future programs, are described. Past programs discussed include the Olympus solar array ambient and thermal vacuum deployments, and the Anik-E array and reflector deployments. The proposed MSAT reflector and boom ambient deployment tests, as well as the proposed RADARSAT Synthetic Aperture Radar (SAR) ambient and thermal vacuum deployment tests will also be presented. A series of tests relating to various component parts of the offloading equipment systems was required. These tests included the characterization and understanding of linear bearings and large (180 in-lbf) constant force spring motors in a thermal vacuum environment, and the results from these tests are presented.

Caravaggio, Levino↗

Test Frame for Gravity Offload Systems

Advances in space telescope and aperture technology have created a need to launch larger structures into space. Traditional truss structures will be too heavy and bulky to be effectively used in the next generation of space-based structures. Large deployable structures are a possible solution. By packaging deployable trusses, the cargo volume of these large structures greatly decreases. The ultimate goal is to three dimensionally measure a boom's deployment in simulated microgravity. This project outlines the construction of the test frame that supports a gravity offload system. The test frame is stable enough to hold the gravity offload system and does not interfere with deployment of, or vibrations in, the deployable test boom. The natural frequencies and stability of the frame were engineered in FEMAP. The test frame was developed to have natural frequencies that would not match the first two modes of the deployable beam. The frame was then modeled in Solidworks and constructed. The test frame constructed is a stable base to perform studies on deployable structures.

Murray, Alexander R.↗

Active Response Gravity Offload System

The Active Response Gravity Offload System (ARGOS) provides the ability to simulate with one system the gravity effect of planets, moons, comets, asteroids, and microgravity, where the gravity is less than Earth fs gravity. The system works by providing a constant force offload through an overhead hoist system and horizontal motion through a rail and trolley system. The facility covers a 20 by 40-ft (approximately equals 6.1 by 12.2m) horizontal area with 15 ft (approximately equals4.6 m) of lifting vertical range.

Valle, Paul↗

Development and Evaluation of the Active Response Gravity Offload System as a Lunar and Martian EVA Simulation Environment

In preparation for future exploration missions, NASA seeks the ability to simulate partial-gravity operations for use in ground-based research, crew training, and engineering design evaluations. The Active Response Gravity Offload System (ARGOS) at the Johnson Space Center (JSC) is designed to simulate reduced gravity environments, such as lunar, Martian, or microgravity, using a robotic system similar to an overhead bridge crane. ARGOS continuously offloads a portion of a suited human’s weight during all dynamic motions within the test facility, which can include basic functional movements such as walking, running, and jumping, as well as a wide range of planetary surface activities. This system will be used as part of a metabolic-rate task characterization study to determine the workload associated with partial-gravity extravehicular activity (EVA). Pilot testing was conducted using the MKIII prototype planetary space suit and two gimbal designs to determine the ability of the ARGOS test environment to simulate planetary EVA operations. This paper will describe the lessons learned from the feasibility testing, simulation-environment mockup design, and the results from the pilot tests and their influence into the final study design. Being able to effectively simulate partial-gravity environments and characterize the performance of crewmembers will have an impact on multiple domains including suit design, task design, thermal models, and life-support-system capacity verification plans, among others.

Omar S Bekdash↗

LSMS–L35, Miniature Crane for Payload Offloading and Manipulation: Development, and Application

The Lightweight Surface Manipulation System (LSMS) is a robotic agent for autonomous surface construction activities on planetary surfaces, that was designed at NASA Langley Research Center and has over a decade of research and development. The LSMS is a key component to achieving many goals of the NASA Artemis program. The LSMS is lightweight, structurally efficient system that can be easily packaged for launch and deployment on-surface, capable of a suite of surface activities enabled by modular end-effectors at the wrist. The focus of recent development work has been on using the LSMS for payload offloading and handling from lunar landers. Discussed in the paper is the development of the LSMS-L35 hardware (35 kg wrist lifting capacity on the lunar surface), designed to integrate with a Commercial Lunar Payload Services (CLPS) lander to offload payloads to the surface. The LSMS-L35 hardware development is part of a larger effort to enable autonomous payload handling and manipulation.

Iok M. Wong↗

LANDO: Developing Autonomous Payload Offloading Capabilities for Lunar Surface Operations

Introduction: The Lightweight Surface Manipulation System (LSMS) AutoNomy capabilities Development for surface Operations and construction (LANDO) project is an Early Career Initiative selected for funding by NASA Space Technology Mission Directorate. LANDO is developing a general-purpose autonomy framework applicable to serial and tension-actuated manipulation agents, that will be validated using an existing prototype of the LSMS-L35 (35-kg wrist lift capacity on the lunar surface [Fig. 1], sized for a Commercial Lunar Pay-load Services (CLPS) mission). The autonomous LSMS-L35 will be used to demonstrate autonomous payload handling capabilities for Lunar and other planetary surfaces, directly addressing STMD capability gaps in autonomous excavation and construction operations, advanced robotics and spacecraft autonomy technologies, and technologies supporting emerging space industries including the In-Space Servicing, Assembly and Manufacturing national strategy. LSMS: The LSMS is a tension actuated robotic agent that is scalable (reach and lifting capacity in different gravity environments), versatile (types of surface operations), and reusable. Compared to serial arms, the LSMS provides significantly higher structural efficiency and mechanical advantage, enabling a greater payload lift capacity at a lower system mass. The LSMS is envisioned to be a crucial part of the excavation and construction portfolio, capable of supporting a variety of activities on the lunar surface. Autonomous payload handling is one of the first activities the LSMS can support that develops capabilities that are extensible to other surface operations. Payload handling is required to: remove payloads from a lander; place payloads on mobile agents for transport from a lander to construction site/assembly point; emplace payloads in their operational configuration, and aggregate components to create an asset. As an example of this critical gap, manifested CLPS missions do not currently have a ubiquitous payload offloading capability; payloads (excluding rovers) are designed to remain on the lander. Why Autonomy? Autonomous robotic systems capable of carrying out excavation and construction operations are a fundamental and critical capability required for realizing the NASA Artemis program vision to “emplace and build the infrastructure, systems, and robotic missions that can enable a sustained lunar surface presence.” While teleoperation is still feasible for lunar surface operations, increased latency at Mars will require validated supervised autonomous technologies capable of operating with minimal human involvement (human-on-the-loop) unless an unexpected event occurs requiring human intervention. Autonomy reduces operator burden, allows operations to continue during uncrewed periods, in-creases the safety of operations by automatically detecting and handling faults, and allows operating in high latency environments. Development Activities: LANDO is extending critical autonomous operations to the manipulation domain and creating an integrated system, based on reusable software modules, that is capable of planning and executing payload handling and autonomous surface operations without requiring hu-man intervention beyond a supervisory role. The priority features under development are 1) autonomously offload payloads from a tilted lander deck without buckling the LSMS; 2) sensing whether a payload is safe to lift and handle; and 3) integrate with Astrobotic’s CLPS lander. The poster presentation will highlight current development activities over the past year on LSMS-L35 prototype hard-ware design, and autonomy software.

in-space assembly↗

ATHLETE Offloader Limb as a High-capacity Crane

A new concept for the NASA Jet Propulsion Laboratory (JPL) All-Terrain Hex-Limbed Extra-Terrestrial Explorer (ATHLETE) robotic constructor / mobility system employs tendon-driven actuation of individual limbs, similar to high-capacity cranes in terrestrial work environments, that can offload cargo from tall landers. While maintaining mechanical joints that allow each limb to function as a highly dexterous multi-Degree-Of-Freedom (DOF) robotic arm, tendon-driven truss sections increase the capacity of moment loads for extended configurations of the limb. The tendon-driven crane-like limb will be capable of offloading cargo from tall landers such as the SpaceX Starship. This paper describes the structural analysis, operations, calibration, and actuation of limbs designed to address specific target load cases that might be required for human exploration missions on planetary surfaces.

Wilcox, Brian↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

Queen bees offload pesticide burden to eggs when social buffering is overwhelmed

Honey bee colonies pollinate about one-third of the world’s food crops, and their rapid decline directly threatens agricultural productivity and ecosystem stability. Understanding how colony-level social defenses influence pesticide fate and the circumstances under which they fail is therefore a crucial question in pollinator biology. We used biological accelerator mass spectrometry (BioAMS), a sensitive radiotracer technique, to track the movement of a model pesticide through a small honey bee colony under laboratory conditions. We tested the hypothesis that social buffering protects honey bees from toxic accumulation and that this protection can be overcome, leading to maternal offloading of the pesticide to developing eggs. Consistent with this hypothesis, our results identified three key mechanisms governing chemical movement within a social insect colony: (1) worker bees initially decrease dietary pesticide levels by 95% through diet filtering and deposition in honeycombs, though this declines to 86% by day 10; (2) queen bees maintain markedly lower pesticide levels than workers but, over time, they accumulate the pesticide in their ovaries and transfer it into developing eggs, revealing a previously undocumented protective mechanism in reproductive individuals; and (3) the presence of a queen bee shifts colony-wide chemical distribution by concentrating worker exposure and increasing pesticide deposition in wax. Our findings show that honey bee colonies function as integrated detoxification networks, in which chemical fate depends on complex social behaviors and caste-specific physiology. When social buffering is overwhelmed, reproductive queens may survive by transferring their chemical burden to their offspring.

Biological and medical sciences↗

An OpenMP GPU-offload implementation of a non-equilibrium solidification cellular automata model for additive manufacturing

Here, in this paper, performance strategies on GPU-based HPC platforms of a cellular automata (CA) simulation code for non-equilibrium solidification, including nucleation, grain growth, solute partitioning and transport for the metal additive manufacturing (AM) process are investigated using OpenMP 4.5. To accurately report the speed-up for multicore CPUs and GPUs, a rigorous performance analysis employed optimizations appropriate for both CPU-only code (baseline) and GPU offload codes for an isothermal test problem. The performance results on Summit at the Oak Ridge Leadership Computing Facility indicate that using a precomputed list of interface cells significantly decreased the wall-clock time on GPUs. The speedup due to GPU acceleration was evaluated for a full Summit node and measured to be 1.8X when comparing a 6 MPI tasks run with 6 GPUs versus 36 MPI tasks on the CPU only. That speed-up was found to be 7.9X when comparing 6 MPI tasks with 6 GPUs versus the 6 MPI tasks running on the CPU only. Performance measurements showed that system total time is almost constant for runs with more than 96 MPI tasks (or GPUs), indicating that the GPU-accelerated code showed an excellent weak scaling performance. Finally, a rapid directional solidification problem was considered to demonstrate the CA code capability on Summit. It was found that a mesh size of at least 0.05 μm is recommended for the AM-like simulations in order to obtain accurate elongated grain microstructure and elongated subgrain features, which are in qualitative good agreement with experimental data. The results presented in this study indicate that the performance strategies on GPU-based HPC platforms for the CA code are appropriate for novel HPC exascale platforms.

36 MATERIALS SCIENCE↗

Porting Fragmentation Methods to Graphical Processing Units Using an OpenMP Application Programming Interface: Offloading the Fock Build for Low Angular Momentum Functions

Here, a framework to offload four-index two-electron repulsion integrals to graphical processing units (GPUs) using OpenMP is discussed. The method has been applied to the Fock build for low angular momentum s and p functions in both the restricted Hartree–Fock (RHF) and in the effective fragment molecular orbital (EFMO) framework. Benchmark calculations for the GPU code for the pure RHF method show an increasing speedup relative to the existing OpenMP CPU code in GAMESS from 1.04 to 52× for clusters of 70–569 water molecules. The parallel efficiency on 24 NVIDIA V100 GPU boards also increases when increasing the system size: from 75 to 94% for water clusters that contain 303–1120 molecules. In the EFMO framework, the GPU Fock build shows a high linear scalability up to 4608 V100s with a parallel efficiency of 96% for calculations on a solvated mesoporous silica nanoparticle system with ~67,000 basis functions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee

Currently, the Weather Research and Forecasting model (WRF) utilizes shared memory (OpenMP) and distributed memory (MPI) parallelisms. To take advantage of GPU resources on the Perlmutter supercomputer at NERSC, we port parts of the computationally expensive routine Fast Spectral Bin Microphysics (FSBM) to NVIDIA GPUs using OpenMP device offloading directives. To facilitate this process, we explore a workflow for optimization which uses both runtime profilers and a static code inspection tool Codee to refactor the subroutine. We observe an 2.24x overall speedup for the CONUS-12km storm test case.

Wichitrnithed, Chayanon (Namo) [Odin Institute]↗

Design of a Single Layer Metamaterial for Pressure Offloading of Transtibial Amputees

While using a prosthesis, transtibial amputees may experience pain and discomfort brought on by large pressure gradients at the interface between the residual limb and prosthetic socket. Current prosthetic interface solutions attempt to alleviate these pressure gradients by using soft homogenous liners to reduce and distribute pressures. This research investigates an additively manufactured metamaterial inlay with a tailored mechanical response to reduce peak pressure gradients around the limb. The inlay uses a hyperelastic behaving metamaterial (US10244818) comprised of triangular pattern unit cells which can be 3D printed with walls of various thicknesses controlled by draft angles. The hyperelastic material properties are modeled using a Yeoh 3rd order model. The 3rd order coefficients can be adjusted and optimized, which corresponds to a change in the unit cell wall thickness to create an inlay that can meet the unique offloading needs of an amputee. Finite element analysis simulations evaluated the pressure gradient reduction from: 1) a common homogenous silicone liner, 2) a prosthetist's inlay prescription that utilizes three variations of the metamaterial, and 3) a metamaterial solution with optimized Yeoh 3rd order coefficients. When compared to a traditional homogenous silicone liner for two unique limb loading scenarios, the prosthetist prescribed inlay and the optimized material inlay can achieve equal or greater pressure gradient reduction capabilities. These results show the potential feasibility of implementing this metamaterial as a method of personalized medicine for transtibial amputees by creating a customizable interface solution to meet the unique performance needs of an individual patient.

60 APPLIED LIFE SCIENCES↗

Computational Offload with BlueField Smart NICs

The recent introduction of a new generation of "smart NICs" have provided new accelerator platforms that include CPU cores or reconfigurable fabric in addition to traditional networking hardware and packet offloading capabilities. While there are currently several proposals for using these smartNICs for low-latency, in-line packet processing operations, there remains a gap in knowledge as to how they might be used as computational accelerators for traditional high-performance applications. This work aims to look at benchmarks and mini-applications to evaluate possible benefits of using a smartNIC as a compute accelerator for HPC applications. We investigate NVIDIA's current-generation BlueField-2 card, which includes eight Arm CPUs along with a small amount of storage, and we test the networking and data movement performance of these cards compared to a standard Intel server host. We then detail how two different applications, YASK and miniMD can be modified to make more efficient use of the BlueField-2 device with a focus on overlapping computation and communication for operations like neighbor building and halo exchanges. Our results show that while the overall compute performance of these devices is limited, using them with a modified miniMD algorithm allows for potential speedups of 5 to 20% over the host CPU baseline with no loss in simulation accuracy.

97 MATHEMATICS AND COMPUTING↗