Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computation offloading”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Performance Analysis of Data Processing in Distributed File Systems with Near Data Processing

In the era of big data, the escalating volume and velocity of data generation pose significant challenges in data processing. Traditional systems like Spark and Hadoop manage the increasing amount and velocity of data by improving data placement and processing speeds. However, they face inherent limitations due to the essential data movement required for processing. In this paper, we explore the Skyhook framework, a novel extension of the Ceph distributed system, which significantly reduces the need for data movement. We present an extensive case study using the Skyhook framework, applying it with the TPC-H and K-means clustering algorithms. More specifically, we leverage the TPC-H benchmark to distinguish between CPU-intensive and I/O-intensive tasks. We explore the integration of K-means clustering into SQL, coupled with a near-data processing system to offload the computational burden of the K-means clustering algorithm to storage nodes. We conduct a comprehensive performance evaluation of distributed data processing applications across three processing approaches: traditional layout (baseline), optimized layout, and near-data processing. Additionally, we introduce the use of the FIO tool to simulate real-world system workloads, enabling the measurement of performance metrics such as average latency and CPU utilization. Our research is a significant advance in understanding how to optimize data processing systems to meet the demands of the modern data landscape.

Hou, Shiyue↗

An Offload NIC for NASA, NLR, and Grid Computing

This work addresses distributed data management and access dynamically configurable high-speed access to data distributed and shared over wide-area high-speed network environments. An offload engine NIC (network interface card) is proposed that scales at nX10-Gbps increments through 100-Gbps full duplex. The Globus de facto standard was used in projects requiring secure, robust, high-speed bulk data transport. Novel extension mechanisms were derived that will combine these technologies for use by GridFTP, bandwidth management resources, and host CPU (central processing unit) acceleration. The result will be wire-rate encrypted Globus grid data transactions through offload for splintering, encryption, and compression. As the need for greater network bandwidth increases, there is an inherent need for faster CPUs. The best way to accelerate CPUs is through a network acceleration engine. Grid computing data transfers for the Globus tool set did not have wire-rate encryption or compression. Existing technology cannot keep pace with the greater bandwidths of backplane and network connections. Present offload engines with ports to Ethernet are 32 to 40 Gbps f-d at best. The best of ultra-high-speed offload engines use expensive ASICs (application specific integrated circuits) or NPUs (network processing units). The present state of the art also includes bonding and the use of multiple NICs that are also in the planning stages for future portability to ASICs and software to accommodate data rates at 100 Gbps. The remaining industry solutions are for carrier-grade equipment manufacturers, with costly line cards having multiples of 10-Gbps ports, or 100-Gbps ports such as CFP modules that interface to costly ASICs and related circuitry. All of the existing solutions vary in configuration based on requirements of the host, motherboard, or carriergrade equipment. The purpose of the innovation is to eliminate data bottlenecks within cluster, grid, and cloud computing systems, and to add several more capabilities while reducing space consumption and cost. Provisions were designed for interoperability with systems used in the NASA HEC (High-End Computing) program. The new acceleration engine consists of state-ofthe- art FPGA (field-programmable gate array) core IP, C, and Verilog code; novel communication protocol; and extensions to the Globus structure. The engine provides the functions of network acceleration, encryption, compression, packet-ordering, and security added to Globus grid or for cloud data transfer. This system is scalable in nX10-Gbps increments through 100-Gbps f-d. It can be interfaced to industry-standard system-side or network-side devices or core IP in increments of 10 GigE, scaling to provide IEEE 40/100 GigE compliance.

Awrach, James↗

Runtime performance of a GAMESS quantum chemistry application offloaded to GPUs

Summary Computational chemistry is at the forefront of solving urgent societal problems, such as polymer upcycling and carbon capture. The complexity of modeling these processes at appropriate length and time scales is mainly manifested in the number and types of chemical species involved in the reactions and may require models of several thousand atoms and large basis sets to accurately capture the chemical complexity and heterogeneity in the physical and chemical processes. The quantum chemistry package General Atomic and Molecular Electronic Structure System (GAMESS) has a wide array of methods that can efficiently and accurately treat complex chemical systems. In this work, we have used the GAMESS Effective Fragment Molecule Orbital (EFMO) method for electronic structure calculation of a challenging mesoporous silica nanoparticle (MSN) model surrounded by about 4700 water molecules to investigate the strong scaling and GPU offloading on hybrid CPU‐GPU nodes. Experiments were performed on the Perlmutter platform at the National Energy Research Scientific Computing Center. Good strong scaling and load balancing have been observed on up to 88 hybrid nodes for different settings of the execution parameters for the calculation considered here. When GPUs are oversubscribed by offloading work from multiple CPU processes, using the NVIDIA multi‐process service (MPS) has consistently reduced time to solution and energy consumed. Additionally, for some configuration parameter settings, oversubscription with MPS improved performance by up to 5.8% over the case without oversubscription.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Performance Characteristics of the BlueField-2 SmartNIC

High-performance computing (HPC) researchers have long envisioned scenarios where application workflows could be improved through the use of programmable processing elements embedded in the network fabric. Recently, vendors have introduced programmable Smart Network Interface Cards (SmartNICs) that enable computations to be offloaded to the edge of the network. There is great interest in both the HPC and high-performance data analytics (HPDA) communities in understanding the roles these devices may play in the data paths of upcoming systems. This paper focuses on characterizing both the networking and computing aspects of NVIDIA’s new BlueField-2 SmartNIC when used in a 100Gb/s Ethernet environment. For the networking evaluation we conducted multiple transfer experiments between processors located at the host, the SmartNIC, and a remote host. These tests illuminate how much effort is required to saturate the network and help estimate the processing headroom available on the SmartNIC during transfers. For the computing evaluation we used the stress-ng benchmark to compare the BlueField-2 to other servers and place realistic bounds on the types of offload operations that are appropriate for the hardware. Our findings from this work indicate that while the BlueField-2 provides a flexible means of processing data at the network’s edge, great care must be taken to not overwhelm the hardware. While the host can easily saturate the network link, the SmartNIC’s embedded processors may not have enough computing resources to sustain more than half the expected bandwidth when using kernel-space packet processing. From a computational perspective, encryption operations, memory operations under contention, and on-card IPC operations on the SmartNIC perform significantly better than the general-purpose servers used for comparisons in our experiments. Therefore, applications that mainly focus on these operations may be good candidates for offloading to the SmartNIC.

97 MATHEMATICS AND COMPUTING↗

Collision Avoidance Approach Using Deep Reinforcement Learning

A method to enable autonomous robots moving in a 2D space collision free motivates the purposed approach for collision avoidance for autonomous UAM vehicles. Challenges of autonomous collision free navigation for both problems are similar. Agents in each environment do not know the intent, or goal, of the other. Finding the time efficient paths require some level of anticipation with neighboring agents which is computationally expensive. In the original work, these obstacles were overcome with a novel application of deep reinforcement learning which offloads the online computation to an offline learning algorithm. A value network that encodes the estimated time to the goal given the agent’s state and the observable portion of the other agent’s state is trained on a baseline policy and further refined with reinforcement learning to promote time efficient collision free navigation. Online, the value network efficiently informs the agent’s decision making in the face of uncertainty of the other agent’s next move. In this paper, challenges extending this methodology to the 3D environment of autonomous UAM vehicles with kinematic constraints are discussed and initial results shown.

Collision Avoidance↗

Collision Avoidance Approach Using Deep Reinforcement Learning

A method to enable autonomous robots moving in a 2D space collision free motivates the purposed approach for collision avoidance for autonomous UAM vehicles. Challenges of autonomous collision free navigation for both problems are similar. Agents in each environment do not know the intent, or goal, of the other. Finding the time efficient paths require some level of anticipation with neighboring agents which is computationally expensive. In the original work, these obstacles were overcome with a novel application of deep reinforcement learning which offloads the online computation to an offline learning algorithm. A value network that encodes the estimated time to the goal given the agent’s state and the observable portion of the other agent’s state is trained on a baseline policy and further refined with reinforcement learning to promote time efficient collision free navigation. Online, the value network efficiently informs the agent’s decision making in the face of uncertainty of the other agent’s next move. In this paper, challenges extending this methodology to the 3D environment of autonomous UAM vehicles with kinematic constraints are discussed and initial results shown.

Collision Avoidance↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

System Decommutes And Displays Telemetry Data

TDPlus computer program software system for decommutation of pulse-code-modulation (PCM) telemetry signals. Provides synchronization, conversion into engineering units, and display of serial bit streams. Transforms IBM PC-compatible computer into PCM-telemetry-decommutation system. Synchronizes telemetric signals data and enables conversion back into such meaningful forms as voltage, current, pressure, and the like. Also controls operation of digital-to-analog converters to ship data to paper strip charts or to parallel digital ports for offloading to other computers. Software used to process actual data only when telemetry-data-processing computer modified in accordance with specifications contained in "TDPlus TM Data Processor" (GSC-13291). Written in Turbo C and 8088 Assembly language.

Massey, D. E.↗

The Instrumented Walking and Turning Test to Evaluate Suited Gait Dynamics and Performance in Extravehicular Activity Training Environments

Background and aims: Walking will be required for many exploration tasks on the Moon during the Artemis program. Walking in a straight line on the confined floorspace of a testing area, and repetitive treadmill walking that requires no change in direction may not adequately reflect the balance and coordination required during ambulation. Also, performance of turning maneuvers may be affected differently in different extravehicular activity (EVA) training facilities that simulate partial gravity. For example, the Neutral Buoyancy Lab (NBL) simulates lunar gravity by adding weight to underwater subjects to alter buoyancy and achieve the equivalent ground reaction force of 1/6 of Earth’s gravity (1/6G), whereas the Active Response Gravity Offload System (ARGOS) uses a computer controlled overhead suspension system programmed to continuously offload a percentage of a subject’s weight to simulate 1/6G. The degree to which dynamic movements such as turning are comparable across these EVA training facilities has not yet been evaluated. The instrumented gait test helps NASA scientists and engineers evaluate gait dynamics and performance in suited conditions, and this test demonstrates the unique characteristics and limitations of EVA training facilities. We developed an instrumented walking and turning test using inertial measurement units (IMUs) and conducted the test at NASA’s EVA training facilities. Results were used to compare suited walking and turning characteristics in the ARGOS and the NBL. Methods: Subjects donned the Mark III space suit during offloading with the ARGOS spreader bar gimbal and donned the Z2.5 space suit while underwater in the NBL with weights and floatation added to achieve realistic suit center of gravity. The test team securely attached three Opal (APDM, OR, USA) wireless IMUs on the space suit for each test run: one on the middle of the hard upper torso, and one on the left and on the right ankle bearings. During the NBL tests, the IMUs were encased in a waterproof housing (GoPro) with foam added to create a tighter fit. At both testing facilities, 6.3 m x 1.0 m (LxW) walking lines were marked, and a cone for turning or walking around was located at the end of the walking path with another line on the other side of the cone to indicate the stopping point after walking around the cone. Under simulated 1/6G, subjects began by standing at the marked line with their arms folded across the chest, they then walked at a preferred speed along the straight walking path until they reached the end, turned 180 degrees around the cone, and finally stopped at the marked stopping point. All IMU data recorded during testing were automatically saved to the internal memory. Then, raw IMU signals were processed using custom MATLAB (Mathworks, MA, USA) code to compare gait parameters during both the walking and the turning components of the task. These parameters included time (s), speed (m/s for walking and rad/s for turning), step number (n) and walk:turn time ratio (% time spent straight walking versus turning). Results: Less time, faster gait, fewer steps, and higher walk:turn ratio during both walking and turning components were exhibited during tests performed at the ARGOS versus those performed at the NBL. During the NBL tests, the slower walking speed continued at the same rate throughout a U-shape turn. During the ARGOS tests, the subjects performed shorter and tighter turns at 4 times the speed of the NBL turns because they walked 30% faster and the vertical offloading system gave them more support. Conclusion: Our data show that the differences in walking and turning parameters during the NBL tests may be due to the high viscosity in the water environment where the motion of the lower limbs was slow and did not reach full flexion and extension. These tests improve the current knowledge of testing environments in preparation for EVAs on the lunar surface.

Kyoung Jae Kim↗

GPU acceleration of hybrid functional calculations in the SPARC electronic structure code

We present a Graphics Processing Unit (GPU)-accelerated version of the real-space SPARC electronic structure code for performing hybrid functional calculations in generalized Kohn–Sham density functional theory. In particular, we develop a batch variant of the recently formulated Kronecker product-based linear solver for the simultaneous solution of multiple linear systems. We then develop a modular, math kernel based implementation for hybrid functionals on NVIDIA architectures, where computationally intensive operations are offloaded to the GPUs, while the remaining workload is handled by the central processing units (CPUs). Considering bulk and slab examples, we demonstrate that GPUs enable up to 8× speedup in node-hours and 80× in core-hours compared to CPU-only execution, reducing the time to solution on V100 GPUs to around 300 s for a metallic system with over 6000 electrons, and significantly reducing the computational resources required for a given wall time.

Kohn-Sham density functional theory↗

Numerical eigen-spectrum slicing, accurate orthogonal eigen-basis, and mixed-precision eigenvalue refinement using OpenMP data-dependent tasks and accelerator offload

Performing a variety of numerical computations efficiently and, at the same time, in a portable fashion requires both an overarching design followed by a number of implementation strategies. All of these are exemplified below as we present transitioning the PLASMA numerical library from relying on dependence-driven large tasks to achieving utilization of fine grain tasking and offload to hardware accelerators while keeping its core dependence sets: OpenMP source code pragmas and runtime for most system-level functionality and basic low-level numerical kernels provided directly by hardware vendors or open source projects with vendor contributions. We also present new algorithmic methods and their efficient parallel implementations including fine grained tasking for eigen-spectrum slicing and offload for mixed-precision eigenvalue refinement. We provide performance, scaling, and numerical results showing sizable gains over the available solutions from either the open source and vendor-provided packages.

Luszczek, Piotr↗

The Transient Multi-Level method for Monte Carlo reactor statics calculations

The Transient Multi-Level (TML) method is applied to a time-dependent Monte Carlo transport solver to offload some of the computational burden of the expensive Monte Carlo solve to lower-order Coarse Mesh Finite Difference (CMFD) and Exact Point Kinetics Equations (EPKE) solvers via factorization of the neutron flux at the transport and CMFD levels using the Predictor Corrector Quasi-Static Method (PCQM). The Monte Carlo transient is solved by a modified fission source iteration scheme that introduces a single transient source bank. The method is implemented in the production-level Monte Carlo code, Shift, and verified with prescribed reactivity ramps from the two-dimensional version of the C5G7-TD reactor benchmark. The results show that, as compared to other quasi-static methods, the TML reduces the stochastic noise inherent to the transient Monte Carlo solver by factors of ~2 to 6 for various norm comparisons of the reactor power amplitude. Finally, the TML additionally reduces the number of Monte Carlo evaluations needed to simulate the transient, leading to roughly an order of magnitude improvement in CPU time relative to the standard PCQM for the problems tested.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Methods of downloading to user institutions

The Pilot Climate Data System (PCDS) not only supports the ability to output data in a uniform structure via the Common Data Format (CDF, but also supports the ability to provide data in native format for any data set supported by the PCDS. Methods are discussed for acquiring data in either format from the PCDS for further work at remote sites. Four levels of remote utilization are defined, based on the extent of offloading the National Space Science Data Center (NSSDC) computer and local PCDS processing. Characteristics of each level are thoroughly explained, including details of information and data transfers, downloading, uploading, and offloading of the NSSDC computer. Only the levels themselves are specified here. The first level defined is that of a network-based distributed PCDS. A subset of the PCDS software is ported to another VAX and made available on a network (i.e., SPAN) node. There is no subset of the PCDS at the second level, but it is also network based. Non-network utilization of the PCDS, requiring dial-up log on, is denoted as a third level. Finally, at the fourth level, personal computer utilization of the PCDS through dial-up log on with proper terminal emulation is defined.

Treinish, L.↗

Smart Payload Development for High Data Rate Instrument Systems

This slide presentation reviews the development of smart payloads instruments systems with high data rates. On-board computation has become a bottleneck for advanced science instrument and engineering capabilities. In order to improve the computation capability on board, smart payloads have been proposed. A smart payload is a Localized instrument, that can offload the flight processor of extensive computing cycles, simplify the interfaces, and minimize the dependency of the instrument on the flight system. This has been proposed for the Mars mission, Mars Atmospheric Trace Molecule Spectroscopy (MATMOS). The design of this system is discussed; the features of the Virtex-4, are discussed, and the technical approach is reviewed. The proposed Hybrid Field Programmable Gate Array (FPGA) technology has been shown to deliver breakthrough performance by tightly coupling hardware and software. Smart Payload designs for instruments such as MATMOS can meet science data return requirements with more competitive use of available on-board resources and can provide algorithm acceleration in hardware leading to implementation of better (more advanced) algorithms in on-board systems for improved science data return

Virtex-4↗

Final Technical Report for CMSC 838L

This paper describes Rahul Vishnoi’s final project supporting in his Graduate School curriculum CMSC 838L, Advanced Topics in Programming Languages and Computer Architecture. This project was selected to intersect with his work as a Pathways Intern supporting Code 583, the Ground Software Systems Branch, at NASA’s Goddard Space Flight Center (GSFC). In this project, Field Programmable Gate Array (FPGA) hardware from Xilinx is used to replace and offload processor and memory-intensive computations from a microcontroller/Processing System (PS) to the FPGA Programmable Logic (PL). An interface between the PL and PS in the form of a C library allows for this bridging of capability.

Microcontroller, FPGA, Embedded Development, Xilin↗

SMART: The Future of Spaceflight Avionics

A novel avionics approach is necessary to meet the future needs of low cost space and lunar missions that require low mass and low power electronics. The current state of the art for avionics systems are centralized electronic units that perform the required spacecraft functions. These electronic units are usually custom-designed for each application and the approach compels avionics designers to have in-depth system knowledge before design can commence. The overall design, development, test and evaluation (DDT&E) cycle for this conventional approach requires long delivery times for space flight electronics and is very expensive. The Small Multi-purpose Advanced Reconfigurable Technology (SMART) concept is currently being developed to overcome the limitations of traditional avionics design. The SMART concept is based upon two multi-functional modules that can be reconfigured to drive and sense a variety of mechanical and electrical components. The SMART units are key to a distributed avionics architecture whereby the modules are located close to or right at the desired application point. The drive module, SMART-D, receives commands from the main computer and controls the spacecraft mechanisms and devices with localized feedback. The sensor module, SMART-S, is used to sense the environmental sensors and offload local limit checking from the main computer. There are numerous benefits that are realized by implementing the SMART system. Localized sensor signal conditioning electronics reduces signal loss and overall wiring mass. Localized drive electronics increase control bandwidth and minimize time lags for critical functions. These benefits in-turn reduce the main processor overhead functions. Since SMART units are standard flight qualified units, DDT&E is reduced and system design can commence much earlier in the design cycle. Increased production scale lowers individual piece part cost and using standard modules also reduces non-recurring costs. The benefit list continues, but the overall message is already evident: the SMART concept is an evolution in spacecraft avionics. SMART devices have the potential to change the design paradigm for future satellites, spacecraft and even commercial applications.

Alhorn, Dean C.↗

Enabling Fortran Standard Parallelism in GAMESS for Accelerated Quantum Chemistry Calculations

The performance of Fortran 2008 DO CONCURRENT (DC) relative to OpenACC and OpenMP target offloading (OTO) with different compilers is studied for the GAMESS quantum chemistry application. Specifically, DC and OTO are used to offload the Fock build, which is a computational bottleneck in most quantum chemistry codes, to GPUs. The DC Fock build performance is studied on NVIDIA A100 and V100 accelerators and compared with the OTO versions compiled by the NVIDIA HPC, IBM XL, and Cray Fortran compilers. The results show that DC can speed up the Fock build by 3.0× compared with that of the OTO model. Finally, with similar offloading efforts, DC is a compelling programming model for offloading Fortran applications to GPUs.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗