Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computation offloading”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Quantum optimization algorithms: Energetic implications

Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. Here in this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least 2x to 4x more energy efficient than their classical counterparts; why feedback-based quantum optimization is energy-inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of ≥2 0 x. Finally, we use the NchooseK high-level programming model to run optimization problems on both gate-based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate-based counterparts.

97 MATHEMATICS AND COMPUTING↗

RISE: Reducing I/O Contention in Staging-based Extreme-Scale In-situ Workflows

While in-situ workflow formulations have addressed some of the data-related challenges associated with extreme-scale scientific workflows, these workflows involve complex interactions and different modes of data exchange. In the context of increasing system complexity, such workflows present significant resource management challenges, requiring complex cost-performance tradeoffs. This paper presents RISE, an intelligent staging-based data management middleware, which builds on the DataSpaces framework and performs intelligent scheduling of data management operations to reduce I/O contention. In RISE, data are always written immediately to local buffers to reduce the effect of the transfer impact upon application performance. RISE identifies applications’ data access patterns and moves data towards data consumers only when the network is expected to be idle, reducing the impact of asynchronous background data movement upon critical data read/write requests. Here, we experimentally demonstrate that RISE can take advantage of staging nodes to offload data during writes without degrading application data movement performance.

97 MATHEMATICS AND COMPUTING↗

thornado+FLASH-X: A Hybrid Discontinuous Galerkin–Implicit-explicit and Finite-volume Framework for Neutrino-radiation Hydrodynamics in Core-collapse Supernovae

We present neutrino-transport algorithms implemented in the toolkit for high-order neutrino-radiation hydrodynamics (thornado) and their coupling to self-gravitating hydrodynamics within the adaptive mesh refinement–based multiphysics simulation framework FLASH-X. thornado, developed primarily for simulations of core-collapse supernovae (CCSNe), employs a spectral, six-species two-moment formulation with algebraic closure and special-relativistic observer corrections accurate to $\mathcal{O}(v/c)$, and uses discontinuous Galerkin (DG) methods for phase-space discretization combined with implicit-explicit time stepping. A key development is a nonlinear neutrino–matter coupling algorithm based on nested fixed-point iteration with Anderson acceleration, enabling fully implicit treatment of collisional processes, including energy-coupling interactions such as neutrino–electron scattering and pair production. Coupling to finite-volume (FV) hydrodynamics is achieved through a hybrid DG-FV representation of the fluid variables and operator-split evolution within FLASH-X. The implementation is verified using basic transport tests with idealized opacities and relaxation and deleptonization problems with tabulated microphysics. Spherically symmetric CCSN simulations demonstrate accuracy and robustness of the coupled scheme, including close agreement with the CCSN simulation code Chimera. An axisymmetric CCSN simulation further demonstrates the viability of DG-based neutrino transport for multidimensional supernova modeling within FLASH-X. thornado’s neutrino-transport solver is GPU-enabled using OpenMP offloading or OpenACC, and all CCSN applications included in this work use the GPU implementation. Together, these results establish a foundation for future enhancements in physics fidelity, numerical algorithms, and computational performance, for increasingly realistic large-scale CCSN simulations.

Endeve, Eirik [Oak Ridge National Laboratory (ORNL↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

PYSEQM

PYSEQM is a package for performing semi-empirical quantum mechanical (SEQM) simulations on molecular systems utilizing PyTorch. SEQM simulations determine molecular properties (energy, electron density, dipole moment, ect.) by solving an approximate Schrödinger equation for the motions of electrons in a molecule. The use of PyTorch provides three specific advantages. First, it allows the calculations to be offloaded to GPU accelerators, giving an order of magnitude increase in speed. Second, back propagation is used to get atomic forces (derivative of the total energy with respect to atomic position) at the same computational cost as the energy calculation itself. Finally, the use of PyTorch makes for a natural interface to modern machine learning methods, which can be used to adjust the semi-empirical parameters build into SEQM methods. Additionally, PYSEQM implements various other optimizations for performing quantum mechanics based molecular dynamics, including SP2 for rapid GPU based solution of the self-consistent field algorithm, and the extended Lagrangian method for rapid QM-MD.

Nebgen, Benjamin↗

The Portals 4.3 Network Programming Interface

This report presents a specification for the Portals 4 network programming interface. Portals 4 is intended to allow scalable, high-performance network communication between nodes of a parallel computing system. Portals 4 is well suited to massively parallel processing and embedded systems. Portals 4 represents an adaption of the data movement layer developed for massively parallel processing platforms, such as the 4500-node Intel TeraFLOPS machine. Sandia's Cplant cluster project motivated the development of Version 3.0, which was later extended to Version 3.3 as part of the Cray Red Storm machine and XT line. Version 4 is targeted to the next generation of machines employing advanced network interface architectures that support enhanced offload capabilities.

97 MATHEMATICS AND COMPUTING↗

Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUs

Computational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications.

Thavappiragasam, Mathialakan↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

Celeritas: Accelerating Geant4 with GPUs

Celeritas [1] is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications (specifically, High Lumi- nosity Large Hadron Collider (HL-LHC) simulation) on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a GPU-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.3, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling integration of Celeritas into high energy physics (HEP) experimental frameworks such as CMSSW. On the Summit supercomputer, Celeritas performs EM physics between 6 and 32 faster using the machine’s Nvidia GPUs compared to using only CPUs. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.8–2.3× on GPU and 1.2× on CPU.

Johnson, Seth R.↗

Space station autonomy requirements

Several concepts of autonomy have emerged from space station technology and mission studies. Unifying these concepts is important to enabling the needs for and requirements of autonomous systems for a space station to be explored. One purpose of autonomy related to a space station is to offload routine, demonstrated, precisely specifiable tasks and functions from humans to machines in order to increase habitability, increase human-machine system productivity, or to decrease operational costs. Defining incremental roles, functions, and technical capabilities for autonomy leads to identification of computing systems technologies needed to enable various degrees of autonomy.

Anderson, J. L.↗

NASA Space Environment Analog for Training, Engineering, Science, and Technology (SEATEST) 6 Detailed Final Report

After more than 50 years since the last crewed lunar landing, plans for more missions to the moon are in development. For these missions, efficient and sustainable logistics will be critical. Additionally, innovative methods of cargo transfer to and from a lunar outpost should be considered for successfully establishing a permanent presence on the moon. SEATEST (Space Environment Analog for Training, Engineering, Science, and Technology) is an immersive mission-analogous operational atmosphere where buoyancy effects and supplemental weights can simulate partial gravity conditions similar to those astronauts will experience on the moon. SEATEST 6 took place at the University of Southern California (USC) Wrigley Marine Science Center on Santa Catalina Island from July 18-30, 2023. The analog was used to collect preliminary logistics data on two different offloading conceptual methods (a davit and a zipline) during a simulated lunar mission. Pre-test analysis indicated for a crew of two on a 14-day mission, approximately three Medium Pressurized Logistics Containers (MPLC) sized logistics containers (or a total of 37.5 single Cargo Transfer Bag Equivalents (CTBE)) would be needed to support a mission. A Computer-Aided Design (CAD) analysis was employed on the SEATEST airlock mockup to determine how many logistic containers would fit with two suited crewmembers, don/doff stands, and hatch operations. It was determined that for SEATEST, a total of 15 1.0 Small Pressurized Logistics Containers (SPLCs) and 8 2.0 SPLCs would adequately fit into the approximate 9.5 cubic meter airlock volume. This does not fully represent a complete 14-day logistic supply; however, it does provide a preliminary estimate to initiate design conversations between logistics teams and crew at this early stage of development. Data were collected in eight logistics transfer scenarios over two days with four scenarios per day. Five test subject crew participated in scenarios as pairs. Scenarios included two sizes of logistics containers – 1.0 SPLC (equivalent to a single Cargo Transfer Bag (CTB) and 2.0 SPLC (equivalent to two CTBs). Planed evaluations included the use of a logistics port compared to transfer through an Airlock hatch, offloading methods based on either a davit or a zipline system, choreography of cargo in the airlock to permit ingress and suit doffing, and dust removal protocols for an understanding of the overall impact to transfer ops. Data collected included objective data (task times for conducting overall tasks and subtasks, full audio/video of test activities, and inadvertent “dings” on hardware) and subjective data (crew consensus of: task acceptability and capability assessment ratings related to best practices, considerations, and constraints for EVA-driven logistics transfer ConOps, sim quality of the test environment, and more general debrief comments). The two logistic offloading transfer concepts (davit, zipline) presented both advantages and limitations. The davit’s flexibility in allowing the crew to pick up the containers without physical interaction was well regarded by the crew. Some limitations of select davit hardware components were noted, but the overall concept was acceptable. The zipline system proved to be the most efficient way of moving logistics from the lander to the airlock and eliminated the need for dust operations. However, extended and repetitive lifting of containers to the line could be fatiguing. In conclusion, logistics transfer could hypothetically be achieved without an offloading method; however, the time requirement for such operations would be prohibitive. Results of crew subjective feedback proposed a combined or hybrid davit/zipline method to increase efficiency.

Logistics↗

Systems Architecture for Fully Autonomous Space Missions

The NASA Goddard Space Flight Center is working to develop a revolutionary new system architecture concept in support of fully autonomous missions. As part of GSFC's contribution to the New Millenium Program (NMP) Space Technology 7 Autonomy and on-Board Processing (ST7-A) Concept Definition Study, the system incorporates the latest commercial Internet and software development ideas and extends them into NASA ground and space segment architectures. The unique challenges facing the exploration of remote and inaccessible locales and the need to incorporate corresponding autonomy technologies within reasonable cost necessitate the re-thinking of traditional mission architectures. A measure of the resiliency of this architecture in its application to a broad range of future autonomy missions will depend on its effectiveness in leveraging from commercial tools developed for the personal computer and Internet markets. Specialized test stations and supporting software come to past as spacecraft take advantage of the extensive tools and research investments of billion-dollar commercial ventures. The projected improvements of the Internet and supporting infrastructure go hand-in-hand with market pressures that provide continuity in research. By taking advantage of consumer-oriented methods and processes, space-flight missions will continue to leverage on investments tailored to provide better services at reduced cost. The application of ground and space segment architectures each based on Local Area Networks (LAN), the use of personal computer-based operating systems, and the execution of activities and operations through a Wide Area Network (Internet) enable a revolution in spacecraft mission formulation, implementation, and flight operations. Hardware and software design, development, integration, test, and flight operations are all tied-in closely to a common thread that enables the smooth transitioning between program phases. The application of commercial software development techniques lays the foundation for delivery of product-oriented flight software modules and models. Software can then be readily applied to support the on-board autonomy required for mission self-management. An on-board intelligent system, based on advanced scripting languages, facilitates the mission autonomy required to offload ground system resources, and enables the spacecraft to manage itself safely through an efficient and effective process of reactive planning, science data acquisition, synthesis, and transmission to the ground. Autonomous ground systems in turn coordinate and support schedule contact times with the spacecraft. Specific autonomy software modules on-board include mission and science planners, instrument and subsystem control, and fault tolerance response software, all residing within a distributed computing environment supported through the flight LAN. Autonomy also requires the minimization of human intervention between users on the ground and the spacecraft, and hence calls for the elimination of the traditional operations control center as a funnel for data manipulation. Basic goal-oriented commands are sent directly from the user to the spacecraft through a distributed internet-based payload operations "center". The ensuing architecture calls for the use of spacecraft as point extensions on the Internet. This paper will detail the system architecture implementation chosen to enable cost-effective autonomous missions with applicability to a broad range of conditions. It will define the structure needed for implementation of such missions, including software and hardware infrastructures. The overall architecture is then laid out as a common thread in the mission life cycle from formulation through implementation and flight operations.

Esper, Jamie↗

CCAMP: An Integrated Translation and Optimization Framework for OpenACC and OpenMP

Heterogeneous computing and exploration into specialized accelerators are inevitable in current and future supercomputers. Although this diversity of devices is promising for performance, the array of architectures presents programming challenges. High-level programming strategies have emerged to face these challenges, such as the OpenMP offloading model and OpenACC. The varying levels of support for these standards, however, within vendor-specific and open-source tools, as well as the lack of performance portability across devices, have prevented the standards from achieving their goals. To address these shortcomings, we present CCAMP, an OpenMP and OpenACC interoperable framework. CCAMP provides two primary facilities: language translation between the two standards and device-specific directive optimization within each standard. We show that by using the CCAMP framework, programmers can easily transplant non-portable code into new ecosystems for new architectures. Additionally, by using CCAMP device-specific directive optimizations, users can achieve optimized performance across architectures using a single source code.

Lambert, Jacob↗

COMPOFF: A Compiler Cost model using Machine Learning to predict the Cost of OpenMP Offloading

The HPC industry is inexorably moving towards an era of extremely heterogeneous architectures, with more devices configured on any given HPC platform and potentially more kinds of devices, some of them highly specialized. Writing a separate code suitable for each target system for a given HPC application is not practical. The better solution is to use directive-based parallel programming models such as OpenMP. OpenMP provides a number of options for offloading a piece of code to devices like GPUs. To select the best option from such options during compilation, most modern compilers use analytical models to estimate the cost of executing the original code and the different offloading code variants. Building such an analytical model for compilers is a difficult task that necessitates a lot of effort on the part of a compiler engineer. Recently, machine learning techniques have been successfully applied to build cost models for a variety of compiler optimization problems. In this paper, we present COMPOFF, a cost model which uses the multi-layer perceptrons to statically estimates the Cost of OpenMP OFFloading. We used six different transformations on a parallel code of Wilson Dslash Operator to support GPU offloading, and we predicted their cost of execution on different GPUs using COMPOFF during compile time. Our results show that this model can predict offloading costs with a root mean squared error in prediction of less than 0.5 seconds. Our preliminary findings indicate that this work will make it much easier and faster for scientists and compiler developers to port legacy HPC applications that use OpenMP to new heterogeneous computing environment.

97 MATHEMATICS AND COMPUTING↗

Application Experiences on a GPU-Accelerated Arm-based HPC Testbed

This paper assesses and reports the experience of ten teams working to port, validate, and benchmark several High Performance Computing applications on a novel GPU-accelerated Arm testbed system. The testbed consists of eight NVIDIA Arm HPC Developer Kit systems, each one equipped with a server-class Arm CPU from Ampere Computing and two data center GPUs from NVIDIA Corp. The systems are connected together using InfiniBand interconnect. The selected applications and mini-apps are written using several programming languages and use multiple accelerator-based programming models for GPUs such as CUDA, OpenACC, and OpenMP offloading. Working on application porting requires a robust and easy-to-access programming environment, including a variety of compilers and optimized scientific libraries. The goal of this work is to evaluate platform readiness and assess the effort required from developers to deploy well-established scientific workloads on current and future generation Arm-based GPU-accelerated HPC systems. The reported case studies demonstrate that the current level of maturity and diversity of software and tools is already adequate for large-scale production deployments.

Elwasif, Wael↗

Finite Element Analysis and Test Correlation of a 10-Meter Inflation-Deployed Solar Sail

Under the direction of the NASA In-Space Propulsion Technology Office, the team of L Garde, NASA Jet Propulsion Laboratory, Ball Aerospace, and NASA Langley Research Center has been developing a scalable solar sail configuration to address NASA's future space propulsion needs. Prior to a flight experiment of a full-scale solar sail, a comprehensive phased test plan is currently being implemented to advance the technology readiness level of the solar sail design. These tests consist of solar sail component, subsystem, and sub-scale system ground tests that simulate the vacuum and thermal conditions of the space environment. Recently, two solar sail test articles, a 7.4-m beam assembly subsystem test article and a 10-m four-quadrant solar sail system test article, were tested in vacuum conditions with a gravity-offload system to mitigate the effects of gravity. This paper presents the structural analyses simulating the ground tests and the correlation of the analyses with the test results. For programmatic risk reduction, a two-prong analysis approach was undertaken in which two separate teams independently developed computational models of the solar sail test articles using the finite element analysis software packages: NEiNastran and ABAQUS. This paper compares the pre-test and post-test analysis predictions from both software packages with the test data including load-deflection curves from static load tests, and vibration frequencies and mode shapes from vibration tests. The analysis predictions were in reasonable agreement with the test data. Factors that precluded better correlation of the analyses and the tests were uncertainties in the material properties, test conditions, and modeling assumptions used in the analyses.

Sleight, David W.↗

Structural Analysis of an Inflation-Deployed Solar Sail With Experimental Validation

Under the direction of the NASA In-Space Propulsion Technology Office, the team of L Garde, NASA Jet Propulsion Laboratory, Ball Aerospace, and NASA Langley Research Center has been developing a scalable solar sail configuration to address NASA s future space propulsion needs. Prior to a flight experiment of a full-scale solar sail, a comprehensive phased test plan is currently being implemented to advance the technology readiness level of the solar sail design. These tests consist of solar sail component, subsystem, and sub-scale system ground tests that simulate the vacuum and thermal conditions of the space environment. Recently, two solar sail test articles, a 7.4-m beam assembly subsystem test article and a 10-m four-quadrant solar sail system test article, were tested in vacuum conditions with a gravity-offload system to mitigate the effects of gravity. This paper presents the structural analyses simulating the ground tests and the correlation of the analyses with the test results. For programmatic risk reduction, a two-prong analysis approach was undertaken in which two separate teams independently developed computational models of the solar sail test articles using the finite element analysis software packages: NEiNastran and ABAQUS. This paper compares the pre-test and post-test analysis predictions from both software packages with the test data including load-deflection curves from static load tests, and vibration frequencies and mode shapes from structural dynamics tests. The analysis predictions were in reasonable agreement with the test data. Factors that precluded better correlation of the analyses and the tests were uncertainties in the material properties, test conditions, and modeling assumptions used in the analyses.

Sleight, David W.↗