Piper: Pipelining OpenMP Offloading Execution Through Compiler Optimization For Performance
Explore the source record for details and available documents.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Celeritas is a GPU-optimized Monte Carlo (MC) particle transport code designed to meet the growing computational demands of next-generation high energy physics (HEP) experiments. It provides efficient simulation of electromagnetic (EM) physics processes in complex geometries with magnetic fields, detector hit scoring, and seamless integration into Geant4-driven applications to offload EM physics to GPUs. Recent efforts have focused on performance optimizations and expanding profiling capabilities. This paper presents some key advancements, including the integration of the Perfetto system profiling tool for detailed performance analysis and the development of track-sorting methods to improve computational efficiency.
The Exascale Computing Project (ECP) focuses on the development of future exascale-capable applications. Most ECP applications use the message passing interface (MPI) as their parallel programming model with mini-apps serving as proxies. This paper explores the explicit usage of MPI in such ECP proxy applications. We empirically analyze 14 proxy applications from the ECP Proxy Apps Suite. We use the MPI profiling interface (PMPI) to collect MPI usage patterns in ECP proxy apps. Our analysis shows that a small subset of features from MPI is commonly used in the proxies of exascale-capable applications, even when they reference third-party libraries. Overall, this study is intended to provide a better understanding of the use of MPI in current exascale applications. The findings can help focus software investments made for exascale systems in the MPI middleware including optimization, fault-tolerance, tuning, and hardware-offload.
Computing resources in the Worldwide LHC Computing Grid (WLCG) have been based entirely on the x86 architecture for more than two decades. In the near future, however, heterogeneous non-x86 resources, such as ARM, POWER and Risc-V, will become a substantial fraction of the resources that will be provided to the LHC experiments, due to their presence in existing and planned world-class HPC installations. The CMS experiment, one of the four large detectors at the LHC, has started to prepare for this situation, with the CMS software stack (CMSSW) already compiled for multiple architectures. In order to allow for a production use, the tools for workload management and job distribution need to be extended to be able to exploit heterogeneous architectures. Profiting from the opportunity to exploit the first sizable IBM Power9 allocation available on Marconi100 HPC system at CINECA, CMS developed all the needed modifications to the CMS workload management system. After a successful proof of concept, a full physics validation has been performed in order to bring the system in production. The experiences are of very high value, when it comes to commissioning of the similar (even larger) Summit HPC system at Oak Ridge, where CMS is also expecting a resource allocation. Moreover the compute power of those systems is being provided also via GPUs and this represents an extremely valuable opportunity to exploit the offloading capability already implemented in CMSSW. The status of the current integration including the exploitation of the GPUs, the results of the validation as well as the future plans will be shown and discussed.
BACKGROUND: The Active Response Gravity Offload System (ARGOS) provides an analog environment for extravehicular activity (EVA) testing and training. Discomfort has been observed during longer suited test sessions. While the subject’s core is offloaded during surface EVA evaluations, his/her arms experience full Earth gravity and can become overly fatigued, especially during suited tests which involve reaching and prolonged arm extensions. A device (ARGOS Negation of Gravitational Effects on the Limbs: ANGEL) to offload the weight of the arms and suit sleeves is being developed by JSC’s Flight Systems Branch of the Software, Robotics, and Simulation Division. Previously we have shared preliminary modeling of that device and kinematics based on motion capture data. Here we present an alternative approach to determine device kinematics by calculating ANGEL component angles with an OpenSim plugin. We compare calculated angles to inverse kinematics (IK) derived ones with the goal of validating the model. This new method can be further informative for device design and analytically testing different configurations to achieve desired reduced gravity conditions (e.g., lunar gravity (Lg) or Martian gravity (Mg)). We have compared calculated angles with IK-derived angles in tests with a shirt-sleeve subject positioned in a test stand with a Mark-III Hard Upper Torso (HUT) and Portable Life Support System (PLSS) mockup and arm weights to emulate the weight of the suit sleeve as well as a suited subject in ARGOS with a Mark-III suit. A variety of upper body tasks were completed in the former and full-body tasks in the latter. METHODS AND RESULTS: To model the offload device, we augment the OpenSim human model topology with the offload mechanism components and joints, using CAD models to represent the mechanism graphically. The joint angles of the device are calculated in the OpenSim plugin by modeling how the components configure themselves under the offloading spring tension given a particular IK-derived arm position. There are four ANGEL components with a total of 5 degrees of freedom (DOFs), each component has a single DOF except for the cuff which is modeled as 2 DOFs. The sickle/yaw bracket and cuff rotation angles are determined statically based on the assumptions that the sickle will track the attachment point of the cuff and that the cuff will rotate such that the attachment point is at its highest point. The cuff tilt, linker and V-bracket angles are then determined by optimizing their positions to approach a mechanical equilibrium. The calculated linker angle is compared to three different methods of determining the linker line-of-force kinematically (from V-bracket to center-cuff, cuff highest point or marker-derived position). Given the joint angles of the device, the spring force and resulting force on the arm is computed by the plugin and applied as an external load in inverse dynamics (ID) to enable study of overall shoulder joint torques as well as offload achieved. We verify the calculated joint angles by using the inverse kinematic data. The average difference in angles is the smallest for the V-bracket and linker, around 1 to 5 degrees for most trials. The resulting offload and shoulder torque are comparable between calculated and IK-derived angles. In summary, we have developed a method to calculate the joint angles of an exoskeleton-like upper limb offloading device currently in development. We have also developed a custom plugin which will be a valuable tool to optimize device configurations for a desired gravitational environment, probe the offload achieved for motions recorded outside of our test suite, and inform future design improvements.
BACKGROUND: The Active Response Gravity Offload System (ARGOS) provides an analog environment for extravehicular activity (EVA) testing and training. Discomfort has been observed during longer suited test sessions. While the subject’s core is offloaded during surface EVA evaluations, his/her arms experience full Earth gravity and can become overly fatigued, especially during suited tests which involve reaching and prolonged arm extensions. A device (ARGOS Negation of Gravitational Effects on the Limbs: ANGEL) to offload the weight of the arms and suit sleeves is being developed by JSC’s Flight Systems Branch of the Software, Robotics, and Simulation Division. Previously we have shared preliminary modeling of that device and kinematics based on motion capture data. Here we present an alternative approach to determine device kinematics by calculating ANGEL component angles with an OpenSim plugin. We compare calculated angles to inverse kinematics (IK) derived ones with the goal of validating the model. This new method can be further informative for device design and analytically testing different configurations to achieve desired reduced gravity conditions (e.g., lunar gravity (Lg) or Martian gravity (Mg)). We have compared calculated angles with IK-derived angles in tests with a shirt-sleeve subject positioned in a test stand with a Mark-III Hard Upper Torso (HUT) and Portable Life Support System (PLSS) mockup and arm weights to emulate the weight of the suit sleeve as well as a suited subject in ARGOS with a Mark-III suit. A variety of upper body tasks were completed in the former and full-body tasks in the latter. METHODS AND RESULTS: To model the offload device, we augment the OpenSim human model topology with the offload mechanism components and joints, using CAD models to represent the mechanism graphically. The joint angles of the device are calculated in the OpenSim plugin by modeling how the components configure themselves under the offloading spring tension given a particular IK-derived arm position. There are four ANGEL components with a total of 5 degrees of freedom (DOFs), each component has a single DOF except for the cuff which is modeled as 2 DOFs. The sickle/yaw bracket and cuff rotation angles are determined statically based on the assumptions that the sickle will track the attachment point of the cuff and that the cuff will rotate such that the attachment point is at its highest point. The cuff tilt, linker and V-bracket angles are then determined by optimizing their positions to approach a mechanical equilibrium. The calculated linker angle is compared to three different methods of determining the linker line-of-force kinematically (from V-bracket to center-cuff, cuff highest point or marker-derived position). Given the joint angles of the device, the spring force and resulting force on the arm is computed by the plugin and applied as an external load in inverse dynamics (ID) to enable study of overall shoulder joint torques as well as offload achieved. We verify the calculated joint angles by using the inverse kinematic data. The average difference in angles is the smallest for the V-bracket and linker, around 1 to 5 degrees for most trials. The resulting offload and shoulder torque are comparable between calculated and IK-derived angles. In summary, we have developed a method to calculate the joint angles of an exoskeleton-like upper limb offloading device currently in development. We have also developed a custom plugin which will be a valuable tool to optimize device configurations for a desired gravitational environment, probe the offload achieved for motions recorded outside of our test suite, and inform future design improvements.
Science gateways are altering the manner in which people interact with high performance computing (HPC) by providing a web browser based interface to advanced computing platforms. In particular, science gateways lower the barrier to using HPC by simplifying the process of submitting workloads to such systems and by offloading the efforts required to use HPC to the maintainers of the system. While science gateways decrease the time-to-science that comes with using such advanced systems, progress can still be made in improving the user's experience. In this paper we explore two strategies for integrating artificial intelligence tools commonly found in non-HPC service workflows: voice activated assistants and chatbots. Since August 2021, the HPC group at Idaho National Laboratory answers an average of 581 support tickets per month of which a large percentage could be addressed via these two strategies. This work defines the key capabilities that an HPC voice activated assistant and chatbot would need to address for a userbase consisting of largely non-expert users as well as a design for integration into the Open OnDemand science gateway.
The current light-water reactor fleet uses time-based maintenance strategies to achieve high-capacity factors. But to make nuclear more competitive in the energy market, these reactors could utilize emerging artificial intelligence (AI) and cloud computing technologies to achieve a cost-effective, predictive-maintenance strategy. This paper presents discussion and results on the application of cloud computing in the nuclear industry. The technical viability of cloud computing was analyzed using data from a boiling-water reactor’s safety relief valve. The models were hosted on three different systems: a local personal computer, Idaho National Laboratory’s high-performance computer system, and Microsoft Azure. The data were loaded and processed, and two types of models were trained in an A/B fashion. Based on the speed at which these actions were completed, it was determined that cloud computing affords adequate computing resources. Additionally, the computing power can scale with the demanded load. To enable cloud computing in the existing fleet, additional sensors, networks, and other requirements must be implemented to ensure a smooth transition from current maintenance strategies. However, the benefit is that the plants no longer need to manage their own servers, software, cybersecurity, and information technology support staff for in-house data analytics purpose. Many of these features can be offloaded to the cloud provider for a potential cost savings. Demonstrating how AI can improve the maintenance and operation of non-safety-related systems seems the likely path forward for implementing AI and cloud computing resources inside nuclear power plants.
We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.
This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.
This paper reports on experiences gained and practices adopted when using the latest features of OpenMP to port a variety of HPC applications and mini-apps based on different computational motifs (BerkeleyGW, WDMApp/XGC, GAMESS, GESTS, and GridMini) to accelerator-based, leadership-class, high-performance supercomputer systems at the Department of Energy. As recent enhancements to OpenMP become available in implementations, there is a need to share the results of experimentation with them in order to better understand their behavior in practice, to identify pitfalls, and to learn how they can be effectively deployed in scientific codes. Additionally, we identify best practices from these experiences that we can share with the rest of the OpenMP community.
FastCaloSim is a parameterized simulation of the particle energy response and of the energy distribution in the ATLAS calorimeter. It is a relatively small and self-contained package with massive inherent parallelism and captures the essence of GPU offloading via important operations like data transfer, memory initialization, floating point operations, and reduction. It was identified by the High Energy Physics Center for Computational Excellence project as a good testbed for evaluating the performance and ease of portability of programming models. In this paper, we will discuss the results of our evaluation of the porting process to Kokkos, SYCL, Alpaka, OpenMP and std::par (nvc++), and compare performance on NVIDIA, AMD and Intel GPUs, as well as multicore CPUs.
In the past years the landscape of tools for expressing parallel algorithms in a portable way across various compute accelerators has continued to evolve significantly. There are many technologies on the market that provide portability between CPU, GPUs from several vendors, and in some cases even FPGAs. These technologies include C++ libraries such as Alpaka and Kokkos, compiler directives such as OpenMP, the SYCL open specification that can be implemented as a library or in a compiler, and standard C++ where the compiler is solely responsible for the offloading. Given this developing landscape, users have to choose the technology that best fits their applications and constraints. For example, in the CMS experiment the experience so far in heterogeneous reconstruction algorithms suggests that the full application contains a large number of relatively short computational kernels and memory transfer operations. In this work we use a stand-alone version of the CMS heterogeneous pixel reconstruction code as a realistic use case of HEP reconstruction software that is capable of leveraging GPUs effectively. We summarize the experience of porting this code base from CUDA to Alpaka, Kokkos, SYCL, std::par, and OpenMP offloading. We compare the event processing throughput achieved by each version on NVIDIA and AMD GPUs as well as on a CPU, and compare those to what a native version of the code achieves on each platform.
Supercomputing systems are used for a wide range of computationally demanding tasks in many fields of science and engineering. They play a key role in numerical simulation, in which mathematical models are computed in order to simulate the behavior of physical systems. Scientists and engineers that use supercomputers for numerical simulation often have their productivity limited by the need to manually organize and manage extremely large amounts of data that are often produced and consumed by the software programs run on these systems. Recognizing these limitations, Kitware Inc. (Clifton Park, NY) and SLAC National Accelerator Laboratory (Menlo Park, CA) are developing an advanced software platform that can reduce the cognitive overhead required by knowledge workers when using supercomputers for numerical simulation. Phase I of the project is complete and includes the development of new capabilities for organizing simulation project files, improvements to the user interface and overall usability, and deployment of a “middle tier” server to sit between user desktop machines and supercomputers to offload much of the data management workload. The project also developed prototype software for executing sequences of numerical simulations, and a prototype for migrating supercomputing software to cloud-based computing systems to provide a potential alternative to supercomputers with different logistical and price-to-performance tradeoffs.
Lunar robotic functions include: 1. Transport of crew and payloads on the surface of the moon; 2. Offloading payloads from a lunar lander; 3. Handling the deployment of surface systems; with 4. Human commanding of these functions from inside a lunar vehicle, habitat, or extravehicular (space walk), with Earth-based supervision. The systems that will perform these functions may not look like robots from science fiction. In fact, robotic functions may be automated trucks, cranes and winches. Use of this equipment prior to the crew s arrival or in the potentially long periods without crews on the surface, will require that these systems be computer controlled machines. The public release of NASA's Exploration plans at the 2nd Space Exploration Conference (Houston, December 2006) included a lunar outpost with as many as four unique mobility chassis designs. The sequence of lander offloading tasks involved as many as ten payloads, each with a unique set of geometry, mass and interface requirements. This plan was refined during a second phase study concluded in August 2007. Among the many improvements to the exploration plan were a reduction in the number of unique mobility chassis designs and a reduction in unique payload specifications. As the lunar surface system payloads have matured, so have the mobility and offloading functional requirements. While the architecture work continues, the community can expect to see functional requirements in the areas of surface mobility, surface handling, and human-systems interaction as follows: Surface Mobility 1. Transport crew on the lunar surface, accelerating construction tasks, expanding the crew s sphere of influence for scientific exploration, and providing a rapid return to an ascent module in an emergency. The crew transport can be with an un-pressurized rover, a small pressurized rover, or a larger mobile habitat. 2. Transport Extra-Vehicular Activity (EVA) equipment and construction payloads. 3. Transport habitats and power modules over long distances, pre-positioning them for the arrival of crew on a subsequent lander. Surface Handling 1. Offload surface system payloads from the lander, breaking launch restraints and power/data connections. Payloads may be offloaded to a wheeled vehicle for transport. 2. Deploy payloads from a wheeled vehicle at a field site, placing the payloads in their final use site on the ground or mating them with existing surface systems. 3. Support regolith collection, site preparation, berm construction, or other civil engineering tasks using tools and implements attached to rovers. Human-Systems Interaction 1. Provide a safe command and control interface for suited EVA to ride on and drive the vehicles, making sure that the systems are also safe for working near dismounted crew. 2. Provide an effective control system for IV crew to tele-operate vehicles, cranes and other equipment from inside the surface habitats with evolving independence from Earth. .. Provide a supervisory system that allows machines to be commanded from the ground, working across the Earth-Lunar time delays on the order of 5-10 seconds (round trip) to support operations when crew are not resident on the surface. Technology Development Needs 1. Surface vehicles that can dock, align and mate with outpost equipment such as landers, habitats and fluid/power interfaces. 2. Long life motors, drive trains, seals, motor electronics, sensors, processors, cable harnesses, and dash board displays. 3. Active suspension control, localization, high speed obstacle avoidance, and safety systems for operating near dismounted crew. 4. High specific energy and specific power batteries that are safe, rechargeable, and long lived.
We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.
The main objective of the Holodeck Testbed is to create a cost effective, realistic, and highly immersive environment that can be used to train astronauts, carry out engineering analysis, develop procedures, and support various operations tasks. Currently, the Holodeck testbed allows to step into a simulated ISS (International Space Station) and interact with objects; as well as, perform Extra Vehicular Activities (EVA) on the surface of the Moon or Mars. The Holodeck Testbed is using the products being developed in the Hybrid Reality Lab (HRL). The HRL is combining technologies related to merging physical models with photo-realistic visuals to create a realistic and highly immersive environment. The lab also investigates technologies and concepts that are needed to allow it to be integrated with other testbeds; such as, the gravity offload capability provided by the Active Response Gravity Offload System (ARGOS). My main two duties were to develop and animate models for use in the HRL environments and work on a new way to interface with computers using Brain Computer Interface (BCI) technology. On my first task, I was able to create precise computer virtual tool models (accurate down to the thousandths or hundredths of an inch). To make these tools even more realistic, I produced animations for these tools so they would have the same mechanical features as the tools in real life. The computer models were also used to create 3D printed replicas that will be outfitted with tracking sensors. The sensor will allow the 3D printed models to align precisely with the computer models in the physical world and provide people with haptic/tactile feedback while wearing a VR (Virtual Reality) headset and interacting with the tools. Getting close to the end of my internship the lab bought a professional grade 3D Scanner. With this, I was able to replicate more intricate tools at a much more time-effective rate. The second task was to investigate the use of BCI to control objects inside the hybrid reality ISS environment. This task looked at using an Electroencephalogram (EEG) headset to collect brain state data that could be mapped to commands that a computer could execute. On this Task, I had a setback with the hardware, which stopped working and was returned to the vendor for repair. However, I was still able to collect some data, was able to process it, and started to create correlation algorithms between the electrical patterns in the brain and the commands we wanted the computer to carry out. I also carried out a test to investigate the comfort of the headset if it is worn for a long time. The knowledge gained will benefit me in my future career. I learned how to use various modeling and programming tools that included Blender, Maya, Substance Painter, Artec Studio, Github, and Unreal Engine 4. I learned how to use a professional grade 3D scanner and 3D printer. On the BCI Project I learned about data mining and how to create correlation algorithms. I also supported various demos including a live demo of the hybrid reality lab capabilities at ComicPalooza. This internship has given me a good look into engineering at NASA. I developed a more thorough understanding of engineering and my overall confidence has grown. I have also realized that any problem can be fixed, if you try hard enough, and as an engineer it is your job to not only fix problems but to embrace coming up with solutions to those problems.
High-throughput structure-based screening of drug-like molecules has become a common tool in biomedical research. Recently, acceleration with graphics processing units (GPUs) has provided a large performance boost for molecular docking programs. Both cloud and high-performance computing (HPC) resources have been used for large screens with molecular docking programs; while NVIDIA GPUs have dominated cloud and HPC resources, new vendors such as AMD and Intel are now entering the field, creating the problem of software portability across different GPUs. Ideally, software productivity could be maximized with portable programming models that are able to maintain high performance across architectures. While in many cases compiler directives have been used as an easy way to offload parallel regions of a CPU-based program to a GPU accelerator, they may also be an attractive programming model for providing portability across different GPU vendors, in which case the porting process may proceed in the reverse direction: from low-level, architecture-specific code to higher-level directive-based abstractions. MiniMDock is a new mini-application (miniapp) designed to capture the essential computational kernels found in molecular docking calculations, such as are used in phar-maceutical drug discovery efforts, in order to test different solutions for porting across GPU architectures. Here we extend MiniMDock to GPU offloading with OpenMP directives, and compare to performance of kernels using CUDA and HIP on NVIDIA and AMD GPUs, respectively, as well as across different compilers, exploring performance bottlenecks. We document this reverse-porting process, from highly optimized device code to a higher-level version using directives, compare code structure, and describe barriers that were overcome in this effort.