Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “offloading”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

BASIC HARTREE-FOCK PROXY APPLICTION

This proxy application simulates the compute load and data-movement of the kernel of the Hartree-Fock method in quantum chemistry. The proxy features a simplified algorithm for computing electron repulsion integrals that is easily offloaded to GPUs.

FLETCHER, GRAHAMD↗

PYSEQM

PYSEQM is a package for performing semi-empirical quantum mechanical (SEQM) simulations on molecular systems utilizing PyTorch. SEQM simulations determine molecular properties (energy, electron density, dipole moment, ect.) by solving an approximate Schrödinger equation for the motions of electrons in a molecule. The use of PyTorch provides three specific advantages. First, it allows the calculations to be offloaded to GPU accelerators, giving an order of magnitude increase in speed. Second, back propagation is used to get atomic forces (derivative of the total energy with respect to atomic position) at the same computational cost as the energy calculation itself. Finally, the use of PyTorch makes for a natural interface to modern machine learning methods, which can be used to adjust the semi-empirical parameters build into SEQM methods. Additionally, PYSEQM implements various other optimizations for performing quantum mechanics based molecular dynamics, including SP2 for rapid GPU based solution of the self-consistent field algorithm, and the extended Lagrangian method for rapid QM-MD.

Nebgen, Benjamin↗

Exascale-capable cellular automaton for nucleation and grain growth

This model is a cellular automata-based algorithm used to predict grain-scale microstructure development. Nucleation and growth of existing grains are considered in an initially liquid domain, where the temperature field evolution is either driven through heat transport simulation data or a frozen temperature approximation (i.e., a constant cooling rate and a fixed thermal gradient). The growth of grains is based on the decentered octahedron algorithm published by Gandin and Rappaz (https://doi.org/10.1016/S1359-6454(96)00303-5), where the grain envelopes are tracked as approximations of the evolving dendritic growth front. The input heat transport data, which can come from simulation of an arbitrary processing technique, consists of coordinates and times at which said coordinates fall below the liquidus and solidus temperatures for the final time. As a result, only the final solidification of a given region is modeled (no remelting). The code is designed to be compatible with pre-Exascale architectures, and uses MPI and the Kokkos library to decompose the computational domain and offload solidification calculations to GPUs.

Rolchigo, MatthewR.↗

Multiscale Ecosystem for solving Maxwell-Schrodinger equations of open quantum systems (OpenMS)

Light-matter interactions play an important role in many branches of physics, chemistry, energy, and materials science. In the strong coupling regime, light-matter interactions are able to tune the materials properties via the formation of new quasiparticles, such as plasmons and polaritons. However, theoretical and numerical modeling of the light-matter interaction-mediated processes remain a big challenge because light-matter interactions are fundamentally multiscale and multiphysics problems involving multiple interactions between electrons, nuclei, and photons at different time/length scales. In the current community, light-matter interactions were treated at different levels of theoretical complexity in quantum chemistry and quantum optics. In quantum optics or quantum photonics, the matter is usually simplified as a few-level system, and the light is treated quantum-mechanically. On the other hand, quantum chemistry explores first-principles methods, including both single-particle and many-body-based techniques, to describe the electronic properties of matter in detail. However, light is usually prescribed as a classical electromagnetic field, and the light-matter interaction is taken into account as an external potential via classical approximations. This software is designed to fill current modeling shortcomings by delivering a first-ever scalable multiscale platform for simulating light-matter interactions in realistic electromagnetic environments. The software solves Maxwell and Schrodinger equations self-consistent on the heterogeneous platforms. It implements HF/DFT, TDDFT, and coupled-cluster counterparts for light-matter interactions and adopts modular programming to offload massively parallel algorithms on a large number of CPU/GPUs.

Zhang, Yu↗

Designing and prototyping extensions to the Message Passing Interface in MPICH

As HPC system architectures and the applications running on them continue to evolve, the MPI standard itself must evolve. The trend in current and future HPC systems toward powerful nodes with multiple CPU cores and multiple GPU accelerators makes efficient support for hybrid programming critical for applications to achieve high performance. However, the support for hybrid programming in the MPI standard has not kept up with recent trends. The MPICH implementation of MPI provides a platform for implementing and experimenting with new proposals and extensions to fill this gap and to gain valuable experience and feedback before the MPI Forum can consider them for standardization. Here, in this work, we detail six extensions implemented in MPICH to increase MPI interoperability with other runtimes, with a specific focus on heterogeneous architectures. First, the extension to MPI generalized requests lets applications integrate asynchronous tasks into MPI’s progress engine. Second, the iovec extension to datatypes lets applications use MPI datatypes as a general-purpose data layout API beyond just MPI communications. Third, a new MPI object, MPIX_Stream, can be used by applications to identify execution contexts beyond MPI processes, including threads and GPU streams. MPIX stream communicators can be created to make existing MPI functions thread-aware and GPU-aware, thus providing applications with explicit ways to achieve higher performance. Fourth, MPIX Streams are extended to support the enqueue semantics for offloading MPI communications onto a GPU stream context. Fifth, thread communicators allow MPI communicators to be constructed with individual threads, thus providing a new level of interoperability between MPI and on-node runtimes such as OpenMP. Lastly, we present an extension to invoke MPI progress, which lets users spawn progress threads with fine-grained control to adapt the communication performance to their application designs. We describe the design and implementation of these extensions, provide usage examples, and highlight their expected benefits with performance results.

97 MATHEMATICS AND COMPUTING↗

Integrating ytopt and libEnsemble to autotune OpenMC

Ytopt is a Python machine-learning-based autotuning software package developed within the ECP PROTEAS-TUNE project. The ytopt software adopts an asynchronous search framework that consists of sampling a small number of input parameter configurations and progressively fitting a surrogate model over the input-output space until exhausting the user-defined maximum number of evaluations or the wall-clock time. libEnsemble is a Python toolkit for coordinating workflows of asynchronous and dynamic ensembles of calculations across massively parallel resources developed within the ECP PETSc/TAO project. libEnsemble helps users take advantage of massively parallel resources to solve design, decision, and inference problems and expands the class of problems that can benefit from increased parallelism. In this paper we present our methodology and framework to integrate ytopt and libEnsemble to take advantage of massively parallel resources to accelerate the autotuning process. Specifically, we focus on using the proposed framework to autotune the ECP ExaSMR application OpenMC, an open source Monte Carlo particle transport code. OpenMC has seven tunable parameters some of which have large ranges such as the number of particles in-flight, which is in the range of 100,000 to 8 million, with its default setting of 1 million. Setting the proper combination of these parameter values to achieve the best performance is extremely time-consuming. Therefore, we apply the proposed framework to autotune the MPI/OpenMP offload version of OpenMC based on a user-defined metric such as the figure of merit (FoM) (particles/s) or energy efficiency energy-delay product (EDP) on Crusher at Oak Ridge Leadership Computing Facility. In conclusion, the experimental results show that we achieve the improvement up to 29.49% in FoM and up to 30.44% in EDP.

Autotuning↗

Investigating Scientific Workload Acceleration using BlueField SmartNICs [Slides]

Modern computing platforms whose workloads generate large amounts of network traffic, such as cloud and HPC systems, often suffer from performance bottlenecks associated with the network interface. In order to alleviate the effects of this obstacle, a new generation of accelerators known as ‘SmartNICs’, which are designed to offload low level networking tasks from the processor into the NIC, have emerged.

42 ENGINEERING↗

INTEGRATED WORKFLOW MANAGEMENT FOR PARTICLE ACCELERATOR SIMULATION

Supercomputing systems are used for a wide range of computationally demanding tasks in many fields of science and engineering. They play a key role in numerical simulation, in which mathematical models are computed in order to simulate the behavior of physical systems. Scientists and engineers that use supercomputers for numerical simulation often have their productivity limited by the need to manually organize and manage extremely large amounts of data that are often produced and consumed by the software programs run on these systems. Recognizing these limitations, Kitware Inc. (Clifton Park, NY) and SLAC National Accelerator Laboratory (Menlo Park, CA) are developing an advanced software platform that can reduce the cognitive overhead required by knowledge workers when using supercomputers for numerical simulation. Phase I of the project is complete and includes the development of new capabilities for organizing simulation project files, improvements to the user interface and overall usability, and deployment of a “middle tier” server to sit between user desktop machines and supercomputers to offload much of the data management workload. The project also developed prototype software for executing sequences of numerical simulations, and a prototype for migrating supercomputing software to cloud-based computing systems to provide a potential alternative to supercomputers with different logistical and price-to-performance tradeoffs.

Tourtellott, John↗

OpenSNAPI: Toward a Unified API for SmartNICs

The end of Moore’s Law and Dennard Scaling has produced a renaissance in the field of computer architecture. Unable to continue leveraging silicon-level processor improvements to further enhance performance and scalability, system architects have been forced to explore other options. In this new era of heterogeneous architectures and hardware/software codesign, a new class of devices known as “accelerators” has emerged. Independently designed for optimized execution of distinct workloads, these devices have proven critical to the continued advancement of application performance. SmartNICs, accelerator devices integrated with a network controller, have conventionally been utilized to offload low-level networking functionality. However, newer SmartNIC variants, which incorporate a system-on-chip (SoC) with traditional designs, are challenging this precedent. Leveraging significantly augmented resources, these new devices offer increased versatility and the potential to more effectively complement a given architecture’s CPU. In this talk, we introduce the motivation underlying acceleration, explore the fundamentals of SmartNICs, and discuss traditional use cases. We also detail our initial efforts to investigate the feasibility and benefits of SmartNICs as general-purpose accelerators. We present the OpenSNAPI project created to define a uniform application programming interface (API) for this emerging class of devices. Finally, we provide a brief tutorial regarding development of SmartNIC-accelerated applications on Los Alamos National Laboratory’s SmartNIC-enabled platforms.

97 MATHEMATICS AND COMPUTING↗

Accelerate M-TIP on GPUs and deploy to Summit and NERSC-9 (against simulated data) WBS 2.2.4.05 ExaFEL, Milestone ADSE13-199

We evaluated the orientation matching step in the M-TIP SPI workflow for potential offloading to accelerators. We ported the code to GPUs, benchmarked it, optimized and down-selected the best versions. The accelerated version of the orientation matching code that was developed at LANL (LANL GPU v3) is 34-55X faster than sequential, 2.4-4.9X faster than the fastest OpenMP open source version we found (FAISS OpenMP) and 1.5-4X faster than the fastest GPU open source version we found (FAISS GPU). Summit single-node GPU versions were somewhat faster than Cori GPU. Image size plays a role; mid-range image sizes take more time. The LANL CUDA multi-node, multi-GPU implementation shows mostly linear strong scaling. I/O also plays a large role; splitting data into parts improves read time and burst buffers dramatically improve read times. This work will be integrated into the M-TIP workflow as part of the next milestone ADSE13-193.

97 MATHEMATICS AND COMPUTING↗

The Portals 4.3 Network Programming Interface

This report presents a specification for the Portals 4 network programming interface. Portals 4 is intended to allow scalable, high-performance network communication between nodes of a parallel computing system. Portals 4 is well suited to massively parallel processing and embedded systems. Portals 4 represents an adaption of the data movement layer developed for massively parallel processing platforms, such as the 4500-node Intel TeraFLOPS machine. Sandia's Cplant cluster project motivated the development of Version 3.0, which was later extended to Version 3.3 as part of the Cray Red Storm machine and XT line. Version 4 is targeted to the next generation of machines employing advanced network interface architectures that support enhanced offload capabilities.

97 MATHEMATICS AND COMPUTING↗

Processing Particle Data Flows with SmartNICs

Many distributed applications implement complex data flows and need a flexible mechanism for routing data between producers and consumers. Recent advances in programmable network interface cards, or SmartNICs, represent an opportunity to offload data-flow tasks into the network fabric, thereby freeing the hosts to perform other work. System architects in this space face multiple questions about the best way to leverage SmartNICs as processing elements in data flows. In this paper, we advocate the use of Apache Arrow as a foundation for implementing data-flow tasks on SmartNICs. We report on our experiences adapting a partitioning algorithm for particle data to Apache Arrow and measure the on-card processing performance for the BlueField-2 SmartNIC. Our experiments confirm that the BlueField-2’s (de)compression hardware can have a significant impact on in-transit workflows where data must be unpacked, processed, and repacked.

97 MATHEMATICS AND COMPUTING↗

Assessment of Cloud-Based Applications Enabling a Scalable Risk-Informed Predictive Maintenance Strategy Across the Nuclear Fleet

The current light water reactor fleet uses time-based or failure-based maintenance strategies to achieve high-capacity factors. But to make nuclear more competitive in the energy market, these reactors could utilize emerging technologies in terms of artificial intelligence (AI) and cloud computing to enable a cost-effective, predictive maintenance strategy. This report examines the feasibility of cloud computing for the nuclear industry’s needs in terms of the cloud’s computing capabilities, feasibility, and regulatory concerns. The technical viability of cloud computing was analyzed using one year worth of data from a boiling water reactor’s safety relief valve. Models were hosted on a local desktop, Idaho National Laboratory’s high-performance computer, and Microsoft Azure. Data was loaded, processed, and two types of models were trained in an A/B fashion. Based on the speed at which these actions were completed, it was used to determined that cloud computing has adequate computing resources. Additionally, the computing power can scale with the demanded load. To enable cloud computing in the existing fleet, additional sensors, networks, and other requirements must be implemented to ensure a smooth transition from current maintenance strategies. However, there is a benefit as the plant no longer needs manage their own servers, software, cybersecurity, and IT support staff. Many of these features can be offloaded on to the cloud provider. A comprehensive analysis was completed that showed the current annual cost of operating is more expensive than using cloud computing resources. Lastly, the regulatory framework does not explicitly address AI or autonomous control. Currently, the NRC and other regulatory bodies are evaluating providing guidance to address gaps rather than new regulations to address the use of AI and ML. But since many of the AI applications are focused on non-safety related applications, such as balance-of-plant components, they will likely have little or no regulatory restrictions or necessary approvals. Demonstrating how AI can improve maintenance and operation of these non-safety related systems seems like the likely path forward for implementing AI and cloud computing resources inside nuclear power plants (NPPs).

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Celeritas R&D Report: Accelerating Geant4

Celeritas is a new Monte Carlo (MC) detector simulation code designed for computationally intensive applications on high-performance heterogeneous architectures. In the past two years Celeritas has advanced from prototyping a Graphics Processing Unit (GPU)-based single physics model in infinite medium to implementing a full set of electromagnetic (EM) physics processes in complex geometries. The current release of Celeritas, version 0.4, has incorporated full device-based navigation, an event loop in the presence of magnetic fields, and detector hit scoring. New functionality incorporates a scheduler to offload electromagnetic physics to the GPU within a Geant4-driven simulation, enabling straightforward integration of Celeritas into the high energy physics (HEP) experimental frameworks CMSSW and ATLAS FullSimLight. On the Perlmutter supercomputer, Celeritas performs EM physics between 3× and 18× faster using the machine’s Nvidia GPUs compared to using only CPUs, corresponding to an electrical power efficiency up to a factor of 5. When running a multithreaded Geant4 ATLAS test beam application with full hadronic physics, using Celeritas to accelerate the EM physics results in an overall simulation speedup of 1.7–2.2× on GPU and 1.2× on CPU. In a CMS test application using tt¯ events and the prototype Run 4 configuration, compared to Geant4 CPU, Celeritas with a Nvidia A100 improves overall throughput up to a factor of 2.7× but cannot be efficiently shared with more than 8 cores.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Deployment of inference as a service at the US CMS Tier-2 data centers

Coprocessors, especially GPUs, will be a vital ingredient of data production workflows at the HL-LHC. At CMS, the GPU-as-a-service approach for production workflows is implemented by the SONIC project (Services for Optimized Network Inference on Coprocessors). SONIC provides a mechanism for outsourcing computationally demanding algorithms, such as neural network inference, to remote servers, where requests from multiple clients are intelligently distributed across multiple GPUs by a load-balancing service. This talk highlights the recent progress in deploying SONIC at selected U.S. CMS Tier-2 data centers. Using realistic CMS Run3 data processing workflows, such as those containing transformer-based algorithms, we demonstrate how SONIC is integrated into the production-like environment to enable accelerated inference offloading. We will present developments from both the client and server sides, including production job and data center configurations for NVIDIA and AMD GPUs. We will also present performance scaling benchmarks and discuss the challenges of operating SONIC in CMS production, such as server discovery, GPU saturation, fallback server logic, etc.

Holzman, Burt↗

Geant4 Event Biasing and Fast Simulation

Geant4 offers advanced event biasing techniques to significantly accelerate simulations involving rare events. Various biasing methods, such as leading particle selection, cross-section biasing, radioactive decay enhancement, and bremsstrahlung splitting, enable efficient event sampling, though they require careful handling. Additionally, Geant4 provides a Fast Simulation Interface, allowing the replacement of standard processes in specific region and for selected particles, enabling faster execution or external code integration. Applications of fast simulation include electromagnetic shower modeling in calorimeters, machine learning inference, and offloading tasks to specialized hardware like GPUs, making Geant4 a powerful tool for computationally demanding simulations.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Heterogeneous Computing

To leverage the increasing heterogeneity in modern computing resources, Geant4 incorporates advanced software tools and a task-based framework (G4Tasking) that enables efficient parallelism at event, sub-event, and track levels. Ongoing R&D efforts focus on integrating GPUs into high-energy physics (HEP) simulations, including optical photon simulation with Opticks/NVIDIA OptiX, offloading electromagnetic particle transport using G4HepEM/AdePT and Celeritas, and employing advanced surface-based geometry models such as VecGeom2.0 and ORANGE. As Geant4 continues evolving toward high-performance computing (HPC) and heterogeneous architectures, it remains a key tool for large-scale simulations in HEP and beyond.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗