Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Global Pathway Selection with Zero-RK

Global Pathway Selection (GPS) is an algorithm to effectively generates reduced (skeletal) chemistry mechanisms, which speeds up simulations and can be used as a systematic analytics tool to extract insights from complex reacting system. This release is an extension of the original code to run in parallel and to use LLNL's Zero-RK solver for fast solution of chemical problems.

Whitesides, RussellA↗

Clang UPC2C Translator (Clang UPC2C) v9.0.1-1

Clang Unified Parallel C 2 C (Clang UPC2C) translator compiles programs written in the UPC (Unified Parallel C) language to ISO C99, with calls to the Berkeley UPC runtime system. The Clang UPC2C compiler extends the capabilities of the Clang LLVM C frontend to comply with the UPC Language Specification version 1.3. The compiler generates programs that run on a wide variety of systems ranging from workstations to leadership-class supercomputers, in conjunction with the Berkeley UPC runtime and GASNet communication system. In addition to the standard UPC libraries, Clang UPC2C also provides access to Berkeley UPC library extensions.

Hargrove, Paul↗

Global Pathway Selection with Zero-RK v0.5

Global Pathway Selection (GPS) is an algorithm to effectively generates reduced (skeletal) chemistry mechanisms, which speeds up simulations and can be used as a systematic analytics tool to extract insights from complex reacting system. This release is an extension of the original code to run in parallel and to use LLNL's Zero-RK solver for fast solution of chemical problems.

Whitesides, RussellA [Lawrence Livermore National ↗

Adorym: a multi-platform generic X-ray image reconstruction framework based on automatic differentiation

We describe and demonstrate an optimization-based X-ray image reconstruction framework called Adorym. Our framework provides a generic forward model, allowing one code framework to be used for a wide range of imaging methods ranging from near-field holography to fly-scan ptychographic tomography. By using automatic differentiation for optimization, Adorym has the flexibility to refine experimental parameters including probe positions, multiple hologram alignment, and object tilts. It is written with strong support for parallel processing, allowing large datasets to be processed on high-performance computing systems. We demonstrate its use on several experimental datasets to show improved image quality through parameter refinement.

36 MATERIALS SCIENCE↗

Asynchronous Iterative Solvers for Extreme-Scale Computing

The Asynchronous Iterative Solvers for Extreme-Scale Computing (AsyncIS) project aims to explore more efficient numerical algorithms by decreasing their overhead. AsyncIS does this by replacing the outer Krylov subspace solver with an asynchronous optimized Schwarz method, thereby removing the global synchronization and bulk synchronous operations typically used in numerical codes. AsyncIS—a U.S. Department of Energy (DOE)-funded collaboration between Georgia Tech, the University of Tennessee, Knoxville, Temple University, and Sandia National Laboratories—also focuses on the development and optimization of asynchronous preconditioners (i.e., preconditioners that are generated and/or applied in an asynchronous fashion). The novel preconditioning algorithms that provide fine-grained parallelism enable preconditioned Krylov solvers to run efficiently on large-scale distributed systems and manycore accelerators like GPUs.

97 MATHEMATICS AND COMPUTING↗

Elastic Solutions to 2D Plane Strain Problems: Nonlinear Contact and Settlement Analysis for Shallow Foundations

The classical Neumann boundary value problem of an isotropic, homogeneous elastic half-plane under plane strain conditions is readdressed as the limiting case of the fully three-dimensional problem. Analytical solutions of the stress and strain tensors are obtained by taking the limit from known three-dimensional solutions. It is shown that the displacement fields for the plane strain problem are not well defined. A small number of simple expressions are developed, which provide a general solution for linearly-varying traction over arbitrary regions on the boundary. A simple, efficient, and rapidly convergent algorithm is developed which uses these solutions as analytic elements and provides a solution approach to the general boundary value problem. The method is verified against known solutions for Hertzian contact between parallel cylinders. Two numerical examples are presented for the analysis of shallow foundation systems. In the first, the boundary conditions are informed by analytical elastoplastic calculations and a strain influence analysis is performed and compared with the Schmertmann method. Subsequently, empirical laboratory contact traction distributions measured by Bauer et al., in both the normal and tangential directions are employed as boundary conditions for an analysis of the underlying stress field.

42 ENGINEERING↗

AthenaK: A Performance-portable Version of the Athena++ Adaptive Mesh Refinement Framework

We describe AthenaK: a new implementation of the Athena++ block-based adaptive mesh refinement framework using the Kokkos programming model. Finite volume methods for Newtonian, special relativistic, and general relativistic (GR) hydrodynamics and magnetohydrodynamics (MHD), and GR-radiation hydrodynamics and MHD, as well as a module for evolving Lagrangian tracer or charged test particles (e.g., cosmic rays) are implemented using the framework. In two companion papers, we describe (1) a new solver for the Einstein equations based on the Z4c formalism, and (2) a GRMHD solver in dynamical spacetimes also implemented using the framework, enabling new applications in numerical relativity. By adopting Kokkos, the code can be run on virtually any hardware, including CPUs, GPUs from multiple vendors, and emerging Advanced RISC Machine processors. AthenaK shows excellent performance and weak scaling, achieving over 1 billion cell updates per second for hydrodynamics in three dimensions on a single NVIDIA Grace Hopper processor. It does this with a typical parallel efficiency of 80% on 65,536 AMD GPUs on the OLCF Frontier system. Such performance portability enables AthenaK to leverage modern exascale computing systems for challenging applications in astrophysical fluid dynamics, numerical relativity, and multimessenger astrophysics.

79 ASTRONOMY AND ASTROPHYSICS↗

Modeling commercial-scale CO 2 storage in the gas hydrate stability zone with PFLOTRAN v6.0

Abstract. Safe and secure carbon dioxide (CO2) storage is likely to be critical for mitigating some of the most dangerous effects of climate change. In the last decade, there has been a significant increase in activity associated with reservoir characterization and site selection for large-scale CO2 storage projects across the globe. These prospective storage sites tend to be selected for their optimal structural, petrophysical, and geochemical trapping potential. However, it has also been suggested that storing CO2 in reservoirs within the CO2 hydrate stability zone (GHSZ), characterized by high pressures and low temperatures (e.g., Arctic or marine environments), could provide a natural thermodynamic barrier to gas leakage. Evaluating the prospect of commercial-scale, long-term storage of CO2 in the GHSZ requires reservoir-scale modeling capabilities designed to account for the unique physics and thermodynamics associated with these systems. We have developed the HYDRATE flow mode and the accompanying fully implicit parallel well model in the massively parallel subsurface flow and reactive transport simulator PFLOTRAN to model CO2 injection into the marine GHSZ. We have applied these capabilities to a set of CO2 injection scenarios designed to reveal the challenges and opportunities for commercial-scale CO2 storage in the GHSZ.

carbon storage↗

Locality-Aware Scheduling for Scalable Heterogeneous Environments

Heterogeneous computing promise boost performance of scientific applications by allowing massively parallel execution of computational tasks. However, manually managing extremely heterogeneous, multi-device systems is complicated and may result in sub-optimal performance. Specifically, data management is an extremely challenging problem on multi-device systems. In this work, we introduce two locality-aware schedulers for the Minos Computing Library (MCL), an asynchronous, task-based programming model and runtime for extremely heterogeneous systems. The first scheduler implements a pure locality-aware algorithm to maximize data reuse, though it might incur in ”hot-spots” that limit system utilization. The second scheduler mitigates this drawback by dynamically targeting between locality-awareness and system utilization based on the current workload and available computing devices. Our results show that locality-awareness greatly benefit applications that exhibit data reuse, providing up to 6.9x and 7.9x over the original MCL scheduler and equivalent OpenCL implementations, respectively. Moreover, our schedulers introduce negligible overhead compared with the original MCL scheduler and achieve similar performance for applications that don’t benefit from data locality.

Architecture, co-design, Task-based programming mo↗

Triple bubbler system, fast-bubbling approach, and related methods

A triple bubbler system includes a first fluid probe, a second fluid probe, a third fluid probe, a gas source operably coupled to the first fluid probe, the second fluid probe, and the third fluid probe and configured to meter gas through the first fluid probe, the second fluid probe, and the third fluid probe to form bubbles at tips of each of the first fluid probe, the second fluid probe, and the third fluid probe, and a cover member disposed over the tips of the first, second, and third fluid probes and configured to at least partially prevent bubbles formed and escaping the tips of the first, second, and third fluid probes from interfering with other bubbles formed at each other tips. The bubbler system includes a thermocouple having a plurality of junctions disposed along an axis parallel to longitudinal axes of the first, second, and third fluid probes.

Galbreth, Gregory G.↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

VTK-m User's' Guide (V.1.7)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.This book includes contributions from the VTK-m community including the VTK-m development team and the user community.

97 MATHEMATICS AND COMPUTING↗

The VTK-m Users' Guide (V.1.9)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

The VTK-m Users' Guide (V.2.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

The VTK-m User's Guide (V. 2.2)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction.

97 MATHEMATICS AND COMPUTING↗

Supercomputer-Based Ensemble Docking Drug Discovery Pipeline with Application to Covid-19

In this work, we present a supercomputer-driven pipeline for in silico drug discovery using enhanced sampling molecular dynamics (MD) and ensemble docking. Ensemble docking makes use of MD results by docking compound databases into representative protein binding-site conformations, thus taking into account the dynamic properties of the binding sites. We also describe preliminary results obtained for 24 systems involving eight proteins of the proteome of SARS-CoV-2. The MD involves temperature replica exchange enhanced sampling, making use of massively parallel supercomputing to quickly sample the configurational space of protein drug targets. Using the Summit supercomputer at the Oak Ridge National Laboratory, more than 1 ms of enhanced sampling MD can be generated per day. We have ensemble docked repurposing databases to 10 configurations of each of the 24 SARS-CoV-2 systems using AutoDock Vina. Comparison to experiment demonstrates remarkably high hit rates for the top scoring tranches of compounds identified by our ensemble approach. We also demonstrate that, using Autodock-GPU on Summit, it is possible to perform exhaustive docking of one billion compounds in under 24 h. Finally, we discuss preliminary results and planned improvements to the pipeline, including the use of quantum mechanical (QM), machine learning, and artificial intelligence (AI) methods to cluster MD trajectories and rescore docking poses.

60 APPLIED LIFE SCIENCES↗

Cluster perturbation theory X: A parallel implementation of Lagrangian perturbation series for the coupled cluster singles and doubles ground-state energy through fifth order

We describe an efficient implementation of cluster perturbation and Møller-Plesset Lagrangian energy series through fifth order that target the coupled cluster singles and doubles energy utilizing the resolution of the identity approximation. We illustrate the computational performance of the implementation by performing ground state energy calculations on systems with up to 1200 basis functions using a single node and by comparison to conventional CCSD calculations. We further show that our hybrid MPI/OMP parallel implementation that also utilizes graphical processing units can be used to obtain fifth order energies on systems with almost 1200 basis functions with a 90 minute "time to solution" running on Frontier at Oak Ridge National Laboratory.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗