Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Toward exascale whole-device modeling of fusion devices: Porting the GENE gyrokinetic microturbulence code to GPU

GENE solves the five-dimensional gyrokinetic equations to simulate the development and evolution of plasma microturbulence in magnetic fusion devices. The plasma model used is close to first principles and computationally very expensive to solve in the relevant physical regimes. In order to use the emerging computational capabilities to gain new physics insights, several new numerical and computational developments are required. Here, we focus on the fact that it is crucial to efficiently utilize GPUs (graphics processing units) that provide the vast majority of the computational power on such systems. In this paper, we describe the various porting approaches considered and given the constraints of the GENE code and its development model, justify the decisions made, and describe the path taken in porting GENE to GPUs. We introduce a novel library called gtensor that was developed along the way to support the process. Performance results are presented for the ported code, which in a single node of the Summit supercomputer achieves a speed-up of almost 15× compared to running on central processing unit (CPU) only. Typical GPU kernels are memory-bound, achieving about 90% of peak. Our analysis shows that there is still room for improvement if we can refactor/fuse kernels to achieve higher arithmetic intensity. We also performed a weak parallel scalability study, which shows that the code runs well on a massively parallel system, but communication costs start becoming a significant bottleneck.

Germaschewski, K. (ORCID:0000000284956354)↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

EXAGRAPH: Graph and combinatorial methods for enabling exascale applications

Combinatorial algorithms in general and graph algorithms in particular play a critical enabling role in numerous scientific applications. However, the irregular memory access nature of these algorithms makes them one of the hardest algorithmic kernels to implement on parallel systems. With tens of billions of hardware threads and deep memory hierarchies, the exascale computing systems in particular pose extreme challenges in scaling graph algorithms. The codesign center on combinatorial algorithms, ExaGraph, was established to design and develop methods and techniques for efficient implementation of key combinatorial (graph) algorithms chosen from a diverse set of exascale applications. Algebraic and combinatorial methods have a complementary role in the advancement of computational science and engineering, including playing an enabling role on each other. In this paper, we survey the algorithmic and software development activities performed under the auspices of ExaGraph from both a combinatorial and an algebraic perspective. In particular, we detail our recent efforts in porting the algorithms to manycore accelerator (GPU) architectures. We also provide a brief survey of the applications that have benefited from the scalable implementations of different combinatorial algorithms to enable scientific discovery at scale. We believe that several applications will benefit from the algorithmic and software tools developed by the ExaGraph team.

97 MATHEMATICS AND COMPUTING↗

Priority-BF: A Task Manager for Priority-Based Scheduling

The increasing demand for computational resources, particularly in High-Performance Computing environments, necessitates to rethink how we handle job scheduling strategies. This work addresses the challenge of managing concurrent jobs with differing priorities on overloaded parallel systems, where strict QoS constraints are often difficult for users to define. Our solution relies on a qualitative description of priorities and pulls from two key approaches: the Easy-BF algorithm and the Conservative Backfilling algorithms. This solution improves the response time for high-priority jobs by 50% without affecting the overall system utilization. We show its applicability in several critical scenarios such as High-Performance Computing (HPC) resource management and in-situ computing.

Gainaru, Ana [ORNL]↗

PUMIPic: A mesh-based approach to unstructured mesh Particle-In-Cell on GPUs

Unstructured mesh particle-in-cell, PIC, simulations executing on the current and next generation of massively parallel systems require new methods for both the mesh and particles to achieve performance and scalability on GPUs. The traditional approach to implementing PIC simulations defines data structures and algorithms in terms of particles with a full copy of the unstructured mesh on every process. To effectively scale the unstructured mesh and particles, mesh-based PIC uses the unstructured mesh as the predominant data structure with the particles stored in terms of the mesh entities. Here, this paper details the PUMIPic library, a framework for developing efficient and performance-portable mesh-based PIC simulations on GPU systems. A pseudo physics simulation based on a five-dimensional gyro-kinetic code for modeling plasma physics is used to examine the performance of PUMIPic. Scaling studies of the unstructured mesh partition and number of particles are performed up to 4096 nodes of the Summit system at Oak Ridge National Laboratory. The studies show that mesh-based PIC can utilize a partitioned mesh and maintain scaling up to system limitations.

97 MATHEMATICS AND COMPUTING↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Modeling Data Movement Performance on Heterogeneous Architectures

The cost of data movement on parallel systems varies greatly with machine architecture, job partition, and nearby jobs. Performance models that accurately capture the cost of data movement provide a tool for analysis, allowing for communication bottlenecks to be pinpointed. Modern heterogeneous architectures yield increased variance in data movement as there are a number of viable paths for inter-GPU communication. In this paper, we present performance models for the various paths of inter-node communication on modern heterogeneous architectures, including the trade-off between GPUDirect communication and copying to CPUs. Furthermore, we present a novel optimization for inter-node communication based on these models, utilizing all available CPU cores per node. Finally, we show associated performance improvements for MPI collective operations.

97 MATHEMATICS AND COMPUTING↗

ADG: automated generation and evaluation of many-body diagrams

The goal of the present paper is twofold. First, a novel expansion many-body method applicable to superfluid open-shell nuclei, the so-called Bogoliubov in-medium similarity renormalization group (BIMSRG) theory, is formulated. This generalization of standard single-reference IMSRG theory for closed-shell systems parallels the recent extensions of coupled cluster, self-consistent Green’s function or many-body perturbation theory. Within the realm of IMSRG theories, BIMSRG provides an interesting alternative to the already existing multi-reference IMSRG (MR-IMSRG) method applicable to open-shell nuclei. The algebraic equations for low-order approximations, i.e., BIMSRG(1) and BIMSRG(2), can be derived manually without much difficulty. However, such a methodology becomes already impractical and error prone for the derivation of the BIMSRG(3) equations, which are eventually needed to reach high accuracy. Based on a diagrammatic formulation of BIMSRG theory, the second objective of the present paper is thus to describe the third version (v3.0) of the code that automatically (1) generates all valid BIMSRG(n) diagrams and (2) evaluates their algebraic expressions in a matter of seconds. This is achieved in such a way that equations can easily be retrieved for both the flow equation and the Magnus expansion formulations of BIMSRG. Expanding on this work, the first future objective is to numerically implement BIMSRG(2) (eventually BIMSRG(3)) equations and perform ab initio calculations of mid-mass open-shell nuclei.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Management and Storage of Scientific Data

Scientific discoveries rely heavily on efficient access, search, and management of massive data sets. Data management technologies have, for decades, provided foundational capabilities for scientific computing. Just as storage, input/output (I/O), and data management have been fundamental to simulation-based science for many years, so too are capable data-management technologies key to the success of today’s scientific workflows utilizing data intensive and machine learning (ML) techniques. The Department of Energy, Office of Science, Advanced Scientific Computing Research (ASCR) program has invested broadly in data-management research focused on high-performance computing (HPC) systems, from parallel file systems that store data to application software that makes these systems more productive. Still, advances in technology combined with growing diversity of supported science strongly motivate continued investment in this area. In January 2022, ASCR convened a workshop to identify priority research directions in the area of data management for high-performance and scientific computing. Attendees were challenged to identify promising approaches that would support the breadth of the DOE mission, including the explosion of artificial intelligence (AI) uses and the growing needs of experimental and observational science. Technological and science drivers were identified and considered as they relate to key aspects of data management such as interfaces, architectural design, and FAIR principles (Findable, Accessible, Interoperable, and Reusable). The thoughts of the workshop participants were distilled into a set of four priority research directions with the potential for high impact on DOE science. These research directions are summarized in the following pages.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Charon User Manual: v. 2.2 (revision1)

This manual gives usage information for the Charon semiconductor device simulator. Charon was developed to meet the modeling needs of Sandia National Laboratories and to improve on the capabilities of the commercial TCAD simulators; in particular, the additional capabilities are running very large simulations on parallel computers and modeling displacement damage and other radiation effects in significant detail. The parallel capabilities are based around the MPI interface which allows the code to be ported to a large number of parallel systems, including linux clusters and proprietary “big iron” systems found at the national laboratories and in large industrial settings.

42 ENGINEERING↗

CompLaB v1.0: a scalable pore-scale model for flow, biogeochemistry, microbial metabolism, and biofilm dynamics

Abstract. Microbial activity and chemical reactions in porous media depend on the local conditions at the pore scale and can involve complex feedback with fluid flow and mass transport. We present a modeling framework that quantitatively accounts for the interactions between the bio(geo)chemical and physical processes and that can integrate genome-scale microbial metabolic information into a dynamically changing, spatially explicit representation of environmental conditions. The model couples a lattice Boltzmann implementation of Navier–Stokes (flow) and advection–diffusion-reaction (mass conservation) equations. Reaction formulations can include both kinetic rate expressions and flux balance analysis, thereby integrating reactive transport modeling and systems biology. We also show that the use of surrogate models such as neural network representations of in silico cell models can speed up computations significantly, facilitating applications to complex environmental systems. Parallelization enables simulations that resolve heterogeneity at multiple scales, and a cellular automaton module provides additional capabilities to simulate biofilm dynamics. The code thus constitutes a platform suitable for a range of environmental, engineering and – potentially – medical applications, in particular ones that involve the simulation of microbial dynamics.

58 GEOSCIENCES↗

A Computationally Improved Heuristic Algorithm for Transmission Switching Using Line Flow Thresholds for Load Shed Reduction

We present a computationally improved heuristic algorithm for transmission switching (TS) to recover load shed. Research from the past showed that changing power system topology may control power flows and remove line congestion. Hence, TS may reduce the required load shed. One of the main challenges is to find a potential TS candidate in a suitable time. Here, we propose a novel heuristic method that is capable of finding the potential TS candidate faster than existing algorithms in literature. The proposed method is compatible with both the AC and DC optimal power flows (OPF). Three metrics are used to compare the proposed algorithm with the state-of-the-art from literature to show the speedup and accuracy achieved. The proposed method is implemented on the IEEE 30-bus system, PEGASE 89-bus system, IEEE 118-bus system, and Polish 2383- bus system. The results on the large-scale Polish 2383-bus system shows that the proposed algorithm is scalable to large real-world systems. Parallel computing is implemented to further improve the computational performance of the proposed algorithm.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Bringing OpenCL to Commodity RISC-V CPUs

The importance of open-source hardware has been increasing in recent years with the introduction of the RISC-V Open ISA. This has also accelerated the push for support of the open-source software stack from compiler tools to full-blown operating systems. Parallel computing with today’s Application Programming Interfaces such as OpenCL has proven to be effective at leveraging the parallelism in commodity multi-core processors and programmable parallel accelerators. However, to the best of our knowledge, there is currently no publicly available implementation of OpenCL targeting commodity RISC-V processors that is accessible to the open-source community. Besides opening RISC-V to the existing rich variety of scientific parallel applications, OpenCL also provides access to a unique genre of benchmarks useful in computer architecture research. In this work, we extended an Open-source implementation of OpenCL to target RISC-V CPUs. Our work not only cover commodity multi-core RISC-V processors, but also plethora of low- profile embedded RISC-V CPUs that often do not support atomic instructions or multi-threading.

Tine, Blaise↗

MOOSE: A Modular Platform for Fission and Fusion Multiphysics

The Multiphysics Object-Oriented Simulation Environment (MOOSE) Framework, as well as MOOSE-based simulation tools, have accelerated the development of fission energy and advanced reactor technologies through the United States Department of Energy, Office of Nuclear Science, Nuclear Energy Advanced Modeling & Simulation (NEAMS) Program. MOOSE contains a complete platform of multiphysics simulation capabilities, capable of running on massively parallel systems, and is developed in an open-source manner with great attention paid to high-quality software quality assurance practices. This overall approach could greatly benefit the fusion energy community, which requires rapid design iteration and improvement in order to facilitate the successful development of fusion as an alternative energy source to fossil fuels. In the first half of this talk, applications of MOOSE and MOOSE-based tools for advanced reactor designs will be showcased, as well as MOOSE ecosystem infrastructure (such as the NEAMS Virtual Test Bed) that enables and accelerates fission reactor design. In the second half, a discussion of how the MOOSE approach to modeling and simulation is currently being applied internationally in fusion energy research and development at the United Kingdom Atomic Energy Authority will be discussed, and ongoing/future domestic research efforts will be highlighted.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

PRISMA: PARALLEL REFINEMENT AND INTEGRATION SYSTEM FOR MULTI-AZIMUTHAL ANALYSIS

The Parallel Refinement and Integration System for Multi-azimuthal Analysis (PRISMA, version 1.1.0) is a Python application for processing X-ray diffraction (XRD) image data. PRISMA wraps GSAS-II to perform azimuthally-binned peak refinement, computes per-frame strain and d-spacing from those fits, and provides three PyQt5 graphical interfaces: (1) a Recipe Builder for selecting GSAS-II control (.imctrl) files, optional mask (.immask) files or threshold-ased masking, reference and experiment image sets, peaks, zimuthal range and bin size, and an optional ceria-based auto-calibration; (2) a Batch Processor that uses Dask on local workstations and pure MPI (mpi4py.futures.MPICommExecutor) on HPC to distribute GSAS-II refinement across cores or compute nodes and write results to a 4-dimensional (peaks x frames x azimuths x measurements) Zarr dataset; and (3) a Data Analyzer that renders heatmaps of fit parameters, strain, frame-to-frame deltas, and percent-change-vs-reference, and exports user-defined subsections to CSV or Excel. The peak-refinement algorithm is deterministic. Benchmark on ALCF Crux: a 20,000-image set, single-peak fit in frame mode with 44 azimuthal bins on 128 nodes x 128 workers, 48 seconds total wall time.

Lorenzo Martin, Maria De La Cinta [Argonne Nationa↗

Parallelized multiple nozzle system and method to produce layered droplets and fibers for microencapsulation

The present disclosure relates to a nozzle system for use in a microfluidic production application for producing at least one of particles, capsules or fibers. The system has a main body portion having a compressed fluid inlet and a core fluid inlet, and a plurality of parallel arranged core fluid nozzles that receive the core fluid and create a plurality of core fluid streams. At least one compressed fluid inlet associated with the main body channels compressed fluid to areas adjacent ends of the core fluid nozzles. An apertured plate having a plurality of apertures is arranged near the ends of the core fluid nozzles, with each aperture being uniquely associated with a single one of the core fluid nozzles. The compressed fluid acts on the core fluid streams exiting the core fluid nozzles to help create, with the apertures, at least one of core fluid droplets or core fluid fibers from the core fluid streams.

Ye, Congwang↗

Stability Analysis of Parallel Connected Bidirectional WPT System

This paper presents a stability analysis of parallel-connected bi-directional series-series resonant network wireless power transfer (WPT), optimized for Electric Vehicle (EV) charging and vehicle-to-grid (V2G) applications. The study addresses critical stability challenges in systems integrated with diverse distributed energy resources (DERs), including photovoltaics, fuel cells, wind turbines, energy storage systems, and the AC grid. The stability of such integrated DC grid systems is paramount for ensuring reliable operation, particularly under varying power flow conditions and dynamic interactions between parallel WPT systems. The analysis included system impedance characterization, state-space modeling, and open and closed-loop stability evaluations. The results demonstrated that the integration of a robust control architecture effectively mitigates instability risks and supports scalable, efficient operation. This work underscores the converter's adaptability and its potential for large-scale deployment in wireless EV charging infrastructures and integrated DC grid systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗