Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel and high performance computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Topanga: A kinetic ion plasma code for large-scale ionospheric simulations on magnetohydrodynamic timescales

Topanga is a kinetic ion code developed for simulating large-scale plasma phenomena in the Earth's ionosphere on magnetohydrodynamic timescales. It is a domain-decomposed parallel code that runs on high-performance computing platforms. Features of Topanga include spherical geometry for simplified boundary conditions and computational efficiency; a hybrid plasma model with inertia-less fluid electrons, kinetic ions, and an electric field specified via an Ohm's law; a Maxwell-FDTD (finite difference time domain) plasma model which retains the displacement current in Maxwell's equations and models electron currents in the ionosphere with a tensor conductivity; sponge-layer boundary conditions for absorption of electromagnetic and plasma waves incident on the domain boundaries; and a novel mixed-implicit algorithm for evolving the EM fields inside the Maxwell-FDTD region that is stable over many orders of magnitude in the electron–ion collision frequency. We verify the numerical methods used in Topanga on a pair of test problems. The first test involves modeling a three-dimensional collisionless shock using the hybrid set of equations. The second test involves modeling a spherical TEM mode in vacuum using the Maxwell-FDTD set of equations. Finally, we demonstrate how using the combined set of hybrid and Maxwell-FDTD equations to model the Starfish Prime high-altitude nuclear test recovers a “missing” EM signal on the ground that is not present when using only the hybrid set of equations. The magnitude of this signal in the simulation containing the Maxwell-FDTD region agrees well with the E3a portion of the magnetohydrodynamic electromagnetic pulse from Starfish Prime.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Supercomputing Pipelines Search for Therapeutics Against COVID-19

The urgent search for drugs to combat SARS-CoV-2 has included the use of supercomputers. The use of general-purpose graphical processing units (GPUs), massive parallelism, and new software for high-performance computing (HPC) has allowed researchers to search the vast chemical space of potential drugs faster than ever before. We developed a new drug discovery pipeline using the Summit supercomputer at Oak Ridge National Laboratory to help pioneer this effort, with new platforms that incorporate GPU-accelerated simulation and allow for the virtual screening of billions of potential drug compounds in days compared to weeks or months for their ability to inhibit SARS-COV-2 proteins. Here, this effort will accelerate the process of developing drugs to combat the current COVID-19 pandemic and other diseases.

60 APPLIED LIFE SCIENCES↗

It’s Time to Talk About HPC Storage: Perspectives on the Past and Future

High-performance computing (HPC) storage systems are a key component of the success of HPC to date. Recently, we have seen major developments in storage-related technologies, as well as changes to how HPC platforms are used, especially in relation to artificial intelligence and experimental data analysis workloads. Additionally, these developments merit a revisit of HPC storage system architectural designs. In this article, we discuss the drivers, identify key challenges to status quo posed by these developments, and discuss directions future research might take to unlock the potential of new technologies for the breadth of HPC applications.

97 MATHEMATICS AND COMPUTING↗

VAN-DAMME: GPU-accelerated and symmetry-assisted quantum optimal control of multi-qubit systems

We present an open-source software package, VAN-DAMME (Versatile Approaches to Numerically Design, Accelerate, and Manipulate Magnetic Excitations), for massively-parallelized quantum optimal control (QOC) calculations of multi-qubit systems. To enable large QOC calculations, the VAN-DAMME software package utilizes symmetry-based techniques with custom GPU-enhanced algorithms. This combined approach allows for the simultaneous computation of hundreds of matrix exponential propagators that efficiently leverage the intra-GPU parallelism found in high-performance GPUs. In addition, to maximize the computational efficiency of the VAN-DAMME code, we carried out several extensive tests on data layout, computational complexity, memory requirements, and performance. These extensive analyses allowed us to develop computationally efficient approaches for evaluating complex-valued matrix exponential propagators based on Padé approximants. To assess the computational performance of our GPU-accelerated VAN-DAMME code, we carried out QOC calculations of systems containing 10 - 15 qubits, which showed that our GPU implementation is 18.4× faster than the corresponding CPU implementation. Our GPU-accelerated enhancements allow efficient calculations of multi-qubit systems, which can be used for the efficient implementation of QOC applications across multiple domains.

97 MATHEMATICS AND COMPUTING↗

Planetary normal mode computation: Parallel algorithms, performance and reproducibility

This report is an extension of work entitled “Computing planetary interior normal modes with a highly parallel polynomial filtering eigensolver.” by Shi et al., [1] originally presented at the SC18 conference. A highly parallel polynomial filtered eigensolver was developed and exploited to calculate the planetary normal modes. The proposed method is ideally suited for computing interior eigenpairs for large-scale eigenvalue problems as it greatly enhances memory and computational efficiency. In this article, the second-order finite element method is used to further improve the accuracy as only the first-order finite element method was deployed in the previous work. The parallel algorithm, its parallel performance up to 20k processors, and the great computational accuracy are illustrated. The reproducibility of the previous work was successfully performed on the Student Cluster Competition at the SC19 conference by several participant teams using a completely different Mars-model dataset on different clusters. Both weak and strong scaling performances of the reproducibility by the participant teams were impressive and encouraging. The analysis and reflection of their results are demonstrated and future direction is discussed.

97 MATHEMATICS AND COMPUTING↗

ChatMPI: LLM-Driven MPI Code Generation for HPC Workloads

The Message Passing Interface (MPI) standard plays a crucial role in enabling scientific applications for parallel computing and is an essential component in high-performance computing (HPC). However, implementing MPI code manually—especially applying a proper domain decomposition and communication pattern—is a challenging and error-prone task. We present ChatMPI, an AI assistant for MPI parallelization of sequential C codes. In our analysis, we focus on testing six essential HPC workloads, which are based on Basic Linear Algebra Subprograms levels 1, 2, and 3 as well as sparse, stencil, and iterative operations. We analyze the process of creating ChatMPI by using the ChatHPC library. This lightweight large language model (LLM)–based infrastructure enables HPC experts to efficiently create and supervise trustworthy AI capabilities for critical HPC software tasks. We study the data required for training (fine-tuning) ChatMPI to generate parallel codes that not only use MPI syntax correctly but also apply HPC techniques to reduce memory communication and maximize performance by using proper work decomposition. With a relatively small training dataset composed of a few dozen prompts and fewer than 15 minutes of fine-tuning on one node equipped with two NVIDIA H100 GPUs, ChatMPI elevates trustworthiness for MPI code generation of current LLMs (e.g., Code Llama, ChatGPT-4o and ChatGPT 5). Additionally, we evaluate the performance of the MPI codes generated by ChatMPI in comparison with the ones generated by ChatGPT-4o and ChatGPT-5. The codes generated by ChatMPI provide up to a 4 × boost in performance by using better problem decomposition, communication patterns, and HPC techniques (e.g., communication avoiding).

Valero Lara, Pedro [ORNL] (ORCID:0000000214794310)↗

TAO Users Manual (Rev. 3.15)

The Toolkit for Advanced Optimization (TAO) focuses on the development of algorithms and software for the solution of large-scale optimization problems on high-performance architectures. Areas of interest include unconstrained and bound-constrained optimization, nonlinear least squares problems, optimization problems with partial differential equation constraints, and variational inequalities and complementarity constraints. The development of TAO was motivated by the scattered support for parallel computations and the lack of reuse of external toolkits in current optimization software. Our aim is to produce high-quality optimization software for computing environments ranging from workstations and laptops to massively parallel high-performance architectures. Our design decisions are strongly motivated by the challenges inherent in the use of large-scale distributed memory architectures and the reality of working with large, often poorly structured legacy codes for specific applications.

97 MATHEMATICS AND COMPUTING↗

Enabling Low-Overhead HT-HPC Workflows at Extreme Scale using GNU Parallel

GNU Parallel is a versatile and powerful tool for process parallelization widely used in scientific computing. This paper demonstrates its effective application in high-performance computing (HPC) environments, particularly focusing on its scalability and efficiency in executing large-scale high-throughput high-performance computing (HT-HPC) workflows. Through real-world examples, we highlight GNU Parallel’s performance across various HPC workloads, including GPU computing, container-based workloads, and node-local NVMe storage. Our results on two leading supercomputers, OLCF’s Frontier and NERSC’s Perlmutter, showcase GNU Parallel’s rapid process dispatching ability and its capacity to maintain low overhead even at extreme scales. We explore GNU Parallel’s application in massive parallel file transfers using a scheduled Data Transfer Node (DTN) cluster, emphasizing its broad utility in diverse scientific workflows. Beyond its direct application as a viable workflow manager, GNU Parallel can be employed in conjunction with other workflow systems as a "last-mile" parallelizing driver and as a quick prototyping tool to design and extract parallel profiles from application executions. We then argue that the potential for GNU Parallel to transform workflow management at extreme scales is substantial, paving the way for more efficient and effective scientific discoveries.

Maheshwari, Ketan↗

OpenMP Target Task: Tasking and Target Offloading on Heterogeneous Systems

This work evaluated the use of OpenMP tasking with target GPU offloading as a potential solution for programming productivity and performance on heterogeneous systems. Also, it is proposed a new OpenMP specification to make the implementation of heterogeneous codes simpler by using OpenMP target task, which integrates both OpenMP tasking and target GPU offloading in a single OpenMP pragma. As a test case, the authors used one of the most popular and widely used Basic Linear Algebra Subprogram Level-3 routines: triangular solver (TRSM). To benefit from the heterogeneity of the current high-performance computing systems, the authors propose a different parallelization of the algorithm by using a nonuniform decomposition of the problem. This work used target GPU offloading inside OpenMP tasks to address the heterogeneity found in the hardware. This new approach can outperform the state-of-the-art algorithms, which use a uniform decomposition of the data, on both the CPU-only and hybrid CPU-GPU systems, reaching speedups of up to one order of magnitude. The performance that this approach achieves is faster than the IBM ESSL math library on CPU and competitive relative to a highly optimized heterogeneous CUDA version. One node of Oak Ridge National Laboratory’s supercomputer, Summit, was used for performance analysis.

Valero Lara, Pedro↗

Toward a persistent event-streaming system for high-performance computing applications

High-performance computing (HPC) applications have traditionally relied on parallel file systems and file transfer services to manage data movement and storage. Alternative approaches have been proposed that use direct communications between application components, trading persistence and fault tolerance for speed. Event-driven architectures, as popularized in enterprise contexts, present a compelling middle ground, avoiding the performance cost and API constraints of parallel file systems while retaining persistence and offering impedance matching between application components. However, adapting streaming frameworks to HPC workloads requires addressing challenges unique to HPC systems. This paper investigates the potential for a streaming framework designed for HPC infrastructures and use cases. We introduce Mofka, a persistent event-streaming framework designed specifically for HPC environments. Mofka combines the capabilities of a traditional streaming service with optimizations tailored to the HPC context, such as support for massively multicore nodes, efficient scaling for large producer-consumer workflows, RDMA-enabled high-performance network communications, specialized network fabrics with multiple links per node, and efficient handling of large scientific data payloads. Built using the Mochi suite of HPC data service components, Mofka provides a lightweight, modular, and high-performance solution for persistent streaming in HPC systems. We present the architecture of Mofka and evaluate its performance against Kafka and Redpanda using benchmarks on diverse platforms, including Argonne's Polaris and Oak Ridge's Frontier supercomputers, showing up to 8× improvement in throughput in some scenarios. We then demonstrate its utility in several real-world applications: a tomographic reconstruction pipeline, a workflow for the discovery of metal-organic frameworks for carbon capture, and the instrumentation of Dask workflows for provenance tracking and performance analysis.

HPC↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers part I: Computational models and workflow

Computational simulations have become central to the seismic analysis and design of major infrastructure over the past several decades. Most major structures are now “proof tested” virtually through representative simulations of earthquake-induced response. More recently, with the advancement of high-performance computing (HPC) platforms and the associated massively parallel computational ecosystems, simulation is beginning to play a role in increased understanding and prediction of ground motions for earthquake hazard assessments. However, the computational requirements for regional-scale geophysics-based ground motion simulations are extreme, which has restricted the frequency resolution of direct simulations and limited the ability to perform the large number of simulations required to numerically explore the problem parametric space. In this article, recent developments toward an integrated, multidisciplinary earth science-engineering computational framework for the regional-scale simulation of both ground motions and resulting structural response are described with a particular emphasis on advancing simulations to frequencies relevant to engineered systems. This multidisciplinary computational development is being carried out as part of the US Department of Energy (DOE) Exascale Computing Project with the goal of achieving a computational framework poised to exploit emerging DOE exaflop computer platforms scheduled for the 2022–2023 timeframe.

58 GEOSCIENCES↗

High-fidelity parallel entangling gates on a neutral-atom quantum computer

The ability to perform entangling quantum operations with low error rates in a scalable fashion is a central element of useful quantum information processing. Neutral-atom arrays have recently emerged as a promising quantum computing platform, featuring coherent control over hundreds of qubits and any-to-any gate connectivity in a flexible, dynamically reconfigurable architecture. The main outstanding challenge has been to reduce errors in entangling operations mediated through Rydberg interactions. Here we report the realization of two-qubit entangling gates with 99.5% fidelity on up to 60 atoms in parallel, surpassing the surface-code threshold for error correction. Our method uses fast, single-pulse gates based on optimal control, atomic dark states to reduce scattering and improvements to Rydberg excitation and atom cooling. We benchmark fidelity using several methods based on repeated gate applications, characterize the physical error sources and outline future improvements. Finally, we generalize our method to design entangling gates involving a higher number of qubits, which we demonstrate by realizing low-error three-qubit gates. By enabling high-fidelity operation in a scalable, highly connected system, these advances lay the groundwork for large-scale implementation of quantum algorithms, error-corrected circuits and digital simulations.

97 MATHEMATICS AND COMPUTING↗

ZERNIPAX: A fast and accurate Zernike polynomial calculator in Python

Zernike polynomials serve as an orthogonal basis on the unit disc, and have proven to be effective in optics simulations, astrophysics, and more recently in plasma simulations. Unlike Bessel functions, Zernike polynomials are inherently finite and smooth at the disc center (r=0), ensuring continuous differentiability along the axis. This property makes them particularly suitable for simulations, requiring no additional handling at the origin. We developed ZERNIPAX, an open-source Python package capable of utilizing CPU/GPUs, leveraging Google's JAX package and available on GitHub as well as the Python software repository PyPI. Furthermore, our implementation of the recursion relation between Jacobi polynomials significantly improves computation time compared to alternative methods by use of parallel computing while still performing more accurately for high-mode numbers.

Astrophysics↗

Insights From Dayflow: A Historical Streamflow Reanalysis Dataset for the Conterminous United States

Abstract Reconstructed historical streamflow time series can supplement limited streamflow gauge observations. However, there are common challenges of typical modeling approaches: process‐based hydrologic models can be data/computation‐intensive, and statistics‐based models can be region/stream‐specific. Here we present a nationally scalable modeling framework integrating the simulated runoff from the Variable Infiltration Capacity (VIC) model with the Routing Application for Parallel computatIon of Discharge (RAPID) routing model leveraging high‐performance computing. We demonstrate an efficient method of assimilating streamflow at US Geological Survey (USGS) streamflow monitoring sites using a simple hierarchical approach in the VIC‐RAPID framework. The result is a reconstructed 36‐year (1980–2015) daily and monthly streamflow dataset (Dayflow) at ∼2.7 million NHDPlusV2 stream reaches in the conterminous US (CONUS). We perform a comprehensive evaluation at 7,526 USGS sites and characterize their error statistics. The results demonstrate that 49% of the USGS sites demonstrate Kling–Gupta Efficiency (KGE) > 0.5 and 58% of the sites show percentage bias within ±20% for the daily naturalized streamflow. Streamflow data assimilation across CONUS shows an overall improvement over naturalized streamflow, notably in the western semiarid‐to‐arid regions. Comparison to other national and global streamflow reanalysis datasets such as the National Water Model and Global Reach‐scale A priori Discharge Estimates for SWOT demonstrates improved KGE, reduced bias, and directions for Dayflow improvements. Investigations of error statistics with key hydrologic, hydroclimatic, and geomorphologic basin characteristics reveal region‐specific patterns which may help improve future framework applications. Overall, Dayflow may enable a better understanding of hydrologic conditions in a changing environment, especially in locations currently not represented by streamflow monitoring networks.

54 ENVIRONMENTAL SCIENCES↗

Using OpenMP for HEP framework algorithm scheduling

The OpenMP standard is the primary mechanism used at high performance computing facilities to allow intra-process parallelization. In contrast, many HEP specific software packages (such as CMSSW, GaudiHive, and ROOT) make use of Intel’s Threading Building Blocks (TBB) library to accomplish the same goal. In these proceedings we will discuss our work to compare TBB and OpenMP when used for scheduling algorithms to be run by a HEP style data processing framework. This includes both scheduling of different interdependent algorithms to be run concurrently as well as scheduling concurrent work within one algorithm. As part of the discussion we present an overview of the OpenMP threading model. We also explain how we used OpenMP when creating a simplified HEP-like processing framework. Using that simplified framework, and a similar one written using TBB, we will present performance comparisons between TBB and different compiler versions of OpenMP.

97 MATHEMATICS AND COMPUTING↗

MILK : a Python scripting interface to MAUD for automation of Rietveld analysis

Modern diffraction experiments ( e.g. in situ parametric studies) present scientists with many diffraction patterns to analyze. Interactive analyses via graphical user interfaces tend to slow down obtaining quantitative results such as lattice parameters and phase fractions. Furthermore, Rietveld refinement strategies ( i.e. the parameter turn-on-off sequences) tend to be instrument specific or even specific to a given dataset, such that selection of strategies can become a bottleneck for efficient data analysis. Managing multi-histogram datasets such as from multi-bank neutron diffractometers or caked 2D synchrotron data presents additional challenges due to the large number of histogram-specific parameters. To overcome these challenges in the Rietveld software Material Analysis Using Diffraction ( MAUD ), the MAUD Interface Language Kit ( MILK ) is developed along with an updated text batch interface for MAUD . The open-source software MILK is computer-platform independent and is packaged as a Python library that interfaces with MAUD . Using MILK , model selection ( e.g. various texture or peak-broadening models), Rietveld parameter manipulation and distributed parallel batch computing can be performed through a high-level Python interface. A high-level interface enables analysis workflows to be easily programmed, shared and applied to large datasets, and external tools to be integrated with MAUD . Through modification to the MAUD batch interface, plot and data exports have been improved. The resulting hierarchical folders from Rietveld refinements with MILK are compatible with Cinema: Debye–Scherrer , a tool for visualizing and inspecting the results of multi-parameter analyses of large quantities of diffraction data. In this manuscript, the combined Python scripting and visualization capability of MILK is demonstrated with a quantitative texture and phase analysis of data collected at the HIPPO neutron diffractometer.

97 MATHEMATICS AND COMPUTING↗

GentenMPI: Distributed Memory Sparse Tensor Decomposition

GentenMPl is a toolkit of sparse canonical polyadic (CP) tensor decomposition algorithms that is designed to run effectively on distributed-memory high-performance computers. Its use of distributed-memory parallelism enables it to efficiently decompose tensors that are too large for a single compute node's memory. GentenMPl leverages Sandia's decades-long investment in the Trilinos solver framework for much of its parallel-computation capability. Trilinos contains numerical algorithms and linear algebra classes that have been optimized for parallel simulation of complex physical phenomena. This work applies these tools to the data science problem of sparse tensor decomposition. In this report, we describe the use of Trilinos in GentenMPl, extensions needed for sparse tensor decomposition, and implementations of the CP-ALS (CP via alternating least squares) and GCP-SGD (generalized CP via stochastic gradient descent) sparse tensor decomposition algorithms. We show that GentenMPl can decompose sparse tensors of extreme size, e.g., a 12.6-terabyte tensor on 8192 computer cores. We demonstrate that the Trilinos backbone provides good strong and weak scaling of the tensor decomposition algorithms.

97 MATHEMATICS AND COMPUTING↗