Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel computer architecture”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

A Block-Based Triangle Counting Algorithm on Heterogeneous Environments

Triangle counting is a fundamental building block in graph algorithms. In this article, we propose a block-based triangle counting algorithm to reduce data movement during both sequential and parallel execution. Our block-based formulation makes the algorithm naturally suitable for heterogeneous architectures. The problem of partitioning the adjacency matrix of a graph is well-studied. Our task decomposition goes one step further: it partitions the set of triangles in the graph. By streaming these small tasks to compute resources, we can solve problems that do not fit on a device. We demonstrate the effectiveness of our approach by providing an implementation on a compute node with multiple sockets, cores and GPUs. The current state-of-the-art in triangle enumeration processes the Friendster graph in 2.1 seconds, not including data copy time between CPU and GPU. Using that metric, our approach is 20 percent faster. When copy times are included, our algorithm takes 3.2 seconds. This is 5.6 times faster than the fastest published CPU-only time.

97 MATHEMATICS AND COMPUTING↗

A Block-Based Triangle Counting Algorithm on Heterogeneous Environments

Triangle counting is a fundamental building block in graph algorithms. In this paper, we propose a block-based triangle counting algorithm to reduce data movement during both sequential and parallel execution. Our block-based formulation makes the algorithm naturally suitable for heterogeneous architectures. The problem of partitioning the adjacency matrix of a graph is well-studied. Our task decomposition goes one step further: it partitions the set of triangles in the graph. By streaming these small tasks to compute resources, we can solve problems that do not fit on a device. We demonstrate the effectiveness of our approach by providing an implementation on a compute node with multiple sockets, cores and GPUs. The current state-of-the-art in triangle enumeration processes the Friendster graph in 2.1 seconds, not including data copy time between CPU and GPU. Using that metric, our approach is 20 percent faster. When copy times are included, our algorithm takes 3.2 seconds. This is 5.6 times faster than the fastest published CPU-only time.

97 MATHEMATICS AND COMPUTING↗

Artificial Intelligence for Multiphysics Nuclear Design Optimization with Additive Manufacturing

The geometric flexibility of additively manufactured metals and ceramics generates a very large and open design space that requires advanced modeling and simulation tools for physics simulations and the rigorous definition of design problems. This effort deploys artificial intelligence (AI) and machine learning (ML) algorithms to understand the design space, evaluate potential designs, and more efficiently generate optimized results. The Transformational Challenge Reactor (TCR) program is leveraging advances in several scientific areas—including materials, manufacturing, sensors and control systems, data analytics, and high-fidelity modeling and simulation—to accelerate the design, manufacturing, qualification, and deployment of advanced nuclear energy systems. Through a manufacturing-informed design approach, the TCR program seeks to integrate digital data for rapid nuclear innovation; accelerate the adoption of advances in manufacturing, materials, and computational sciences for nuclear applications; and dramatically reduce deployment costs and timelines for new nuclear reactor technologies. This report documents efforts under the TCR program to leverage advanced modeling and simulation techniques driven by AI/ML algorithms on high-performance computing (HPC) systems to yield more optimized TCR core designs. A multiphysics ML surrogate model was developed to run on the HPC architectures. The surrogate model is trained on high-fidelity simulation data of coupled neutronics and thermofluidics and is used to quickly evaluate thousands of candidate core designs in parallel, which drives the evolution of the cooling channel shapes to minimize temperature peaking and material stress. Outcomes from these activities provide design information and feedback into the core design efforts.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Machine Committee Framework for Power Grid Disturbances Analysis Using Synchrophasors Data

Events detection is a key challenge in power grid frequency disturbances analysis. Accurate detection of events is crucial for situational awareness of the power system. In this paper, we study the problem of events detection in power grid frequency disturbance analysis using synchrophasors data streams. Current events detection approaches for power grid rely on individual detection algorithm. This study integrates some of the existing detection algorithms using the concept of machine committee to develop improved detection approaches for grid disturbance analysis. Specifically, we propose two algorithms—an Event Detection Machine Committee (EDMC) algorithm and a Change-Point Detection Machine Committee (CPDMC) algorithm. Both algorithms use parallel architecture to fuse detection knowledge of its individual methods to arrive at an overall output. The EDMC algorithm combines five individual event detection methods, while the CPDMC algorithm combines two change-point detection methods. Each method performs the detection task separately. The overall output of each algorithm is then computed using a voting strategy. The proposed algorithms are evaluated using three case studies of actual power grid disturbances. Compared with the individual results of the various detection methods, we found that the EDMC algorithm is a better fit for analyzing synchrophasors data; it improves the detection accuracy; and it is suitable for practical scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decoding the protein–ligand interactions using parallel graph neural networks

Abstract Protein–ligand interactions (PLIs) are essential for biochemical functionality and their identification is crucial for estimating biophysical properties for rational therapeutic design. Currently, experimental characterization of these properties is the most accurate method, however, this is very time-consuming and labor-intensive. A number of computational methods have been developed in this context but most of the existing PLI prediction heavily depends on 2D protein sequence data. Here, we present a novel parallel graph neural network (GNN) to integrate knowledge representation and reasoning for PLI prediction to perform deep learning guided by expert knowledge and informed by 3D structural data. We develop two distinct GNN architectures: $$\hbox {GNN}_{\mathrm{F}}$$ GNN F is the base implementation that employs distinct featurization to enhance domain-awareness, while $$\hbox {GNN}_{\mathrm{P}}$$ GNN P is a novel implementation that can predict with no prior knowledge of the intermolecular interactions. The comprehensive evaluation demonstrated that GNN can successfully capture the binary interactions between ligand and protein’s 3D structure with 0.979 test accuracy for $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and 0.958 for $$\hbox {GNN}_{\mathrm{P}}$$ GNN P for predicting activity of a protein–ligand complex. These models are further adapted for regression tasks to predict experimental binding affinities and $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 crucial for compound’s potency and efficacy. We achieve a Pearson correlation coefficient of 0.66 and 0.65 on experimental affinity and 0.50 and 0.51 on $$\hbox {pIC}_{\mathrm{50}}$$ pIC 50 with $$\hbox {GNN}_{\mathrm{F}}$$ GNN F and $$\hbox {GNN}_{\mathrm{P}}$$ GNN P , respectively, outperforming similar 2D sequence based models. Our method can serve as an interpretable and explainable artificial intelligence (AI) tool for predicted activity, potency, and biophysical properties of lead candidates. To this end, we show the utility of $$\hbox {GNN}_{\mathrm{P}}$$ GNN P on SARS-Cov-2 protein targets by screening a large compound library and comparing the prediction with the experimentally measured data.

59 BASIC BIOLOGICAL SCIENCES↗

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

TChem v2.0 - A Software Toolkit for the Analysis of Complex Kinetic Models

TChem is an open source software library for solving complex computational chemistry problems and analyzing detailed chemical kinetic models. The software provides support for: complex kinetic models for gas-phase and surface chemistry; thermodynamic properties based on NASA polynomials; species production/consumption rates; stable time integrator for solving stiff time ordinary differential equations; and, reactor models such as homogenous gas-phase ignition (with analytical Jacobian matrices), continuously stirred tank reactor, plug-flow reactor. This toolkit builds upon earlier versions that were written in C and featured tools for gas-phase chemistry only. The current version of the software was completely refactored in C++, uses an object-oriented programming model, and adopts Kokkos as its portability layer to make it ready for the next generation computing architectures i.e., multi/many core computing platforms with GPU accelerators. We have expanded the range of kinetic models to include surface chemistry and have added examples pertaining to Continuously Stirred Tank Reactors (CSTR) and Plug Flow Reactor (PFR) models to complement the homogenous ignition examples present in the earlier versions. To exploit the massive parallelism available from modern computing platforms, the current software interface is designed to evaluate samples in parallel, which enables large scale parametric studies, e.g. for sensitivity analysis and model calibration.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Advancement of hybrid fluid-kinetic modeling for HEDP and ICF science

We report on the development progress of a hybrid fluid-kinetic code for simulating fluids and plasmas in a wide range of environments, such as laser–matter interactions, inertial confinement fusion, magnetic confinement fusion, and pulsed power. The suite of numerical tools under development utilizes heterogeneous computer architectures and leverages the benefits of particle–based simulation techniques. By working to combine the kinetic particle-in-cell (PIC) model with a particle-based fluid simulation technique, such as smoothed particle hydrodynamics, we are developing a flexible framework capable of accurately modeling complex flows within and between kinetic and fluid regimes. The TriForce code is under development as a C++ framework for parallel, 3D, particle-based, hybrid fluid-kinetic plasma simulations. The fluid half of TriForce will be based upon the meshless smoothed-particle-hydrodynamics (SPH) approach, well-suited for shear, mixing, and turbulence, whereas the kinetic half resembles a traditional particle-in-cell (PIC) code; other particle-based approaches to fluid modeling that do use a mesh are also possible to use and are under investigation. Maxwell’s electromagnetic field equations are solved either via explicit or implicit algorithms or approximated via resistive magnetohydrodynamics (MHD) using an Ohm’s law and resulting induction equation (extended MHD is under development). A primary goal of enabling direct comparisons, from the same code, between results from the variants of MHD and implicit electromagnetic solutions is to improve our fundamental understanding of systems with magnetic fields. The code is under development to recover results from both radiation-MHD and fully kinetic codes in those limits, and is continuing to be developed from other follow-on grants to operate in between where both descriptions may co-exist and interact. For certain applications, it is desired for a simulation to contain fluid ions and electrons as well as kinetic ions and electrons. Typically, it is too computationally intensive to model a full-scale ICF or HEDP experiment fully kinetically since many cycles are expended with very small time steps on modeling the fluid part of a material that is well treated by the fluid approximation. In this case, many traditional PIC particles can be replaced with a single fluid particle representing the thermal part of the distribution function, and there are fewer needed kinetic particles, which describe the non-thermal part and can be sub-cycled relative to the fluid particle advance. Furthermore, a pure fluid code may, depending on the problem, simply lack many physically important details that are beyond the scope of its reduced approximations and assumptions. In this report, we summarize the objectives achieved in the development of the collisional and kinetic half of the code, and the physics problems to which the code has been applied in the areas of advanced and innovative fusion concepts, pulsed power, and magneto-inertial fusion.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hybrid programming-model strategies for GPU offloading of electronic structure calculation kernels

To address the challenge of performance portability and facilitate the implementation of electronic structure solvers, we developed the basic matrix library (BML) and Parallel, Rapid O(N), and Graph-based Recursive Electronic Structure Solver (PROGRESS) library. The BML implements linear algebra operations necessary for electronic structure kernels using a unified user interface for various matrix formats (dense and sparse) and architectures (CPUs and GPUs). Focusing on density functional theory and tight-binding models, PROGRESS implements several solvers for computing the single-particle density matrix and relies on BML. In this paper, we describe the general strategies used for these implementations on various computer architectures, using OpenMP target functionalities on GPUs, in conjunction with third-party libraries to handle performance critical numerical kernels. In this study, we demonstrate the portability of this approach and its performance in benchmark problems.

36 MATERIALS SCIENCE↗

Weighted relaxation for multigrid reduction in time

Current trends in computer architectures now mean that faster computation speed must come primarily from increased concurrency, not faster clock speeds, which are stagnating. Thus, this situation creates bottlenecks for serial algorithms, including the well-known bottleneck for sequential time-integration, where each individual time-value (i.e., time-step) is computed sequentially. One approach to alleviate this and achieve parallelism in time is with multigrid. Here, in this work, we consider multigrid-reduction-in-time (MGRIT), a multilevel method applied to the time dimension that computes multiple time-steps in parallel. Like all multigrid methods, MGRIT relies on the complementary relationship between relaxation on a fine-grid and a correction from the coarse grid to solve the problem. All current MGRIT implementations are based on unweighted-Jacobi relaxation; here we introduce the concept of weighted relaxation to MGRIT. We derive new convergence bounds for weighted relaxation, and use this analysis to guide the selection of relaxation weights. Numerical results then demonstrate that by choosing appropriate non-unitary relaxation weights, one can achieve faster convergence rates and lower iteration counts for MGRIT when compared with unweighted relaxation. In most cases, weighted relaxation yields a 10%–20% saving in iterations, which is significant when using large high-performance computers. For A-stable integration schemes, results also illustrate that under-relaxation can restore convergence in some cases where unweighted relaxation is not convergent.

97 MATHEMATICS AND COMPUTING↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Toward a 2D Local Implementation of Quantum Low-Density Parity-Check Codes

Geometric locality is an important theoretical and practical factor for quantum low-density parity-check (qLDPC) codes that affects code performance and ease of physical realization. For device architectures restricted to two-dimensional (2D) local gates, naively implementing the high-rate codes suitable for low-overhead fault-tolerant quantum computing incurs prohibitive overhead. In this work, we present an error-correction protocol built on a bilayer architecture that aims to reduce operational overheads when restricted to 2D local gates by measuring some generators less frequently than others. We investigate the family of bivariate-bicycle qLDPC codes and show that they are well suited for a parallel syndrome-measurement scheme using fast routing with local operations and classical communication (LOCC). Through circuit-level simulations, we find that in some parameter regimes, bivariate-bicycle codes implemented with this protocol have logical error rates comparable to the surface code while using fewer physical qubits. Published by the American Physical Society 2025

Berthusen, Noah (ORCID:0000000275862786)↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

Towards a more general understanding of the algorithmic utility of recurrent connections

Lateral and recurrent connections are ubiquitous in biological neural circuits. Yet while the strong computational abilities of feedforward networks have been extensively studied, our understanding of the role and advantages of recurrent computations that might explain their prevalence remains an important open challenge. Foundational studies by Minsky and Roelfsema argued that computations that require propagation of global information for local computation to take place would particularly benefit from the sequential, parallel nature of processing in recurrent networks. Such “tag propagation” algorithms perform repeated, local propagation of information and were originally introduced in the context of detecting connectedness, a task that is challenging for feedforward networks. Here, we advance the understanding of the utility of lateral and recurrent computation by first performing a large-scale empirical study of neural architectures for the computation of connectedness to explore feedforward solutions more fully and establish robustly the importance of recurrent architectures. In addition, we highlight a tradeoff between computation time and performance and construct hybrid feedforward/recurrent models that perform well even in the presence of varying computational time limitations. We then generalize tag propagation architectures to propagating multiple interacting tags and demonstrate that these are efficient computational substrates for more general computations of connectedness by introducing and solving an abstracted biologically inspired decision-making task. Our work thus clarifies and expands the set of computational tasks that can be solved efficiently by recurrent computation, yielding hypotheses for structure in population activity that may be present in such tasks.

59 BASIC BIOLOGICAL SCIENCES↗

The VTK-m Users' Guide (V.2.0)

High-performance computing relies on ever finer threading. Advances in processor technology include ever greater numbers of cores, hyperthreading, accelerators with integrated blocks of cores, and special vectorized instructions, all of which require more software parallelism to achieve peak performance. Traditional visualization solutions cannot support this extreme level of concurrency. Extreme scale systems require a new programming model and a fundamental change in how we design algorithms. To address these issues we created VTK-m: the visualization toolkit for multi-/many-core architectures. VTK-m supports a number of algorithms and the ability to design further algorithms through a top-down design with an emphasis on extreme parallelism. VTK-m also provides support for finding and building links across topologies, making it possible to perform operations that determine manifold surfaces, interpolate generated values, and find adjacencies. Although VTK-m provides a simplified high-level interface for programming, its template-based code removes the overhead of abstraction. VTK-m simplifies the development of parallel scientific visualization algorithms by providing a framework of supporting functionality that allows developers to focus on visualization operations. Consider the listings in Figure 1.1 that compares the size of the implementation for the Marching Cubes algorithm in VTK-m with the equivalent reference implementation in the CUDA software development kit. Because VTK-m internally manages the parallel distribution of work and data, the VTK-m implementation is shorter and easier to maintain. Additionally, VTK-m provides data abstractions not provided by other libraries that make code written in VTK-m more versatile.

97 MATHEMATICS AND COMPUTING↗

Distributed Training for High Resolution Images: A Domain and Spatial Decomposition Approach

In this work we developed two Pytorch libraries using the PyTorch RPC interface for distributed deep learning approaches on high resolution images. The spatial decomposition library allows for distributedtraining on very large images, which otherwise won’t be possible on a single GPU. The domain parallelism library allows for distributed training across multiple domain unlabeled data, by leveraging the domain separation architecture. Both of those libraries were tested on the Summit supercomputer at a moderate scale, and we are releasing the code for both of them.

97 MATHEMATICS AND COMPUTING↗

Portage: A Modular Data Remap Library for Multiphysics Applications on Advanced Architectures

Portage is a scalable and extensible remap library for numerical simulations. It supports state-of-the-art remap schemes for meshes and particles in 2D and 3D up to a second-order accuracy. Portage ensures critical properties such as local/global conservation and bounds preservation for mesh remap. It enables multi-material field remap through a dedicated plugin, and leverages the hybrid parallelism exposed by advanced architectures using multi-processing and multi-threading.

97 MATHEMATICS AND COMPUTING↗

HPDR: High-Performance Portable Scientific Data Reduction Framework

The rapid growth in scientific data generation is outpacing advancements in computing systems necessary for efficient storage, transfer, and analysis, particularly in the context of exascale computing. With the deployment of first-generation exascale computing systems and next-generation experimental facilities, this gap is widening and necessitates effective data reduction techniques to manage enormous data volumes. Over the past decade, various data reduction methods, including lossless compression, error-controlled lossy compression, and data refactoring, have been developed to accelerate I/O in scientific workflows. Despite significant reductions in data volume, these methods introduce considerable computational overhead, which can become the new bottleneck in data processing. To mitigate this, GPU-accelerated data reduction algorithms have been introduced. However, challenges remain in their integration into exascale workflows, including limited portability across different GPU architectures, substantial memory transfer overhead, and reduced scalability on dense multi-GPU systems. To address these challenges, we propose HPDR, a high-performance and portable data reduction framework. HPDR is designed to enable the execution of state-of-the-art reduction algorithms across diverse processor architectures while reducing memory transfer overhead to 2.3 % of the original, resulting in up to 3.5× faster throughput compared to existing solutions. It also achieves up to 96% of the theoretical speedup in multi-GPU settings. In addition, evaluations on accelerating I/O operations at scale up to 1,024 nodes of the Frontier supercomputer demonstrate that HPDR can achieve up to 103 TB/s reduction throughput, providing up to 4× acceleration in parallel I/O performance compared to existing data reduction routines. This work highlights the potential of HPDR to significantly enhance data reduction efficiency in exascale computing environments.

Chen, Jieyang [University of Oregon]↗