Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195↗

Biochemical parallels between catabolic pathways for lignin-associated aromatic dimers

Lignin is one of the most common biopolymers on Earth. In nature, lignin is primarily deconstructed by fungi into mixtures of aromatic compounds that are then assimilated by bacteria and fungi. Industrially, lignin is primarily generated as a byproduct of pulp and paper production and burned for process heat. However, if the appropriate assimilatory pathways were identified, deconstructed lignin could be funneled into value-added products using engineered bacteria. Foundational work has described pathways for assimilation of diverse monomeric aromatic compounds such as protocatechuate, ferulate, and syringate, as well as select dimers including those with β-O-4 and 5-5 interunit linkages. Recent advances have elucidated additional pathways for dimer assimilation, including pathways for new substrates as well as parallel pathways for previously characterized substrates. Comparing these dimer assimilation pathways can illuminate the underlying biochemical logic of assimilation for lignin-associated aromatic dimers and provide opportunities for metabolic engineering to enhance lignin valorization.

Sphingomonas↗

Efficient Parallel Sparse Symmetric Tucker Decomposition for High-Order Tensors

Tensor based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making it an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTMc) in general sparse tensors, the Sparse Symmetric Tensor Times Same Matrix chain (S3TTMc) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. In this work, we present a novel compressed storage format CSS for sparse symmetric tensors, along with an efficient parallel algorithm for the S3TTMc operation. We theoretically establish that S3TTMc on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations. We demonstrate experimental findings that confirm these results and achieve up to 2.9× speedup on synthetic and real datasets.

Shivakumar, Shruti↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗

Determining Levels of Detail for Simulators of Parallel and Distributed Computing Systems via Automated Calibration

There are two sources of inaccuracy when simulating parallel and distributed computing systems: (i) a simulator implemented at an insufficient level of detail; and (ii) incorrectly calibrated simulation parameter values. Increasing the simulator’s level of detail can improve accuracy, but at the cost of higher space, time, and/or software complexity. Furthermore, evaluating the intrinsic accuracy of a simulator requires that its parameters be well-calibrated. Making decisions regarding the level of detail is thus challenging. We propose a methodology for instantiating the simulation calibration process and a framework for automating this process, which makes it possible to pick appropriate levels of detail for any simulator. We demonstrate the usefulness of our approach via two case studies for two different domains.

McDonald, Jessie [University of Hawaii at Manoa, H↗

A Parallel Kinetic Model for Surface and Bulk Charge Storage in ε -MnO 2 Pseudocapacitors

Pseudocapacitive materials such as manganese dioxide (MnO 2 ) are attractive for energy storage applications due to their ability to combine the fast kinetics of capacitors with the higher energy density of battery-type systems. However, the electrochemical behavior of MnO 2 remains difficult to interpret mechanistically, in part because existing models often fail to distinguish between surface-based redox processes and bulk intercalation mechanisms. In this work, we develop a physics-based model that represents MnO 2 pseudocapacitance as a linear combination of two independent, parallel electrochemical processes: (i) the surface or near-surface redox storage and (ii) lithium ion intercalation into the bulk material. These two processes are treated with distinct kinetic and thermodynamic parameters and are assumed to proceed independently. The total measured current is assumed to be the sum of these two partial currents. We validate the model using rate-dependent cyclic voltammetry experiments, demonstrating that it captures key trends and provides physically interpretable parameters reflecting the relative contributions of surface and bulk processes. By enabling a clear separation between these mechanisms, the model offers a useful framework for analyzing pseudocapacitive materials and can guide the rational design of high-performance energy storage electrodes.

Energy - Storage↗

Massively Parallel and Portable Genomic Sequence Analysis

Massively Parallel and Portable Genomic Sequence Analysis (mappgene) is a sequencing analysis workflow for high performance computing. It incorporates novel technologies to simplify and accelerate genetics research.

Moon, Josephy↗

Parallel Dislocation Simulator

ParaDiS, or Parallel Dislocation Simulator, is a simulation tool that performs direct numerical simulation of dislocation ensembles, the carriers of plasticity, to predict the strength in crystalline materials from the fundamental physics of defect motion, evolution, and interaction. The code has been successfully deployed on high performance computing architectures and used to study the origins of strength and strain hardening for cubic crystals, the strength of micro-pillars, and irradiated materials at LLNL. The ParaDiS code has been successfully deployed on more than one hundred thousand CPU's with over ten million active degrees of freedom.

Bulatov, Vasily↗

4P (Plastic Parallel Pathways Platform) [SWR 23-84]

The Plastic Parallel Pathways Platform (4P) combines life cycle assessment, agent-based modeling within a dynamic material flow analysis structure to compute the environmental impacts of different recycling options under various behavioral interventions.

Walzberg, Julien↗

Phlex: Parallel, Hierarchical, and Layered EXecution of data-processing algorithms

Phlex is a computing framework supporting the parallel, hierarchical, and layered execution of data-processing algorithms. It is based on the functional-programming paradigm, thus guaranteeing thread-safety when invoking user-defined pure functions. Phlex allows users to specify arbitrary graph-based hierarchies of data organization, enabling more flexible processing of data as required by the constraints of the program.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗

PRISMA: PARALLEL REFINEMENT AND INTEGRATION SYSTEM FOR MULTI-AZIMUTHAL ANALYSIS

The Parallel Refinement and Integration System for Multi-azimuthal Analysis (PRISMA, version 1.1.0) is a Python application for processing X-ray diffraction (XRD) image data. PRISMA wraps GSAS-II to perform azimuthally-binned peak refinement, computes per-frame strain and d-spacing from those fits, and provides three PyQt5 graphical interfaces: (1) a Recipe Builder for selecting GSAS-II control (.imctrl) files, optional mask (.immask) files or threshold-ased masking, reference and experiment image sets, peaks, zimuthal range and bin size, and an optional ceria-based auto-calibration; (2) a Batch Processor that uses Dask on local workstations and pure MPI (mpi4py.futures.MPICommExecutor) on HPC to distribute GSAS-II refinement across cores or compute nodes and write results to a 4-dimensional (peaks x frames x azimuths x measurements) Zarr dataset; and (3) a Data Analyzer that renders heatmaps of fit parameters, strain, frame-to-frame deltas, and percent-change-vs-reference, and exports user-defined subsections to CSV or Excel. The peak-refinement algorithm is deterministic. Benchmark on ALCF Crux: a 20,000-image set, single-peak fit in frame mode with 44 azimuthal bins on 128 nodes x 128 workers, 48 seconds total wall time.

Lorenzo Martin, Maria De La Cinta [Argonne Nationa↗

ConnectIt: a framework for static and incremental parallel graph connectivity algorithms

Connected components is a fundamental kernel in graph applications. The fastest existing multicore algorithms for solving graph connectivity are based on some form of edge sampling and/or linking and compressing trees. However, many combinations of these design choices have been left unexplored. In this paper, we design the ConnectIt framework, which provides different sampling strategies as well as various tree linking and compression schemes. ConnectIt enables us to obtain several hundred new variants of connectivity algorithms, most of which extend to computing spanning forest. In addition to static graphs, we also extend ConnectIt to support mixes of insertions and connectivity queries in the concurrent setting. We present an experimental evaluation of ConnectIt on a 72-core machine, which we believe is the most comprehensive evaluation of parallel connectivity algorithms to date. Compared to a collection of state-of-the-art static multicore algorithms, we obtain an average speedup of 12.4x (2.36x average speedup over the fastest existing implementation for each graph). Using ConnectIt, we are able to compute connectivity on the largest publicly-available graph (with over 3.5 billion vertices and 128 billion edges) in under 10 seconds using a 72-core machine, providing a 3.1x speedup over the fastest existing connectivity result for this graph, in any computational setting. For our incremental algorithms, we show that our algorithms can ingest graph updates at up to several billion edges per second. To guide the user in selecting the best variants in ConnectIt for different situations, we provide a detailed analysis of the different strategies. Finally, we show how the techniques in ConnectIt can be used to speed up two important graph applications: approximate minimum spanning forest and SCAN clustering.

Computer Science↗

Nyx: A Massively Parallel AMR Code for Computational Cosmology

Nyx is a highly parallel, adaptive mesh, finite-volume N-body compressible hydrodynamics solver for cosmological simulations. It has been used to simulate different cosmological scenarios with a recent focus on the intergalactic medium and Lyman alpha forest. Together, Nyx, the compressible astrophysical simulation code, Castro, and the low Mach number code MAESTROeX, make up the AMReX-Astrophysics Suite of open-source, adaptive mesh, performance-portable astrophysical simulation codes. Other examples of cosmological simulation research codes include Enzo, Enzo-P/Cello, RAMSES, ART, FLASH, Cholla, as well as Gadget, Gasoline, Arepo, Gizmo, and SWIFT.

79 ASTRONOMY AND ASTROPHYSICS↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

¬¬Integration of Quantification of Margins and Uncertainties Methodology into Parallel Discrete Event Simulator Framework

Parallel Discrete Events Simulation (PDES) is becoming increasingly important to lab efforts in security and intelligence. It is used to model complex asynchronous systems such as computer networks, satellite systems, vehicular traffic, and human performance. Uncertainty Quantification (UQ) techniques have become a mainstay of Verification and Validation (V&V) efforts on physics simulations. However, PDES models are very different from traditional physics simulations, and research into UQ techniques for PDES is in its infancy. It is not clear which traditional UQ techniques can be applied to PDES, or what new techniques will need to be developed. The goal of this project was to identify existing UQ techniques that can be applied to PDES, and to develop new techniques as necessary. The project implemented or developed techniques to handle issues that do not appear in traditional physics simulations, but are common among PDES, including sampling techniques for high-dimensional homogenous inputs, response surfaces for high-variance heteroskedastic output, and characterization of skewed output distributions. These are foundational UQ techniques that must be used in any complete UQ analysis (Tong 2018).

97 MATHEMATICS AND COMPUTING↗