Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21

CCM vs. CRM Design Optimization of a Boost-derived Parallel Active Power Decoupler for Microinverter Applications

Single-phase inverter or rectifier systems often make use of an auxiliary active power decoupler (APD) to balance the mismatch between steady DC power and fluctuating AC power. This paper deals with efficiency and size optimization of a parallel boost-type APD circuit for PV microinverter applications. Specifically, design of an eGaNFET-based, 400 W APD circuit, employing planar inductor and operating in either continuous conduction mode (CCM) or critical conduction mode (CRM) is considered. Available design variables including inductance value, inductor core geometry, capacitor voltage, switching frequency, and modulation scheme (CCM vs. CRM) are explored to identify Pareto-optimal configurations, which can achieve low California Energy Commission (CEC) efficiency drop while also reducing the footprint area of the inductor. The theoretical study predicts that the optimal CRM design can achieve 37% reduced inductor size, while operating with similar efficiency drop, compared to the optimal CCM design. Experimental results, obtained using two separate 40 V, 400 W hardware prototypes for CCM and CRM, are presented to verify the analyses.

14 SOLAR ENERGY↗

Multiterminal High-Voltage dc Systems with Series-Parallel Valve Group-Based High-Voltage dc Substations

To transfer large amount of power over long distances, multiterminal direct current (MTdc) system based on bipole high-voltage direct current (HVdc) technology is a viable option. However, such system results in large dc transmission loss. The same can be reduced by increasing the dc voltage level. This paper introduces a new MTdc system architecture comprising of series (for increasing dc voltage level) and parallel (for increasing dc current capability) connected HVdc converters. The new architecture is compared with the bipole MTdc architecture in terms of equipment needed and dc transmission loss. The control modifications needed for the MTdc system are identified and the performance of the developed control is verified through electromagnetic transient (EMT) simulations.

Jaldanki, Sreenivasa↗

Fast Parallel Tensor Times Same Vector for Hypergraphs

Hypergraphs are a popular paradigm to rep- resent complex real-world networks exhibiting multi-way relationships of varying sizes. Mining centrality in hyper- graphs via symmetric adjacency tensors has only recently become computationally feasible for large and complex datasets. To enable scalable computation of these and related hypergraph analytics, here we focus on the Sparse Symmetric Tensor Times Same Vector (S3TTVC) oper- ation. We introduce the Compound Compressed Sparse Symmetric (CCSS) format, an extension of the compact CSS format for hypergraphs of varying hyperedge sizes and present a shared-memory parallel algorithm to compute S3TTVC. We experimentally show S3TTVC computation using the CCSS format achieves better performance than the naive baseline, and is subsequently more performant for hypergraph H-eigenvector centrality.

Shivakumar, Shruti↗

DeepThermo: Deep Learning Accelerated Parallel Monte Carlo Sampling for Thermodynamics Evaluation of High Entropy Alloys

Since the introduction of Metropolis Monte Carlo (MC) sampling, it and its variants have become standard tools used for thermodynamics evaluations of physical systems. However, a long-standing problem that hinders the effectiveness and efficiency of MC sampling is the lack of a generic method (a.k.a. MC proposal) to update the system configurations. Consequently, current practices are not scalable. Here we propose a parallel MC sampling framework for thermodynamics evaluation—DeepThermo. By using deep learning–based MC proposals that can globally update the system configurations, we show that DeepThermo can effectively evaluate the phase transition behaviors of high entropy alloys, which have an astronomical configuration space. For the first time, we directly evaluate a density of states expanding over a range of ~e 10,000 for a real material. We also demonstrate DeepThermo’s performance and scalability up to 3,000 GPUs on both NVIDIA V100 and AMD MI250X-based supercomputers.

Yin, Junqi↗

Toward Automated Detection of Portability Bugs in Kokkos Parallel Programs

Performance-portable programming frameworks provide abstractions for parallel execution to allow easily porting an application to multiple backend programming models, such as CUDA, HIP, and OpenMP. However, programs may still have portability bugs that manifest only on specific backends. Traditional testing is ineffective in discovering these bugs, as it would require concrete execution on all supported hardware configurations for a potentially infinite set of inputs. To mitigate this issue, we focused on a specific programming framework, Kokkos, and identified several categories of common portability bugs. We then developed Klokkos, a static analysis approach based on symbolic execution that can run on commodity hardware, before execution on supercomputers. As a proof-of-concept, we ran Klokkos on examples encoding the identified bugs. Our results show that Klokkos is effective, efficient, and precise: it detected all the considered bugs, quickly, and without any false positives. Although preliminary, our results motivate further research and development in this direction.

Kale, Vivek↗

Decentralized Carrier Phase Shifting for Optimal Harmonic Minimization in Asymmetric Parallel-Connected Inverters

This paper presents a carrier phase shifting technique for minimizing the aggregate harmonics in networks of asymmetric parallel-connected inverters for distributed power generation system applications. The proposed technique is: 1) implemented in a decentralized manner, relying only on local voltage and current measurements, and 2) optimal in the sense that it minimizes a cost function representing the carrier-frequency current harmonics. The analysis indicates that the proposed optimal carrier phase shifting technique can enable order-of-magnitude reductions in harmonic power, and also universal improvements compared to symmetric carrier interleaving for asymmetric inverter networks. Moreover, compared to existing methods that require either centralized communication or information exchange between inverters to coordinate carriers, the proposed technique is completely decentralized, which provides important practical benefits for implementation, including improved robustness and reduced cost. The technique is experimentally validated on a network of three single-phase 2-kW inverters and demonstrates a 36.5% reduction in the weighted total harmonic distortion factor of the aggregate inverter current, and the ability to converge to the optimal carrier phase spacing dynamically in less than one line frequency cycle (16.7 ms) in steady state and transient operating conditions.

42 ENGINEERING↗

Comparison of CCM- and CRM-Based Boost Parallel Active Power Decoupler for PV Microinverter

Single-phase inverter or rectifier systems often make use of an active power decoupler (APD) to balance the mismatch between constant dc power and fluctuating ac power. This article deals with the comparison of continuous conduction mode (CCM) and critical conduction mode (CRM) operation-based design of a parallel boost-type APD for photovoltaic microinverter applications. From a design perspective, multiobjective analysis of efficiency, volume, and cost is explored within a decision space including planar inductors, gallium nitride based devices, film capacitors, switching frequency, and modulation (CCM vs. CRM). The theoretical study analyzes all possible design configurations within CCM and CRM and identifies Pareto-optimal designs, from which the selected CRM design can achieve reduced system volume and lower cost with the use of smaller inductor core, while operating with similar California Energy Commission efficiency drop as the selected CCM design. From a control perspective, a pulsewidth modulation based control strategy is proposed to implement closed-loop CRM modulation that does not rely on zero-crossing detection. Furthermore, closed-loop systems are designed for the optimal CCM and CRM realizations, and the final system characteristics are compared. Experimental results, obtained using two separate 40-V, 400-W hardware prototypes for CCM and CRM, are presented to verify the analyses.

42 ENGINEERING↗

Memory-Aware External Facelist Calculation: A Data-Parallel Atomic Hash Counting Approach

Unstructured volumetric meshes serve as fundamental data representations in various scientific simulations and analyses. They play a crucial role in representing complex computational domains and are essential for important numerical techniques, such as finite element analysis. Whenever such a mesh is read from a file, streamed in-situ, or generated by algorithms, scientific visualization libraries rely on calculating the external surface of a geometry, named “external facelist”, to produce a polygonal mesh for rendering. Consequently, external facelist calculation has become one of the most widely used algorithms in the scientific visualization domain, necessitating optimal performance. In this paper, we explore relevant work on external facelist calculation algorithms in two common visualization libraries, VTK and Viskores, assess their performance and memory constraints, and introduce a novel memory-aware external facelist calculation algorithm employing an atomic hash counting approach. This algorithm fully leverages Viskores' data-parallel primitive operations, facilitating its execution across diverse many-core architectures. Our algorithm features the lowest memory footprint on the GPU and the second-lowest on the CPU among all evaluated methods, and it also delivers the fastest performance on both CPU and GPU. It has been made available under an open-source license in the VTK and Viskores visualization systems.

Tsalikis, Spiros [Kitware] (ORCID:0000000151137195↗

Biochemical parallels between catabolic pathways for lignin-associated aromatic dimers

Lignin is one of the most common biopolymers on Earth. In nature, lignin is primarily deconstructed by fungi into mixtures of aromatic compounds that are then assimilated by bacteria and fungi. Industrially, lignin is primarily generated as a byproduct of pulp and paper production and burned for process heat. However, if the appropriate assimilatory pathways were identified, deconstructed lignin could be funneled into value-added products using engineered bacteria. Foundational work has described pathways for assimilation of diverse monomeric aromatic compounds such as protocatechuate, ferulate, and syringate, as well as select dimers including those with β-O-4 and 5-5 interunit linkages. Recent advances have elucidated additional pathways for dimer assimilation, including pathways for new substrates as well as parallel pathways for previously characterized substrates. Comparing these dimer assimilation pathways can illuminate the underlying biochemical logic of assimilation for lignin-associated aromatic dimers and provide opportunities for metabolic engineering to enhance lignin valorization.

Sphingomonas↗

Efficient Parallel Sparse Symmetric Tucker Decomposition for High-Order Tensors

Tensor based methods are receiving renewed attention in recent years due to their prevalence in diverse real-world applications. There is considerable literature on tensor representations and algorithms for tensor decompositions, both for dense and sparse tensors. Many applications in hypergraph analytics, machine learning, psychometry, and signal processing result in tensors that are both sparse and symmetric, making it an important class for further study. Similar to the critical Tensor Times Matrix chain operation (TTMc) in general sparse tensors, the Sparse Symmetric Tensor Times Same Matrix chain (S3TTMc) operation is compute and memory intensive due to high tensor order and the associated factorial explosion in the number of non-zeros. In this work, we present a novel compressed storage format CSS for sparse symmetric tensors, along with an efficient parallel algorithm for the S3TTMc operation. We theoretically establish that S3TTMc on CSS achieves a better memory versus run-time trade-off compared to state-of-the-art implementations. We demonstrate experimental findings that confirm these results and achieve up to 2.9× speedup on synthetic and real datasets.

Shivakumar, Shruti↗

A massively parallel and scalable multi-CPU material point method

Harnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid (G2P2G) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array (AoSoA) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 10243 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation.

Wang, Xinlei↗

CEAZ: Accelerating Parallel I/O Via Hardware-Algorithm Co-Designed Adaptive Lossy Compression

As supercomputers continue to grow to exa-scale, the amount of data that needs to be saved or transmitted is exploding. To this end, many previous works have studied using error-bounded lossy compressors to reduce the data size and improve the I/O performance. However, little work has been done for effectively offloading lossy compression onto FPGA-based SmartNICs to reduce the compression overhead. In this paper, we propose a hardware-algorithm co-design of efficient and adaptive lossy compressor for scientific data on FPGAs (called CEAZ) to accelerate parallel I/O. Our contribution is fourfold: (1) We propose an efficient Huffman coding approach that can adaptively update Huffman codewords online based on codewords generated offline (from a variety of representative scientific datasets). (2) We derive a theoretical analysis to support a precise control of compression ratio under an error-bounded compression mode, enabling accurate offline Huffman codewords generation. This also help us create a fixed-ratio compression mode for consistent throughput. (3) We develop an efficient compression pipeline by adopting cuSZ’s dual-quantization algorithm to our hardware use case. (4) We evaluate CEAC on five real-world datasets with both a single FPGA board and 256 nodes from Bridges2 supercomputer. Experiments show that CEAZ outperforms the second-best FPGA-based lossy compressor by 2× of throughput and 9.6× of compression ratio. It also improves MPI_File_write and MPI_Gather throughputs by up to 32.7× and 31.4×, respectively.

Zhang, Chengming↗

Accelerated Constrained Sparse Tensor Factorization on Massively Parallel Architectures

This study presents the first constrained sparse tensor factorization (cSTF) framework that optimizes and fully offloads computation to massively parallel GPU architectures, and the first performance characterization of cSTF on GPU architectures. In contrast to prior work on tensor factorization, where the matricized tensor times Khatri-Rao product (MTTKRP) is the primary performance bottleneck, our systematic analysis of the cSTF algorithm on GPUs reveals that adding constraints creates an additional bottleneck in the update operation for many real-world sparse tensors. While executing the update operation on the GPU brings significant speedup over its CPU counterpart, it remains a significant bottleneck. To further accelerate the update operation, we propose cuADMM, a new update algorithm that leverages algorithmic and code optimization strategies to minimize both computation and data movement on GPUs. As a result, our framework delivers significantly improved performance compared to prior state-of-the-art. On 10 real-world sparse tensors, our framework achieves geometric mean speedup of 5.1 × (max 41.59 ×) and 7.01 × (max 58.05 ×) on the NIVIDA A100 and H100 GPUs, respectively, over the state-of-the-art SPLATT library running on a 26-core Intel Ice Lake Xeon CPU.

Soh, Yongseok↗

Determining Levels of Detail for Simulators of Parallel and Distributed Computing Systems via Automated Calibration

There are two sources of inaccuracy when simulating parallel and distributed computing systems: (i) a simulator implemented at an insufficient level of detail; and (ii) incorrectly calibrated simulation parameter values. Increasing the simulator’s level of detail can improve accuracy, but at the cost of higher space, time, and/or software complexity. Furthermore, evaluating the intrinsic accuracy of a simulator requires that its parameters be well-calibrated. Making decisions regarding the level of detail is thus challenging. We propose a methodology for instantiating the simulation calibration process and a framework for automating this process, which makes it possible to pick appropriate levels of detail for any simulator. We demonstrate the usefulness of our approach via two case studies for two different domains.

McDonald, Jessie [University of Hawaii at Manoa, H↗

A Parallel Kinetic Model for Surface and Bulk Charge Storage in ε -MnO 2 Pseudocapacitors

Pseudocapacitive materials such as manganese dioxide (MnO 2 ) are attractive for energy storage applications due to their ability to combine the fast kinetics of capacitors with the higher energy density of battery-type systems. However, the electrochemical behavior of MnO 2 remains difficult to interpret mechanistically, in part because existing models often fail to distinguish between surface-based redox processes and bulk intercalation mechanisms. In this work, we develop a physics-based model that represents MnO 2 pseudocapacitance as a linear combination of two independent, parallel electrochemical processes: (i) the surface or near-surface redox storage and (ii) lithium ion intercalation into the bulk material. These two processes are treated with distinct kinetic and thermodynamic parameters and are assumed to proceed independently. The total measured current is assumed to be the sum of these two partial currents. We validate the model using rate-dependent cyclic voltammetry experiments, demonstrating that it captures key trends and provides physically interpretable parameters reflecting the relative contributions of surface and bulk processes. By enabling a clear separation between these mechanisms, the model offers a useful framework for analyzing pseudocapacitive materials and can guide the rational design of high-performance energy storage electrodes.

Energy - Storage↗

Massively Parallel and Portable Genomic Sequence Analysis

Massively Parallel and Portable Genomic Sequence Analysis (mappgene) is a sequencing analysis workflow for high performance computing. It incorporates novel technologies to simplify and accelerate genetics research.

Moon, Josephy↗

Parallel Dislocation Simulator

ParaDiS, or Parallel Dislocation Simulator, is a simulation tool that performs direct numerical simulation of dislocation ensembles, the carriers of plasticity, to predict the strength in crystalline materials from the fundamental physics of defect motion, evolution, and interaction. The code has been successfully deployed on high performance computing architectures and used to study the origins of strength and strain hardening for cubic crystals, the strength of micro-pillars, and irradiated materials at LLNL. The ParaDiS code has been successfully deployed on more than one hundred thousand CPU's with over ten million active degrees of freedom.

Bulatov, Vasily↗

4P (Plastic Parallel Pathways Platform) [SWR 23-84]

The Plastic Parallel Pathways Platform (4P) combines life cycle assessment, agent-based modeling within a dynamic material flow analysis structure to compute the environmental impacts of different recycling options under various behavioral interventions.

Walzberg, Julien↗