Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Parallel exponential time differencing methods for geophysical flow simulations

Two ocean models are considered for geophysical flow simulations: the multilayer shallow water equations and the multilayer primitive equations. For the former, we investigate the parallel performance of exponential time differencing (ETD) methods, including exponential Rosenbrock–Euler, ETD2wave, and B-ETD2wave. For the latter, we take advantage of the splitting of barotropic and baroclinic modes and propose a new two-level method in which an ETD method is applied to solve the fast barotropic mode. Furthermore, these methods could improve the computational efficiency of numerical simulations because ETD methods allow for much larger time step sizes than traditional explicit time-stepping techniques that are commonly used in existing computational ocean models. Several standard benchmark tests for ocean modeling are performed and comparison of the numerical results demonstrates a great potential of applying the parallel ETD methods for simulating real-world geophysical flows.

54 ENVIRONMENTAL SCIENCES↗

Parallel derivative-free optimization for simulation-based design of behind-the-meter energy systems

In this work, the integrated design and dispatch of behind-the-meter or distributed resources (e.g. stationary battery storage and solar PV generation) is considered. A simulation-based framework is employed, generating high-fidelity results with closed-loop predictive control at a fine resolution, at the expense of high computational cost (several minutes to a few hours per design point). To address this challenge, parallel derivative-free design methods are considered. Four methods are compared, including state-of-the-art surrogate-based methods (Radial-Basis Functions and Gaussian processes) and sampling strategies, an evolutionary-based method, and a simple sequential grid refinement method. As a case study, two types of design problem with increasing complexity are considered, namely, the design of behind-the-meter resources (three design variables) and the inclusion of grid capacity (four design variables). The second yields a constrained design problem for which violations can only be determined after solving the computationally expensive simulation. For the three-dimensional case, all methods present a good performance, achieving a solution within 1% of the optimum after the first iteration, with the sequential grid refinement exhibiting the fastest convergence and achieving the best final objective value. This indicates that the parallel evaluation of multiple sampling points may be more important than the choice of method for small decision spaces. For the four-dimensional constrained case, the Genetic Algorithm presents the best tradeoff between performance and computational effort, while the rough objective function terrain generated by constraint violation penalties reduces the performance of surrogate-based methods. Contour plots with flat regions indicate flexibility in the optimal design and highlight the importance of characterizing the solution space.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A parallel and performance portable implementation of a full-field crystal plasticity model

We have developed a parallel implementation of an Elasto-Viscoplastic Fast Fourier Transform-based (EVPFFT) micromechanical solver to enable computationally efficient crystal plasticity modeling for polycrystalline materials. Our primary focus lies in achieving performance portability, allowing a single EVPFFT implementation to run optimally on various homogeneous architectures, including multi-core Central Processing Units (CPUs), as well as on heterogeneous computer architectures comprising multi-core CPUs and Graphics Processing Units (GPUs) from different vendors. To accomplish this goal, we have leveraged MATAR, a C++ software library that simplifies the creation and utilization of multidimensional dense or sparse matrix and array data structures. These data structures are designed to be portable across diverse architectures through the use of Kokkos, a performance-portable library. Additionally, we have employed the Message Passing Interface (MPI) to efficiently distribute the computational workload among processors. The heFFTe (Highly Efficient FFT for Exascale) library is used to facilitate the performance portability of the fast Fourier transforms (FFTs) computation. The computational performance of EVPFFT is evaluated and presented in terms of parallel scalability and simulation runtime on different high-performance computing (HPC) architectures. As a result, the utility of the developed framework to efficiently simulate the micro-mechanical fields in polycrystalline microstructures in engineering applications is discussed.

36 MATERIALS SCIENCE↗

OpenSn: A massively parallel, open-source simulation environment for discrete ordinates radiation transport

OpenSn is an open-source, massively parallel deterministic radiation transport code for solving the discrete-ordinates ( S N ) form of the Boltzmann transport equation on unstructured, arbitrary polyhedral meshes. It supports high-fidelity simulations involving steady-state, eigenvalue, and adjoint problems for neutral particles (e.g., neutrons, photons, multi-particles), using the multigroup approximation in energy. OpenSn combines angular discretization via discrete ordinates with a discontinuous Galerkin finite element method (DGFEM) in space, enabling accurate resolution of transport physics on arbitrary polyhedral cells, included locally refined spatial grids. It includes multiple angular quadrature types, including locally refined angular quadratures. Written in modern C++ with a Python API, OpenSn runs efficiently on platforms ranging from laptops to supercomputers. The transport sweep algorithm is implemented using a task-based, directed-acyclic-graph (DAG) approach for each angle and supports asynchronous parallelism across thousands of MPI ranks. Group-set aggregation improves compute intensity, and synthetic acceleration techniques (e.g., diffusion synthetic acceleration, second-moment method) enhance solver convergence. OpenSn has been verified on reactor physics problems and demonstrated excellent weak and strong scaling performance on more than 32,768 processes, making it a versatile and robust platform for large-scale transport simulations in complex geometries.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An Open-Source Parallel EMT Simulation Framework

As the integration level of inverter-based resources (IBRs) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗

Parallel transport dynamics for mixed quantum states with applications to time-dependent density functional theory

Direct simulation of the von Neumann dynamics for a general (pure or mixed) quantum state can often be expensive. One prominent example is the real-time time-dependent density functional theory (rt-TDDFT), a widely used framework for the first principle description of many-electron dynamics in chemical and materials systems. Practical rt-TDDFT calculations often avoid the direct simulation of the von Neumann equation, and solve instead a set of Schrödinger equations, of which the dynamics is equivalent to that of the von Neumann equation. However, the time step size employed by the Schrödinger dynamics is often much smaller. Here, in order to improve the time step size and the overall efficiency of the simulation, we generalize a recent work of the parallel transport (PT) dynamics for simulating pure states [An, Lin, Multiscale Model. Simul. 18, 612, 2020] to general quantum states. The PT dynamics provides the optimal gauge choice, and can employ a time step size comparable to that of the von Neumann dynamics. Going beyond the linear and near adiabatic regime in previous studies, we find that the error of the PT dynamics can be bounded by certain commutators between Hamiltonians, density matrices, and their derived quantities. Such a commutator structure is not present in the Schrödinger dynamics. We demonstrate that the parallel transport-implicit midpoint (PT-IM) method is a suitable method for simulating the PT dynamics, especially when the spectral radius of the Hamiltonian is large. The commutator structure of the error bound, and numerical results for model rt-TDDFT calculations in both linear and nonlinear regimes, confirm the advantage of the PT dynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A second-order distributed memory parallel fast sweeping method for the Eikonal equation

The Eikonal equation is used to calculate wave propagation and distance fields, and due to its complexity requires numerical treatment for its solution. In this work, we present a second-order distributed memory parallel fast sweeping method. The second-order solution switches on a two-point stencil when two upwind points are available, and reverts to first-order otherwise. In all examples, the second-order method improves the solution over the first-order, allowing for significant savings in memory while achieving the same accuracy. Parallelization over distributed memory saw good weak scaling with optimal convergence. The computational time for second-order was approximately 2.5 times slower than first-order, where the largest amount of mesh points ran on 144 cores (512 GB) was ≈20 billion. The savings in memory from the second-order method combined with the distributed memory algorithm result in the ability to solve problems much larger than are possible with the serial first-order method.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Improving scalability of parallel CNN training by adaptively adjusting parameter update frequency

Synchronous SGD with data parallelism, the most popular parallelization strategy for CNN training, suffers from the expensive communication cost of averaging gradients among all workers. The iterative parameter updates of SGD cause frequent communications and it becomes the performance bottleneck. In this paper, we propose a lazy parameter update algorithm that adaptively adjusts the parameter update frequency to address the expensive communication cost issue. Our algorithm accumulates the gradients if the difference of the accumulated gradients and the latest gradients is sufficiently small. Here, the less frequent parameter updates reduce the per-iteration communication cost while maintaining the model accuracy. Our experimental results demonstrate that the lazy update method remarkably improves the scalability while maintaining the model accuracy. For ResNet50 training on ImageNet, the proposed algorithm achieves a significantly higher speedup (739.6 on 2048 Cori KNL nodes) as compared to the vanilla synchronous SGD (276.6) while the model accuracy is almost not affected (<0.2% difference).

97 MATHEMATICS AND COMPUTING↗

Aperture size distribution, length, and preferential location of bed-parallel veins in shale

Bed-parallel, calcite-filled veins (BPVs) are common in shale formations, and although they have been widely described in other studies, little is known about their population aperture size distribution. To address this knowledge gap, we analyzed BPV sizes in outcrops and cores from the Vaca Muerta Formation, Neuquén Basin, Argentina; in two cores from the Marcellus Formation, Appalachian Basin, northeast Pennsylvania; and in one core from the Wolfcamp Shale, Delaware Basin, West Texas. Nine out of ten aperture size populations follow a negative exponential distribution, with one following a weak power law. Bed-parallel vein size distribution and intensity vary among formations and within the same shale. We define three groups of distributions: (1) Vaca Muerta outcrops, with the highest BPV intensity and the largest BPVs (cumulative frequency of 4.9 BPVs per meter [BPVs/m] for apertures 0.265 mm to 8.7 cm); (2) Vaca Muerta cores with a similar BPV intensity overall but with no apertures wider than 1.2 cm; and (3) Vaca Muerta, Wolfcamp, and Marcellus cores with the fewest BPVs (cumulative frequency up to 0.63 BPVs/m) and very few wider than 1 cm. Aperture and length in two outcrop data sets are weakly positively correlated and follow power laws with exponents of 0.44 and 0.49. Mechanical interfaces at boundaries between different lithologies exert a strong control on BPV location, with 65–75% of observed interfaces having BPVs along them. Only 25–30% of the BPVs occur at observed material interfaces, however, and unless subtle, unobserved mechanical layering is present, other factors must also control location. BPV intensity and organic richness (TOC) from Vaca Muerta well logs are correlated in some instances but not in others, indicating TOC is not always a good proxy for BPV location or intensity. Furthermore, these findings provide useful information for modeling of hydraulic fracture treatments where BPVs may influence development of the stimulated fracture network, for example by limiting height growth.

58 GEOSCIENCES↗

Scalable line and plane relaxation in a parallel structured multigrid solver

The efficient solution of sparse, linear systems that arise through the discretization of partial differential equations remains a key challenge for a range of high performance scientific simulations. One approach for reducing data movement and improving performance is by exposing and exploiting structure in a problem through the use of robust structured multilevel solvers. By choosing coarsening that preserves the structure of the problem, these methods maintain efficient structured computation and communication throughout the multigrid hierarchy. However, when coarsening is not permitted to be dependent on the operator, anisotropy must be addressed by the smoother — producing error compatible for coarse-grid correction with structured coarsening. Here, the components required in a scalable parallel structured solver are described with a focus on memory and communication efficiency of robust smoothers. While the implementation of communication and memory reduction techniques in smoothers integrated in a complete 3D solver present a significant engineering challenge, a novel approach is proposed that addresses these challenges systematically through a change to the solver’s execution model. Enabled by user-level threading paired with a set of data and communication abstractions, this approach permits seamless aggregation of communication in plane smoothers — directly reusing code for a 2D distributed multilevel cycle. Results show an effective reduction in communication costs for coarse-grid problems, and result in a speedup of 8.7x in smoothing routines shown in Fig. 12 using this approach. This produces a significant improvement to strong scalability while maintaining favorable weak scaling behavior. Finally, a parallel scaling study using a series of refined meshes is included that demonstrates the effectiveness of this approach in an application of interest.

97 MATHEMATICS AND COMPUTING↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parallel Quantum-Enhanced Sensing

Quantum metrology takes advantage of quantum correlations to enhance the sensitivity of sensors and measurement techniques beyond their fundamental classical limit, given by the shot-noise limit. The use of both temporal and spatial correlations present in quantum states of light can extend quantum-enhanced sensing to a parallel configuration that can simultaneously probe an array of sensors or independently measure multiple parameters. To this end, we use multispatial-mode bright twin beams of light, which are characterized by independent quantum-correlated spatial subregions in addition to quantum temporal correlations, to probe a four-sensor quadrant plasmonic array. We show that it is possible to independently and simultaneously measure local changes in refractive index for all four sensors with a quantum enhancement in sensitivity in the range of 22% to 24% over the corresponding classical configuration. Finally, these results provide a first step toward highly parallel spatially resolved quantum-enhanced sensing techniques and pave the way toward more complex quantum sensing and quantum imaging platforms.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel algorithms for hyperdynamics and local hyperdynamics

Hyperdynamics (HD) is a method for accelerating the timescale of standard molecular dynamics (MD). It can be used for simulations of systems with an energy potential landscape that is a collection of basins, separated by barriers, where transitions between basins are infrequent. HD enables the system to escape from a basin more quickly while enabling a statistically accurate renormalization of the simulation time, thus effectively boosting the timescale of the simulation. In [Kim, Perez, Voter, J Chem Phys, 139:144110, 2013)1, a local version of HD was formulated, which exploits the intrinsic locality characteristic typical of most systems to mitigate the poor scaling properties of standard HD as the system size is increased. In this paper, we discuss how both HD and local HD can be formulated to run efficiently in parallel. We have implemented these ideas in the LAMMPS MD code, which means HD can be used with any interatomic potential LAMMPS supports. Together, these parallel methods allow simulations of any size to achieve the time acceleration offered by HD (which can be orders of magnitude), at a cost 3-5x that of standard MD. As examples, we performed two simulations of a million-atom system to model the diffusion and clustering of Pt adatoms on a large patch of Pt(100) surface for 80 and 160 μs.

74 ATOMIC AND MOLECULAR PHYSICS↗

Cross-field and parallel dynamics of SOL filaments in TCV

Using recently installed scrape-off layer diagnostics on the tokamak à configuration variable, we characterise the poloidal and parallel properties of turbulent filaments. We access both attached and detached divertor conditions across a wide range of core densities (f G ϵ [0.09, 0.66]) in diverted L-mode plasma configurations. With a gas puff imaging (GPI) diagnostic at the outer midplane we observed filaments with a monotonic increase in radial velocity (from 390 m s –1 to 800 m s –1 ) and cross-field radii (from 8.5 mm to 13.4 mm) with increasing core density. Interpreting the filament behaviour in the context of the two-region model, we find that they populate the ideal-interchange regime (C i ) in discharges at very low densities, and the resistive X (RX)-point regime for all other discharges. The scaling of filament velocity versus size shows good agreement with this interpretation. These results are discussed and compared with previous probe-based measurements for similar conditions, which mostly placed filaments in TCV in the resistive ballooning (RB) regime. In addition, for the first time in TCV, the parallel filament extension is studied by magnetically aligning the GPI measurements at the outboard midplane with a reciprocating probe in the divertor. In agreement with the filaments being in the ideal-interchange and the RX-point regimes, they are found to extend beyond the X-point into the outer divertor leg.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Contingency Analysis Based on Partitioned and Parallel Holomorphic Embedding

In the steady-state contingency analysis, the traditional Newton-Raphson method suffers from non-convergence issues when solving post-outage power flow problems, which hinders the integrity and accuracy of security assessment. In this paper, we propose a novel robust contingency analysis approach based on holomorphic embedding (HE). Here, the HE-based simulator provides theoretical convergence guarantee, which is desirable because it avoids the influence of numerical issues and provides a credible security assessment conclusion. In addition, based on the multi-area characteristics of real-world power systems, a partitioned HE (PHE) method is proposed with an interfacebased partitioning of HE formulation. The PHE method does not undermine the numerical robustness of HE and significantly reduces the computation burden in large-scale contingency analysis. The PHE method is further enhanced by parallel or distributed computation to become parallel PHE (P2HE). Tests on a 458-bus system, a synthetic 419-bus system and a large-scale 21447-bus system demonstrate the advantages of the proposed methods in robustness and efficiency.

42 ENGINEERING↗

Advanced Shuttle Strategies for Parallel QCCD Architectures

Trapped ions (TIs) are at the forefront of quantum computing implementation, offering unparalleled coherence, fidelity, and connectivity. However, the scalability of TI systems is hampered by the limited capacity of individual ion traps, necessitating intricate ion shuttling for advanced computational tasks. The quantum charge-coupled device (QCCD) framework has emerged as a promising solution, facilitating ion mobility for universal quantum computation. Current QCCD architectures predominantly feature a linear topology, which is increasingly recognized as inefficient for complex quantum operations. Anticipating the shift toward more efficacious designs, this article introduces an innovative quantum scheduling strategy optimized for parallel QCCD topologies. Our strategy proposes a probabilistic formula for ion movement, alongside ingenious methods for local layer generation and layer compression, yielding a significant reduction in ion shuttle times. Through simulations, we demonstrate that our strategy not only substantially outstrips the linear model but also exhibits better performance over other parallel strategies that employ greedy algorithms. This is achieved through our nuanced resolution of complexities, such as traffic blocks and trap capacity limitations. The consequent reduction in shuttle operations leads to lower energy consumption and an enhancement in the quantum computer's fidelity, ultimately accelerating program execution times.

43 PARTICLE ACCELERATORS↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗