Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Pairwise‐Parallel Entangling Gates on Orthogonal Modes in a Trapped‐Ion Chain

Abstract Parallel operations are important for both near‐term quantum computers and larger‐scale fault‐tolerant machines because they reduce execution time and qubit idling. This study proposes and implements a pairwise‐parallel gate scheme on a trapped‐ion quantum computer. The gates are driven simultaneously on different sets of orthogonal motional modes of a trapped‐ion chain. This work demonstrates the utility of this scheme by creating a Greenberger‐Horne‐Zeilinger (GHZ) state in one step using parallel gates with one overlapping qubit. It also shows its advantage for circuits by implementing a digital quantum simulation of the dynamics of an interacting spin system, the transverse‐field Ising model. This method effectively extends the available gate depth by up to two times with no overhead when no overlapping qubit is involved, apart from additional initial cooling. This scheme can be easily applied to different trapped‐ion qubits and gate schemes, broadly enhancing the capabilities of trapped‐ion quantum computers.

Optics↗

Twelve Ways to Fool the Masses When Giving Parallel-in-Time Results

Getting good speedup—let alone high parallel efficiency—for parallel-in-time (PinT) integration examples can be frustratingly difficult. The high complexity and large number of parameters in PinT methods can easily (and unintentionally) lead to numerical experiments that overestimate the algorithm’s performance. In the tradition of Bailey’s article “Twelve ways to fool the masses when giving performance results on parallel computers”, we discuss and demonstrate pitfalls to avoid when evaluating the performance of PinT methods. Despite being written in a light-hearted tone, this paper is intended to raise awareness that there are many ways to unintentionally fool yourself and others and that by avoiding these fallacies more meaningful PinT performance results can be obtained.

97 MATHEMATICS AND COMPUTING↗

A parallel evolutionary multiple-try metropolis Markov chain Monte Carlo algorithm for sampling spatial partitions

We develop an Evolutionary Markov Chain Monte Carlo (EMCMC) algorithm for sampling spatial partitions that lie within a large, complex, and constrained spatial state space. Our algorithm combines the advantages of evolutionary algorithms (EAs) as optimization heuristics for state space traversal and the theoretical convergence properties of Markov Chain Monte Carlo algorithms for sampling from unknown distributions. Local optimality information that is identified via a directed search by our optimization heuristic is used to adaptively update a Markov chain in a promising direction within the framework of a Multiple-Try Metropolis Markov Chain model that incorporates a generalized Metropolis-Hastings ratio. We further expand the reach of our EMCMC algorithm by harnessing the computational power afforded by massively parallel computing architecture through the integration of a parallel EA framework that guides Markov chains running in parallel.

97 MATHEMATICS AND COMPUTING↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Osmotic control of the spacing of parallel shear cracks in shale growing subcritically in geologic past

The geological genesis of natural cracks in sedimentary rocks such as shale is a problem that needs to be understood to improve the technology of hydraulic fracturing as well as deep sequestration of harmful fluids. Why are the vertical natural cracks roughly parallel and equidistant, and why is the spacing roughly 10 cm rather than 1 cm or 100 cm? Fracture mechanics of critical cracks cannot answer this question. Neither can the material heterogeneity. The growth of critical parallel cracks is impossible because the relative crack face displacements would immediately localize into one crack, leading to an earthquake. The cracks must have formed, on the tectonic time scale, by a slow growth of subcritical shear cracks governed by the Charles-Evans law. The idea advanced here is that what controls the crack spacing is the balance between the reduction, due to shear dilatancy, of the concentration of ions such as Na + and Cl - in each fracture process zone (PFZ), which decelerates the cracks, and the restoration of ion concentration by diffusion of ions from the space between the cracks into the FPZ. This diffusion of water is driven mainly by the osmotic pressure gradient, which offsets the deceleration and depends strongly on the crack spacing. A simple analytical solution of the steady state is rendered possible by approximating the ion concentration profiles between adjacent cracks by parabolic arcs. Applying this theory to Woodford shale yields the approximate crack spacing of 10 cm, which is realistic. Furthermore, the stability of unlimited parallel mode II frictional crack growth is proven by examining the second variation of the free energy. Water concentration drop in the FPZ due to shear dilatancy and its restoration by water diffusion from the inter-crack space have similar effect, although probably much weaker.

42 ENGINEERING↗

A parallel-kinetic-perpendicular-moment model for magnetised plasmas

We describe a new model for the study of weakly collisional, magnetised plasmas derived from exploiting the separation of the dynamics parallel and perpendicular to the magnetic field. This unique system of equations retains the particle dynamics parallel to the magnetic field while approximating the perpendicular dynamics through a spectral expansion in the perpendicular degrees of freedom, analogous to moment-based fluid approaches. In so doing, a hybrid approach is obtained that is computationally efficient enough to allow for larger-scale modelling of plasma systems while eliminating a source of difficulty in deriving fluid equations applicable to magnetised plasmas. We connect this system of equations to historical asymptotic models and discuss advantages and disadvantages of this approach, including the extension of this parallel-kinetic-perpendicular moment beyond the typical region of validity of these more traditional asymptotic models. This paper forms the first of a multi-part series on this new model, covering the theory and derivation, alongside demonstration benchmarks of this approach that include shocks and magnetic reconnection.

astrophysical plasmas↗

High-throughput synthesis of high-entropy alloys via parallelized electric field assisted sintering

Materials discovery and design is an expensive and time-consuming process, though necessary to advance many engineering fields. In this work, a novel tooling design is utilized in conjunction with electric field assisted sintering (EFAS) to effectively create a new high-throughput synthesis technique: parallelized EFAS. Through this technique, a wide range of material compositions and geometries can be synthesized in parallel as isolated samples or as part of contiguous arrays. Multiple tooling designs are explored to examine both the flexibility and limitations of the technique. A series of increasing complex alloys is produced simultaneously using in situ alloying, beginning with pure Ni and adding equimolar constituents up to the septenary high-entropy alloy AlCoCrCuFeMnNi. Microstructural characterization reveals each sample is effectively fully dense and chemically homogenous while exhibiting phases in agreement with CALPHAD predictions. Scalability of parallelized EFAS is then experimentally demonstrated and the implications for materials discovery and automation are discussed.

36 - MATERIALS SCIENCE↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Design and performance of parallel-channel nanocryotrons in magnetic fields

We introduce a design modification to conventional geometry of the cryogenic three-terminal switch, the nanocryotron (nTron). The conventional geometry of nTrons is modified by including parallel current-carrying channels, an approach aimed at enhancing the device's performance in magnetic field environments. The common challenge in nTron technology is to maintain efficient operation under varying magnetic field conditions. Here, we show that the adaptation of parallel channel configurations leads to an enhanced gate signal sensitivity, an increase in operational gain, and a reduction in the impact of superconducting vortices on nTron operation within magnetic fields up to 1 T. Contrary to traditional designs that are constrained by their effective channel width, the parallel nanowire channels permits larger nTron cross sections, further bolstering the device's magnetic field resilience while improving electro-thermal recovery times due to reduced local inductance. This advancement in nTron design not only augments its functionality in magnetic fields but also broadens its applicability in technological environments, offering a simple design alternative to existing nTron devices.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Understanding cold electron impact on parallel-propagating whistler chorus waves via moment-based quasilinear theory

Earth's magnetosphere hosts a wide range of collisionless particle populations that interact through various wave-particle processes. Among these, cold electrons, with energies below 100 eV, often dominate the plasma density but remain poorly characterized due to measurement challenges such as spacecraft charging and photoelectron contamination. Understanding the contribution of these cold populations to wave–particle interaction is of significant interest. Recent kinetic simulations identified a secondary drift-driven instability, in which parallel-propagating whistler-mode chorus waves excite oblique electrostatic whistler waves near the resonance cone and Bernstein-mode turbulence. These secondary modes enable a new channel of energy transfer from the parallel-propagating whistler wave to the cold electrons. In this work, we develop a moment-based quasilinear theory of the secondary instabilities to quantify such energy exchange. Our results show that these secondary instabilities persist for a wide range of parameters and, in many cases, lead to nearly complete damping of the primary wave. Such secondary instability might limit the amplitude of parallel-propagating whistler waves in Earth's magnetosphere and might explain why high-amplitude oblique whistler or electron Bernstein waves are rarely observed simultaneously with high-amplitude field-aligned whistler waves in the inner magnetosphere.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel simulation via SPPARKS of on-lattice kinetic and Metropolis Monte Carlo models for materials processing

Abstract SPPARKS is an open-source parallel simulation code for developing and running various kinds of on-lattice Monte Carlo models at the atomic or meso scales. It can be used to study the properties of solid-state materials as well as model their dynamic evolution during processing. The modular nature of the code allows new models and diagnostic computations to be added without modification to its core functionality, including its parallel algorithms. A variety of models for microstructural evolution (grain growth), solid-state diffusion, thin film deposition, and additive manufacturing (AM) processes are included in the code. SPPARKS can also be used to implement grid-based algorithms such as phase field or cellular automata models, to run either in tandem with a Monte Carlo method or independently. For very large systems such as AM applications, the Stitch I/O library is included, which enables only a small portion of a huge system to be resident in memory. In this paper we describe SPPARKS and its parallel algorithms and performance, explain how new Monte Carlo models can be added, and highlight a variety of applications which have been developed within the code.

36 MATERIALS SCIENCE↗

A phase-shift-periodic parallel boundary condition for low-magnetic-shear scenarios

Abstract We formulate a generalized periodic boundary condition as a limit of the standard twist-and-shift parallel boundary condition that is suitable for simulations of plasmas with low magnetic shear. This is done by applying a phase shift in the binormal direction when crossing the parallel boundary. While this phase shift can be set to zero without loss of generality in the local flux-tube limit when employing the twist-and-shift boundary condition, we show that this is not the most general case when employing periodic parallel boundaries, and may not even be the most desirable. A non-zero phase shift can be used to avoid the convective cells that plague simulations of the three-dimensional Hasegawa–Wakatani system, and is shown to have measurable effects in periodic low-magnetic-shear gyrokinetic simulations. We propose a numerical program where a sampling of periodic simulations at random pseudo-irrational flux surfaces are used to determine physical observables in a statistical sense. This approach can serve as an alternative to applying the twist-and-shift boundary condition to low-magnetic-shear scenarios, which, while more straightforward, can be computationally demanding.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Optimizing the hit finding algorithm for liquid argon TPC neutrino detectors using parallel architectures

Neutrinos are particles that interact rarely, so identifying them requires large detectors which produce lots of data. Processing this data with the computing power available is becoming even more difficult as the detectors increase in size to reach their physics goals. Liquid argon time projection chamber (LArTPC) neutrino experiments are expected to grow in the next decade to have 100 times more wires than in currently operating experiments, and modernization of LArTPC reconstruction code, including parallelization both at data- and instruction-level, will help to mitigate this challenge. The LArTPC hit finding algorithm is used across multiple experiments through a common software framework. In this paper we discuss a parallel implementation of this algorithm. Using a standalone setup we find speedup factors of two times from vectorization and 30–100 times from multi-threading on Intel architectures. The new version has been incorporated back into the framework so that it can be used by experiments. On a serial execution, the integrated version is about 10 times faster than the previous one and, once parallelization is enabled, more speedups comparable to the standalone program are achieved.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Drosophila melanogaster pigmentation demonstrates adaptive phenotypic parallelism over multiple spatiotemporal scales

Abstract Populations are capable of responding to environmental change over ecological timescales via adaptive tracking. However, the translation from patterns of allele frequency change to rapid adaptation of complex traits remains unresolved. We used abdominal pigmentation in Drosophila melanogaster as a model phenotype to address the nature, genetic architecture, and repeatability of rapid adaptation in the field. We show that D. melanogaster pigmentation evolves as a highly parallel and deterministic response to shared environmental variation across latitude and season in natural North American populations. We then experimentally evolved replicate, genetically diverse fly populations in field mesocosms to remove any confounding effects of demography and/or cryptic structure that may drive patterns in wild populations; we show that pigmentation rapidly responds, in parallel, in fewer than 15 generations. Thus, pigmentation evolves concordantly in response to spatial and temporal climatic axes. We next examined whether phenotypic differentiation was associated with allele frequency change at loci with established links to genetic variance in pigmentation in natural populations. We found that across all spatial and temporal scales, phenotypic patterns were associated with variation at pigmentation-related loci, and the sets of genes we identified at each scale were largely nonoverlapping. Therefore, our findings suggest that parallel phenotypic evolution is associated with distinct components of the polygenic architecture shifting across each environmental axis to produce redundant adaptive patterns.

Evolutionary Biology↗

Random fields from quenched disorder in an archetype for correlated electrons: The parallel spin stripe phase of La 1.6 – x Nd 0.4 Sr x CuO 4 at the 1/8 anomaly

The parallel stripe phase is remarkable both in its own right, and in relation to the other phases with which it coexists. Its inhomogeneous nature makes such states susceptible to random fields from quenched magnetic vacancies. Here we argue this is the case by introducing low concentrations of nonmagnetic Zn impurities (0%–10%) into La 1.6–x ⁢Nd 0.4⁢ Sr x ⁢CuO 4 (Nd-LSCO) with x=0.125 in single-crystal form, well below the percolation threshold of ~41% for a two-dimensional square lattice. Elastic neutron scattering measurements on these crystals show clear magnetic quasi-Bragg peaks at all Zn dopings. While all the Zn-doped crystals display order parameters that merge into each other and the background at ~68 K, the temperature dependence of the order parameter as a function of Zn concentration is drastically different. This result is consistent with meandering charge stripes within the parallel stripe phase, which are pinned in the presence of quenched magnetic vacancies. In turn it implies vacancies that preferentially occupy sites within the charge stripes, and hence that can be very effective at disrupting superconductivity in Nd-LSCO (x=0.125), and, by extension, in all systems exhibiting parallel stripes.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Unbalanced Parallel I/O: An Often-Neglected Side Effect of Lossy Scientific Data Compression

Lossy compression techniques have demonstrated promising results in significantly reducing the scientific data size while guaranteeing the compression error bounds. However, one important yet often neglected side effect of lossy scientific data compression is its impact on the performance of parallel I/O. Our key observation is that the compressed data size is often highly skewed across processes in lossy scientific compression. To understand this behavior, we conduct extensive experiments where we apply three lossy compressors MGARD, ZFP, and SZ, which are specifically designed and optimized for scientific data, to three real-world scientific applications Gray-Scott simulation, WarpX, and XGC. Our analysis result demonstrates that the size of the compressed data is always skewed even if the original data is evenly decomposed among processes. Such skewness widely exists in different scientific applications using different compressors as long as the information density of the data varies across processes. We then systematically study how this side effect of lossy scientific data compression impacts the performance of parallel I/O. We observe that the skewness in the sizes of the compressed data often leads to I/O imbalance, which can significantly reduce the efficiency of I/O bandwidth utilization if not properly handled. In addition, writing data concurrently to a single shared file through MPI-IO library is more sensitive to the unbalanced I/O loads. Therefore, we believe our research community should pay more attention to the unbalanced parallel I/O caused by lossy scientific data compression.

Wang, Xinying↗

A Parallel Cut-Cell Algorithm for the Free-Boundary Grad--Shafranov Problem

A parallel cut-cell algorithm is described to solve the free-boundary problem of the Grad--Shafranov equation. The algorithm reformulates the free-boundary problem in an irregular bounded domain and its important aspects include a searching algorithm for the magnetic axis and separatrix, a surface integral along the irregular boundary to determine the boundary values, an approach to optimize the coil current based on a targeting plasma shape, Picard iterations with Aitken's acceleration for the resulting nonlinear problem, and a Cartesian grid embedded boundary method to handle the complex geometry. Here the algorithm is implemented in parallel using a standard domain-decomposition approach and a good parallel scaling is observed. Numerical results verify the accuracy and efficiency of the free-boundary Grad--Shafranov solver.

97 MATHEMATICS AND COMPUTING↗

Parallel Memory-Independent Communication Bounds for SYRK

In this paper, we focus on the parallel communication cost of multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK). SYRK requires half the computation of general matrix multiplication because of the symmetry of the output matrix. Recent work (Beaumont et al., SPAA '22) has demonstrated that the sequential I/O complexity of SYRK is also a constant factor smaller than that of general matrix multiplication. Inspired by this progress, we establish memory-independent parallel communication lower bounds for SYRK with smaller constants than general matrix multiplication, and we show that these constants are tight by presenting communication-optimal algorithms. The crux of the lower bound proof relies on extending a key geometric inequality to symmetric computations and analytically solving a constrained nonlinear optimization problem. Here, the optimal algorithms use a triangular blocking scheme for parallel distribution of the symmetric output matrix and corresponding computation.

Communication costs↗