Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27

Spontaneous Hot Flow Anomalies at Quasi-Parallel Shocks: 2. Hybrid Simulations

Motivated by recent THEMIS observations, this paper uses 2.5-D electromagnetic hybrid simulations to investigate the formation of Spontaneous Hot Flow Anomalies (SHFA) upstream of quasi-parallel bow shocks during steady solar wind conditions and in the absence of discontinuities. The results show the formation of a large number of structures along and upstream of the quasi-parallel bow shock. Their outer edges exhibit density and magnetic field enhancements, while their cores exhibit drops in density, magnetic field, solar wind velocity and enhancements in ion temperature. Using virtual spacecraft in the simulation, we show that the signatures of these structures in the time series data are very similar to those of SHFAs seen in THEMIS data and conclude that they correspond to SHFAs. Examination of the simulation data shows that SHFAs form as the result of foreshock cavitons interacting with the bow shock. Foreshock cavitons in turn form due to the nonlinear evolution of ULF waves generated by the interaction of the solar wind with the backstreaming ions. Because foreshock cavitons are an inherent part of the shock dissipation process, the formation of SHFAs is also an inherent part of the dissipation process leading to a highly non-uniform plasma in the quasi-parallel magnetosheath including large scale density and magnetic field cavities.

Quasi-Parallel↗

Parallel Aircraft Trajectory Optimization with Analytic Derivatives

Trajectory optimization is an integral component for the design of aerospace vehicles, but emerging aircraft technologies have introduced new demands on trajectory analysis that current tools are not well suited to address. Designing aircraft with technologies such as hybrid electric propulsion and morphing wings requires consideration of the operational behavior as well as the physical design characteristics of the aircraft. The addition of operational variables can dramatically increase the number of design variables which motivates the use of gradient based optimization with analytic derivatives to solve the larger optimization problems. In this work we develop an aircraft trajectory analysis tool using a Legendre-Gauss-Lobatto based collocation scheme, providing analytic derivatives via the OpenMDAO multidisciplinary optimization framework. This collocation method uses an implicit time integration scheme that provides a high degree of sparsity and thus several potential options for parallelization. The performance of the new implementation was investigated via a series of single and multi-trajectory optimizations using a combination of parallel computing and constraint aggregation. The computational performance results show that in order to take full advantage of the sparsity in the problem it is vital to parallelize both the non-linear analysis evaluations and the derivative computations themselves. The constraint aggregation results showed a significant numerical challenge due to difficulty in achieving tight convergence tolerances. Overall, the results demonstrate the value of applying analytic derivatives to trajectory optimization problems and lay the foundation for future application of this collocation based method to the design of aircraft with where operational scheduling of technologies is key to achieving good performance.

aircraft↗

Massively Parallel Algorithms for Real-Time Wavefront Control of a Dense Adaptive Optics System

In this paper massively parallel algorithms and architectures for real-time wavefront control of a dense adaptive optic system (SELENE) are presented. We have already shown that the computation of a near optimal control algorithm for SELENE can be reduced to the solution of a discrete Poisson equation on a regular domain. Although this represents an optimal computation, due the large size of the system and the high sampling rate requirement, the implementation of this control algorithm poses a computationally challenging problem since it demands a sustained computational throughput of the order of 10 GFlops. We develop a novel algorithm, designated as Fast Invariant Imbedding algorithm, which offers a massive degree of parallelism with simple communication and synchronization requirements. Due to these features, our algorithm is significantly more efficient than other Fast Poisson Solvers for implementation on massively parallel architectures.

massively↗

Improved CDMA Performance Using Parallel Interference Cancellation

This paper considers a general parallel interference cancellation scheme that significantly reduces the degradation effect of user interference but with a lesser implementation complexity than the maximum-likelihood technique. The scheme operates on the fact that parallel processing simultaneously removes from each user the total interference produced by the remaining most reliably received users accessing the channel. The parallel processing can be done in multiple stages. The proposed scheme uses tentative decision devices with different optimum thresholds at the multiple stages to produce the most reliably received data for generation and cancellation of user interference.

CDMA↗

Massively Parallel Neurocomputing for Aerospace Applications

An innovative hybrid, analog-digital charge-domain technology, for the massively parallel VLSI implementation of certain large scale matrix-vector operations, has recently been introduced. It employs arrays of Charge Coupled/Charge Injection Device cells holding an analog matrix of charge, which process digital vectors in parallel by means of binary, non-destructive charge transfer operations. The impact of this technology on massively parallel processing is discussed.

massively↗

Massively Parallel Algorithms for Solution of Schrodinger Equation

In this paper massively parallel algorithms for solution of Schrodinger equation are developed. Our results clearly indicate that the Crank-Nicolson method, in addition to its excellent numerical properties, is also highly suitable for massively parallel computation.

parallel algorithms MIMD parallel architectures ti↗

Pairwise‐Parallel Entangling Gates on Orthogonal Modes in a Trapped‐Ion Chain

Abstract Parallel operations are important for both near‐term quantum computers and larger‐scale fault‐tolerant machines because they reduce execution time and qubit idling. This study proposes and implements a pairwise‐parallel gate scheme on a trapped‐ion quantum computer. The gates are driven simultaneously on different sets of orthogonal motional modes of a trapped‐ion chain. This work demonstrates the utility of this scheme by creating a Greenberger‐Horne‐Zeilinger (GHZ) state in one step using parallel gates with one overlapping qubit. It also shows its advantage for circuits by implementing a digital quantum simulation of the dynamics of an interacting spin system, the transverse‐field Ising model. This method effectively extends the available gate depth by up to two times with no overhead when no overlapping qubit is involved, apart from additional initial cooling. This scheme can be easily applied to different trapped‐ion qubits and gate schemes, broadly enhancing the capabilities of trapped‐ion quantum computers.

Optics↗

Twelve Ways to Fool the Masses When Giving Parallel-in-Time Results

Getting good speedup—let alone high parallel efficiency—for parallel-in-time (PinT) integration examples can be frustratingly difficult. The high complexity and large number of parameters in PinT methods can easily (and unintentionally) lead to numerical experiments that overestimate the algorithm’s performance. In the tradition of Bailey’s article “Twelve ways to fool the masses when giving performance results on parallel computers”, we discuss and demonstrate pitfalls to avoid when evaluating the performance of PinT methods. Despite being written in a light-hearted tone, this paper is intended to raise awareness that there are many ways to unintentionally fool yourself and others and that by avoiding these fallacies more meaningful PinT performance results can be obtained.

97 MATHEMATICS AND COMPUTING↗

A parallel evolutionary multiple-try metropolis Markov chain Monte Carlo algorithm for sampling spatial partitions

We develop an Evolutionary Markov Chain Monte Carlo (EMCMC) algorithm for sampling spatial partitions that lie within a large, complex, and constrained spatial state space. Our algorithm combines the advantages of evolutionary algorithms (EAs) as optimization heuristics for state space traversal and the theoretical convergence properties of Markov Chain Monte Carlo algorithms for sampling from unknown distributions. Local optimality information that is identified via a directed search by our optimization heuristic is used to adaptively update a Markov chain in a promising direction within the framework of a Multiple-Try Metropolis Markov Chain model that incorporates a generalized Metropolis-Hastings ratio. We further expand the reach of our EMCMC algorithm by harnessing the computational power afforded by massively parallel computing architecture through the integration of a parallel EA framework that guides Markov chains running in parallel.

97 MATHEMATICS AND COMPUTING↗

Parallel quantum computing simulations via quantum accelerator platform virtualization

Quantum circuit execution is a central task in quantum computation. Due to inherent quantum-mechanical constraints, quantum computing workflows often involve a considerable number of independent measurements over a large set of slightly different quantum circuits. Here we discuss a simple model for parallelizing such quantum circuit executions that is based on introducing a large array of virtual quantum processing units (mapped to HPC nodes in our case) as a parallel quantum computing platform. Implemented within the XACC framework, the model can readily take advantage of its backend-agnostic features, enabling parallel quantum computing/simulation over any target backend supported by XACC. We illustrate the performance of this approach by demonstrating strong scaling in two pertinent domain science problems, namely in computing the gradients for the multi-contracted variational quantum eigensolver and in data-driven quantum circuit learning, where we vary the number of qubits and the number of circuit layers. Here, the latter simulation leverages the cuQuantum library to run efficiently on GPU-accelerated HPC platforms.

97 MATHEMATICS AND COMPUTING↗

Osmotic control of the spacing of parallel shear cracks in shale growing subcritically in geologic past

The geological genesis of natural cracks in sedimentary rocks such as shale is a problem that needs to be understood to improve the technology of hydraulic fracturing as well as deep sequestration of harmful fluids. Why are the vertical natural cracks roughly parallel and equidistant, and why is the spacing roughly 10 cm rather than 1 cm or 100 cm? Fracture mechanics of critical cracks cannot answer this question. Neither can the material heterogeneity. The growth of critical parallel cracks is impossible because the relative crack face displacements would immediately localize into one crack, leading to an earthquake. The cracks must have formed, on the tectonic time scale, by a slow growth of subcritical shear cracks governed by the Charles-Evans law. The idea advanced here is that what controls the crack spacing is the balance between the reduction, due to shear dilatancy, of the concentration of ions such as Na + and Cl - in each fracture process zone (PFZ), which decelerates the cracks, and the restoration of ion concentration by diffusion of ions from the space between the cracks into the FPZ. This diffusion of water is driven mainly by the osmotic pressure gradient, which offsets the deceleration and depends strongly on the crack spacing. A simple analytical solution of the steady state is rendered possible by approximating the ion concentration profiles between adjacent cracks by parabolic arcs. Applying this theory to Woodford shale yields the approximate crack spacing of 10 cm, which is realistic. Furthermore, the stability of unlimited parallel mode II frictional crack growth is proven by examining the second variation of the free energy. Water concentration drop in the FPZ due to shear dilatancy and its restoration by water diffusion from the inter-crack space have similar effect, although probably much weaker.

42 ENGINEERING↗

A parallel-kinetic-perpendicular-moment model for magnetised plasmas

We describe a new model for the study of weakly collisional, magnetised plasmas derived from exploiting the separation of the dynamics parallel and perpendicular to the magnetic field. This unique system of equations retains the particle dynamics parallel to the magnetic field while approximating the perpendicular dynamics through a spectral expansion in the perpendicular degrees of freedom, analogous to moment-based fluid approaches. In so doing, a hybrid approach is obtained that is computationally efficient enough to allow for larger-scale modelling of plasma systems while eliminating a source of difficulty in deriving fluid equations applicable to magnetised plasmas. We connect this system of equations to historical asymptotic models and discuss advantages and disadvantages of this approach, including the extension of this parallel-kinetic-perpendicular moment beyond the typical region of validity of these more traditional asymptotic models. This paper forms the first of a multi-part series on this new model, covering the theory and derivation, alongside demonstration benchmarks of this approach that include shocks and magnetic reconnection.

astrophysical plasmas↗

High-throughput synthesis of high-entropy alloys via parallelized electric field assisted sintering

Materials discovery and design is an expensive and time-consuming process, though necessary to advance many engineering fields. In this work, a novel tooling design is utilized in conjunction with electric field assisted sintering (EFAS) to effectively create a new high-throughput synthesis technique: parallelized EFAS. Through this technique, a wide range of material compositions and geometries can be synthesized in parallel as isolated samples or as part of contiguous arrays. Multiple tooling designs are explored to examine both the flexibility and limitations of the technique. A series of increasing complex alloys is produced simultaneously using in situ alloying, beginning with pure Ni and adding equimolar constituents up to the septenary high-entropy alloy AlCoCrCuFeMnNi. Microstructural characterization reveals each sample is effectively fully dense and chemically homogenous while exhibiting phases in agreement with CALPHAD predictions. Scalability of parallelized EFAS is then experimentally demonstrated and the implications for materials discovery and automation are discussed.

36 - MATERIALS SCIENCE↗

A parallel, distributed memory implementation of the adaptive sampling configuration interaction method

The many-body simulation of quantum systems is an active field of research that involves several different methods targeting various computing platforms. Many methods commonly employed, particularly coupled cluster methods, have been adapted to leverage the latest advances in modern high-performance computing. Selected configuration interaction (sCI) methods have seen extensive usage and development in recent years. However, the development of sCI methods targeting massively parallel resources has been explored only in a few research works. Here, we present a parallel, distributed memory implementation of the adaptive sampling configuration interaction approach (ASCI) for sCI. In particular, we will address the key concerns pertaining to the parallelization of the determinant search and selection, Hamiltonian formation, and the variational eigenvalue calculation for the ASCI method. Load balancing in the search step is achieved through the application of memory-efficient determinant constraints originally developed for the ASCI-PT2 method. The presented benchmarks demonstrate near optimal speedup for ASCI calculations of Cr 2 (24e, 30o) with 10 6 , 10 7 , and 3 × 10 8 variational determinants on up to 16 384 CPUs. Importantly, to the best of the authors’ knowledge, this is the largest variational ASCI calculation to date.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Design and performance of parallel-channel nanocryotrons in magnetic fields

We introduce a design modification to conventional geometry of the cryogenic three-terminal switch, the nanocryotron (nTron). The conventional geometry of nTrons is modified by including parallel current-carrying channels, an approach aimed at enhancing the device's performance in magnetic field environments. The common challenge in nTron technology is to maintain efficient operation under varying magnetic field conditions. Here, we show that the adaptation of parallel channel configurations leads to an enhanced gate signal sensitivity, an increase in operational gain, and a reduction in the impact of superconducting vortices on nTron operation within magnetic fields up to 1 T. Contrary to traditional designs that are constrained by their effective channel width, the parallel nanowire channels permits larger nTron cross sections, further bolstering the device's magnetic field resilience while improving electro-thermal recovery times due to reduced local inductance. This advancement in nTron design not only augments its functionality in magnetic fields but also broadens its applicability in technological environments, offering a simple design alternative to existing nTron devices.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Understanding cold electron impact on parallel-propagating whistler chorus waves via moment-based quasilinear theory

Earth's magnetosphere hosts a wide range of collisionless particle populations that interact through various wave-particle processes. Among these, cold electrons, with energies below 100 eV, often dominate the plasma density but remain poorly characterized due to measurement challenges such as spacecraft charging and photoelectron contamination. Understanding the contribution of these cold populations to wave–particle interaction is of significant interest. Recent kinetic simulations identified a secondary drift-driven instability, in which parallel-propagating whistler-mode chorus waves excite oblique electrostatic whistler waves near the resonance cone and Bernstein-mode turbulence. These secondary modes enable a new channel of energy transfer from the parallel-propagating whistler wave to the cold electrons. In this work, we develop a moment-based quasilinear theory of the secondary instabilities to quantify such energy exchange. Our results show that these secondary instabilities persist for a wide range of parameters and, in many cases, lead to nearly complete damping of the primary wave. Such secondary instability might limit the amplitude of parallel-propagating whistler waves in Earth's magnetosphere and might explain why high-amplitude oblique whistler or electron Bernstein waves are rarely observed simultaneously with high-amplitude field-aligned whistler waves in the inner magnetosphere.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel simulation via SPPARKS of on-lattice kinetic and Metropolis Monte Carlo models for materials processing

Abstract SPPARKS is an open-source parallel simulation code for developing and running various kinds of on-lattice Monte Carlo models at the atomic or meso scales. It can be used to study the properties of solid-state materials as well as model their dynamic evolution during processing. The modular nature of the code allows new models and diagnostic computations to be added without modification to its core functionality, including its parallel algorithms. A variety of models for microstructural evolution (grain growth), solid-state diffusion, thin film deposition, and additive manufacturing (AM) processes are included in the code. SPPARKS can also be used to implement grid-based algorithms such as phase field or cellular automata models, to run either in tandem with a Monte Carlo method or independently. For very large systems such as AM applications, the Stitch I/O library is included, which enables only a small portion of a huge system to be resident in memory. In this paper we describe SPPARKS and its parallel algorithms and performance, explain how new Monte Carlo models can be added, and highlight a variety of applications which have been developed within the code.

36 MATERIALS SCIENCE↗

A phase-shift-periodic parallel boundary condition for low-magnetic-shear scenarios

Abstract We formulate a generalized periodic boundary condition as a limit of the standard twist-and-shift parallel boundary condition that is suitable for simulations of plasmas with low magnetic shear. This is done by applying a phase shift in the binormal direction when crossing the parallel boundary. While this phase shift can be set to zero without loss of generality in the local flux-tube limit when employing the twist-and-shift boundary condition, we show that this is not the most general case when employing periodic parallel boundaries, and may not even be the most desirable. A non-zero phase shift can be used to avoid the convective cells that plague simulations of the three-dimensional Hasegawa–Wakatani system, and is shown to have measurable effects in periodic low-magnetic-shear gyrokinetic simulations. We propose a numerical program where a sampling of periodic simulations at random pseudo-irrational flux surfaces are used to determine physical observables in a statistical sense. This approach can serve as an alternative to applying the twist-and-shift boundary condition to low-magnetic-shear scenarios, which, while more straightforward, can be computationally demanding.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗