Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44

Commuting embeddings for parallel strategies in non-local games

Non-local games provide a versatile framework for probing quantum correlations and for benchmarking the power of entanglement. In finite dimensions, the standard method for playing several games in parallel requires a tensor product of the local Hilbert spaces, which scales additively in the number of qubits. In this work, we show that this additive cost can be reduced by exploiting algebraic embeddings. We introduce two forms of compressions. First, when a referee selects one game from a finite collection of games at random, the game quantum strategy can be implemented using a maximally entangled state of dimension equal to the largest individual game, thereby eliminating the need for repeated state preparations. Second, we establish conditions under which several games can be played simultaneously in parallel on fewer qubits than the tensor product baseline. These conditions are expressed in terms of commuting embeddings of the game algebras. Moreover, we provide a constructive framework for building such embeddings. Using tools from Lie theory, we show that aligning the various game algebras into a common Cartan decomposition enables such a qubit reduction. Beyond the theoretical contribution, our framework casts NLGs as algebraic primitives for distributed and resource-constrained quantum computations and suggested NLGs as a comparable device-independent dimension witness.

Commuting embeddings↗

Extending TOUGH + HYDRATE with a parallel particle transport simulator: numerical investigation of sand production during gas production from hydrate deposits

A new parallel code for simulating particle transport in porous media is integrated with the TOUGH + HYDRATE simulator to investigate sand production associated with gas production from unconsolidated gas hydrate-bearing sediments (HBS). Here, the parallel coupled simulator is named THMPT and uses the integral finite difference method to describe the Darcian and non-Darcian flow of fluids and heat transport, the finite element method to describe the associated geomechanical changes, and the discrete element method to track the trajectory of individual sand particles within the HBS. The THMPT simulator is written in Fortran, incorporates multiple optimized algorithms, and can comprehensively address the coupled flow, thermal, chemical, geomechanical, and particle transport processes that characterize the system behaviors during gas production from HBS. The simulator can capture all processes involved in sand particle transport in porous media, including sand detachment, collision, clogging (i.e., bridging), and migration. A benchmark case study of sand production in the course of depressurization-induced gas production from a representative HBS reveals various distinct microscopic particle migration mechanisms and the adverse impact of sand particle detachment, transport, and clogging. The numerical investigation also examines the effect of bottomhole pressure on mitigating sand production. The simulation results indicate that sand clogging near the wellbore significantly reduces permeability, decreasing gas production by at least 50%. Lastly, the efficiency of gravel packing in mitigating sand production is numerically evaluated, revealing that the structure of the porous media appears to profoundly influence the macroscopic motion behavior of sand particles and sand clogging characteristics.

discrete element method↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

RF-transpond: A 1D coupled cold plasma wave and plasma transport model for ponderomotive force driven density modification parallel to B

The RF-Transpond code couples a fluid plasma transport solver with a frequency domain cold plasma RF wave solver in a 1D domain parallel to a strong background magnetic field. A ponderomotive force term proportional to parallel gradients in the electric field strength is included in the transport model in order to describe ponderomotive effects in the scrape-off layer (SOL) of fusion plasmas. The transport and wave codes are verified independently and a coupled case corresponding to experimental parameters from the LArge Plasma Device (LAPD) is presented. The density perturbation ratio R n , calculated to describe ponderomotive force driven modifications, is up to 20% for the simulation inputs used.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

SPACE: 3D parallel solvers for Vlasov-Maxwell and Vlasov-Poisson equations for relativistic plasmas with atomic transformations

A parallel, relativistic, three-dimensional particle-in-cell code SPACE has been developed for the simulation of electromagnetic fields, relativistic particle beams, and plasmas. In addition to the standard second-order Particle-in-Cell (PIC) algorithm, SPACE includes efficient novel algorithms to resolve atomic physics processes such as multi-level ionization of plasma atoms, recombination, and electron attachment to dopants in dense neutral gases. SPACE also contains a highly adaptive particle-based method, called Adaptive Particle-in-Cloud (AP-Cloud), for solving the Vlasov-Poisson problems. It eliminates the traditional Cartesian mesh of PIC and replaces it with an adaptive octree data structure. The code's algorithms, structure, capabilities, parallelization strategy, and performance have been discussed. Additionally, typical examples of SPACE applications to accelerator science and engineering problems are described.

43 PARTICLE ACCELERATORS↗

Parallel diffusion operator for magnetized plasmas with improved spectral fidelity

Diffusive transport processes in magnetized plasmas are highly anisotropic, with fast parallel transport along the magnetic field lines sometimes faster than perpendicular transport by orders of magnitude. This constitutes a major challenge for describing non-grid-aligned magnetic structures in Eulerian (grid-based) simulations. Here, the present paper describes and validates a new method for parallel diffusion in magnetized plasmas based on the anti-symmetry representation [Halpern and Waltz, Phys. Plasmas 25, 060703 (2018)]. In the anti-symmetry formalism, diffusion manifests as a flow operator involving the logarithmic derivative of the transported quantity. Qualitative plane wave analysis shows that the new operator naturally yields better discrete spectral resolution compared to its conventional counterpart. Numerical simulations comparing the new method against existing finite difference methods are carried out, showing significant improvement. In particular, we find that combining anti-symmetry with finite differences in diagonally staggered grids essentially eliminates the so-called “artificial numerical diffusion” that affects conventional finite difference and finite volume methods.

Anisotropic diffusion↗

Design-time performance modeling of compositional parallel programs

Performance models are powerful instruments for understanding the performance of parallel systems and uncovering their bottlenecks. Already during system design, performance models can help ponder alternative development options. However, creating a performance model – whether theoretically or empirically – for an entire application that does not exist yet is challenging. In this paper, we propose to generate performance models of full programs from performance models of their components using formal composition operators derived from parallel design patterns. As long as the design of the overall system follows such a pattern, its performance model can be predicted with reasonable accuracy without an actual implementation. In conclusion, we demonstrate our approach with design patterns of varying complexity, including pipeline, task pool, and eventually MapReduce, which is representative of a broad class of data-analytics applications.

97 MATHEMATICS AND COMPUTING↗

Dynamics of an inertially collapsing gas bubble between two parallel, rigid walls

The collapse of cavitation bubbles in channel flows can give rise to structural damage along neighbouring walls. Although the collapse of a bubble near a single wall has been studied extensively, less is known about bubble collapse between two walls, e.g. as in a channel. We conduct highly resolved, direct simulations of the Navier–Stokes equations to investigate the bubble dynamics and pressures produced by the collapse of a bubble between two parallel rigid walls. We examine the dependence of the dynamics and pressures on the initial bubble location, confinement and driving pressure. For a fixed initial stand-off distance, as the channel width increases the bubble volume, migration distance and re-entrant jet speed approach their single-wall counterparts. We obtain an expression for the minimum channel width at which the confinement does not affect the bubble dynamics depending on the driving pressure difference and initial stand-off distance. For a fixed channel width, varying stand-off distance reduced the maximum wall pressures in the channel relative to the single wall; the trend was consistent for three different driving pressures. Two different jetting behaviours are seen when the bubble is centred in the channel, depending on the channel width. Under significant confinement, wall-parallel re-entrant jets impinge upon each other and further intensify the collapse of the vortex ring.

Rodriguez, Jr, Mauro (ORCID:0000000305450265)↗

Pedestal origin and extrapolation of high-density small edge-localised-modes peak parallel energy fluence in ITER and SPARC

Experimental analysis and simulations with the BOUT++ code show that small edge-localised modes (ELMs) in reactor-relevant high-density regimes originate in a region close to the separatrix and only marginally perturb the pedestal structure. The measured divertor peak parallel energy fluence (ε ∥,peak ) for a database of small ELM scenarios in DIII-D and ASDEX Upgrade can be reproduced, within 40 % accuracy on average, if an ad hoc modification of the Eich peak parallel ELM energy fluence model is applied to account for the small ELM pedestal birth location. This allows for first-order extrapolation of small-ELM divertor ε ∥,peak to ITER and SPARC, resulting in values that satisfy the nominal melting threshold of tungsten monoblocks of 12 MJ m −2 . The findings reported in this study, both via modelling and direct measurements, constitute a step forward in assessing small ELMs in high edge-collisionality scenarios as a viable plasma regime for the operation of next-generation fusion machines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Cyclotron breaking: a mechanism for parallel ion cyclotron waves to heat the fast solar wind

The Parker Solar Probe mission has observed near-continuous power in parallel ion cyclotron waves (PICWs) in the young, fast solar wind. These waves are unlikely to be directly produced by the turbulent cascade and are likely born of a local instability; yet, they are observed to both cool – and heat – the plasma. We propose that these observations can be self-consistently explained as the natural consequence of PICWs propagating in the inhomogeneous solar wind after they have been driven unstable. In this work, we argue that strong proton heating by a turbulent cascade of oblique ICWs will result in PICWs being driven unstable in a process known as quasi-linear focusing. Because the power in the turbulent cascade is concentrated at scales above the turbulent transition region, PICWs will be driven unstable within a range of wavenumbers parallel to the background magnetic field, 𝑘 ∥ , that is bounded from above by 𝑘$^{∗}_{∥P}$, corresponding to the start of the transition region. As unstable PICWs propagate away from the Sun to regions of lower proton density, their 𝑘 ∥ , multiplied by the proton inertial length 𝑑 p , increases. Eventually, 𝑘$^{∗}_{∥P}$ of the PICWs becomes larger than 𝑘$^{∗}_{∥P}$⁢𝑑 p and the waves damp, heating the solar wind. We call this effect ‘cyclotron breaking’, in analogy with ocean waves breaking on the shore. We then discuss the testable predictions of the theory, including a distinct heating signature in which PICWs cool fast protons and heat slow protons at any given heliocentric distance 𝑟. Finally, we conjecture that cyclotron breaking can lead to net heating by PICWs if the power emitted as PICWs decreases sufficiently rapidly with 𝑟 that local emission of PICWs is overwhelmed by the local damping of PICWs generated closer to the Sun.

plasma heating↗

Implementation of Relativistic Coupled Cluster Theory for Massively Parallel GPU-Accelerated Computing Architectures

In this paper, we report reimplementation of the core algorithms of relativistic coupled cluster theory aimed at modern heterogeneous high-performance computational infrastructures. The code is designed for parallel execution on many compute nodes with optional GPU coprocessing, accomplished via the new ExaTENSOR back end. The resulting ExaCorr module is primarily intended for calculations of molecules with one or more heavy elements, as relativistic effects on the electronic structure are included from the outset. In the current work, we thereby focus on exact two-component methods and demonstrate the accuracy and performance of the software. The module can be used as a stand-alone program requiring a set of molecular orbital coefficients as the starting point, but it is also interfaced to the DIRAC program that can be used to generate these. We therefore also briefly discuss an improvement of the parallel computing aspects of the relativistic self-consistent field algorithm of the DIRAC program.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Substituent and Heteroatom Effects on π–π Interactions: Evidence That Parallel-Displaced π-Stacking is Not Driven by Quadrupolar Electrostatics

Stacking interactions are a recurring motif in supramolecular chemistry and biochemistry, where a persistent theme is a preference for parallel-displaced aromatic rings rather than face-to-face π-stacking. This is usually explained in terms of quadrupole–quadrupole interactions between the arene moieties but that interpretation is inconsistent with accurate calculations, which reveal that the quadrupolar picture is qualitatively wrong. At typical π-stacking distances, quadrupolar electrostatics may differ in sign from an exact calculation based on charge densities of the interacting arenes. We apply symmetry-adapted perturbation theory to dimers composed of substituted benzene and various aromatic heterocycles, which display a wide range of electrostatic interactions, and we investigate the interplay of Pauli repulsion, dispersion, and electrostatics as it pertains to parallel-displaced π-stacking. Profiles of energy components along cofacial slip-stacking coordinates support a prominent role for the “van der Waals model” (dispersion in competition with Pauli repulsion), even for polar monomers where electrostatic interactions are significant. While electrostatic interactions are necessary to explain the optimal face-to-face π-stacking distance and to account for the relative orientation of one polar arene with respect to another, we find no evidence to support continued invocation of quadrupolar electrostatics as a basis for π-stacking. Our results suggest that a driving force for offset-stacking exists even in the absence of electrostatic interactions. Consequently, tuning electrostatics via functionalization does not guarantee that slip-stacking can be avoided. This has implications for rational design of soft materials and other supramolecular architectures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel IO Libraries for Managing HEP Experimental Data

The computing and storage requirements of the energy and intensity frontiers will grow significantly during the Run 4 & 5 and the HL-LHC era. Similarly, in the intensity frontier, with larger trig ger readouts during supernovae explosions, the Deep Underground Neutrino Experiment (DUNE) will have unique computing challenges that could be addressed by the use of parallel and accelerated dataprocessing capabilities. Most of the requirements of the energy and intensity frontier experiments rely on increasing the role of high performance computing (HPC) in the HEP community. In this presentation, we will describe our ongoing efforts that are focused on using HPC resources for the next generation HEP experiments. The HEPCCE (High Energy Physics-Center for Computational Excellence) IOS (Input/Output and Storage) group has been developing approaches to map HEP data to the HDF5 , an IO library optimized for the HPC platforms to store the intermediate HEP data. The complex HEP data products are serialized using ROOT to allow for experiment independent general mapping approaches of the HEP data to the HDF5 format. The mapping approaches can be optimized for high performance parallel IO. Similarly, simpler data can be directly mapped into the HDF5, which can also be suitable for offloading into the GPUs directly. We will present our works on both complex and simple data model models.

Bashyal, Amit↗

Massively Parallel Quantum Chemistry: A high-performance research platform for electronic structure

The Massively Parallel Quantum Chemistry (MPQC) program is a 30-year-old project that enables facile development of electronic structure methods for molecules for efficient deployment to massively parallel computing architectures. Here, we describe the historical evolution of MPQC’s design into its latest (fourth) version, the capabilities and modular architecture of today’s MPQC, and how MPQC facilitates rapid composition of new methods as well as its state-of-the-art performance on a variety of commodity and high-end distributed-memory computer platforms.

Peng, Chong↗

Parallel high-frequency magnetic sensing with an array of flux transformers and multi-channel optically pumped magnetometer for hand MRI application

Here, we investigate an approach for parallel high-frequency magnetic sensing based on a multi-channel radio frequency (RF) optically pumped magnetometer (OPM) coupled to multiple flux transformers (FTs) with a focus on hand magnetic resonance imaging (MRI) application at ultra-low field (ULF). Multiple RF OPM sensing channels are realized by using a single large-area alkali-metal vapor cell and two laser beams for pumping and probing, shared for all the channels. This design leads to significant cost reduction when multi-channel sensing is desirable, as in the case of ULF MRI. The FT, composed of two connected coils, serves as a transmitter of a target magnetic field to the OPM, while decoupling the OPM from untargeted magnetic fields in the sensing area that can limit the OPM performance. For hand MRI application, theoretical and numerical analysis is performed to determine an optimal geometry for the FT array that could improve signal-to-noise ratio (SNR) and sufficiently reduce crosstalk between FTs. We estimate that the optimized multi-channel FT-OPM sensor can achieve a magnetic field sensitivity of the order of 1 fT / Hz 1 / 2 above 100 kHz, which would be sufficient for 1 mm resolution MRI. In general, the multi-channel capability enables simultaneous magnetic measurements, thus reducing the sensing time and improving the SNR, and we anticipate many applications of the multi-channel FT-OPM sensor beyond the targeted here hand MRI: anatomical parallel ULF MRI of the human brain and other parts of the body, airport security screening, magnetic material imaging, and many others.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Structured illumination with thermal imaging (SI-TI): A dynamically reconfigurable metrology for parallelized thermal transport characterization

The recent push for the “materials by design” paradigm requires synergistic integration of scalable computation, synthesis, and characterization. Among these, techniques for efficient measurement of thermal transport can be a bottleneck limiting the experimental database size, especially for diverse materials with a range of roughness, porosity, and anisotropy. Traditional contact thermal measurements have challenges with throughput and the lack of spatially resolvable property mapping, while non-contact pump-probe laser methods generally need mirror smooth sample surfaces and also require serial raster scanning to achieve property mapping. Here, we present structured illumination with thermal imaging (SI-TI), a new thermal characterization tool based on parallelized all-optical heating and thermometry. Experiments on representative dense and porous bulk materials as well as a 3D printed thermoelectric thick film (~50 μm) demonstrate that SI-TI (1) enables paralleled measurement of multiple regions and samples without raster scanning; (2) can dynamically adjust the heating pattern purely in software, to optimize the measurement sensitivity in different directions for anisotropic materials; and (3) can tolerate rough (~3 μm) and scratched sample surfaces. Here, this work highlights a new avenue in adaptivity and throughput for thermal characterization of diverse materials.

42 ENGINEERING↗

On parallel laser beam merger in plasmas

Self-focusing instability is a well-known phenomenon of nonlinear optics, which is of great importance in the field of laser–plasma interactions. Self-focusing instability leads to beam focusing and, consequently, breakup into multiple laser filaments. The majority of applications tend to avoid a laser filamentation regime due to its detrimental role on laser spot profile and peak intensity. In our work, using nonlinear Schrödinger equation solver and particle-in-cell simulations, we address the problem of interaction of multiple parallel beams in plasmas. We consider both non-relativistic and moderately relativistic regimes and demonstrate how the physics of parallel beam interaction transitions from the familiar self- and mutual-focusing instabilities in the non-relativistic regime to a moderately relativistic regime, where an analytical description of filament interaction is not available.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Kelvin–Helmholtz instability under stabilizing parallel magnetic field in nonhomogeneous compressible MHD flows

We study the Kelvin–Helmholtz instability (KHI) for the general case of a compressible, nonhomogeneous, magnetized plasma flow. The study is limited to a vortex sheet interface with an imposed parallel magnetic field. We introduce a new formalism based on a convective Mach number M c , a convective Alfvénic Mach number M Ac , and a total convective Mach number that combines the two. We derive an analytic expression of the KHI growth rate for a homogeneous flow (i.e., zero Atwood number, A=0) that converges toward both the expression for unmagnetized compressible flow and Chandrasekhar's expression for magnetized incompressible flow. Otherwise, the dispersion relation is solved numerically and allows deriving general stability diagrams of magnetized KHI for the triplet (A, M c , β −plasma) parameters. We show these parameters uniquely define all configurations for a parallel magnetic field. We also construct diagrams with respect to the convective Alfvénic Mach number, the β − plasma parameter, or the magnetic field showing which magnetic field strength is required for stabilizing a given shear flow. The theoretical growth rates are compared with 18 simulations made with the GAMERA code, currently used for 3D magnetospheric simulations. Finally, we apply our results to the analysis of a past KHI experiment performed at the OMEGA laser facility, showing linear theory succeeds to provide accurate estimates of the growth rate at early times. We further discuss how our results can inform future experiments in the high-Mach magnetized regime at the National Ignition Facility. Possible limitations of the study due to resistive, mixing, or turbulence effects are discussed.

compressible flows↗