Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45

Semi-implicit continuum kinetic modeling of weakly collisional parallel transport in a magnetic mirror

We present implicit-explicit (IMEX) kinetic simulations of weakly collisional parallel plasma transport in magnetic mirror configurations using the continuum code COGENT. The numerical scheme employs a Jacobian-free Newton–Krylov method with algebraic multigrid preconditioning to overcome the severe time step limitations imposed by strong mirror forces in fully explicit schemes. Applied to parameters relevant to the Wisconsin HTS Axisymmetric Mirror experiment, the IMEX approach enables time steps up to 2.5×10 4 times larger than those permitted by explicit methods, resulting in a 2500× speedup in 1D–2V simulations of parallel transport with kinetic ions and Boltzmann electrons. Additionally, a reduced bounce-averaged model for a square mirror is implemented to support the computationally intensive fully kinetic simulations. The bounce-averaged formulation is used to evaluate the numerical convergence of the velocity-space discretization algorithms and to assess the role of the collision model by comparing simulations employing the nonlinear Fokker–Planck and the simplified Lenard–Bernstein–Dougherty collision operators.

Collision theories↗

Phasing of seven-channel fibre laser radiation with dynamic turbulent phase distortions using a stochastic parallel gradient algorithm at a bandwidth of 450 kHz

We have demonstrated an experimental setup for the coherent phasing of a seven-channel fibre laser system ( λ = 1064 nm) in a scheme comprising a master oscillator and a set of parallel amplifiers with lithium niobate-based fibre-optic phase modulators. Using a stochastic parallel gradient algorithm, an instrumental phase modulator control unit ensures a bandwidth of the system up to 450 kHz. The effectiveness of phasing of light transmitted through a turbulent medium with a characteristic time scale τ{sub turb} has been studied experimentally as a function of phasing time τ{sub ph}. The results demonstrate that the average Strehl ratio begins to rise at τ{sub turb}/τ{sub ph} ⩾ 2 and that the effectiveness of compensation for dynamic phase distortions in the beam propagation path rises sharply at τ{sub turb}/τ{sub ph} ≈ 20. For τ{sub turb}/τ{sub ph} ⩾ 30 – 40, the average Strehl ratio remains constant at the level reached. (control of laser radiation parameters)

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Ion heat and parallel momentum transport by stochastic magnetic fields and turbulence

In this work, the theory of turbulent transport of parallel momentum and ion heat by the interaction of stochastic magnetic fields and turbulence is presented. Attention is focused on determining the kinetic stress and the compressive energy flux. A critical parameter is identified as the ratio of the turbulent scattering rate to the rate of parallel acoustic dispersion. For the parameter large, the kinetic stress takes the form of a viscous stress. For the parameter small, the quasilinear residual stress is recovered. In practice, the viscous stress is the relevant form, and the quasilinear limit is not observable. This is the principal prediction of this paper. A simple physical picture is developed and shown to recover the results of the detailed analysis.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel convection and E × B drifts in the TCV snowflake divertor and their effects on target heat-fluxes

Parallel convection and E × B drifts act together to redistribute heat between the strike-points in the low field side snowflake minus (LFS SF–). The cumulative heat convection from both mechanisms is enhanced near the secondary X-point and is shown to dominate over heat conduction, partly explaining why the LFS SF– distributes power more evenly than the single null (SN) or other snowflake (SF) configurations. Pressure profiles at the entrance of the divertor are strongly affected by the position of the secondary X-point and magnetic field direction indicating the importance of E × B drifts. Pressure drops of up to 50% appear between the outer-midplane (OMP) and the divertor entrance enhancing the role of parallel heat convection. The electron temperature and density profiles and the radial turbulent fluxes measured at the OMP are largely unaffected by the changes in divertor geometry, even on flux surfaces where the connection length is infinite.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

New insights on divertor parallel flows,E × B drifts, and fluctuations from in situ, two-dimensional probe measurement in the Tokamak à Configuration Variable

Abstract In situ , two-dimensional (2D) Langmuir probe measurements across a large part of the TCV outer divertor are reported in L-mode discharges with and without divertor baffles. This provides detailed insights into time averaged profiles, particle fluxes, and fluctuation behavior in different divertor regimes. The presence of the baffles is shown to substantially increase the divertor neutral pressure for a given upstream density and to facilitate the access to detachment, an effect that increases with plasma current. The detailed, 2D probe measurements allow for a divertor particle balance, including ion flux contributions from parallel flows and E × B drifts. The poloidal flux contribution from the latter is often comparable or even larger than the former, and the divertor parallel flow direction reverses in some conditions, pointing away from the target. In most conditions, the integrated particle flux at the outer target can be predominantly ascribed to ionization along the outer divertor leg, consistent with a closed-box approximation of the divertor. The exception is a strongly detached divertor, achieved here only with baffles, where the total poloidal ion flux even decreases towards the outer target, indicative of significant plasma recombination. The most striking observation from relative density fluctuation measurements along the outer divertor leg is the transition from poloidally uniform fluctuation levels in attached conditions to fluctuations strongly peaking near the X-point when approaching detachment.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Speeding up particle track reconstruction using a parallel Kalman filter algorithm

One of the most computationally difficult problems expected for the High-Luminosity Large Hadron Collider (HL-LHC) is determining the trajectory of charged particles during event reconstruction. Algorithms used at the LHC today rely on Kalman filtering, which builds physical trajectories incrementally while incorporating material effects and error estimation. Recognizing the need for faster computational throughput, we have adapted Kalman-filter-based methods for highly parallel, many-core SIMD architectures that are now prevalent in high-performance hardware. In this paper, we discuss the design and performance of the improved tracking algorithm, referred to as mkFit. A key piece of the algorithm is the Matriplex library, containing dedicated code to optimally vectorize operations on small matrices. The physics performance of the mkFit algorithm is comparable to the nominal CMS tracking algorithm when reconstructing tracks from simulated proton-proton collisions within the CMS detector. We study the scaling of the algorithm as a function of the parallel resources utilized and find large speedups both from vectorization and multi-threading. mkFit achieves a speedup of a factor of 6 compared to the nominal algorithm when run in a single-threaded application within the CMS software framework.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Turbulence-level dependence of cosmic ray parallel diffusion

ABSTRACT Understanding the transport of energetic cosmic rays belongs to the most challenging topics in astrophysics. Diffusion due to scattering by electromagnetic fluctuations is a key process in cosmic ray transport. The transition from a ballistic to a diffusive-propagation regime is presented in direct numerical calculations of diffusion coefficients for homogeneous magnetic field lines subject to turbulent perturbations. Simulation results are compared with theoretical derivations of the parallel diffusion coefficient’s dependences on the energy and the fluctuation amplitudes in the limit of weak turbulence. The present study shows that the widely used extrapolation of the energy scaling for the parallel diffusion coefficient to high turbulence levels predicted by quasi-linear theory does not provide a universally accurate description in the resonant-scattering regime. It is highlighted here that the numerically calculated diffusion coefficients can be polluted for low energies due to missing resonant interaction possibilities of the particles with the turbulence. Five reduced-rigidity regimes are established, which are separated by analytical boundaries derived in this work. Consequently, a proper description of cosmic ray propagation can only be achieved by using a turbulence-level-dependent diffusion coefficient and can contribute to solving the Galactic cosmic ray gradient problem.

79 ASTRONOMY AND ASTROPHYSICS↗

The first crystal structures of hybrid and parallel four-tetrad intramolecular G-quadruplexes

Abstract G-quadruplexes (GQs) are non-canonical DNA structures composed of stacks of stabilized G-tetrads. GQs play an important role in a variety of biological processes and may form at telomeres and oncogene promoters among other genomic locations. Here, we investigate nine variants of telomeric DNA from Tetrahymena thermophila with the repeat (TTGGGG)n. Biophysical data indicate that the sequences fold into stable four-tetrad GQs which adopt multiple conformations according to native PAGE. Excitingly, we solved the crystal structure of two variants, TET25 and TET26. The two variants differ by the presence of a 3′-T yet adopt different GQ conformations. TET25 forms a hybrid [3 + 1] GQ and exhibits a rare 5′-top snapback feature. Consequently, TET25 contains four loops: three lateral (TT, TT, and GTT) and one propeller (TT). TET26 folds into a parallel GQ with three TT propeller loops. To the best of our knowledge, TET25 and TET26 are the first reported hybrid and parallel four-tetrad unimolecular GQ structures. The results presented here expand the repertoire of available GQ structures and provide insight into the intricacy and plasticity of the 3D architecture adopted by telomeric repeats from T. thermophila and GQs in general.

59 BASIC BIOLOGICAL SCIENCES↗

Optimizing temperature distributions for training neural quantum states using parallel tempering

Parametrized artificial neural networks (ANNs) can be very expressive ansatzes for variational algorithms, reaching state-of-the-art energies on many quantum many-body Hamiltonians. Nevertheless, the training of the ANN can be slow and stymied by the presence of local minima in the parameter landscape. One approach to mitigate this issue is to use parallel tempering methods, and in this work, we focus on the role played by the temperature distribution of the parallel tempering replicas. Using an adaptive method that adjusts the temperatures in order to equate the exchange probability between neighboring replicas, we show that this temperature optimization can significantly increase the success rate of the variational algorithm with negligible computational cost by eliminating bottlenecks in the replicas' random walk. Furthermore, we demonstrate this using two different neural networks, a restricted Boltzmann machine and a feedforward network, which we use to study a toy problem based on a permutation invariant Hamiltonian with a pernicious local minimum and the 𝐽 1 −𝐽 2 model on a rectangular lattice.

Neural network simulations↗

A High-Performance Design for Hierarchical Parallelism in the QMCPACK Monte Carlo code

We introduce a new high-performance design for parallelism within the Quantum Monte Carlo code QMCPACK. We demonstrate that the new design is better able to exploit the hierarchical parallelism of heterogeneous architectures compared to the previous GPU implementation. The new version is able to achieve higher GPU occupancy via the new concept of crowds of Monte Carlo walkers, and by enabling more host CPU threads to effectively offload to the GPU. The higher performance is expected to be achieved independent of the underlying hardware, significantly improving developer productivity and reducing code maintenance costs. Scientific productivity is also improved with full support for fallback to CPU execution when GPU implementations are not available or CPU execution is more optimal.

Luo, Ye↗

Parallel String Graph Construction and Transitive Reduction for De Novo Genome Assembly

One of the most computationally intensive tasks in computational biology is de novo genome assembly, the decoding of the sequence of an unknown genome from redundant and erroneous short sequences. A common assembly paradigm identifies overlapping sequences, simplifies their layout, and creates consensus. Despite many algorithms developed in the literature, the efficient assembly of large genomes is still an open problem. In this work, we introduce new distributed-memory parallel algorithms for overlap detection and layout simplification steps of de novo genome assembly, and implement them in the diBELLA 2D pipeline. Our distributed memory algorithms for both overlap detection and layout simplification are based on linear-algebra operations over semirings using 2D distributed sparse matrices. Our layout step consists of performing a transitive reduction from the overlap graph to a string graph. We provide a detailed communication analysis of the main stages of our new algorithms. diBELLA 2D achieves near linear scaling with over 80% parallel efficiency for the human genome, reducing the runtime for overlap detection by 1.2-1.3× for the human genome and 1.5-1.9× for C.elegans compared to the state-of-the-art. Our transitive reduction algorithm outperforms an existing distributed-memory implementation by 10.5-13.3× for the human genome and 18-29× for the C. elegans. Our work paves the way for efficient de novo assembly of large genomes using long reads in distributed memory.

59 BASIC BIOLOGICAL SCIENCES↗

Automated Calibration of Parallel and Distributed Computing Simulators: A Case Study

Many parallel and distributed computing research results are obtained in simulation, using simulators that mimic real-world executions on some target system. Each such simulator is configured by picking values for parameters that define the behavior of the underlying simulation models it implements. The main concern for a simulator is accuracy: simulated behaviors should be as close as possible to those observed in the real-world target system. This requires that values for each of the simulator's parameters be carefully picked, or “calibrated,” based on ground-truth real-world executions. Examining the current state of the art shows that simulator calibration, at least in the field of parallel and distributed computing, is often undocumented (and thus perhaps often not performed) and, when documented, is described as a labor-intensive, manual process. In this work we evaluate the benefit of automating simulation calibration using simple algorithms. Specifically, we use a real-world case study from the field of High Energy Physics and compare automated calibration to calibration performed by a domain scientist. Our main finding is that automated calibration is on par with or significantly outperforms the calibration performed by the domain scientist. Furthermore, automated calibration makes it straightforward to operate desirable tradeoffs between simulation accuracy and simulation speed.

Mc donald, Jesse↗

A Tradeoff Analysis of Series / Parallel Three-Phase Converter Topologies for Wireless Extreme Chargers

In this paper, extreme fast charging (XFC) technology is studied considering the charge rates of 300 kW for wireless power transfer (WPT) applications. Tradeoff analysis of series and parallel connection of three-phase WPT system are presented by comparisons of voltage and current stresses on power electronics active / passive components. In addition, star (Y) / delta (Δ) connection configurations for three-phase wireless power transfer coupling coils are analyzed with series and LCC resonant compensation circuits. The system series and parallel connection controllability is also reviewed considering voltage and current balance techniques with output control. In a conclusion of evaluation analysis, it is revealed that each component of 300 kW wireless charging network must be designed for high fast charging system and the overall system operation need to be strategically planned for high power charging and infrastructure deployment.

Asa, Erdem↗

A Compact 50kW High Power Density, Hybrid 3-Level Paralleled T-type Inverter for More Electric Aircraft Applications

The demand for high performing, lightweight, reliable inverters, increased the scope of wide bandgap and high-switching frequency based solutions. To achieve such high efficiency inverters, it is vital to focus on the system level design considerations to maximize the benefits of these advanced technologies. This paper presents an improved design based on considerations to further reap the benefits of choosing the right inverter topology; increased capabilities through paralleling devices, with reduced total number of switches; and designing a planarized inverter with PCB based busbar. Appropriate thermal analysis and heatsink design has aided in increased system power density along with the overall efficiency. Demonstration of a 50kW 3-phase 3-level paralleled T-type SiC inverter operating at 40kHz switching frequency for aircraft applications is shown to evaluate the benefits of proposed design methodology. Here, the prototype achieves a high power density of 11kW/L.

42 ENGINEERING↗

Stability Analysis of Parallel Connected Bidirectional WPT System

This paper presents a stability analysis of parallel-connected bi-directional series-series resonant network wireless power transfer (WPT), optimized for Electric Vehicle (EV) charging and vehicle-to-grid (V2G) applications. The study addresses critical stability challenges in systems integrated with diverse distributed energy resources (DERs), including photovoltaics, fuel cells, wind turbines, energy storage systems, and the AC grid. The stability of such integrated DC grid systems is paramount for ensuring reliable operation, particularly under varying power flow conditions and dynamic interactions between parallel WPT systems. The analysis included system impedance characterization, state-space modeling, and open and closed-loop stability evaluations. The results demonstrated that the integration of a robust control architecture effectively mitigates instability risks and supports scalable, efficient operation. This work underscores the converter's adaptability and its potential for large-scale deployment in wireless EV charging infrastructures and integrated DC grid systems.

Asa, Erdem [ORNL] (ORCID:0000000190884812)↗

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

Image Gradient Decomposition for Parallel and Memory-Efficient Ptychographic Reconstruction

Ptychography is a popular microscopic imaging modality for many scientific discoveries and sets the record for highest image resolution. Unfortunately, the high image resolution for ptychographic reconstruction requires significant amount of memory and computations, forcing many applications to compromise their image resolution in exchange for a smaller memory footprint and a shorter reconstruction time. In this paper, we propose a novel image gradient decomposition method that significantly reduces the memory footprint for ptychographic reconstruction by tessellating image gradients and diffraction measurements into tiles. In addition, we propose a parallel image gradient decomposition method that enables asynchronous point-to-point communications and parallel pipelining with minimal overhead on a large number of GPUs. Our experiments on a Titanate material dataset (PbTiO3) with 16632 probe locations show that our Gradient Decomposition algorithm reduces memory footprint by 51 times. In addition, it achieves time-to-solution within 2.2 minutes by scaling to 4158 GPUs with a super-linear strong scaling efficiency at 364% compared to runtimes at 6 GPUs. This performance is 2.7 times more memory efficient, 9 times more scalable and 86 times faster than the state-of-the-art algorithm.

Wang, Xiao↗

Distributed-Memory Parallel Symmetric Nonnegative Matrix Factorization

We develop the first distributed-memory parallel implementation of Symmetric Nonnegative Matrix Factorization (SymNMF), a key data analytics kernel for clustering and dimensionality reduction. Our implementation includes two different algorithms for SymNMF, which give comparable results in terms of time and accuracy. The first algorithm is a parallelization of an existing sequential approach that uses solvers for non symmetric NMF. The second algorithm is a novel approach based on the Gauss-Newton method. It exploits second-order information without incurring large computational and memory costs. We evaluate the scalability of our algorithms on the Summit system at Oak Ridge National Laboratory, scaling up to 128 nodes (4,096 cores) with 70% efficiency. Additionally, we demonstrate our software on an image segmentation task.

Eswar, Srinivas↗