Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Parallel Randomized Tucker Decomposition Algorithms

The Tucker tensor decomposition is a natural extension of the singular value decomposition (SVD) to multiway data. Here, we propose to accelerate Tucker tensor decomposition algorithms by using randomization and parallelization. We present two algorithms that scale to large data and many processors, significantly reduce both computation and communication cost compared to previous deterministic and randomized approaches, and obtain nearly the same approximation errors. The key idea in our algorithms is to perform randomized sketches with Kronecker-structured random matrices, which reduces computation compared to unstructured matrices and can be implemented using a fundamental tensor computational kernel. We provide probabilistic error analysis of our algorithms and implement a new parallel algorithm for the structured randomized sketch. Our experimental results demonstrate that our combination of randomization and parallelization achieves accurate Tucker decompositions much faster than alternative approaches. We observe up to a 16X speedup over the fastest deterministic parallel implementation on 3D simulation data.

Tucker decompositions↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

NESSi : The N on- E quilibrium S ystems S imulation package

The nonequilibrium dynamics of correlated many-particle systems is of interest in connection with pump–probe experiments on molecular systems and solids, as well as theoretical investigations of transport properties and relaxation processes. Nonequilibrium Green’s functions are a powerful tool to study interaction effects in quantum many-particle systems out of equilibrium, and to extract physically relevant information for the interpretation of experiments. Here, we present the open-source software package NESSi (The Non-Equilibrium Systems Simulation package) which allows to perform many-body dynamics simulations based on Green’s functions on the L-shaped Kadanoff–Baym contour. NESSi contains the library libcntr which implements tools for basic operations on these nonequilibrium Green’s functions, for constructing Feynman diagrams, and for the solution of integral and integro-differential equations involving contour Green’s functions. The library employs a discretization of the Kadanoff–Baym contour into time points and a high-order implementation of integration routines. The total integrated error scales up to $\mathcal{O}(N^{-7})$, which is important since the numerical effort increases at least cubically with the simulation time. A distributed-memory parallelization over reciprocal space allows large-scale simulations of lattice systems. We provide a collection of example programs ranging from dynamics in simple two-level systems to problems relevant in contemporary condensed matter physics, including Hubbard clusters and Hubbard or Holstein lattice models. The libcntr library is the basis of a follow-up software package for nonequilibrium dynamical mean-field theory calculations based on strong-coupling perturbative impurity solvers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Polarized fusion and potential in situ tests of fuel polarization survival in a tokamak plasma

Abstract The use of spin-polarized fusion fuels would provide a significant boost towards the ignition of a burning plasma. The cross section for D + T → α + n, would be increased by 1.5 if the fuels were injected with parallel polarization. Furthermore, our simulations demonstrate additional non-linear power gains in large-scale machines such as ITER, due to increased alpha heating. Such benefits require the survival of spin polarizations for periods comparable to the particle confinement time. During the 1980s, calculations predicted that polarizations could survive a plasma environment, although concerns persisted regarding the cumulative impacts of wall recycling. In that era, technical challenges prevented direct tests and left the large scale fueling of a power reactor beyond reach. Over the last decades, this situation has changed dramatically. Detailed simulations of ITER have predicted negligible wall recycling in a high-power reactor, and recent advances in laser-driven sources project the capability of producing large quantities of ∼100% polarized D and T. The remaining crucial step is an in-situ demonstration of polarization survival in a plasma. For this, we outline a measurement strategy using the isospin-mirror reaction, D + 3 He → α + p. Polarized 3 He avoids the complexities of handling tritium, while encompassing the same spin-physics. We evaluate two methods of delivering deuterium, using dynamically polarized Lithium-Deuteride (with vector polarization P V D of 70%) or frozen-spin Hydrogen-Deuteride (with P V D of 40%), together with a method of injecting optically-pumped 3 He (with 65% polarization). Pellets of these materials all have long polarization decay times (∼6 min for LiD at 2 K, ∼2 months for HD at 2 K, and ∼3 d for 3 He at 77 K), all far greater than a plasma shot in a research tokamak such as DIII-D (∼20 s). Both species can be propelled from a single cryogenic injection gun. We review plasma requirements and strategies for detecting polarization survival. Polarization alters both fusion yields and the angular distribution of fusion products, and each of these provides a potential signal. In this paper we simulate a selection of shots with similar characteristics in a future high-T ion H plasma, and find ratios of yields from shots with fuel spins parallel and antiparallel reaching 1.3 (HD + 3 He) to 1.6 (LiD + 3 He) over a wide range of poloidal angles. (A companion paper finds sensitivity to fusion product angular distributions as reflected in the pitch angles of protons and alphas reaching the plasma facing wall.)

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Laser-produced plasmas as drivers of laboratory collisionless quasi-parallel shocks

The creation of a repeatable collisionless quasi-parallel shock in the laboratory would provide a valuable platform for experimental studies of space and astrophysical shocks. However, conducting such an experiment presents substantial challenges. Scaling the results of hybrid simulations of quasi-parallel shock formation to the laboratory highlights the experimentally demanding combination of dense, fast, and magnetized background and driver plasmas required. One possible driver for such experiments is high-energy laser-produced plasmas (LPPs). Preliminary experiments at the University of California, Los Angeles, have explored LPPs as drivers of quasi-parallel shocks by combining the Phoenix Laser Laboratory [Niemann et al., J. Instrum. 7, P03010 (2012)] with a large plasma device [Gekelman et al., Rev. Sci. Instrum. 87, 025105 (2016)]. Beam instabilities and waves characteristic of the early stages of shock formation are observed, but spatial dispersion of the laser-produced plasma prematurely terminates the process. This result is illustrated by experimental measurements and Monte Carlo calculations of LPP density dispersion. The experimentally validated Monte Carlo model is then applied to evaluate several possible approaches to mitigating LPP dispersion in future experiments.

Heuer, P. V. (ORCID:0000000150506606)↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations [Slides]

A separate MCNP6.3 V&V document reports on all the default calculations for all test suites. This report does not include the subcritical multiplication benchmark suite. After some additional clean-up and finalizing the post-processing and documentation steps, the subcritical multiplication benchmark suite will be released in the next version of our vnvstats repository. We tested the new HDF5 PTRAC feature in MCNP6.3 and found encouraging outcomes. Identical results coming out of the simulation with respect to the legacy PTRAC results. The overall runtime for all simulations is reduced by ~20% with the new HDF5 PTRAC capability. We consider giving the new HDF5 PTRAC features a try and using it for all subcritical multiplication and any other relevant (PTRAC) calculations.

97 MATHEMATICS AND COMPUTING↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations

The MCNP6® code, version 6.3, has several new features that are intended to ultimately replace legacy features that are now marked for deprecation. One of these features is the new particle track output (PTRAC) format and capability, where the legacy PTRAC capability still exists alongside the modern PTRAC capability in MCNP6.3. While the MCNP6.3 code has been extensively verified and validated for many applications, the PTRAC feature is not exercised in any of the typical verification and validation (V&V) applications studied during the course of a typical MCNP code release. The primary goal of this paper is to verify that the legacy and modern PTRAC feature produces equivalent results for subcritical multiplication benchmarks previously studied. In the process of verifying that the simulated benchmark results are equivalent, the computational performance is compared between the legacy and modern PTRAC uses. In addition to verification of the update, which is important to the community as a whole, this effort also supports advances in the simulation of recent subcritical neutron noise measurements that require higher computational effort per second of real-time measurement than that of systems typically measured.

97 MATHEMATICS AND COMPUTING↗

Tokamak disruption simulation

UT contribution in this Tokamak Disruption Simulation (TDS) is to develop a parallel high‐order hybridized Discontinuous Galerkin (HDG) methods for large‐s cale MHD simulations. The following are the major goals: 1) Construction of HDG methods for linearized MHD, 2) Rigorous analysis for the HDG formulations for linearized MHD, 3) 2D and 3D simulations to verify the convergent of the HDG methods for linearized MHD, 4) multigrid solvers/preconditioners for HDG formulations, 5) HDG for reconnection problems, 6) Divergence cleaning with HDG, 7) HDG formulations for nonlinear MHD, 8) Picard fixed point HDG approach, 9) IMEX HDG‐DG for nonlinear HDG; 10) parallel large‐scale HDG for linear and nonlinear MHD simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Beam Dynamics Simulations of Transient Beam Loading Effects in the 10GeV EIC Electron Storage Ring

We report on beam dynamics studies of transient beam loading effects in the 10 GeV EIC electron storage ring [1]. The studies are carried out with time-dependent Vlasov-Fokker-Planck simulations performed with the parallel, particle tracking code SPACE [2], which allows to follow self-consistently the dynamics of h bunches, where h in the number of RF buckets, in arbitrary multi-bunch configurations. The specific goal of the numerical simulations is to determine stable RF cavity settings under heavy beam loading. We also study the option to operate with a passive, third-harmonic cavity (3HC) system for bunch lengthening, addressing both stability and the performance limitation due to a gap in the uniform filling pattern for ion clearing.

43 PARTICLE ACCELERATORS↗

Simulating coupled surface–subsurface flows with ParFlow v3.5.0: capabilities, applications, and ongoing development of an open-source, massively parallel, integrated hydrologic model

Surface flow and subsurface flow constitute a naturally linked hydrologic continuum that has not traditionally been simulated in an integrated fashion. Recognizing the interactions between these systems has encouraged the development of integrated hydrologic models (IHMs) capable of treating surface and subsurface systems as a single integrated resource. IHMs are dynamically evolving with improvements in technology, and the extent of their current capabilities are often only known to the developers and not general users. This article provides an overview of the core functionality, capability, applications, and ongoing development of one open-source IHM, ParFlow. ParFlow is a parallel, integrated, hydrologic model that simulates surface and subsurface flows. ParFlow solves the Richards equation for three-dimensional variably saturated groundwater flow and the two-dimensional kinematic wave approximation of the shallow water equations for overland flow. The model employs a conservative centered finite-difference scheme and a conservative finite-volume method for subsurface flow and transport, respectively. ParFlow uses multigrid-preconditioned Krylov and Newton–Krylov methods to solve the linear and nonlinear systems within each time step of the flow simulations. The code has demonstrated very efficient parallel solution capabilities. ParFlow has been coupled to geochemical reaction, land surface (e.g., the Common Land Model), and atmospheric models to study the interactions among the subsurface, land surface, and atmosphere systems across different spatial scales. This overview focuses on the current capabilities of the code, the core simulation engine, and the primary couplings of the subsurface model to other codes, taking a high-level perspective.

58 GEOSCIENCES↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Effect of neutral interactions on parallel transport and blob dynamics in gyrokinetic scrape-off layer simulations

The effect of neutral interactions on scrape-off layer (SOL) turbulence is investigated in a continuum gyrokinetic code that has been coupled to a continuum kinetic model of neutral transport. This extends the work of a previous paper, which compared two NSTX SOL simulations in simple helical geometry, one with neutrals and one without. The former included electron-impact ionization, charge exchange, and wall recycling. Here, the case with neutrals is compared to a gyrokinetic-only simulation that includes an effective ionization source to separate the effect of sourcing from charge exchange collisions. It is observed that sourcing accounts for many features of the simulated SOL with neutrals, including density and temperature magnitudes and reduced normalized density fluctuations, but differences persist. In particular, a flatter density profile results due to changes in parallel transport when neutral collisions are included, illustrating the importance of neutral drag on global plasma properties. An analysis of coherent turbulent structures, or blobs, in these simulations demonstrates the case with neutrals has slower and larger blobs. Here, a series of seeded blob simulations corroborates the blob velocity observation. In general, the blob motion does not contribute significantly to radial transport in these simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large Eddy Simulations of Turbulence below Antarctic Ice Shelves

Turbulence below ice shelves is of key importance for predicting ice-shelf melt rates and consequently the contribution of ice sheets to sea level rise. In this project we conducted Large-Eddy Simulations (LES) to improve our understanding of sub-ice-shelf ocean turbulence and the relationship between ice-shelf melt rates and ocean conditions. Over the course of the second and final year of this project (FY2020), we accomplished both major code developments for the PArallel Large-eddy simulation Model (PALM; Maronga et al., 2015) and conducted a suite of simulations which form the basis for a publication in preparation.

58 GEOSCIENCES↗

Parallel transport dynamics for mixed quantum states with applications to time-dependent density functional theory

Direct simulation of the von Neumann dynamics for a general (pure or mixed) quantum state can often be expensive. One prominent example is the real-time time-dependent density functional theory (rt-TDDFT), a widely used framework for the first principle description of many-electron dynamics in chemical and materials systems. Practical rt-TDDFT calculations often avoid the direct simulation of the von Neumann equation, and solve instead a set of Schrödinger equations, of which the dynamics is equivalent to that of the von Neumann equation. However, the time step size employed by the Schrödinger dynamics is often much smaller. Here, in order to improve the time step size and the overall efficiency of the simulation, we generalize a recent work of the parallel transport (PT) dynamics for simulating pure states [An, Lin, Multiscale Model. Simul. 18, 612, 2020] to general quantum states. The PT dynamics provides the optimal gauge choice, and can employ a time step size comparable to that of the von Neumann dynamics. Going beyond the linear and near adiabatic regime in previous studies, we find that the error of the PT dynamics can be bounded by certain commutators between Hamiltonians, density matrices, and their derived quantities. Such a commutator structure is not present in the Schrödinger dynamics. We demonstrate that the parallel transport-implicit midpoint (PT-IM) method is a suitable method for simulating the PT dynamics, especially when the spectral radius of the Hamiltonian is large. The commutator structure of the error bound, and numerical results for model rt-TDDFT calculations in both linear and nonlinear regimes, confirm the advantage of the PT dynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗