Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

The formation of quasi-parallel shocks

In a collisionless plasma, the coupling between a piston and the plasma must take place through either laminar or turbulent electromagnetic fields. Of the three types of coupling (laminar, Larmor and turbulent), shock formation in the parallel regime is dominated by the latter and in the quasi-parallel regime by a combination of all three, depending on the piston. In the quasi-perpendicular regime, there is usually a good separation between piston and shock. This is not true in the quasi-parallel and parallel regime. Hybrid numerical simulations for hot plasma pistons indicate that when the electrons are hot, a shock forms, but does not cleanly decouple from the piston. For hot ion pistons, no shock forms in the parallel limit: in the quasi-parallel case, a shock forms, but there is severe contamination from hot piston ions. These results suggest that the properties of solar and astrophysical shocks, such as particle acceleration, cannot be readily separated from their driving mechanism.

Cargill, Peter J.↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations [Slides]

A separate MCNP6.3 V&V document reports on all the default calculations for all test suites. This report does not include the subcritical multiplication benchmark suite. After some additional clean-up and finalizing the post-processing and documentation steps, the subcritical multiplication benchmark suite will be released in the next version of our vnvstats repository. We tested the new HDF5 PTRAC feature in MCNP6.3 and found encouraging outcomes. Identical results coming out of the simulation with respect to the legacy PTRAC results. The overall runtime for all simulations is reduced by ~20% with the new HDF5 PTRAC capability. We consider giving the new HDF5 PTRAC features a try and using it for all subcritical multiplication and any other relevant (PTRAC) calculations.

97 MATHEMATICS AND COMPUTING↗

A Computer Simulation of the System-Wide Effects of Parallel-Offset Route Maneuvers

Most aircraft managed by air-traffic controllers in the National Airspace System are capable of flying parallel-offset routes. This paper presents the results of two related studies on the effects of increased use of offset routes as a conflict resolution maneuver. The first study analyzes offset routes in the context of all standard resolution types which air-traffic controllers currently use. This study shows that by utilizing parallel-offset route maneuvers, significant system-wide savings in delay due to conflict resolution of up to 30% are possible. It also shows that most offset resolutions replace horizontal-vectoring resolutions. The second study builds on the results of the first and directly compares offset resolutions and standard horizontal-vectoring maneuvers to determine that in-trail conflicts are often more efficiently resolved by offset maneuvers.

Lauderdale, Todd A.↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations

The MCNP6® code, version 6.3, has several new features that are intended to ultimately replace legacy features that are now marked for deprecation. One of these features is the new particle track output (PTRAC) format and capability, where the legacy PTRAC capability still exists alongside the modern PTRAC capability in MCNP6.3. While the MCNP6.3 code has been extensively verified and validated for many applications, the PTRAC feature is not exercised in any of the typical verification and validation (V&V) applications studied during the course of a typical MCNP code release. The primary goal of this paper is to verify that the legacy and modern PTRAC feature produces equivalent results for subcritical multiplication benchmarks previously studied. In the process of verifying that the simulated benchmark results are equivalent, the computational performance is compared between the legacy and modern PTRAC uses. In addition to verification of the update, which is important to the community as a whole, this effort also supports advances in the simulation of recent subcritical neutron noise measurements that require higher computational effort per second of real-time measurement than that of systems typically measured.

97 MATHEMATICS AND COMPUTING↗

A nonrecursive order N preconditioned conjugate gradient: Range space formulation of MDOF dynamics

While excellent progress has been made in deriving algorithms that are efficient for certain combinations of system topologies and concurrent multiprocessing hardware, several issues must be resolved to incorporate transient simulation in the control design process for large space structures. Specifically, strategies must be developed that are applicable to systems with numerous degrees of freedom. In addition, the algorithms must have a growth potential in that they must also be amenable to implementation on forthcoming parallel system architectures. For mechanical system simulation, this fact implies that algorithms are required that induce parallelism on a fine scale, suitable for the emerging class of highly parallel processors; and transient simulation methods must be automatically load balancing for a wider collection of system topologies and hardware configurations. These problems are addressed by employing a combination range space/preconditioned conjugate gradient formulation of multi-degree-of-freedom dynamics. The method described has several advantages. In a sequential computing environment, the method has the features that: by employing regular ordering of the system connectivity graph, an extremely efficient preconditioner can be derived from the 'range space metric', as opposed to the system coefficient matrix; because of the effectiveness of the preconditioner, preliminary studies indicate that the method can achieve performance rates that depend linearly upon the number of substructures, hence the title 'Order N'; and the method is non-assembling. Furthermore, the approach is promising as a potential parallel processing algorithm in that the method exhibits a fine parallel granularity suitable for a wide collection of combinations of physical system topologies/computer architectures; and the method is easily load balanced among processors, and does not rely upon system topology to induce parallelism.

Kurdila, Andrew J.↗

Tokamak disruption simulation

UT contribution in this Tokamak Disruption Simulation (TDS) is to develop a parallel high‐order hybridized Discontinuous Galerkin (HDG) methods for large‐s cale MHD simulations. The following are the major goals: 1) Construction of HDG methods for linearized MHD, 2) Rigorous analysis for the HDG formulations for linearized MHD, 3) 2D and 3D simulations to verify the convergent of the HDG methods for linearized MHD, 4) multigrid solvers/preconditioners for HDG formulations, 5) HDG for reconnection problems, 6) Divergence cleaning with HDG, 7) HDG formulations for nonlinear MHD, 8) Picard fixed point HDG approach, 9) IMEX HDG‐DG for nonlinear HDG; 10) parallel large‐scale HDG for linear and nonlinear MHD simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Beam Dynamics Simulations of Transient Beam Loading Effects in the 10GeV EIC Electron Storage Ring

We report on beam dynamics studies of transient beam loading effects in the 10 GeV EIC electron storage ring [1]. The studies are carried out with time-dependent Vlasov-Fokker-Planck simulations performed with the parallel, particle tracking code SPACE [2], which allows to follow self-consistently the dynamics of h bunches, where h in the number of RF buckets, in arbitrary multi-bunch configurations. The specific goal of the numerical simulations is to determine stable RF cavity settings under heavy beam loading. We also study the option to operate with a passive, third-harmonic cavity (3HC) system for bunch lengthening, addressing both stability and the performance limitation due to a gap in the uniform filling pattern for ion clearing.

43 PARTICLE ACCELERATORS↗

Macro Scale Independently Homogenized Subcells for Modeling Braided Composites

An analytical method has been developed to analyze the impact response of triaxially braided carbon fiber composites, including the penetration velocity and impact damage patterns. In the analytical model, the triaxial braid architecture is simulated by using four parallel shell elements, each of which is modeled as a laminated composite. Currently, each shell element is considered to be a smeared homogeneous material. The commercial transient dynamic finite element code LS-DYNA is used to conduct the simulations, and a continuum damage mechanics model internal to LS-DYNA is used as the material constitutive model. To determine the stiffness and strength properties required for the constitutive model, a top-down approach for determining the strength properties is merged with a bottom-up approach for determining the stiffness properties. The top-down portion uses global strengths obtained from macro-scale coupon level testing to characterize the material strengths for each subcell. The bottom-up portion uses micro-scale fiber and matrix stiffness properties to characterize the material stiffness for each subcell. Simulations of quasi-static coupon level tests for several representative composites are conducted along with impact simulations.

Blinzler, Brina J.↗

Modification of a Macromechanical Finite-Element Based Model for Impact Analysis of Triaxially-Braided Composites

A macro level finite element-based model has been developed to simulate the mechanical and impact response of triaxially-braided polymer matrix composites. In the analytical model, the triaxial braid architecture is simulated by using four parallel shell elements, each of which is modeled as a laminated composite. For the current analytical approach, each shell element is considered to be a smeared homogeneous material. The commercial transient dynamic finite element code LS-DYNA is used to conduct the simulations, and a continuum damage mechanics model internal to LS-DYNA is used as the material constitutive model. The constitutive model requires stiffness and strength properties of an equivalent unidirectional composite. Simplified micromechanics methods are used to determine the equivalent stiffness properties, and results from coupon level tests on the braided composite are utilized to back out the required strength properties. Simulations of quasi-static coupon tests of several representative braided composites are conducted to demonstrate the correlation of the model. Impact simulations of a represented braided composites are conducted to demonstrate the capability of the model to predict the penetration velocity and damage patterns obtained experimentally.

Goldberg, Robert K.↗

X-composer: enabling cross-environments in-situ workflows between HPC and cloud

As large-scale scientific simulations and big data analyses become more popular, it is increasingly more expensive to store huge amounts of raw simulation results to perform post-analysis. To minimize the expensive data I/O, "in-situ" analysis is a promising approach, where data analysis applications analyze the simulation generated data on the fly without storing it first. However, it is challenging to organize, transform, and transport data at scales between two semantically different ecosystems due to the distinct software and hardware difference. To tackle these challenges, we design and implement the X-Composer framework. X-Composer connects cross-ecosystem applications to form an "in-situ" scientific workflow, and provides a unified approach and recipe for supporting such hybrid in-situ workflows on distributed heterogeneous resources. X-Composer reorganizes simulation data as continuous data streams and feeds them seamlessly into the Cloud-based stream processing services to minimize I/O overheads. For evaluation, we use X-Composer to set up and execute a cross-ecosystem workflow, which consists of a parallel Computational Fluid Dynamics simulation running on HPC, and a distributed Dynamic Mode Decomposition analysis application running on Cloud. Our experimental results show that X-Composer can seamlessly couple HPC and Big Data jobs in their own native environments, achieve good scalability, and provide high-fidelity analytics for ongoing simulations in real-time.

Wang, Dali↗

Effect of neutral interactions on parallel transport and blob dynamics in gyrokinetic scrape-off layer simulations

The effect of neutral interactions on scrape-off layer (SOL) turbulence is investigated in a continuum gyrokinetic code that has been coupled to a continuum kinetic model of neutral transport. This extends the work of a previous paper, which compared two NSTX SOL simulations in simple helical geometry, one with neutrals and one without. The former included electron-impact ionization, charge exchange, and wall recycling. Here, the case with neutrals is compared to a gyrokinetic-only simulation that includes an effective ionization source to separate the effect of sourcing from charge exchange collisions. It is observed that sourcing accounts for many features of the simulated SOL with neutrals, including density and temperature magnitudes and reduced normalized density fluctuations, but differences persist. In particular, a flatter density profile results due to changes in parallel transport when neutral collisions are included, illustrating the importance of neutral drag on global plasma properties. An analysis of coherent turbulent structures, or blobs, in these simulations demonstrates the case with neutrals has slower and larger blobs. Here, a series of seeded blob simulations corroborates the blob velocity observation. In general, the blob motion does not contribute significantly to radial transport in these simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Large Eddy Simulations of Turbulence below Antarctic Ice Shelves

Turbulence below ice shelves is of key importance for predicting ice-shelf melt rates and consequently the contribution of ice sheets to sea level rise. In this project we conducted Large-Eddy Simulations (LES) to improve our understanding of sub-ice-shelf ocean turbulence and the relationship between ice-shelf melt rates and ocean conditions. Over the course of the second and final year of this project (FY2020), we accomplished both major code developments for the PArallel Large-eddy simulation Model (PALM; Maronga et al., 2015) and conducted a suite of simulations which form the basis for a publication in preparation.

58 GEOSCIENCES↗

A high-order language for a system of closely coupled processing elements

The research reported in this paper was occasioned by the requirements on part of the Real-Time Digital Simulator (RTDS) project under way at NASA Lewis Research Center. The RTDS simulation scheme employs a network of CPUs running lock-step cycles in the parallel computations of jet airplane simulations. Their need for a high order language (HOL) that would allow non-experts to write simulation applications and that could be implemented on a possibly varying network can best be fulfilled by using the programming language Ada. We describe how the simulation problems can be modeled in Ada, how to map a single, multi-processing Ada program into code for individual processors, regardless of network reconfiguration, and why some Ada language features are particulary well-suited to network simulations.

Feyock, S.↗

Parallel transport dynamics for mixed quantum states with applications to time-dependent density functional theory

Direct simulation of the von Neumann dynamics for a general (pure or mixed) quantum state can often be expensive. One prominent example is the real-time time-dependent density functional theory (rt-TDDFT), a widely used framework for the first principle description of many-electron dynamics in chemical and materials systems. Practical rt-TDDFT calculations often avoid the direct simulation of the von Neumann equation, and solve instead a set of Schrödinger equations, of which the dynamics is equivalent to that of the von Neumann equation. However, the time step size employed by the Schrödinger dynamics is often much smaller. Here, in order to improve the time step size and the overall efficiency of the simulation, we generalize a recent work of the parallel transport (PT) dynamics for simulating pure states [An, Lin, Multiscale Model. Simul. 18, 612, 2020] to general quantum states. The PT dynamics provides the optimal gauge choice, and can employ a time step size comparable to that of the von Neumann dynamics. Going beyond the linear and near adiabatic regime in previous studies, we find that the error of the PT dynamics can be bounded by certain commutators between Hamiltonians, density matrices, and their derived quantities. Such a commutator structure is not present in the Schrödinger dynamics. We demonstrate that the parallel transport-implicit midpoint (PT-IM) method is a suitable method for simulating the PT dynamics, especially when the spectral radius of the Hamiltonian is large. The commutator structure of the error bound, and numerical results for model rt-TDDFT calculations in both linear and nonlinear regimes, confirm the advantage of the PT dynamics.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Parallel DSMC Solution of Three-Dimensional Flow Over a Finite Flat Plate

This paper describes a parallel implementation of the direct simulation Monte Carlo (DSMC) method. Runtime library support is used for scheduling and execution of communication between nodes, and domain decomposition is performed dynamically to maintain a good load balance. Performance tests are conducted using the code to evaluate various remapping and remapping-interval policies, and it is shown that a one-dimensional chain-partitioning method works best for the problems considered. The parallel code is then used to simulate the Mach 20 nitrogen flow over a finite-thickness flat plate. It is shown that the parallel algorithm produces results which compare well with experimental data. Moreover, it yields significantly faster execution times than the scalar code, as well as very good load-balance characteristics.

Nance, Robert P.↗

Microprocessor arrays for large scale computation

An important new direction in computer architecture centers around the achievement of very high computational power (capacity, speed and reliability) through the use of tens of thousands of microprocessors, micromemories, and switch modules, all interconnected into a large homogeneous network using one of certain advanced connection schemes. When surrounded and supported by conventional computers and memories, such a machine holds potential for out-performing both conventional and array-based computers of the mid-1980's by one to two orders of magnitude, at least for particular classes of applications amenable to high parallelism, such as aerodynamic simulation. The homogeneous feature of this machine concept also implies size extendability, fault tolerance, and improved flexibility to handle a variety of algorithms of interest. Current work is addressing the design of technologically efficient interconnection configurations and the development of new computation algorithms that are especially efficient for highly parallel computation.

Kautz, W. H.↗

Lowering entry barriers to developing custom simulators of distributed applications and platforms with SimGrid

Researchers in parallel and distributed computing (PDC) often resort to simulation because experiments conducted using a simulator can be for arbitrary experimental scenarios, are less resource-, labor-, and time-consuming than their real-world counterparts, and are perfectly repeatable and observable. Many frameworks have been developed to ease the development of PDC simulators, and these frameworks provide different levels of accuracy, scalability, versatility, extensibility, and usability. Further, the SimGrid framework has been used by many PDC researchers to produce a wide range of simulators for over two decades. Its popularity is due to a large emphasis placed on accuracy, scalability, and versatility, and is in spite of shortcomings in terms of extensibility and usability. Although SimGrid provides sensible simulation models for the common case, it was difficult for users to extend these models to meet domain-specific needs. Furthermore, SimGrid only provided relatively low-level simulation abstractions, making the implementation of a simulator of a complex system a labor-intensive undertaking. In this work we describe developments in the last decade that have contributed to vastly improving extensibility and usability, thus lowering or removing entry barriers for users to develop custom SimGrid simulators.

97 MATHEMATICS AND COMPUTING↗