Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Efficient, massively parallel eigenvalue computation

In numerical simulations of disordered electronic systems, one of the most common approaches is to diagonalize random Hamiltonian matrices and to study the eigenvalues and eigenfunctions of a single electron in the presence of a random potential. An effort to implement a matrix diagonalization routine for real symmetric dense matrices on massively parallel SIMD computers, the Maspar MP-1 and MP-2 systems, is described. Results of numerical tests and timings are also presented.

Huo, Yan↗

Simulation of Unsteady Combustion in a Ramjet Engine Using a Highly Parallel Computer

Combustion instability in ramjets is a complex phenomenon that involve nonlinear interaction between acoustic waves, vortex motion and unsteady heat release in the combustor. To numerically simulate this 3-D, transient phenomenon, enormous computer resources (time, memory and disk storage) are required. Although current generation vector supercomputers are capable of providing adequate resources for simulations of this nature, their high cost and limited availability, makes such machines less than satisfactory for routine use. The primary focus of this study is to assess the feasibility of using highly parallel computer systems as a cost-effective alternative for conducting such unsteady flow simulations. Towards this end, a large-eddy simulation model for combustion instability was implemented on the Intel iPSC/860 and a careful study was conducted to determine the benefits and the problems associated with the use of such machines for transient simulations. Details of this study along with the results obtained from the unsteady combustion simulations carried out on the iPSC/860 are discussed in this paper.

Menon, Suresh↗

Architecture for Co-Simulation of Transportation and Distribution Systems with Electric Vehicle Charging at Scale in the San Francisco Bay Area

This work describes the Grid-Enhanced, Mobility-Integrated Network Infrastructures for Extreme Fast Charging (GEMINI) architecture for the co-simulation of distribution and transportation systems to evaluate EV charging impacts on electric distribution systems of a large metropolitan area and the surrounding rural regions with high fidelity. The current co-simulation is applied to Oakland and Alameda, California, and in future work will be extended to the full San Francisco Bay Area. It uses the HELICS co-simulation framework to enable parallel instances of vetted grid and transportation software programs to interact at every model timestep, allowing high-fidelity simulations at a large scale. This enables not only the impacts of electrified transportation systems across a larger interconnected collection of distribution feeders to be evaluated, but also the feedbacks between the two systems, such as through control systems, to be captured and compared. The findings are that with moderate passenger EV adoption rates, inverter controls combined with some distribution system hardware upgrades can maintain grid voltages within ANSI C.84 range A limits of 0.95 to 1.05 p.u. without smart charging. However, EV charging control may be required for higher levels of charging or to reduce grid upgrades, and this will be explored in future work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Parallel processor engine model program

The Parallel Processor Engine Model Program is a generalized engineering tool intended to aid in the design of parallel processing real-time simulations of turbofan engines. It is written in the FORTRAN programming language and executes as a subset of the SOAPP simulation system. Input/output and execution control are provided by SOAPP; however, the analysis, emulation and simulation functions are completely self-contained. A framework in which a wide variety of parallel processing architectures could be evaluated and tools with which the parallel implementation of a real-time simulation technique could be assessed are provided.

Mclaughlin, P.↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Parallel Randomized Tucker Decomposition Algorithms

The Tucker tensor decomposition is a natural extension of the singular value decomposition (SVD) to multiway data. Here, we propose to accelerate Tucker tensor decomposition algorithms by using randomization and parallelization. We present two algorithms that scale to large data and many processors, significantly reduce both computation and communication cost compared to previous deterministic and randomized approaches, and obtain nearly the same approximation errors. The key idea in our algorithms is to perform randomized sketches with Kronecker-structured random matrices, which reduces computation compared to unstructured matrices and can be implemented using a fundamental tensor computational kernel. We provide probabilistic error analysis of our algorithms and implement a new parallel algorithm for the structured randomized sketch. Our experimental results demonstrate that our combination of randomization and parallelization achieves accurate Tucker decompositions much faster than alternative approaches. We observe up to a 16X speedup over the fastest deterministic parallel implementation on 3D simulation data.

Tucker decompositions↗

On the suitability of the connection machine for direct particle simulation

The algorithmic structure was examined of the vectorizable Stanford particle simulation (SPS) method and the structure is reformulated in data parallel form. Some of the SPS algorithms can be directly translated to data parallel, but several of the vectorizable algorithms have no direct data parallel equivalent. This requires the development of new, strictly data parallel algorithms. In particular, a new sorting algorithm is developed to identify collision candidates in the simulation and a master/slave algorithm is developed to minimize communication cost in large table look up. Validation of the method is undertaken through test calculations for thermal relaxation of a gas, shock wave profiles, and shock reflection from a stationary wall. A qualitative measure is provided of the performance of the Connection Machine for direct particle simulation. The massively parallel architecture of the Connection Machine is found quite suitable for this type of calculation. However, there are difficulties in taking full advantage of this architecture because of lack of a broad based tradition of data parallel programming. An important outcome of this work has been new data parallel algorithms specifically of use for direct particle simulation but which also expand the data parallel diction.

Dagum, Leonard↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

NESSi : The N on- E quilibrium S ystems S imulation package

The nonequilibrium dynamics of correlated many-particle systems is of interest in connection with pump–probe experiments on molecular systems and solids, as well as theoretical investigations of transport properties and relaxation processes. Nonequilibrium Green’s functions are a powerful tool to study interaction effects in quantum many-particle systems out of equilibrium, and to extract physically relevant information for the interpretation of experiments. Here, we present the open-source software package NESSi (The Non-Equilibrium Systems Simulation package) which allows to perform many-body dynamics simulations based on Green’s functions on the L-shaped Kadanoff–Baym contour. NESSi contains the library libcntr which implements tools for basic operations on these nonequilibrium Green’s functions, for constructing Feynman diagrams, and for the solution of integral and integro-differential equations involving contour Green’s functions. The library employs a discretization of the Kadanoff–Baym contour into time points and a high-order implementation of integration routines. The total integrated error scales up to $\mathcal{O}(N^{-7})$, which is important since the numerical effort increases at least cubically with the simulation time. A distributed-memory parallelization over reciprocal space allows large-scale simulations of lattice systems. We provide a collection of example programs ranging from dynamics in simple two-level systems to problems relevant in contemporary condensed matter physics, including Hubbard clusters and Hubbard or Holstein lattice models. The libcntr library is the basis of a follow-up software package for nonequilibrium dynamical mean-field theory calculations based on strong-coupling perturbative impurity solvers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Diodelike response of high-latitude plasma in magnetosphere-ionosphere coupling in the presence of field-aligned currents

The dynamic processes in the plasma along high-latitude field lines plays an important role in ionosphere-magnetosphere coupling process. A time-dependent, large-scale simulation of these dynamics parallel to the geomagnetic field lines from the ionosphere well into the magnetosphere is created. The plasma consists of hot e(-) and H(+) of magnetospheric origin and low-energy e(-), H(+), and O(+) of ionospheric origin. Including multiple electron species, a major improvement to the model, made it possible for the first time to simulate the upward current region properly and to dynamically simulate the diodelike response of the field-line plasma to the parallel currents coupling the ionosphere and magnetosphere. It is shown that return currents flow with small resistance, while upward currents produce kilovolt-sized potential drops along the field, as concluded from satellite observations. The kilovolt potential drops are due to the effect of the converging magnetic field on the high-energy magnetospheric electrons.

Mitchell, H. G., Jr.↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗

Turbomachinery CFD on parallel computers

The role of multistage turbomachinery simulation in the development of propulsion system models is discussed. Particularly, the need for simulations with higher fidelity and faster turnaround time is highlighted. It is shown how such fast simulations can be used in engineering-oriented environments. The use of parallel processing to achieve the required turnaround times is discussed. Current work by several researchers in this area is summarized. Parallel turbomachinery CFD research at the NASA Lewis Research Center is then highlighted. These efforts are focused on implementing the average-passage turbomachinery model on MIMD, distributed memory parallel computers. Performance results are given for inviscid, single blade row and viscous, multistage applications on several parallel computers, including networked workstations.

Blech, Richard A.↗

Turbomachinery CFD on parallel computers

The role of multistage turbomachinery simulation in the development of propulsion system models is discussed. Particularly, the need for simulations with higher fidelity and faster turnaround time is highlighted. It is shown how such fast simulations can be used in engineering-oriented environments. The use of parallel processing to achieve the required turnaround times is discussed. Current work by several researchers in this area is summarized. Parallel turbomachinery CFD research at the NASA Lewis Research Center is then highlighted. These efforts are focused on implementing the average-passage turbomachinery model on MIMD, distributed memory parallel computers. Performance results are given for inviscid, single blade row and viscous, multistage applications on several parallel computers, including networked workstations.

Blech, R. A.↗

Polarized fusion and potential in situ tests of fuel polarization survival in a tokamak plasma

Abstract The use of spin-polarized fusion fuels would provide a significant boost towards the ignition of a burning plasma. The cross section for D + T → α + n, would be increased by 1.5 if the fuels were injected with parallel polarization. Furthermore, our simulations demonstrate additional non-linear power gains in large-scale machines such as ITER, due to increased alpha heating. Such benefits require the survival of spin polarizations for periods comparable to the particle confinement time. During the 1980s, calculations predicted that polarizations could survive a plasma environment, although concerns persisted regarding the cumulative impacts of wall recycling. In that era, technical challenges prevented direct tests and left the large scale fueling of a power reactor beyond reach. Over the last decades, this situation has changed dramatically. Detailed simulations of ITER have predicted negligible wall recycling in a high-power reactor, and recent advances in laser-driven sources project the capability of producing large quantities of ∼100% polarized D and T. The remaining crucial step is an in-situ demonstration of polarization survival in a plasma. For this, we outline a measurement strategy using the isospin-mirror reaction, D + 3 He → α + p. Polarized 3 He avoids the complexities of handling tritium, while encompassing the same spin-physics. We evaluate two methods of delivering deuterium, using dynamically polarized Lithium-Deuteride (with vector polarization P V D of 70%) or frozen-spin Hydrogen-Deuteride (with P V D of 40%), together with a method of injecting optically-pumped 3 He (with 65% polarization). Pellets of these materials all have long polarization decay times (∼6 min for LiD at 2 K, ∼2 months for HD at 2 K, and ∼3 d for 3 He at 77 K), all far greater than a plasma shot in a research tokamak such as DIII-D (∼20 s). Both species can be propelled from a single cryogenic injection gun. We review plasma requirements and strategies for detecting polarization survival. Polarization alters both fusion yields and the angular distribution of fusion products, and each of these provides a potential signal. In this paper we simulate a selection of shots with similar characteristics in a future high-T ion H plasma, and find ratios of yields from shots with fuel spins parallel and antiparallel reaching 1.3 (HD + 3 He) to 1.6 (LiD + 3 He) over a wide range of poloidal angles. (A companion paper finds sensitivity to fusion product angular distributions as reflected in the pitch angles of protons and alphas reaching the plasma facing wall.)

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predictive Models for Semiconductor Device Design and Processing

The device feature size continues to be on a downward trend with a simultaneous upward trend in wafer size to 300 mm. Predictive models are needed more than ever before for this reason. At NASA Ames, a Device and Process Modeling effort has been initiated recently with a view to address these issues. Our activities cover sub-micron device physics, process and equipment modeling, computational chemistry and material science. This talk would outline these efforts and emphasize the interaction among various components. The device physics component is largely based on integrating quantum effects into device simulators. We have two parallel efforts, one based on a quantum mechanics approach and the second, a semiclassical hydrodynamics approach with quantum correction terms. Under the first approach, three different quantum simulators are being developed and compared: a nonequlibrium Green's function (NEGF) approach, Wigner function approach, and a density matrix approach. In this talk, results using various codes will be presented. Our process modeling work focuses primarily on epitaxy and etching using first-principles models coupling reactor level and wafer level features. For the latter, we are using a novel approach based on Level Set theory. Sample results from this effort will also be presented.

Meyyappan, Meyya↗

Investigation of a Macromechanical Approach to Analyzing Triaxially-Braided Polymer Composites

A macro level finite element-based model has been developed to simulate the mechanical and impact response of triaxially-braided polymer matrix composites. In the analytical model, the triaxial braid architecture is simulated by using four parallel shell elements, each of which is modeled as a laminated composite. The commercial transient dynamic finite element code LS-DYNA is used to conduct the simulations, and a continuum damage mechanics model internal to LS-DYNA is used as the material constitutive model. The material stiffness and strength values required for the constitutive model are determined based on coupon level tests on the braided composite. Simulations of quasi-static coupon tests of a representative braided composite are conducted. Varying the strength values that are input to the material model is found to have a significant influence on the effective material response predicted by the finite element analysis, sometimes in ways that at first glance appear non-intuitive. A parametric study involving the input strength parameters provides guidance on how the analysis model can be improved.

Goldberg, Robert K.↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

Flow of a Gas Turbine Engine Low-Pressure Subsystem Simulated

The NASA Lewis Research Center is managing a task to numerically simulate overnight, on a parallel computing testbed, the aerodynamic flow in the complete low-pressure subsystem (LPS) of a gas turbine engine. The model solves the three-dimensional Navier- Stokes flow equations through all the components within the LPS, as well as the external flow around the engine nacelle. The LPS modeling task is being performed by Allison Engine Company under the Small Engine Technology contract. The large computer simulation was evaluated on networked computer systems using 8, 16, and 32 processors, with the parallel computing efficiency reaching 75 percent when 16 processors were used.

Veres, Joseph P.↗