Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Introduction to a system for implementing neural net connections on SIMD architectures

Neural networks have attracted much interest recently, and using parallel architectures to simulate neural networks is a natural and necessary application. The SIMD model of parallel computation is chosen, because systems of this type can be built with large numbers of processing elements. However, such systems are not naturally suited to generalized communication. A method is proposed that allows an implementation of neural network connections on massively parallel SIMD architectures. The key to this system is an algorithm permitting the formation of arbitrary connections between the neurons. A feature is the ability to add new connections quickly. It also has error recovery ability and is robust over a variety of network topologies. Simulations of the general connection system, and its implementation on the Connection Machine, indicate that the time and space requirements are proportional to the product of the average number of connections per neuron and the diameter of the interconnection network.

Tomboulian, Sherryl↗

Introduction to a system for implementing neural net connections on SIMD architectures

Neural networks have attracted much interest recently, and using parallel architectures to simulate neural networks is a natural and necessary application. The SIMD model of parallel computation is chosen, because systems of this type can be built with large numbers of processing elements. However, such systems are not naturally suited to generalized elements. A method is proposed that allows an implementation of neural network connections on massively parallel SIMD architectures. The key to this system is an algorithm permitting the formation of arbitrary connections between the neurons. A feature is the ability to add new connections quickly. It also has error recovery ability and is robust over a variety of network topologies. Simulations of the general connection system, and its implementation on the Connection Machine, indicate that the time and space requirements are proportional to the product of the average number of connections per neuron and the diameter of the interconnection network.

Tomboulian, Sherryl↗

Parallelization of a Six Degree of Freedom Entry Vehicle Trajectory Simulation Using OpenMP and OpenACC

The art and science of writing parallelized software, using methods such as Open Multi-Processing (OpenMP) and Open Accelerators (OpenACC), is dominated by computer scientists. Engineers and non-computer scientists looking to apply these techniques to their project applications face a steep learning curve, especially when looking to adapt their original single threaded software to run multi-threaded on graphics processing units (GPUs). There are significant changes in mindset that must occur; such as how to manage memory, the organization of instructions, and the use of if statements (also known as branching). The purpose of this work is twofold: 1) to demonstrate the applicability of parallelized coding methodologies, OpenMP and OpenACC, to tasks outside of the typical large scale matrix mathematics; and 2) to discuss, from an engineer’s perspective, the lessons learned from parallelizing software using these computer science techniques. This work applies OpenMP, on both multi-core central processing units (CPUs) and Intel® Xeon Phi™ 7210, and OpenACC on GPUs. These parallelization techniques are used to tackle the simulation of thousands of entry vehicle trajectories through the integration of six degree of freedom (DoF) equations of motion (EoM). The forces and moments acting on the entry vehicle, and used by the EoM, are estimated using multiple models of varying levels of complexity. Several benchmark comparisons are made on the execution of six DoF trajectory simulation: single thread Intel® Xeon® E5-2670 CPU, multi-thread CPU using OpenMP, multi-thread Xeon Phi™ 7210 using OpenMP, and multi-thread NVIDIA® Tesla® K40 GPU using OpenACC. These benchmarks are run on the Pleiades Supercomputer Cluster at the National Aeronautics and Space Administration (NASA) Ames Research Center (ARC), and a Xeon Phi™ 7210 node at NASA Langley Research Center (LaRC).

Green, Justin S.↗

Simulation of electrostatic turbulence due to sheared flows parallel and transverse to the magnetic field

A spatially two-dimensional electrostatic particle simulation code is used to examine the stability of a plasma equilibrium characterized by a localized transverse dc electric field and a magnetic-field-aligned electron drift for L much less than Lx, where Lx is the simulation length in the x direction and L is the scale length associated with the dc electric field. It is found that the dc electric field and the field-aligned current can together play a synergistic role to enable the excitation of electrostatic waves even when the threshold values of the field-aligned drift and the E x B drift are individually subcritical. The simulation results indicate that a broadband turbulence is associated with such an equilibrium.

Nishikawa, K.-I.↗

A real-time blade element helicopter simulation for handling qualities analysis

A simulation model which utilizes parallel processing platforms is described in terms of its contributions to improved real-time helicopter simulation. The FLIGHTLAB parallel processing environment is explained, and the relative advantages of the blade element and rotor map models for rigid and elastic articulated blades are discussed. A UH-60 simulation is conducted by means of a rigid model with 14 degrees of freedom, as well as an elastic model with 26 degrees of freedom, to compare trim conditions, longitudinal static margins, and longitudinal and lateral frequency responses. The FLIGHTLAB system is shown to facilitate restructuring for parallel processing as well as the systematic comparison of a variety of models. The system can facilitate the comparison of rigid and elastic blade element rotor models at NASA-Ames and other research facilities.

Du Val, Ronald W.↗

Computational Investigation of Retropropulsion Operating Environments with a GPU-Enabled Detached Eddy Simulation Approach

Human exploration of the surface of Mars will require an extended powered descent phase of flight, during which aerodynamic-propulsive interference effects can be significant. Characterization of these environments to enable implementation of this technology into a flight vehicle will rely heavily on computational simulation. This work advances the understanding of retropropulsion aerodynamics through application of a massively parallel detached eddy simulation approach on a GPU-accelerated computational framework, yielding data that are largely unachievable with conventional high-performance computing resources. This work includes time-dependent and time-averaged forces and moments on a conceptual, full-scale vehicle in environments and operating conditions relevant to human Mars exploration. Conditions are examined where the engine exhaust flow transitions between over-expanded and under-expanded flow structures, and flight operation will require the ability to maintain control of the vehicle during such a transition. These transitions occur as the vehicle decelerates, and as such, this investigation includes supersonic, transonic, and subsonic flight conditions. Options for vehicle control during powered flight include differential throttling of the engines. This paper provides an overview of the computational campaign, approach, and discussion of results in characterizing the resulting aerodynamics for differential throttling with retropropulsion in atmospheric environments.

EDL↗

Pre-Stall Behavior of a Transonic Axial Compressor Stage via Time-Accurate Numerical Simulation

CFD calculations using high-performance parallel computing were conducted to simulate the pre-stall flow of a transonic compressor stage, NASA compressor Stage 35. The simulations were run with a full-annulus grid that models the 3D, viscous, unsteady blade row interaction without the need for an artificial inlet distortion to induce stall. The simulation demonstrates the development of the rotating stall from the growth of instabilities. Pressure-rise performance and pressure traces are compared with published experimental data before the study of flow evolution prior to the rotating stall. Spatial FFT analysis of the flow indicates a rotating long-length disturbance of one rotor circumference, which is followed by a spike-type breakdown. The analysis also links the long-length wave disturbance with the initiation of the spike inception. The spike instabilities occur when the trajectory of the tip clearance flow becomes perpendicular to the axial direction. When approaching stall, the passage shock changes from a single oblique shock to a dual-shock, which distorts the perpendicular trajectory of the tip clearance vortex but shows no evidence of flow separation that may contribute to stall.

Chen, Jen-Ping↗

Developing parallel GeoFEST(P) using the PYRAMID AMR library

The PYRAMID parallel unstructured adaptive mesh refinement (AMR) library has been coupled with the GeoFEST geophysical finite element simulation tool to support parallel active tectonics simulations. Specifically, we have demonstrated modeling of coseismic and postseismic surface displacement due to a simulated Earthquake for the Landers system of interacting faults in Southern California. The new software demonstrated a 25-times resolution improvement and a 4-times reduction in time to solution over the sequential baseline milestone case. Simulations on workstations using a few tens of thousands of stress displacement finite elements can now be expanded to multiple millions of elements with greater than 98% scaled efficiency on various parallel platforms over many hundreds of processors. Our most recent work has demonstrated that we can dynamically adapt the computational grid as stress grows on a fault. In this paper, we will describe the major issues and challenges associated with coupling these two programs to create GeoFEST(P). Performance and visualization results will also be described.

adaptive mesh refinement (AMR)↗

Collecting address traces from parallel computers

Trace driven simulation is a well-established method of performance analysis for single processor computer systems. However, efficient and accurate memory address tracing for parallel computer systems is not well understood. In this paper we present a critical survey of recently implemented approaches to address tracing and highlight the issues specific to collection of traces for both shared and distributed memory parallel computers. These issues include potential distortion of the relative ordering of events by the address tracing activity, realistic interleaving of addresses generated by multiple processors, and I/O and storage problems associated with collecting traces for large parallel systems. The strengths and weaknesses of the parallel tracing approaches are described.

Stunkel, Craig B.↗

Parallel Multiscale Algorithms for Astrophysical Fluid Dynamics Simulations

Our goal is to develop software libraries and applications for astrophysical fluid dynamics simulations in multidimensions that will enable us to resolve the large spatial and temporal variations that inevitably arise due to gravity, fronts and microphysical phenomena. The software must run efficiently on parallel computers and be general enough to allow the incorporation of a wide variety of physics. Cosmological structure formation with realistic gas physics is the primary application driver in this work. Accurate simulations of e.g. galaxy formation require a spatial dynamic range (i.e., ratio of system scale to smallest resolved feature) of 104 or more in three dimensions in arbitrary topologies. We take this as our technical requirement. We have achieved, and in fact, surpassed these goals.

Norman, Michael L.↗

Advances in Parallelization for Large Scale Oct-Tree Mesh Generation

Despite great advancements in the parallelization of numerical simulation codes over the last 20 years, it is still common to perform grid generation in serial. Generating large scale grids in serial often requires using special "grid generation" compute machines that can have more than ten times the memory of average machines. While some parallel mesh generation techniques have been proposed, generating very large meshes for LES or aeroacoustic simulations is still a challenging problem. An automated method for the parallel generation of very large scale off-body hierarchical meshes is presented here. This work enables large scale parallel generation of off-body meshes by using a novel combination of parallel grid generation techniques and a hybrid "top down" and "bottom up" oct-tree method. Meshes are generated using hardware commonly found in parallel compute clusters. The capability to generate very large meshes is demonstrated by the generation of off-body meshes surrounding complex aerospace geometries. Results are shown including a one billion cell mesh generated around a Predator Unmanned Aerial Vehicle geometry, which was generated on 64 processors in under 45 minutes.

O'Connell, Matthew↗

Simulation of electrostatic ion instabilities in the presence of parallel currents and transverse electric fields

A spatially two-dimensional electrostatic PIC simulation code was used to study the stability of a plasma equilibrium characterized by a localized transverse dc electric field and a field-aligned drift for L is much less than Lx, where Lx is the simulation length in the x direction and L is the scale length associated with the dc electric field. It is found that the dc electric field and the field-aligned current can together play a synergistic role to enable the excitation of electrostatic waves even when the threshold values of the field aligned drift and the E x B drift are individually subcritical. The simulation results show that the growing ion waves are associated with small vortices in the linear stage, which evolve to the nonlinear stage dominated by larger vortices with lower frequencies.

Nishikawa, K.-I.↗

Methodology of modeling and measuring computer architectures for plasma simulations

A brief introduction to plasma simulation using computers and the difficulties on currently available computers is given. Through the use of an analyzing and measuring methodology - SARA, the control flow and data flow of a particle simulation model REM2-1/2D are exemplified. After recursive refinements the total execution time may be greatly shortened and a fully parallel data flow can be obtained. From this data flow, a matched computer architecture or organization could be configured to achieve the computation bound of an application problem. A sequential type simulation model, an array/pipeline type simulation model, and a fully parallel simulation model of a code REM2-1/2D are proposed and analyzed. This methodology can be applied to other application problems which have implicitly parallel nature.

Wang, L. P. T.↗

A robot arm simulation with a shared memory multiprocessor machine

A parallel processing scheme for a single chain robot arm is presented for high speed computation on a shared memory multiprocessor. A recursive formulation that is derived from a virtual work form of the d'Alembert equations of motion is utilized for robot arm dynamics. A joint drive system that consists of a motor rotor and gears is included in the arm dynamics model, in order to take into account gyroscopic effects due to the spinning of the rotor. The fine grain parallelism of mechanical and control subsystem models is exploited, based on independent computation associated with bodies, joint drive systems, and controllers. Efficiency and effectiveness of the parallel scheme are demonstrated through simulations of a telerobotic manipulator arm. Two different mechanical subsystem models, i.e., with and without gyroscopic effects, are compared, to show the trade-off between efficiency and accuracy.

Kim, Sung-Soo↗

Re-forming supercritical quasi-parallel shocks. II - Mechanism for wave generation and front re-formation

This paper continues the study of Thomas et al. (1990) in which hybrid simulations of quasi-parallel shocks were performed in one and two spatial dimensions. To identify the wave generation processes, the electromagnetic structure of the shock is examined by performing a number of one-dimensional hybrid simulations of quasi-parallel shocks for various upstream conditions. In addition, numerical experiments were carried out in which the backstreaming ions were removed from calculations to show their fundamental importance in reformation process. The calculations show that the waves are excited before ions can propagate far enough upstream to generate resonant modes. At some later times, the waves are regenerated at the leading edge of the interface, with properties like those of their initial interactions.

Winske, D.↗

The Cirrus Parcel Model Comparison Project

The cirrus Parcel Model Comparison Project involves the systematic comparison of current models of ice crystal nucleation and growth for specified, typical, cirrus cloud environments. In Phase 1 of the project reported here, simulated cirrus cloud microphysical properties are compared for situations of "warm" (-40 C) and "cold" (-60 C) cirrus subject to updrafts of 4, 20 and 100 centimeters per second, respectively. Five models are participating in the project. These models employ explicit microphysical schemes wherein the size distribution of each class of particles (aerosols and ice crystals) is resolved into bins. Simulations are made including both homogeneous and heterogeneous ice nucleation mechanisms. A single initial aerosol population of sulfuric acid particles is prescribed for all simulations. To isolate the treatment of the homogeneous freezing (of haze drops) nucleation process, the heterogeneous nucleation mechanism is disabled for a second parallel set of simulations. Qualitative agreement is found amongst the models for the homogeneous-nucleation-only simulations, e.g., the number density of nucleated ice crystals increases with the strength of the prescribed updraft. However, non-negligible quantitative differences are found. Systematic bias exists between results of a model based on a modified classical theory approach and models using an effective freezing temperature approach to the treatment of nucleation. Each approach is constrained by critical freezing data from laboratory studies. This information is necessary, but not sufficient, to construct consistent formulae for the two approaches. Large haze particles may deviate considerably from equilibrium size in moderate to strong updrafts (20-100 centimeters per second) at -60 C when the commonly invoked equilibrium assumption is lifted. The resulting difference in particle-size-dependent solution concentration of haze particles may significantly affect the ice nucleation rate during the initial nucleation interval. The uptake rate for water vapor excess by ice crystals is another key component regulating the total number of nucleated ice crystals. This rate, the product of ice number concentration and ice crystal diffusional growth rate, partially controls the peak nucleation rate achieved in an air parcel and the duration of the active nucleation time period.

Lin, Ruei-Fong↗

GCSS Cirrus Parcel Model Comparison Project

The Cirrus Parcel Model Comparison Project, a project of GCSS Working Group on Cirrus Cloud Systems (WG2), involves the systematic comparison of current models of ice crystal nucleation and growth for specified, typical, cirrus cloud environments. The goal of this project is to document and understand the factors resulting in significant inter-model differences. The intent is to foment research leading to model improvement and validation. In Phase 1 of the project reported here, simulated cirrus cloud microphysical properties are compared for situations of "warm" (-40 C) and "cold" (-60 C) cirrus subject to updrafts of 4, 20 and 100 cm/s, respectively. Five models participated. These models employ explicit microphysical schemes wherein the size distribution of each class of particles (aerosols and ice crystals) is resolved into bins. Simulations are made including both homogeneous and heterogeneous ice nucleation mechanisms. A single initial aerosol population of sulfuric acid particles is prescribed for all simulations. To isolate the treatment of the homogeneous freezing (of haze drops) nucleation process, the heterogeneous nucleation mechanism is disabled for a second parallel set of simulations. Qualitative agreement is found for the homogeneous-nucleation-only simulations, e.g., the number density of nucleated ice crystals increases with the strength of the prescribed updraft. However, non-negligible quantitative differences are found. Detailed analysis reveals that the homogeneous nucleation formulation, aerosol size, ice crystal growth rate (particularly the deposition coefficient), and water vapor uptake rate are critical components that lead to differences in predicted microphysics. Systematic bias exists between results based on a modified classical theory approach and models using an effective freezing temperature approach to the treatment of nucleation. Each approach is constrained by critical freezing data from laboratory studies, but each includes assumptions that can only be justified by further laboratory data. Consequently, it is not yet clear if the two approaches can be made consistent. Large haze particles may deviate considerably from equilibrium size in moderate to strong updrafts (20-100 cm/s) at -60 C when the commonly invoked equilibrium assumption is lifted. The resulting difference in particle-size-dependent solution concentration of haze particles may significantly affect the ice nucleation rate during the initial nucleation interval. The uptake rate for water vapor excess by ice crystals is another key component regulating the total number of nucleated ice crystals. This rate, the product of ice number concentration and ice crystal diffusional growth rate, which is sensitive to the deposition coefficient when ice particles are small, partially controls the peak nucleation rate achieved in an air parcel and the duration of the active nucleation time period. The effects of heterogeneous nucleation are most pronounced in weak updraft situations. Vapor competition by the nucleated (heterogeneous) ice crystals limits the achieved ice supersaturation and thus suppresses the contribution of homogeneous nucleation. Correspondingly, ice crystal number density is markedly reduced. Definitive laboratory and atmospheric benchmark data are needed for the heterogeneous nucleation process. Inter-model differences are correspondingly greater than in the case of the homogeneous nucleation process acting alone.

Lin, Ruei-Fong↗