Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Automatic Multilevel Parallelization Using OpenMP

In this paper we describe the extension of the CAPO parallelization support tool to support multilevel parallelism based on OpenMP directives. CAPO generates OpenMP directives with extensions supported by the NanosCompiler to allow for directive nesting and definition of thread groups. We report first results for several benchmark codes and one full application that have been parallelized using our system.

Jin, Hao-Qiang↗

TReactMech v4.217

TReactMech couples geomechanical processes (poroelasticity, failure, and inelastic strain) with multiphase nonisothermal flow (derived from TOUGH2) and reactive geochemical transport. At its core is the reactive-transport code TOUGHREACT v4.13. TReactMech is an efficient hybrid parallel simulator, solving the geomechanics using finite elements and MPI/PetSc, the multiphase flow using integrated finite difference and MPI/PETSc, and the reactive chemistry using OpenMP. The advantages of TReactMech are in its multiphase flow capabilities (e.g., supercritical CO2, supercritical water, air) and parallel geomechanics including full 3-D stress tensor, shear and tensile failure, coupled to porosity and permeability changes. It is backwardly compatible with TOUGH2 and TOUGHREACT v4.13, allowing for easier transitions between the codes. TReactMech can be used to simulate many natural and engineered subsurface systems, including geothermal reservoirs, borehole heat exchangers, geologic carbon sequestration, geologic storage of nuclear waste, groundwater resources, weathering, sediment diagenesis, seafloor hydrothermal circulation, hydrofracturing in unconventional reservoirs, and injection/production-induced surface deformation.

Sonnenthal, Eric↗

QuakeSim and the Solid Earth Research Virtual Observatory

We are developing simulation and analysis tools in order to develop a solid Earth science framework for understanding and studying active tectonic and earthquake processes. The goal of QuakeSim and its extension, the Solid Earth Research Virtual Observatory (SERVO), is to study the physics of earthquakes using state-of-the-art modeling, data manipulation, and pattern recognition technologies. We are developing clearly defined accessible data formats and code protocols as inputs to simulations, which are adapted to high-performance computers. The solid Earth system is extremely complex and nonlinear resulting in computationally intensive problems with millions of unknowns. With these tools it will be possible to construct the more complex models and simulations necessary to develop hazard assessment systems critical for reducing future losses from major earthquakes. We are using Web (Grid) service technology to demonstrate the assimilation of multiple distributed data sources (a typical data grid problem) into a major parallel high-performance computing earthquake forecasting code. Such a linkage of Geoinformatics with Geocomplexity demonstrates the value of the Solid Earth Research Virtual Observatory (SERVO) Grid concept, and advances Grid technology by building the first real-time large-scale data assimilation grid.

virtual observatory↗

Hamming and Accumulator Codes Concatenated with MPSK or QAM

In a proposed coding-and-modulation scheme, a high-rate binary data stream would be processed as follows: 1. The input bit stream would be demultiplexed into multiple bit streams. 2. The multiple bit streams would be processed simultaneously into a high-rate outer Hamming code that would comprise multiple short constituent Hamming codes a distinct constituent Hamming code for each stream. 3. The streams would be interleaved. The interleaver would have a block structure that would facilitate parallelization for high-speed decoding. 4. The interleaved streams would be further processed simultaneously into an inner two-state, rate-1 accumulator code that would comprise multiple constituent accumulator codes - a distinct accumulator code for each stream. 5. The resulting bit streams would be mapped into symbols to be transmitted by use of a higher-order modulation - for example, M-ary phase-shift keying (MPSK) or quadrature amplitude modulation (QAM). The novelty of the scheme lies in the concatenation of the multiple-constituent Hamming and accumulator codes and the corresponding parallel architectures of the encoder and decoder circuitry (see figure) needed to process the multiple bit streams simultaneously. As in the cases of other parallel-processing schemes, one advantage of this scheme is that the overall data rate could be much greater than the data rate of each encoder and decoder stream and, hence, the encoder and decoder could handle data at an overall rate beyond the capability of the individual encoder and decoder circuits.

Divsalar, Dariush↗

Verification of electromagnetic simulation capabilities in global gyrokinetic particle-in-cell code GTS

Recently, the numerical scheme presented by Mishchenko et al. enabled explicit gyrokinetic simulations of low-frequency electromagnetic instabilities in tokamaks at experimentally relevant values of plasma β⁠. This scheme resolved the long-standing cancellation problem that previously hindered gyrokinetic particle-in-cell code simulations of magnetohydrodynamic phenomena with inherently small parallel electric fields. Moreover, the scheme did not employ approximations that eliminate critical tearing-type instabilities. Here, we report on the implementation of this numerical scheme in the global gyrokinetic particle-in-cell code GTS. This implementation allows for a more complete and accurate picture of interaction between small scale turbulence and MHD modes in tokamaks. Additionally, we present a comprehensive set of verification simulations of numerous electromagnetic instabilities relevant to present-day tokamaks. These simulations encompass the kinetic ballooning mode, the internal kink mode, the tearing mode, the micro-tearing mode, and the toroidal Alfven eigenmode destabilized by energetic ions, which are all instrumental in understanding tokamak physics. We will also showcase the preliminary nonlinear simulations of kinetic ballooning instabilities and (2,1) island formation due to tearing mode instability. These simulations validate the accuracy of the scheme implementation and pave the way for studying how these instabilities affect plasma confinement and performance.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Portability and Cross-Platform Performance of an MPI-Based Parallel Polygon Renderer

Visualizing the results of computations performed on large-scale parallel computers is a challenging problem, due to the size of the datasets involved. One approach is to perform the visualization and graphics operations in place, exploiting the available parallelism to obtain the necessary rendering performance. Over the past several years, we have been developing algorithms and software to support visualization applications on NASA's parallel supercomputers. Our results have been incorporated into a parallel polygon rendering system called PGL. PGL was initially developed on tightly-coupled distributed-memory message-passing systems, including Intel's iPSC/860 and Paragon, and IBM's SP2. Over the past year, we have ported it to a variety of additional platforms, including the HP Exemplar, SGI Origin2OOO, Cray T3E, and clusters of Sun workstations. In implementing PGL, we have had two primary goals: cross-platform portability and high performance. Portability is important because (1) our manpower resources are limited, making it difficult to develop and maintain multiple versions of the code, and (2) NASA's complement of parallel computing platforms is diverse and subject to frequent change. Performance is important in delivering adequate rendering rates for complex scenes and ensuring that parallel computing resources are used effectively. Unfortunately, these two goals are often at odds. In this paper we report on our experiences with portability and performance of the PGL polygon renderer across a range of parallel computing platforms.

Crockett, Thomas W.↗

Devastator Parallel Discrete Event Simulation Runtime (Devastator) v1.0

The Devastator runtime is a modern C++ implementation of optimistic parallel discrete event simulation methods. Devastator allows simulation application code to productively specify their component and event functionality with C++14 constructs. It utilizes GASNet-EX for distributed memory communication and includes parallel performance optimizations such as light-weight thread message queues and asynchronous GVT. Furthermore, it supports efficient event broadcasts and pause-rewind-resume functionality to support periodic load balancing and outer loop optimization algorithms.

Chan, Cy↗

Thermonuclear Burn in a Multiphysics Code on GPUs

Multiphysics codes links to a library called SINGE for the calculation of thermonuclear (TN) burn rates, but some current multiphysics codes do not attempt to leverage the support for parallel operation that SINGE provides. Our goal is to investigate implementations of the SINGE workflow and analyze how the use of a performance portability layer could reduce run time on CPU archi tectures while also supporting GPU architectures without requiring code modifications. We looked to the Kokkos C++ Performance Portability Ecosystem to implement hardware agnostic parallel patterns.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Automated Concurrent Blackboard System Generation in C++

In his 1992 Ph.D. thesis, "Design and Analysis Techniques for Concurrent Blackboard Systems", John McManus defined several performance metrics for concurrent blackboard systems and developed a suite of tools for creating and analyzing such systems. These tools allow a user to analyze a concurrent blackboard system design and predict the performance of the system before any code is written. The design can be modified until simulated performance is satisfactory. Then, the code generator can be invoked to generate automatically all of the code required for the concurrent blackboard system except for the code implementing the functionality of each knowledge source. We have completed the port of the source code generator and a simulator for a concurrent blackboard system. The source code generator generates the necessary C++ source code to implement the concurrent blackboard system using Parallel Virtual Machine (PVM) running on a heterogeneous network of UNIX(trademark) workstations. The concurrent blackboard simulator uses the blackboard specification file to predict the performance of the concurrent blackboard design. The only part of the source code for the concurrent blackboard system that the user must supply is the code implementing the functionality of the knowledge sources.

Kaplan, J. A.↗

SPEEDES - A multiple-synchronization environment for parallel discrete-event simulation

Synchronous Parallel Environment for Emulation and Discrete-Event Simulation (SPEEDES) is a unified parallel simulation environment. It supports multiple-synchronization protocols without requiring users to recompile their code. When a SPEEDES simulation runs on one node, all the extra parallel overhead is removed automatically at run time. When the same executable runs in parallel, the user preselects the synchronization algorithm from a list of options. SPEEDES currently runs on UNIX networks and on the California Institute of Technology/Jet Propulsion Laboratory Mark III Hypercube. SPEEDES also supports interactive simulations. Featured in the SPEEDES environment is a new parallel synchronization approach called Breathing Time Buckets. This algorithm uses some of the conservative techniques found in Time Bucket synchronization, along with the optimism that characterizes the Time Warp approach. A mathematical model derived from first principles predicts the performance of Breathing Time Buckets. Along with the Breathing Time Buckets algorithm, this paper discusses the rules for processing events in SPEEDES, describes the implementation of various other synchronization protocols supported by SPEEDES, describes some new ones for the future, discusses interactive simulations, and then gives some performance results.

Steinman, Jeff S.↗

A New Code SORD for Simulation of Polarized Light Scattering in the Earth Atmosphere

We report a new publicly available radiative transfer (RT) code for numerical simulation of polarized light scattering in plane-parallel atmosphere of the Earth. Using 44 benchmark tests, we prove high accuracy of the new RT code, SORD (Successive ORDers of scattering). We describe capabilities of SORD and show run time for each test on two different machines. At present, SORD is supposed to work as part of the Aerosol Robotic NETwork (AERONET) inversion algorithm. For natural integration with the AERONET software, SORD is coded in Fortran 90/95. The code is available by email request from the corresponding (first) author or from ftp://climate1.gsfc.nasa.gov/skorkin/SORD/.

AERONET↗

SOMAFOAM: An OpenFOAM based solver for continuum simulations of low-temperature plasmas

Here, we report the development of SOMAFOAM, a finite volume framework for performing continuum simulations of low-temperature plasmas. The primary goal of this work is to discuss the features of SOMAFOAM along with representative results provided as examples for a range of operating conditions and geometries. This includes plasma and plasma–dielectric systems operating in direct current, radio frequency, and microwave regimes from pressures as low as 100 mTorr to atmospheric pressure. The code has several useful features including the ability to run massively parallel simulations using arbitrary geometries, structured/unstructured meshes, choice of models such as drift–diffusion/full-momentum at runtime, and species-dependent timesteps to name a few. The verification/validation studies presented include comparison with previously published continuum simulations (low-pressure direct current and radio frequency plasma), with experiments (Gaseous Electronics Conference Reference Cell and microwave microplasma ignited in a split ring resonator), and previously published kinetic simulations (low-pressure radio frequency plasma). Other examples provided include a direct current atmospheric pressure microplasma bounded by dielectric sidewalls and a helium–nitrogen plasma ignited using a needle electrode facing a dielectric. The performance of the code is also discussed with serial and distributed memory parallel runs demonstrated up to 512 cores. The design and implementation of the code in a modular object-oriented framework allows for easy extension and seamless coupling with other codes and can be expected to play an important role in both academia and industry.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

HPC for Optimizing Process Parameters to Control Material Evolution in Seamless Induction Hardening of Wind Turbine Main Shaft Bearings

Work proposed in this project focused on understanding the effect of martensitic transformation in the steel on the potential for cracking during seamless induction hardening (SIH) as a function of process conditions to allow the process to optimally scale up. Large-scale, three-dimensional phase-field simulations of martensitic transformation were performed using MEUMAPPS-SS (Microstructure Evolution Using Massively Parallel Phase-field Simulations – Solid State) code developed at Oak Ridge National Laboratory. The simulations were guided by location-specific thermal history generated by experimental measurements of time-temperature history generated at The Timken Company. The simulations were able to capture the morphological evolution of the martensite variants in an Fe-1.0C-1.5Cr steel based on the Nishiyama-Wasserman (NW) orientation relationship. The simulations were also able to quantify the stress-state at the interface between impinging martensite variants. The simulations indicated that the magnitude of the various stress and strain components were dependent on the sizes of the impinging plates with a reduction in these quantities with reduced plate size in agreement with experimental findings. The results obtained from the simulations will be used to guide the optimization of the alloy thermal conditions to eliminate quench cracking during SIH of bearing steels.

99 GENERAL AND MISCELLANEOUS↗

Drift kinetic electrostatic simulations of the edge localized mode heat pulse

In the present work, electrostatic drift kinetic simulations of parallel plasma transport within the tokamak scrape-off layer (SOL) are conducted using the COGENT code. The SOL configuration is represented in one-dimensional slab geometry, incorporating a heat source localized in the midplane. The heat source parameters correspond to those characterizing edge-localized modes observed in the Joint European Torus (JET) tokamak. The numerical model includes kinetic treatment of both ions and electrons, a simplified model for the gyrokinetic Poisson equation that allows one to step over short time scales associated with fast electrostatic shear Alfvèn waves, and the logical sheath boundary condition (LSBC) that enforces global system quasineutrality. A third-order accurate LSBC is derived to be consistent with the third-order accurate upwind advection scheme utilized in the code, and it was shown to noticeably impact the simulation results, especially parallel heat flux at the target plate. The findings of this study are in agreement with results from preceding fluid and kinetic simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Performance Evaluation of Different Parallel Programming Models in SCALE-Shift Sequences for Criticality and Shielding Applications [Abstract]

The SCALE code system has been widely used for nuclear criticality safety, reactor physics, radiation shielding, source term generation, and inventory analyses by researchers, industry, and regulatory bodies. Although limited support for shared- and distributed-memory parallel processing was introduced via C++ threading, OpenMP, and MPI, a hybrid parallel programming model with both distributed- and shared-memory parallelism has not been fully supported in the SCALE code system.

Nuclear Criticality Safety Program (NCSP)↗

Structure of the quasi-parallel bow shock - Results of numerical simulations

A one-dimensional, nonperiodic hybrid code in which ion dynamics are treated exactly, while those of electrons are omitted by neglecting electron inertia and pressure, is employed in the simulation of quasi-parallel bow shock structures. It is found that for an Alfven Mach number smaller than 3, the magnetic field profile of the shock is laminar or quasi-laminar, and that the upstream waves are right-hand polarized whistlers. Downstream waves are absent or low in amplitudes. For Alfven Mach numbers greater than 3, the magnetic field profile of the shock is turbulent, and the upstream waves are again right-hand polarized whistlers. The transition from laminar-subcritical to turbulent-supercritical shock structures is shown to be due to the firehose instability which occurs at Alfven Mach numbers greater than about 3.

Kan, J. R.↗

Single and multiple scattering contributions to circumsolar radiation

The contributions to the angular distribution of the almucantar radiance in the forward direction due to multiple scattering are compared to those due to single scattering. The contributions have been calculated by a computer code employing the Gauss-Seidel iterative approach to the solution of the radiative transfer equation for a plane parallel atmosphere composed of air molecules, aerosol particles, and ozone. The code is similar to that of Dave (1972) except in the construction of the source matrix. In the near-forward direction the multiple scattering contributions are significant for optical depths of the order of 0.4. The shape of the angular distribution of almucantar radiance to 10 degrees is less sensitive to multiple scattering.

Box, M. A.↗

Implementation of Pfirsch–Schlüter parallel flow in x-ray imaging crystal spectrometer inversion analysis

The x-ray imaging crystal spectrometer (XICS) tomographic inversion code for Wendelstein 7-X (W7-X) has been modified to consider the effects of parallel flows and has been applied to analyze measurements taken during recent experimental campaigns. Previous analysis neglected the effects of parallel flows due to the primarily perpendicular geometry of the sightlines and the small magnitude predicted by neoclassical theory. To reconsider these effects, the incompressibility condition for plasma flows is used to calculate the parallel Pfirsch–Schlüter flow component for the equilibrium configuration. By incorporating this condition along with the geometry of the sightlines—i.e. the fractional contributions of perpendicular and parallel flows—, an updated expression for the measured flow is used for the profile inversion. Application of this modified inversion code to data from W7-X shows that the magnitude of the radial electric field and the flux surface averaged perpendicular flow are reduced by approximately a factor of 2 and brought into better agreement with neoclassical predictions and charge exchange recombination spectroscopy measurements.

Pfirsch–Schlüter flows↗