Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Error Control Coding Techniques for Space and Satellite Communications

It is well known that the BER performance of a parallel concatenated turbo-code improves roughly as 1/N, where N is the information block length. However, it has been observed by Benedetto and Montorsi that for most parallel concatenated turbo-codes, the FER performance does not improve monotonically with N. In this report, we study the FER of turbo-codes, and the effects of their concatenation with an outer code. Two methods of concatenation are investigated: across several frames and within each frame. Some asymmetric codes are shown to have excellent FER performance with an information block length of 16384. We also show that the proposed outer coding schemes can improve the BER performance as well by eliminating pathological frames generated by the iterative MAP decoding process.

Costello, Daniel J., Jr.↗

Interpreting Broad Double-Peaked Emission Lines in Active Galactic Nuclei

The principal objectives of this project were to probe the inner regions of active galactic nuclei and to test general relativity in the strong-field limit. The approach takes advantage of broad atomic line emission observed from material deep in the potential well of an active galactic nucleus which contains key information as to the physics of the system. Line profiles in a wide range of wavebands from optical to X-ray have provided compelling evidence of the existence of a relativistic accretion disk around a supermassive black hole in a number of galaxies. The simplest model posits a geometrically thin disk in Keplerian orbit, with general relativistic effects in evidence. This model is the point of departure for the proposed work. We developed a high-performance numerical code to calculate photon trajectories in a Schwarzschild or Kerr metric and implemented it on parallel supercomputers. This code includes a general purpose ray tracer that calculates line profiles, light curves, and other observable quantities for a wide variety of emitter configurations. The versatility comes from the fact that the ray tracing algorithm does not depend on any symmetries regarding emitter locations. The speed comes from parallel implementation which enables us to sample hitherto unattainable volumes of disk model parameter space. During the period 1 March 1997 through 28 February 1998, two papers, supported in whole or in part by this grant, were published in refereed journals. They are reproduced in their entirety in the next two sections of this report.

Halpern, Jules↗

Charon Message-Passing Toolkit for Scientific Computations

Charon is a library, callable from C and Fortran, that aids the conversion of structured-grid legacy codes-such as those used in the numerical computation of fluid flows-into parallel, high- performance codes. Key are functions that define distributed arrays, that map between distributed and non-distributed arrays, and that allow easy specification of common communications on structured grids. The library is based on the widely accepted MPI message passing standard. We present an overview of the functionality of Charon, and some representative results.

VanderWijngaart, Rob F.↗

An object-oriented approach for parallel self adaptive mesh refinement on block structured grids

Self-adaptive mesh refinement dynamically matches the computational demands of a solver for partial differential equations to the activity in the application's domain. In this paper we present two C++ class libraries, P++ and AMR++, which significantly simplify the development of sophisticated adaptive mesh refinement codes on (massively) parallel distributed memory architectures. The development is based on our previous research in this area. The C++ class libraries provide abstractions to separate the issues of developing parallel adaptive mesh refinement applications into those of parallelism, abstracted by P++, and adaptive mesh refinement, abstracted by AMR++. P++ is a parallel array class library to permit efficient development of architecture independent codes for structured grid applications, and AMR++ provides support for self-adaptive mesh refinement on block-structured grids of rectangular non-overlapping blocks. Using these libraries, the application programmers' work is greatly simplified to primarily specifying the serial single grid application and obtaining the parallel and self-adaptive mesh refinement code with minimal effort. Initial results for simple singular perturbation problems solved by self-adaptive multilevel techniques (FAC, AFAC), being implemented on the basis of prototypes of the P++/AMR++ environment, are presented. Singular perturbation problems frequently arise in large applications, e.g. in the area of computational fluid dynamics. They usually have solutions with layers which require adaptive mesh refinement and fast basic solvers in order to be resolved efficiently.

Lemke, Max↗

symPACK: A GPU-Capable Fan-Out Sparse Cholesky Solver

Sparse symmetric positive definite systems of equations are ubiquitous in scientific workloads and applications. Parallel sparse Cholesky factorization is the method of choice for solving such linear systems. Therefore, the development of parallel sparse Cholesky codes that can efficiently run on today’s large-scale heterogeneous distributed-memory platforms is of vital importance. Modern supercomputers offer nodes that contain a mix of CPUs and GPUs. To fully utilize the computing power of these nodes, scientific codes must be adapted to offload expensive computations to GPUs. We present symPACK, a GPU-capable parallel sparse Cholesky solver that uses one-sided communication primitives and remote procedure calls provided by the UPC++ library. We also utilize the UPC++ "memory kinds" feature to enable efficient communication of GPU-resident data. We show that on a number of large problems, symPACK outperforms comparable state-of-the-art GPU-capable Cholesky factorization codes by up to 14x on the NERSC Perlmutter supercomputer.

Bellavita, Julian↗

An Evaluation of Structural Analysis Methodologies for Space Deployable Structures

Benchmarks are introduced for evaluating the performance of numerical simulations of space deployable structures. These benchmarks embody the key challenges of interest to future large space deployable structures, including large angle motion, contact between flexible bodies, and the presence of both soft and stiff mechanical components. The benchmarks were used in companion studies to evaluate the ADAMS multibody dynamics code, the LS-Dyna nonlinear finite element code, and the Sierra large-scale parallel nonlinear finite element code. In the past, only multibody codes would have been considered for this application. This study found that all three codes could be used for these benchmarks, a finding that may lead to larger scale, higher fidelity simulations in the future.

Mobrem, Mehran↗

GOES-R Geostationary Lightning Mapper Performance Specifications and Algorithms

The Geostationary Lightning Mapper (GLM) is a single channel, near-IR imager/optical transient event detector, used to detect, locate and measure total lightning activity over the full-disk. The next generation NOAA Geostationary Operational Environmental Satellite (GOES-R) series will carry a GLM that will provide continuous day and night observations of lightning. The mission objectives for the GLM are to: (1) Provide continuous, full-disk lightning measurements for storm warning and nowcasting, (2) Provide early warning of tornadic activity, and (2) Accumulate a long-term database to track decadal changes of lightning. The GLM owes its heritage to the NASA Lightning Imaging Sensor (1997- present) and the Optical Transient Detector (1995-2000), which were developed for the Earth Observing System and have produced a combined 13 year data record of global lightning activity. GOES-R Risk Reduction Team and Algorithm Working Group Lightning Applications Team have begun to develop the Level 2 algorithms and applications. The science data will consist of lightning "events", "groups", and "flashes". The algorithm is being designed to be an efficient user of the computational resources. This may include parallelization of the code and the concept of sub-dividing the GLM FOV into regions to be processed in parallel. Proxy total lightning data from the NASA Lightning Imaging Sensor on the Tropical Rainfall Measuring Mission (TRMM) satellite and regional test beds (e.g., Lightning Mapping Arrays in North Alabama, Oklahoma, Central Florida, and the Washington DC Metropolitan area) are being used to develop the prelaunch algorithms and applications, and also improve our knowledge of thunderstorm initiation and evolution.

Mach, Douglas M.↗

High-Density Plasma Reactors: Simulations for Design

The development of improved and more efficient plasma reactors is a costly process for the semiconductor industry. Until five years ago, the Industry made most of its advancements through a trial and error approach. More recently, the role of computational modeling in the design process has increased. Both conventional computational fluid dynamics (CFD) techniques like Navier-Stokes solvers as well as particle simulation methods are used to model plasma reactor flowfields. However, since high-density plasma reactors generally operate at low gas pressures on the order of 1 to 10 mTorr, a particle simulation may be necessary because of the failure of CFD techniques to model rarefaction effects. The direct simulation Monte Carlo method is the most widely accepted and employed particle simulation tool and has previously been used to investigate plasma reactor flowfields. A plasma DSMC code is currently under development at NASA Ames Research Center with its foundation as the object-oriented parallel Cornell DSMC code, MONACO. The present investigation is a follow up of a neutral flow investigation of the effects of process parameters as well as reactor design on etch rate and etch rate uniformity. The previous work concentrated on silicon etch of a chlorine flow in a configuration typical of electron cyclotron resonance (ECR) or helical resonator type reactors. The effects of the plasma on the dissociation chemistry were modeled by making assumptions about the electron temperature and number density. The electrons or ions themselves were not simulated.The present work extends these results by simulating the charged species.The electromagnetic fields are calculated such that power deposition is modeled self-consistently. Electron impact reactions are modeled along with mechanisms for charge exchange. An bipolar diffusion assumption is made whereby electrons remain tied to the ions. However, the velocities of tile electrons are allowed to be modified during collisions and are not confined to a Maxwellian distribution. The interaction between the neutral flow and plasma is examined, and results for etch rate uniformity from the previous research and the present plasma simulations are compared.

Hash, David B.↗

High-performance computing in water resources hydrodynamics

In this work, we present a vision of future water resources hydrodynamics codes that can fully utilize the strengths of modern high-performance computing. The advances to computing power, formerly driven by the improvement of central processing unit processors, now focus on parallel computing and, in particular, the use of graphics processing units (GPUs). However, this shift to a parallel framework requires refactoring the code to make efficient use of the data as well as changing even the nature of the algorithm that solves the system of equations. These concepts along with other features such as the precision for the computations, dry regions management, and input/output data are analyzed in this paper. A 2D multi-GPU flood code applied to a large-scale test case is used to corroborate our statements and ascertain the new challenges for the next-generation parallel water resources codes.

54 ENVIRONMENTAL SCIENCES↗

Parallel Implicit Algorithms for CFD

The main goal of this project was efficient distributed parallel and workstation cluster implementations of Newton-Krylov-Schwarz (NKS) solvers for implicit Computational Fluid Dynamics (CFD.) "Newton" refers to a quadratically convergent nonlinear iteration using gradient information based on the true residual, "Krylov" to an inner linear iteration that accesses the Jacobian matrix only through highly parallelizable sparse matrix-vector products, and "Schwarz" to a domain decomposition form of preconditioning the inner Krylov iterations with primarily neighbor-only exchange of data between the processors. Prior experience has established that Newton-Krylov methods are competitive solvers in the CFD context and that Krylov-Schwarz methods port well to distributed memory computers. The combination of the techniques into Newton-Krylov-Schwarz was implemented on 2D and 3D unstructured Euler codes on the parallel testbeds that used to be at LaRC and on several other parallel computers operated by other agencies or made available by the vendors. Early implementations were made directly in Massively Parallel Integration (MPI) with parallel solvers we adapted from legacy NASA codes and enhanced for full NKS functionality. Later implementations were made in the framework of the PETSC library from Argonne National Laboratory, which now includes pseudo-transient continuation Newton-Krylov-Schwarz solver capability (as a result of demands we made upon PETSC during our early porting experiences). A secondary project pursued with funding from this contract was parallel implicit solvers in acoustics, specifically in the Helmholtz formulation. A 2D acoustic inverse problem has been solved in parallel within the PETSC framework.

Keyes, David E.↗

Use of advanced computers for aerodynamic flow simulation

The current and projected use of advanced computers for large-scale aerodynamic flow simulation applied to engineering design and research is discussed. The design use of mature codes run on conventional, serial computers is compared with the fluid research use of new codes run on parallel and vector computers. The role of flow simulations in design is illustrated by the application of a three dimensional, inviscid, transonic code to the Sabreliner 60 wing redesign. Research computations that include a more complete description of the fluid physics by use of Reynolds averaged Navier-Stokes and large-eddy simulation formulations are also presented. Results of studies for a numerical aerodynamic simulation facility are used to project the feasibility of design applications employing these more advanced three dimensional viscous flow simulations.

Bailey, F. R.↗

Alfven wave resonances and flow induced by non-linear Alfven waves in a stratified atmosphere

A nonlinear, time-dependent, ideal MHD code has been developed and used to compute the flow induced by nonlinear Alfven waves propagating in an isothermal, stratified, plane-parallel atmosphere. The code is based on characteristic equations solved in a Lagrangian frame and is highly accurate. Results show that resonance behavior of Alfven waves exists in the presence of a continuous density gradient and that the waves with periods corresponding to resonant peaks exert considerably more force on the medium than off-resonance periods; this leads to enhanced flow. If only off-peak periods are considered, the relationship between the wave period and induced longitudinal velocity shows that short period WKB waves push more on the background medium than longer period, non-WKB, waves. The results also show the development of the longitudinal waves produced by the finite amplitude of the Alfven waves. The longitudinal wave becomes strong as the Alfven wave relative amplitude grows above 10 percent and will lead to strong damping of the Alfven waves.

Stark, B. A.↗

Performance Enhancement of APW+lo Calculations by Simplest Separation of Concerns

Full-potential linearized augmented plane wave (LAPW) and APW plus local orbital (APW+lo) codes differ widely in both their user interfaces and in capabilities for calculations and analysis beyond their common central task of all-electron solution of the Kohn–Sham equations. However, that common central task opens a possible route to performance enhancement, namely to offload the basic LAPW/APW+lo algorithms to a library optimized purely for that purpose. To explore that opportunity, we have interfaced the Exciting-Plus (“EP”) LAPW/APW+lo DFT code with the highly optimized SIRIUS multi-functional DFT package. This simplest realization of the separation of concerns approach yields substantial performance over the base EP code via additional task parallelism without significant change in the EP source code or user interface. We provide benchmarks of the interfaced code against the original EP using small bulk systems, and demonstrate performance on a spin-crossover molecule and magnetic molecule that are of size and complexity at the margins of the capability of the EP code itself.

Zhang, Long↗

Production Level CFD Code Acceleration for Hybrid Many-Core Architectures

In this work, a novel graphics processing unit (GPU) distributed sharing model for hybrid many-core architectures is introduced and employed in the acceleration of a production-level computational fluid dynamics (CFD) code. The latest generation graphics hardware allows multiple processor cores to simultaneously share a single GPU through concurrent kernel execution. This feature has allowed the NASA FUN3D code to be accelerated in parallel with up to four processor cores sharing a single GPU. For codes to scale and fully use resources on these and the next generation machines, codes will need to employ some type of GPU sharing model, as presented in this work. Findings include the effects of GPU sharing on overall performance. A discussion of the inherent challenges that parallel unstructured CFD codes face in accelerator-based computing environments is included, with considerations for future generation architectures. This work was completed by the author in August 2010, and reflects the analysis and results of the time.

Duffy, Austen C.↗

Analysis of a parallelized nonlinear elliptic boundary value problem solver with application to reacting flows

A parallelized finite difference code based on the Newton method for systems of nonlinear elliptic boundary value problems in two dimensions is analyzed in terms of computational complexity and parallel efficiency. An approximate cost function depending on 15 dimensionless parameters is derived for algorithms based on stripwise and boxwise decompositions of the domain and a one-to-one assignment of the strip or box subdomains to processors. The sensitivity of the cost functions to the parameters is explored in regions of parameter space corresponding to model small-order systems with inexpensive function evaluations and also a coupled system of nineteen equations with very expensive function evaluations. The algorithm was implemented on the Intel Hypercube, and some experimental results for the model problems with stripwise decompositions are presented and compared with the theory. In the context of computational combustion problems, multiprocessors of either message-passing or shared-memory type may be employed with stripwise decompositions to realize speedup of O(n), where n is mesh resolution in one direction, for reasonable n.

Keyes, David E.↗

A DSMC Study of Low Pressure Argon Discharge

Work toward a self-consistent plasma simulation using the DSMC (Direct Simulation Monte Carlo) method for examination of the flowfields of low-pressure high density plasma reactors is presented. Presently, DSMC simulations for these applications involve either treating the electrons as a fluid or imposing experimentally determined values for the electron number density profile. In either approach, the electrons themselves are not physically simulated. Self-consistent plasma DSMC simulations have been conducted for aerospace applications but at a severe computational cost due in part to the scalar architectures on which the codes were employed. The present work attempts to conduct such simulations at a more reasonable cost using a plasma version of the object-oriented parallel Cornell DSMC code, MONACO, on an IBM SP-2. Due to availability of experimental data, the GEC reference cell is chosen to conduct preliminary investigations. An argon discharge is chosen to conduct preliminary investigations. An argon discharge is examined thus affording a simple chemistry set with eight gas-phase reactions and five species: Ar, Ar(+), Ar(*), Ar(sub 2), and e where Ar(*) is a metastable.

Hash, David B.↗

Quantitative proton radiography and shadowgraphy for arbitrary intensities

Charged-particle radiography and shadowgraphy data can be directly inverted to obtain a line-integrated transverse Lorentz force or a line-integrated transverse refractive index gradient if intensity modulations due to scattering and absorption are negligible, and angular deflections are small. We develop a new direct-inversion algorithm based on plasma physics and compare it to a new Monge–Ampère code and an existing power diagram code. The measured or source intensity is represented by electrons subject to drag, and the other intensity by fixed ions. The decrease in kinetic plus electrostatic energy determines convergence. The displacement of the electrons from their initial to their equilibrium positions determines the line-integrated force or refractive index gradient. We have implemented two approaches: PIC (particle in cell) and Lagrangian fluid, in 1-D and 2-D. The PIC code works for arbitrary intensities, can work efficiently in parallel, and can make use of existing codes. The Lagrangian code requires less memory and is faster than the PIC code without massively parallel processing, but fails in 2-D for large intensity modulations. The Monge–Ampère code is by far the fastest in 2-D, without massively parallel processing, but fails for intensities with large voids, high contrast ratios and large deflections across the boundaries, and could not obtain the degree of convergence possible with the PIC code. As a result, the power diagram code was by far the slowest and most memory intensive, and failed for large peaks in the measured intensity.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗