Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38

Molecular Simulations in Astrobiology

One of the main goals of astrobiology is to understand the origin of cellular life. In the absence of any record of the earliest ancestors of contemporary cells, protocells, the most direct way to test our understanding of their characteristics is to construct laboratory models of protocells. Such efforts, currently underway in the NASA Astrobiology Program, are accompanied by computational studies aimed at explaining self-organization of simple molecules into ordered structures and developing designs of molecules that are capable of performing protocellular functions. Many of these functions, such as importing nutrients, capturing and storing energy, and responding to changes in the environment, are carried out by proteins bound to membranes. We use computer simulations to address the following, questions about these proteins: (1) How do small proteins (peptides) organize themselves into ordered structures at water-membrane interfaces and insert into membranes? (2) How do peptides aggregate to form membrane-spannin(y structures (e.g., channels)? (3) By what mechanisms do such aggregates perform their functions? The simulations are performed using the molecular dynamics (MD) method. In this method, Newton's equations of motion for each atom in the system are solved iteratively. At each time step, the forces exerted on each atom by the remaining atoms are evaluated by dividing them into two parts. Short-range forces are calculated directly in real space while long-range forces are evaluated in reciprocal space, usually using a particle-mesh algorithm which is of order O(NlnN). Currently, a time step of 2 femtoseconds is typically used, thereby making studies of problems occurring on multi-nanosecond time scales (10(exp 6) - 10(exp 8) time steps) accessible. To address a broader range of problems, simulations need to be extended by three orders of magnitude. Such an extension requires both algorithmic improvements and codes scalable to a large number of parallel processors. Work in this direction is in progress. Two specific series of simulations that demonstrate how peptides self-organize and function in membranes are discussed. In one series of simulations, it was shown that nonpolar peptides, disordered in water, translocate to the nonpolar interior of the membrane and, simultaneously, fold into two different helical structures, which remain in equilibrium. Once in the membrane, the peptides can readily change their orientation, especially in response to local electric fields. This structural and orientational flexibility of peptides with changing conditions may have provided a mechanism of transmitting signals between the environment and the interior of the protocell. In another series of simulations, the mechanism by which a simple protein channel efficiently mediates proton transport across membranes was investigated. This process is a key step in cellular bioenergetics. In the channel under study, proton transport is gated by four histidines that occlude the channel pore. The simulations demonstrate that protons move through the gate by a "shuttle" mechanism, wherein one histidine is protonated on the extracellular side and, subsequently, the proton bound on the opposite side is released.

Pohorille, Andrew↗

NAS Parallel Benchmarks I/O Version 2.4

We describe a benchmark problem, based on the Block-Tridiagonal (BT) problem of the NAS Parallel Benchmarks (NPB), which is used to test the output capabilities of high-performance computing systems, especially parallel systems. We also present a source code implementation of the benchmark, called NPBIO2.4-MPI, based on the MPI implementation of NPB, using a variety of ways to write the computed solutions to file.

Wong, Parkson↗

Internal Flow Thermal/Fluid Modeling of STS-107 Port Wing in Support of the Columbia Accident Investigation Board

As part of the aero-thermodynamics team supporting the Columbia Accident Investigation Board (CAB), the Marshall Space Flight Center was asked to perform engineering analyses of internal flows in the port wing. The aero-thermodynamics team was split into internal flow and external flow teams with the support being divided between shorter timeframe engineering methods and more complex computational fluid dynamics. In order to gain a rough order of magnitude type of knowledge of the internal flow in the port wing for various breach locations and sizes (as theorized by the CAB to have caused the Columbia re-entry failure), a bulk venting model was required to input boundary flow rates and pressures to the computational fluid dynamics (CFD) analyses. This paper summarizes the modeling that was done by MSFC in Thermal Desktop. A venting model of the entire Orbiter was constructed in FloCAD based on Rockwell International s flight substantiation analyses and the STS-107 reentry trajectory. Chemical equilibrium air thermodynamic properties were generated for SINDA/FLUINT s fluid property routines from a code provided by Langley Research Center. In parallel, a simplified thermal mathematical model of the port wing, including the Thermal Protection System (TPS), was based on more detailed Shuttle re-entry modeling previously done by the Dryden Flight Research Center. Once the venting model was coupled with the thermal model of the wing structure with chemical equilibrium air properties, various breach scenarios were assessed in support of the aero-thermodynamics team. The construction of the coupled model and results are presented herein.

Sharp, John R.↗

Complexity Computational Environment: Data Assimilation SERVOGrid

We are using Web (Grid) service technology to demonstrate the assimilation of multiple distributed data sources (a typical data grid problem) into a major parallel high-performance computing earthquake forecasting code. Such a linkage of Geoinformatics with Geocomplexity demonstrates the value of the Solid Earth Research Virtual Observatory (SERVO) Grid concept, and advance Grid technology by building the first real-time large-scale data assimilation grid Here we develop the next steps for both the SERVO concept and the identified need for a Solid Earth problem-solving environment. We use a challenging motivating problem of importance to NASA namely integrating NASA space geodetic observations with numerical simulations of a changing earth.

data assimiliation↗

Uncertainty Determination for Aeroheating in Uranus and Saturn Probe Entries by the Monte Carlo Method

The 2013-2022 Decaedal survey for planetary exploration has identified probe missions to Uranus and Saturn as high priorities. This work endeavors to examine the uncertainty for determining aeroheating in such entry environments. Representative entry trajectories are constructed using the TRAJ software. Flowfields at selected points on the trajectories are then computed using the Data Parallel Line Relaxation (DPLR) Computational Fluid Dynamics Code. A Monte Carlo study is performed on the DPLR input parameters to determine the uncertainty in the predicted aeroheating, and correlation coefficients are examined to identify which input parameters show the most influence on the uncertainty. A review of the present best practices for input parameters (e.g. transport coefficient and vibrational relaxation time) is also conducted. It is found that the 2(sigma) - uncertainty for heating on Uranus entry is no more than 2.1%, assuming an equilibrium catalytic wall, with the uncertainty being determined primarily by diffusion and H(sub 2) recombination rate within the boundary layer. However, if the wall is assumed to be partially or non-catalytic, this uncertainty may increase to as large as 18%. The catalytic wall model can contribute over 3x change in heat flux and a 20% variation in film coefficient. Therefore, coupled material response/fluid dynamic models are recommended for this problem. It was also found that much of this variability is artificially suppressed when a constant Schmidt number approach is implemented. Because the boundary layer is reacting, it is necessary to employ self-consistent effective binary diffusion to obtain a correct thermal transport solution. For Saturn entries, the 2(sigma) - uncertainty for convective heating was less than 3.7%. The major uncertainty driver was dependent on shock temperature/velocity, changing from boundary layer thermal conductivity to diffusivity and then to shock layer ionization rate as velocity increases. While radiative heating for Uranus entry was negligible, the nominal solution for Saturn computed up to 20% radiative heating at the highest velocity examined. The radiative heating followed a non-normal distribution, with up to a 3x variation in magnitude. This uncertainty is driven by the H(sub 2) dissociation rate, as H(sub 2) that persists in the hot non-equilibrium zone contributes significantly to radiation.

Palmer, Grant↗

Hypercube matrix computation task

The Hypercube Matrix Computation (Year 1986-1987) task investigated the applicability of a parallel computing architecture to the solution of large scale electromagnetic scattering problems. Two existing electromagnetic scattering codes were selected for conversion to the Mark III Hypercube concurrent computing environment. They were selected so that the underlying numerical algorithms utilized would be different thereby providing a more thorough evaluation of the appropriateness of the parallel environment for these types of problems. The first code was a frequency domain method of moments solution, NEC-2, developed at Lawrence Livermore National Laboratory. The second code was a time domain finite difference solution of Maxwell's equations to solve for the scattered fields. Once the codes were implemented on the hypercube and verified to obtain correct solutions by comparing the results with those from sequential runs, several measures were used to evaluate the performance of the two codes. First, a comparison was provided of the problem size possible on the hypercube with 128 megabytes of memory for a 32-node configuration with that available in a typical sequential user environment of 4 to 8 megabytes. Then, the performance of the codes was anlyzed for the computational speedup attained by the parallel architecture.

Calalo, R.↗

Parallel Methods on Large-Scale Structural Analysis and Physics Applications; Symposium, Hampton, VA, Feb. 5, 6, 1991, Selected Papers

Recent advances in parallel methods and algorithms integrated into large-scale codes are presented. Consideration is given to problem decomposition (substructuring), efficient matrix solution algorithms for shared memory architectures, dynamic and transient analysis algorithms for shared memory architectures, and algorithms for distributed and massively parallel architectures. Particular attention is given to partitioning of unstructured problems for parallel processing, parallel-vector computation for linear-structural analysis and nonlinear unconstraint optimization problems, a parallel-vector equation solver for unsymmetric matrices on supercomputers, parallel nonlinear finite element dynamic response, multigrid algorithms for solving structural mechanics problems on supercomputers, structural analysis on massively parallel computers, explicit finite element methods with contact-impact on SIMD computers, and the impact of mapping and sparsity on parallelized finite element method modules.

Storaasli, Olaf O.↗

Massively Parallel Capability in Sierra/SD for Simulation Vibration with Piezoelectrics

Sierra/SD is an engineering structural dynamics code that provides Sandia and other customers a tool to model structural and acoustic physics on large complex physical systems using massively parallel processing. This report provides a detailed overview on Sierra/SD’s most recent physics package: coupled electro-mechanical physics. This capability uses the finite element method to model coupled electro-mechanical physics exhibited by piezoelectric materials. This report provides an applications overview, theory overview, and verification examples demonstrating the electro-mechanical physics modeling capabilities of Sierra/SD.

97 MATHEMATICS AND COMPUTING↗

LEWICE droplet trajectory calculations on a parallel computer

A parallel computer implementation (128 processors) of LEWICE, a NASA Lewis code used to predict the time-dependent ice accretion process for two-dimensional aerodynamic bodies of simple geometries, is described. Two-dimensional parallel droplet trajectory calculations are performed to demonstrate the potential benefits of applying parallel processing to ice accretion analysis. Parallel performance is evaluated as a function of the number of trajectories and the number of processors. For comparison, similar trajectory calculations are performed on single-processor Cray computers, and the best parallel results are found to be 33 and 23 times faster, respectively, than those of the Cray XMP and YMP.

Caruso, Steven C.↗

Multitasking the INS3D-LU code on the Cray Y-MP

This paper presents the results of multitasking the INS3D-LU code on eight processors. The code is a full Navier-Stokes solver for incompressible fluid in three dimensional generalized coordinates using a lower-upper symmetric-Gauss-Seidel implicit scheme. This code has been fully vectorized on oblique planes of sweep and parallelized using autotasking with some directives and minor modifications. The timing results for five grid sizes are presented and analyzed. The code has achieved a processing rate of over one Gflops.

Fatoohi, Rod↗

A sensitivity analysis of twinning crystal plasticity finite element model using single crystal and poly crystal Zircaloy

The popularity of crystal plasticity finite element method (CPFEM) models is increasing due to their ability to predict the mechanical response of crystalline materials such as metals and metal alloys more accurately than traditional continuum mechanics models. This is since the crystal plasticity models consider the effect of atomic structure, microstructural morphology, and properties of individual grains. These CPFEM models use a large number of material parameters in order to capture the mesoscale physics which comes with the downside of the tedious calibration process. In this paper, a CPFEM code was developed to include the twinning induced grain reorientation and subsequent crystallographic slip for HPC material. The developed code is incorporated in a large-scale, parallelized nonlinear solver WARP3D. Further, a sensitivity analysis with respect to 22 material parameters was then conducted using single crystal and polycrystal representative volume element (RVE) of Zircaloy material. Loading was applied along five different crystallographic orientations for single crystal RVE and along three directions namely, rolling (RD), transverse (TD), and normal (ND) direction for polycrystal RVE. Results obtained from the sensitivity analysis were used for the calibration of material parameters for Zircaloy. Finally, developed code along with calibrated material parameters was used to investigate the effect of the hydride phase formation in Zircaloy which is a typical case observed for nuclear applications. It was found that the volume fraction of the hydride phase has a significant impact on the mechanical properties of Zircaloy.

36 MATERIALS SCIENCE↗

MEDUSA - An overset grid flow solver for network-based parallel computer systems

Continuing improvement in processing speed has made it feasible to solve the Reynolds-Averaged Navier-Stokes equations for simple three-dimensional flows on advanced workstations. Combining multiple workstations into a network-based heterogeneous parallel computer allows the application of programming principles learned on MIMD (Multiple Instruction Multiple Data) distributed memory parallel computers to the solution of larger problems. An overset-grid flow solution code has been developed which uses a cluster of workstations as a network-based parallel computer. Inter-process communication is provided by the Parallel Virtual Machine (PVM) software. Solution speed equivalent to one-third of a Cray-YMP processor has been achieved from a cluster of nine commonly used engineering workstation processors. Load imbalance and communication overhead are the principal impediments to parallel efficiency in this application.

Smith, Merritt H.↗

Testing Models of Resonant Compton Scattering in X-Ray Pulsars

Over the performance period covered by the grant, the principal investigator modified a Monte Carlo Compton scattering code to model the propagation of x-rays through the magnetosphere of accreting neutron stars. These modifications were made to enable the author to compare the observations of x-ray pulsars to theoretical models of the system. The original code was designed to study relativistic plasmas with one of two geometries: a plane parallel plasma with a differential relativistic bulk velocity, and a static spherically symmetric plasma.- This code did not treat gravitational bending or bulk motion in the magnetosphere of a neutron star. Under the grant, the author incorporated code to trace light paths in a Schwarzschild metric. The code was modified to keep track of the photon polarization during propagati on. The investigator also modified the code so that bulk motion in an axisymmetric system is treated properly. An approximate treatment for resonant Compton scattering was added to the code. Finally, code was added that creates model observables that can be compared to observations, such as projected x-ray emission maps and energy-dependent light curves. Comparison to observations is now commencing.

Brainerd, Jerome J.↗

EUPDF-II: An Eulerian Joint Scalar Monte Carlo PDF Module : User's Manual

EUPDF-II provides the solution for the species and temperature fields based on an evolution equation for PDF (Probability Density Function) and it is developed mainly for application with sprays, combustion, parallel computing, and unstructured grids. It is designed to be massively parallel and could easily be coupled with any existing gas-phase CFD and spray solvers. The solver accommodates the use of an unstructured mesh with mixed elements of either triangular, quadrilateral, and/or tetrahedral type. The manual provides the user with an understanding of the various models involved in the PDF formulation, its code structure and solution algorithm, and various other issues related to parallelization and its coupling with other solvers. The source code of EUPDF-II will be available with National Combustion Code (NCC) as a complete package.

Raju, M. S.↗

Evaluating Performance Portability with the CMS Heterogeneous Pixel Reconstruction code

In the past years the landscape of tools for expressing parallel algorithms in a portable way across various compute accelerators has continued to evolve significantly. There are many technologies on the market that provide portability between CPU, GPUs from several vendors, and in some cases even FPGAs. These technologies include C++ libraries such as Alpaka and Kokkos, compiler directives such as OpenMP, the SYCL open specification that can be implemented as a library or in a compiler, and standard C++ where the compiler is solely responsible for the offloading. Given this developing landscape, users have to choose the technology that best fits their applications and constraints. For example, in the CMS experiment the experience so far in heterogeneous reconstruction algorithms suggests that the full application contains a large number of relatively short computational kernels and memory transfer operations. In this work we use a stand-alone version of the CMS heterogeneous pixel reconstruction code as a realistic use case of HEP reconstruction software that is capable of leveraging GPUs effectively. We summarize the experience of porting this code base from CUDA to Alpaka, Kokkos, SYCL, std::par, and OpenMP offloading. We compare the event processing throughput achieved by each version on NVIDIA and AMD GPUs as well as on a CPU, and compare those to what a native version of the code achieves on each platform.

Andriotis, Nikolaos↗

Injector Design Tool Improvements: User's manual for FDNS V.4.5

The major emphasis of the current effort is in the development and validation of an efficient parallel machine computational model, based on the FDNS code, to analyze the fluid dynamics of a wide variety of liquid jet configurations for general liquid rocket engine injection system applications. This model includes physical models for droplet atomization, breakup/coalescence, evaporation, turbulence mixing and gas-phase combustion. Benchmark validation cases for liquid rocket engine chamber combustion conditions will be performed for model validation purpose. Test cases may include shear coaxial, swirl coaxial and impinging injection systems with combinations LOXIH2 or LOXISP-1 propellant injector elements used in rocket engine designs. As a final goal of this project, a well tested parallel CFD performance methodology together with a user's operation description in a final technical report will be reported at the end of the proposed research effort.

Chen, Yen-Sen↗

Constructions for finite-state codes

A class of codes called finite-state (FS) codes is defined and investigated. These codes, which generalize both block and convolutional codes, are defined by their encoders, which are finite-state machines with parallel inputs and outputs. A family of upper bounds on the free distance of a given FS code is derived from known upper bounds on the minimum distance of block codes. A general construction for FS codes is then given, based on the idea of partitioning a given linear block into cosets of one of its subcodes, and it is shown that in many cases the FS codes constructed in this way have a d sub free which is as large as possible. These codes are found without the need for lengthy computer searches, and have potential applications for future deep-space coding systems. The issue of catastropic error propagation (CEP) for FS codes is also investigated.

Pollara, F.↗

Finite-state codes

A class of codes called finite-state (FS) codes is defined and investigated. The codes, which generalize both block and convolutional codes, are defined by their encoders, which are finite-state machines with parallel inputs and outputs. A family of upper bounds on the free distance of a given FS code is derived. A general construction for FS codes is given, and it is shown that in many cases the FS codes constructed in this way have a free distance that is the largest possible. Catastrophic error propagation (CEP) for FS codes is also discussed. It is found that to avoid CEP one must solve the graph-theoretic problem of finding a uniquely decodable edge labeling of the state diagram.

Pollara, Fabrizio↗