Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52

NICS (NASA Instrument Capabilities Study) Instrument Schedule and Cost Study

This paper summarizes work performed on the Flight Projects Directorate Planetary Science Projects Division (PSPD, Code 430) NICS (NASA Instrument Capabilities study) instrument schedule and cost study. Included are a short summary of the original NICS (NASA, 2008), and the design and approach, data collection, analysis, preliminary findings and recommendations from select areas of the current study. The NICS (2008) was chartered by then NASA Chief Engineer Michael Ryschkewitsch and chaired by Goddard Space Flight Center (GSFC) engineer, John Leon. The focus was to identify problem areas in instrument development and, if possible, to offer solutions. In the area of instrument developments, the NICS (2008) identified a lack of resources and authority to successfully manage to instrument cost and schedule requirements; and a lack of critical skills, expertise, and leadership to successfully implement unique (one-of-a-kind) high technology developments (NASA, 2008, pp. 51, 52). Additionally, the NICS (2008) found problems in requirements formulation, reviews and management; unrealistic caps and overly optimistic estimates; and externally directed changes which increased the likelihood of overrunning cost and schedule (NASA, 2008, pp.53, 54). It is noteworthy that NICS findings are consistent with previous studies at the mission level (Robbins, Schmidt & White, 2020). Five years later in 2013, the Instrument Projects Division (IPD) was established to implement and manage instrument projects greater than $20M. The IPD was known as Code 490. Its structure incorporated several of the NICS (2008) recommendations. To see if these incorporated recommendations made a difference, and to identify other potential challenges in instrument developments, two parallel studies were initiated. Originally led by the IPD, now led by the PSPD, and the Instrument and Payload Systems Engineering Branch (IPSE, Code 592), respectively, the instrument schedule and cost study and the instrument technical complexity study began in 2017. Data collection was initiated in 2020 and is on-going. This paper is limited to the IPD/PSPD study. Among other findings, preliminary data indicate IPD/PSPD project management support positively influenced instrument development as related to providing a dedicated level of support staff, including a deputy Instrument Project Manager (dIPM), reducing IPM leadership changes, and providing other project support. Next steps include continued data collection and analysis, and mapping to technical complexity data.

NICS implementation↗

Computational Fluid Dynamic (CFD) analysis of axisymmetric plume and base flow of film/dump cooled rocket nozzle

Film/dump cooling a rocket nozzle with fuel rich gas, as in the National Launch System (NLS) Space Transportation Main Engine (STME), adds potential complexities for integrating the engine with the vehicle. The chief concern is that once the film coolant is exhausted from the nozzle, conditions may exist during flight for the fuel-rich film gases to be recirculated to the vehicle base region. The result could be significantly higher base temperatures than would be expected from a regeneratively cooled nozzle. CFD analyses were conduced to augment classical scaling techniques for vehicle base environments. The FDNS code with finite rate chemistry was used to simulate a single, axisymmetric STME plume and the NLS base area. Parallel calculations were made of the Saturn V S-1 C/F1 plume base area flows. The objective was to characterize the plume/freestream shear layer for both vehicles as inputs for scaling the S-C/F1 flight data to NLS/STME conditions. The code was validated on high speed flows with relevant physics. This paper contains the calculations for the NLS/STME plume for the baseline nozzle and a modified nozzle. The modified nozzle was intended to reduce the fuel available for recirculation to the vehicle base region. Plumes for both nozzles were calculated at 10kFT and 50kFT.

Tucker, P. K.↗

Computational strategy for the solution of large strain nonlinear problems using the Wilkins explicit finite-difference approach

The STEALTH code system, which solves large strain, nonlinear continuum mechanics problems, was rigorously structured in both overall design and programming standards. The design is based on the theoretical elements of analysis while the programming standards attempt to establish a parallelism between physical theory, programming structure, and documentation. These features have made it easy to maintain, modify, and transport the codes. It has also guaranteed users a high level of quality control and quality assurance.

Hofmann, R.↗

Parallel computation with the force

A methodology, called the force, supports the construction of programs to be executed in parallel by a force of processes. The number of processes in the force is unspecified, but potentially very large. The force idea is embodied in a set of macros which produce multiproceossor FORTRAN code and has been studied on two shared memory multiprocessors of fairly different character. The method has simplified the writing of highly parallel programs within a limited class of parallel algorithms and is being extended to cover a broader class. The individual parallel constructs which comprise the force methodology are discussed. Of central concern are their semantics, implementation on different architectures and performance implications.

Jordan, H. F.↗

Automatically parallelizing batch inference on deep neural networks using Fiats and Fortran 2023 `do concurrent`

This paper introduces novel programming strategies that leverage features of the Fortran 2023 standard of the International Standards Organization (ISO) to automatically parallelize computations on deep neural networks. The paper focuses on the interplay of object-oriented, parallel, and functional programming paradigms in the Fiats deep learning library. We demonstrate how several infrequently used language features play a role in enabling efficient, parallel execution. Specifically, the ability to explicitly declare that a procedure is pure facilitates inference in the context of the language’s loop-parallelism construct `do concurrent`. Also, explicitly prohibiting the overriding of a parent type’s type-bound procedures eliminates the need for dynamic dispatch in performance-critical code. Finally, this paper uses batch inference calculations on a neural network surrogate for atmospheric aerosol dynamics to demonstrate that LLVM Flang compiler’s automatic parallelization of `do concurrent` achieves roughly the same performance and scalability as achieved by OpenMP compiler directives. We also demonstrate that double-precision inference costs 37–72% longer runtime than default-real precision with most values in the range 57-60%.

Rouson, Damian↗

On the application of under-decimated filter banks

Maximally decimated filter banks have been extensively studied in the past. A filter bank is said to be under-decimated if the number of channels is more than the decimation ratio in the subbands. A maximally decimated filter bank is well known for its application in subband coding. Another application of maximally decimated filter banks is in block filtering. Convolution through block filtering has the advantages that parallelism is increased and data are processed at a lower rate. However, the computational complexity is comparable to that of direct convolution. More recently, another type of filter bank convolver has been developed. In this scheme, the convolution is performed in the subbands. Quantization and bit allocation of subband signals are based on signal variance, as in subband coding. Consequently, for a fixed rate, the result of convolution is more accurate than is direct convolution. This type of filter bank convolver also enjoys the advantages of block filtering, parallelism, and a lower working rate. Nevertheless, like block filtering, there is no computational saving. In this article, under-decimated systems are introduced to solve the problem. The new system is decimated only by half the number of channels. Two types of filter banks can be used in the under-decimated system: the discrete Fourier transform (DFT) filter banks and the cosine modulated filter banks. They are well known for their low complexity. In both cases, the system is approximately alias free, and the overall response is equivalent to a tunable multilevel filter. Properties of the DFT filter banks and the cosine modulated filter banks can be exploited to simultaneously achieve parallelism, computational saving, and a lower working rate. Furthermore, for both systems, the implementation cost of the analysis or synthesis bank is comparable to that of one prototype filter plus some low-complexity modulation matrices. The individual analysis and synthesis filters have complex coefficients in the DFT filter banks but have real coefficients in the cosine modulated filter banks.

Lin, Y.-P.↗

Source-to-Source Automatic Differentiation of OpenMP Parallel Loops

This article presents our work toward correct and efficient automatic differentiation of OpenMP parallel worksharing loops in forward and reverse mode. Automatic differentiation is a method to obtain gradients of numerical programs, which are crucial in optimization, uncertainty quantification, and machine learning. The computational cost to compute gradients is a common bottleneck in practice. For applications that are parallelized for multicore CPUs or GPUs using OpenMP, one also wishes to compute the gradients in parallel. Here, we propose a framework to reason about the correctness of the generated derivative code, from which we justify our OpenMP extension to the differentiation model. We implement this model in the automatic differentiation tool Tapenade and present test cases that are differentiated following our extended differentiation procedure. Performance of the generated derivative programs in forward and reverse mode is better than sequential, although our reverse mode often scales worse than the input programs.

97 MATHEMATICS AND COMPUTING↗

Parallel Gaussian elimination of a block tridiagonal matrix using multiple microcomputers

The solution of a block tridiagonal matrix using parallel processing is demonstrated. The multiprocessor system on which results were obtained and the software environment used to program that system are described. Theoretical partitioning and resource allocation for the Gaussian elimination method used to solve the matrix are discussed. The results obtained from running 1, 2 and 3 processor versions of the block tridiagonal solver are presented. The PASCAL source code for these solvers is given in the appendix, and may be transportable to other shared memory parallel processors provided that the synchronization outlines are reproduced on the target system.

Blech, Richard A.↗

Vector Radiative Transfer Code SORD: Performance Analysis and Quick Start Guide

We present a new open source polarized radiative transfer code SORD written in Fortran 9095. SORD numerically simulates propagation of monochromatic solar radiation in a plane-parallel atmosphere over a reflecting surface using the method of successive orders of scattering (hence the name). Thermal emission is ignored. We did not improve the method in any way, but report the accuracy and runtime in 52 benchmark scenarios. This paper also serves as a quick start users guide for the code available from ftp:maiac.gsfc.nasa.govpubskorkin, from the JQSRT website, or from the corresponding (first) author.

polarized radiative transfer↗

Turbulence modeling of free shear layers for high performance aircraft

In many flowfield computations, accuracy of the turbulence model employed is frequently a limiting factor in the overall accuracy of the computation. This is particularly true for complex flowfields such as those around full aircraft configurations. Free shear layers such as wakes, impinging jets (in V/STOL applications), and mixing layers over cavities are often part of these flowfields. Although flowfields have been computed for full aircraft, the memory and CPU requirements for these computations are often excessive. Additional computer power is required for multidisciplinary computations such as coupled fluid dynamics and conduction heat transfer analysis. Massively parallel computers show promise in alleviating this situation, and the purpose of this effort was to adapt and optimize CFD codes to these new machines. The objective of this research effort was to compute the flowfield and heat transfer for a two-dimensional jet impinging normally on a cool plate. The results of this research effort were summarized in an AIAA paper titled 'Parallel Implementation of the k-epsilon Turbulence Model'. Appendix A contains the full paper.

Sondak, Douglas↗

Performance Evaluation of Remote Memory Access (RMA) Programming on Shared Memory Parallel Computers

The purpose of this study is to evaluate the feasibility of remote memory access (RMA) programming on shared memory parallel computers. We discuss different RMA based implementations of selected CFD application benchmark kernels and compare them to corresponding message passing based codes. For the message-passing implementation we use MPI point-to-point and global communication routines. For the RMA based approach we consider two different libraries supporting this programming model. One is a shared memory parallelization library (SMPlib) developed at NASA Ames, the other is the MPI-2 extensions to the MPI Standard. We give timing comparisons for the different implementation strategies and discuss the performance.

Jin, Hao-Qiang↗

Rendezvous algorithms for large-scale modeling and simulation

Rendezvous algorithms encode a communication pattern that is useful when processors sending data do not know who the receiving processors should be, or vice versa. The idea is to define an intermediate decomposition where datums from different sending processors can ”rendezvous” to perform a computation, in a manner that both the senders and eventual receivers of the results can identify the appropriate rendezvous processor. Though they were originally designed for interpolating between overlaid grids with independent parallel decompositions (Plimpton et al., 2004), we have recently found rendezvous algorithms useful for a variety of operations in particle- or grid-based simulation codes when running large problems on large numbers of processors. In particular, we show they can perform well when a load-balanced intermediate decomposition is randomized and not spatial, requiring all-to-all communication to move data between processors. In this case rendezvous algorithms leverage the large bisection communication bandwidths which parallel machines provide. We describe how rendezvous algorithms work in a scientific computing context and give specific examples for molecular dynamics and Direct Simulation Monte Carlo codes which result in dramatic performance improvements versus simpler algorithms which do not scale as well. We explain how a generic rendezvous algorithm can be implemented, and also point out similarities with the MapReduce paradigm popularized by Google and Hadoop.

97 MATHEMATICS AND COMPUTING↗

CSRI Summer Proceedings 2020

The Computer Science Research Institute (CSRI) brings university faculty and students to Sandia for focused collaborative research on Department of Energy (DOE) computer and computational science problems. The institute provides an opportunity for university researchers to learn about problems in computer and computational science at DOE laboratories. Participants conduct leading-edge research, interact with scientists and engineers at the laboratories, and help transfer results of their research to programs at the labs. Some specific CSRI research interest areas are: scalable solvers, optimization, adaptivity and mesh refinement, graph-based, discrete, and combinatorial algorithms, uncertainty estimation, mesh generation, dynamic load-balancing, virus and other malicious-code defense, visualization, scalable cluster computers, data-intensive computing, environments for scalable computing, parallel input/output, advanced architectures, and theoretical computer science. The CSRI Summer Program is organized by CSRI and typically includes the organization of a weekly seminar series and the publication of a summer proceedings. In 2020, the CSRI summer program was executed completely virtually; all student interns worked from home, due to the COVID-19 pandemic.

97 MATHEMATICS AND COMPUTING↗

Integrating PGAS and MPI-based Graph Analysis

This project demonstrates that Chapel programs can interface with MPI-based libraries written in C++ without storing multiple copies of shared data. Chapel is a language for productive parallel computing using global address spaces (PGAS). We identified two approaches to interface Chapel code with the MPI-based Grafiki and Trilinos libraries. The first uses a single Chapel executable to call a C function that interacts with the C++ libraries. The second uses the mmap function to allow separate executables to read and write to the same block of memory on a node. We also encapsulated the second approach in Docker/Singularity containers to maximize ease of use. Comparisons of the two approaches using shared and distributed memory installations of Chapel show that both approaches provide similar scalability and performance.

97 MATHEMATICS AND COMPUTING↗

MAESTROeX

MAESTROeX is a massively parallel, finite-volume C++/F90 solver for low Mach number astrophysical flows. The code utilizes a low Mach number equation set allowing for more efficient, long-time integration of highly subsonic flows compared to compressible approaches. The recommended range of applicability is for flows where the Mach number does not exceed~0.1.

Harpole, Alice↗

Three-dimensional Navier-Stokes calculations of multiple interacting vortex rings

Results from a finite-difference Navier-Stokes code for three-dimensional, unsteady, vortical flows in unbounded domains are presented and analyzed in this paper. The vortical flows presented are representative of vortex rings and other closed vortical tubes or structures in fluid mechanics. Such structures are important elements in fluid flows such as jets, atmospheric turbulence, and the far-field wakes of aircraft, and studies of their interaction may aid in an understanding of complex fluid flows. The paper demonstrates that computational methods can be used as a viable alternative or supplement to experimental techniques for studying the physics of vortex flows. The separate visualization of vortex stretching, convection, and diffusion is presented in this paper for a single elliptical vortex ring.The calculations employ a truncated series expansion technique to simulate the unbounded nature of the fluid flow with a finite computational domain, which is a more accurate technique than the conventional freestream boundary specification. The numerical divergence of the three-dimensional vorticity field is considered as a useful estimate of truncation error, and the use of a kinetic energy decay law as a calculation check is demonstrated. Results from the Navier-Stokes code are presented for the unsteady motion of two and four vortex rings along parallel axes, and the results agree qualitatively with experimental flow visualization.

Chamberlain, J. P.↗

Computational simulation of transition to turbulence through inverse modeling

The present investigation has focused on a computational methodology for the fundamental case of transition in channel flow, in which recently published experimental data are utilized both as a stimulus and as a measure of merit of the method. The research has proceeded along three avenues in parallel. The first task has consisted of the development and verification of a computer code which calculates the mean evolution of flow in a channel similar to the one employed experimentally by Blair and Anderson. An analytical test case was created for the dual purposes of code verification and of highlighting the interactions between the Reynolds stress and the mean velocity profile. This test case generated a Reynolds stress by the residue in the momentum equation which is produced by a typical analytical velocity profile. By a substitution of this Reynolds stress into the appropriate code module, the correctness of the code may be verified, along with the accuracy of the computational method. The second task pursued has involved the development of a triple layer model for the Reynolds stress profile, which was suggested and derived from experimental velocity profiles. It is demonstrated that the innermost length scale is based on the local friction velocity, the intermediate layer corresponds to the usual logarithmic law of the wall region in which the normalized Reynolds stress is approximately unity, and the outermost layer is represented by a closed mathematical form depending explicitly on the velocity profile in the wake region. The third task was comprised of scrutiny of the excellent databases developed by Blair and others, and the planning of its incorporation into the transition analysis. These extensive measurements indicate that turbulent statistics in the transition regime may be considered to alternate between laminar and fully turbulent types, the proportions of which are quantified by a measured intermittency function.

Sepri, Paavo↗

Spectral solution of the incompressible Navier-Stokes equations on the Connection Machine2

The issue of solving the time-dependent incompressible Navier-Stokes equations on the Connection Machine 2 is addressed, for the problem of transition to turbulence of the steady flow in a channel. The spectral algorithm used serially requires O(N4) operations when solving the equations on an N x N x N grid; using the massive parallelism of the CM, it becomes an O(N2) problem. Preliminary timings of the code, written in LISP, are included and compared with a corresponding code optimized for the Cray-2 for a 128 x 128 x 101 grid.

Tomboulian, Sherryl↗