Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “coding productivity”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Comparative study of MacCormack and TVD MacCormack schemes for three-dimensional separation at wing/body junctions in supersonic flows

A robust, discontinuity-resolving TVD MacCormack scheme containing no dependent parameters requiring adjustment is presently used to investigate the 3D separation of wing/body junction flows at supersonic speeds. Many production codes employing MacCormack schemes can be adapted to use this method. A numerical simulation of laminar supersonic junction flow is found to yield improved separation location predictions, as well as the axial velocity profiles in the separated flow region.

Lakshmanan, Balakrishnan↗

Algorithms for parallel flow solvers on message passing architectures

The purpose of this project has been to identify and test suitable technologies for implementation of fluid flow solvers -- possibly coupled with structures and heat equation solvers -- on MIMD parallel computers. In the course of this investigation much attention has been paid to efficient domain decomposition strategies for ADI-type algorithms. Multi-partitioning derives its efficiency from the assignment of several blocks of grid points to each processor in the parallel computer. A coarse-grain parallelism is obtained, and a near-perfect load balance results. In uni-partitioning every processor receives responsibility for exactly one block of grid points instead of several. This necessitates fine-grain pipelined program execution in order to obtain a reasonable load balance. Although fine-grain parallelism is less desirable on many systems, especially high-latency networks of workstations, uni-partition methods are still in wide use in production codes for flow problems. Consequently, it remains important to achieve good efficiency with this technique that has essentially been superseded by multi-partitioning for parallel ADI-type algorithms. Another reason for the concentration on improving the performance of pipeline methods is their applicability in other types of flow solver kernels with stronger implied data dependence. Analytical expressions can be derived for the size of the dynamic load imbalance incurred in traditional pipelines. From these it can be determined what is the optimal first-processor retardation that leads to the shortest total completion time for the pipeline process. Theoretical predictions of pipeline performance with and without optimization match experimental observations on the iPSC/860 very well. Analysis of pipeline performance also highlights the effect of uncareful grid partitioning in flow solvers that employ pipeline algorithms. If grid blocks at boundaries are not at least as large in the wall-normal direction as those immediately adjacent to them, then the first processor in the pipeline will receive a computational load that is less than that of subsequent processors, magnifying the pipeline slowdown effect. Extra compensation is needed for grid boundary effects, even if all grid blocks are equally sized.

Vanderwijngaart, Rob F.↗

External Boundary Conditions for Three-Dimensional Problems of Computational Aerodynamics

We consider an unbounded steady-state flow of viscous fluid over a three-dimensional finite body or configuration of bodies. For the purpose of solving this flow problem numerically, we discretize the governing equations (Navier-Stokes) on a finite-difference grid. The grid obviously cannot stretch from the body up to infinity, because the number of the discrete variables in that case would not be finite. Therefore, prior to the discretization we truncate the original unbounded flow domain by introducing some artificial computational boundary at a finite distance of the body. Typically, the artificial boundary is introduced in a natural way as the external boundary of the domain covered by the grid. The flow problem formulated only on the finite computational domain rather than on the original infinite domain is clearly subdefinite unless some artificial boundary conditions (ABC's) are specified at the external computational boundary. Similarly, the discretized flow problem is subdefinite (i.e., lacks equations with respect to unknowns) unless a special closing procedure is implemented at this artificial boundary. The closing procedure in the discrete case is called the ABC's as well. In this paper, we present an innovative approach to constructing highly accurate ABC's for three-dimensional flow computations. The approach extends our previous technique developed for the two-dimensional case; it employs the finite-difference counterparts to Calderon's pseudodifferential boundary projections calculated in the framework of the difference potentials method (DPM) by Ryaben'kii. The resulting ABC's appear spatially nonlocal but particularly easy to implement along with the existing solvers. The new boundary conditions have been successfully combined with the NASA-developed production code TLNS3D and used for the analysis of wing-shaped configurations in subsonic (including incompressible limit) and transonic flow regimes. As demonstrated by the computational experiments and comparisons with the standard (local) methods, the DPM-based ABC's allow one to greatly reduce the size of the computational domain while still maintaining high accuracy of the numerical solution. Moreover, they may provide for a noticeable increase of the convergence rate of multigrid iterations.

Tsynkov, Semyon V.↗

Robust Multigrid Smoothers for Three Dimensional Elliptic Equations with Strong Anisotropies

We discuss the behavior of several plane relaxation methods as multigrid smoothers for the solution of a discrete anisotropic elliptic model problem on cell-centered grids. The methods compared are plane Jacobi with damping, plane Jacobi with partial damping, plane Gauss-Seidel, plane zebra Gauss-Seidel, and line Gauss-Seidel. Based on numerical experiments and local mode analysis, we compare the smoothing factor of the different methods in the presence of strong anisotropies. A four-color Gauss-Seidel method is found to have the best numerical and architectural properties of the methods considered in the present work. Although alternating direction plane relaxation schemes are simpler and more robust than other approaches, they are not currently used in industrial and production codes because they require the solution of a two-dimensional problem for each plane in each direction. We verify the theoretical predictions of Thole and Trottenberg that an exact solution of each plane is not necessary and that a single two-dimensional multigrid cycle gives the same result as an exact solution, in much less execution time. Parallelization of the two-dimensional multigrid cycles, the kernel of the three-dimensional implicit solver, is also discussed. Alternating-plane smoothers are found to be highly efficient multigrid smoothers for anisotropic elliptic problems.

Llorente, Ignacio M.↗

MLP: A Parallel Programming Alternative to MPI for New Shared Memory Parallel Systems

Recent developments at the NASA AMES Research Center's NAS Division have demonstrated that the new generation of NUMA based Symmetric Multi-Processing systems (SMPs), such as the Silicon Graphics Origin 2000, can successfully execute legacy vector oriented CFD production codes at sustained rates far exceeding processing rates possible on dedicated 16 CPU Cray C90 systems. This high level of performance is achieved via shared memory based Multi-Level Parallelism (MLP). This programming approach, developed at NAS and outlined below, is distinct from the message passing paradigm of MPI. It offers parallelism at both the fine and coarse grained level, with communication latencies that are approximately 50-100 times lower than typical MPI implementations on the same platform. Such latency reductions offer the promise of performance scaling to very large CPU counts. The method draws on, but is also distinct from, the newly defined OpenMP specification, which uses compiler directives to support a limited subset of multi-level parallel operations. The NAS MLP method is general, and applicable to a large class of NASA CFD codes.

Taft, James R.↗

Diffuse Galactic Continuum Gamma Rays. A Model Compatible with EGRET Data and Cosmic-ray Measurements

We present a study of the compatibility of some current models of the diffuse Galactic continuum gamma-rays with EGRET data. A set of regions sampling the whole sky is chosen to provide a comprehensive range of tests. The range of EGRET data used is extended to 100 GeV. The models are computed with our GALPROP cosmic-ray propagation and gamma-ray production code. We confirm that the "conventional model" based on the locally observed electron and nucleon spectra is inadequate, for all sky regions. A conventional model plus hard sources in the inner Galaxy is also inadequate, since this cannot explain the GeV excess away from the Galactic plane. Models with a hard electron injection spectrum are inconsistent with the local spectrum even considering the expected fluctuations; they are also inconsistent with the EGRET data above 10 GeV. We present a new model which fits the spectrum in all sky regions adequately. Secondary antiproton data were used to fix the Galactic average proton spectrum, while the electron spectrum is adjusted using the spectrum of diffuse emission it- self. The derived electron and proton spectra are compatible with those measured locally considering fluctuations due to energy losses, propagation, or possibly de- tails of Galactic structure. This model requires a much less dramatic variation in the electron spectrum than models with a hard electron injection spectrum, and moreover it fits the y-ray spectrum better and to the highest EGRET energies. It gives a good representation of the latitude distribution of the y-ray emission from the plane to the poles, and of the longitude distribution. We show that secondary positrons and electrons make an essential contribution to Galactic diffuse y-ray emission.

Strong, Andrew W.↗

Object-Oriented/Data-Oriented Design of a Direct Simulation Monte Carlo Algorithm

Over the past decade, there has been much progress towards improved phenomenological modeling and algorithmic updates for the direct simulation Monte Carlo (DSMC) method, which provides a probabilistic physical simulation of gas Rows. These improvements have largely been based on the work of the originator of the DSMC method, Graeme Bird. Of primary importance are improved chemistry, internal energy, and physics modeling and a reduction in time to solution. These allow for an expanded range of possible solutions In altitude and velocity space. NASA's current production code, the DSMC Analysis Code (DAC), is well-established and based on Bird's 1994 algorithms written in Fortran 77 and has proven difficult to upgrade. A new DSMC code is being developed in the C++ programming language using object-oriented and data-oriented design paradigms to facilitate the inclusion of the recent improvements and future development activities. The development efforts on the new code, the Multiphysics Algorithm with Particles (MAP), are described, and performance comparisons are made with DAC.

Liechty, Derek S.↗

Photoelectrons and Solar Ionizing Radiation at Mars: Predictions Versus MAVEN Observations

Understanding the evolution of the Martian atmosphere requires knowledge of processes transforming solar irradiance into thermal energy well enough to model them accurately. Here we compare Martian photoelectron energy spectra measured at periaps is by Mars Atmosphere and Volatile Evolution MissioN (MAVEN) with calculations made using three photoelectron production codes and three solar irradiance models as well as modeled and measured CO2 densities. We restricted our comparisons to regions where the contribution from solar wind electrons and ions were negligible. The two intervals examined on 19 October 2014 have different observed incident solar irradiance spectra. In spite of the differences in photoionization cross sections and irradiance spectra used, we find the agreement between models to be within the combined uncertainties associated with the observations from the MAVEN neutral density, electron flux, and solar irradiance instruments.

solar ionizing↗

Production Level CFD Code Acceleration for Hybrid Many-Core Architectures

In this work, a novel graphics processing unit (GPU) distributed sharing model for hybrid many-core architectures is introduced and employed in the acceleration of a production-level computational fluid dynamics (CFD) code. The latest generation graphics hardware allows multiple processor cores to simultaneously share a single GPU through concurrent kernel execution. This feature has allowed the NASA FUN3D code to be accelerated in parallel with up to four processor cores sharing a single GPU. For codes to scale and fully use resources on these and the next generation machines, codes will need to employ some type of GPU sharing model, as presented in this work. Findings include the effects of GPU sharing on overall performance. A discussion of the inherent challenges that parallel unstructured CFD codes face in accelerator-based computing environments is included, with considerations for future generation architectures. This work was completed by the author in August 2010, and reflects the analysis and results of the time.

Duffy, Austen C.↗

The Collection 6 'dark-target' MODIS Aerosol Products

Aerosol retrieval algorithms are applied to Moderate resolution Imaging Spectroradiometer (MODIS) sensors on both Terra and Aqua, creating two streams of decade-plus aerosol information. Products of aerosol optical depth (AOD) and aerosol size are used for many applications, but the primary concern is that these global products are comprehensive and consistent enough for use in climate studies. One of our major customers is the international modeling comparison study known as AEROCOM, which relies on the MODIS data as a benchmark. In order to keep up with the needs of AEROCOM and other MODIS data users, while utilizing new science and tools, we have improved the algorithms and products. The code, and the associated products, will be known as Collection 6 (C6). While not a major overhaul from the previous Collection 5 (C5) version, there are enough changes that there are significant impacts to the products and their interpretation. In its entirety, the C6 algorithm is comprised of three sub-algorithms for retrieving aerosol properties over different surfaces: These include the dark-target DT algorithms to retrieve over (1) ocean and (2) vegetated-dark-soiled land, plus the (3) Deep Blue (DB) algorithm, originally developed to retrieve over desert-arid land. Focusing on the two DT algorithms, we have updated assumptions for central wavelengths, Rayleigh optical depths and gas (H2O, O3, CO2, etc.) absorption corrections, while relaxing the solar zenith angle limit (up to 84) to increase pole-ward coverage. For DT-land, we have updated the cloud mask to allow heavy smoke retrievals, fine-tuned the assignments for aerosol type as function of season location, corrected bugs in the Quality Assurance (QA) logic, and added diagnostic parameters such as topographic altitude. For DT-ocean, improvements include a revised cloud mask for thin-cirrus detection, inclusion of wind speed dependence in the retrieval, updates to logic of QA Confidence flag (QAC) assignment, and additions of important diagnostic information. At the same time as we have introduced algorithm changes, we have also accounted for upstream changes including: new instrument calibration, revised land-sea masking, and changed cloud masking. Upstream changes also impact the coverage and global statistics of the retrieved AOD. Although our responsibility is to the DT code and products, we have also added a product that merges DT and DB product over semi-arid land surfaces to provide a more gap-free dataset, primarily for visualization purposes. Preliminary validation shows that compared to surface-based sunphotometer data, the C6, Level 2 (along swath) DT-products compare at least as well as those from C5. C6 will include new diagnostic information about clouds in the aerosol field, including an aerosol cloud mask at 500 m resolution, and calculations of the distance to the nearest cloud from clear pixels. Finally, we have revised the strategy for aggregating and averaging the Level 2 (swath) data to become Level 3 (gridded) data. All together, the changes to the DT algorithms will result in reduced global AOD (by 0.02) over ocean and increased AOD (by 0.02) over land, along with changes in spatial coverage. Changes in calibration will have more impact to Terras time series, especially over land. This will result in a significant reduction in artificial differences in the Terra and Aqua datasets, and will stabilize the MODIS data as a target for AEROCOM studie

Aerosol retrieval algorithms↗

LaRIS: Targeting Portability and Productivity for LAPACK Codes on Extreme Heterogeneous Systems by Using IRIS

In keeping with the trend of heterogeneity in high-performance computing, hardware manufacturers and vendors are developing new architectures and associated software stacks (e.g., libraries) to harness the best possible performance from commonly used kernels (e.g., linear algebra kernels). However, kernels tuned for one architecture are not portable to others. Moreover, the coexistence of different architectures in a single node makes orchestration difficult. To address these challenges, we introduce LaRIS, a portable framework for LAPACK functionalities. LaRIS ensures a separation between linear algebra algorithms and vendor-library kernels by using the IRIS run time and IRIS-BLAS library. Such abstraction at the algorithm level makes the implementation completely agnostic to the vendor library and architecture. LaRIS uses the IRIS run time to dynamically select the vendor-library kernel and suitable processor architecture at run time. Through LU factorization, we demonstrate that LaRIS can fully utilize different heterogeneous systems by launching and orchestrating different vendor-library kernels without any change in the source code.

Monil, M. A. H.↗

Understanding power and energy utilization in large scale production physics simulation codes

Power is an often-cited reason for the move to advanced architectures on the path to Exascale computing. Here, this is due to practical considerations related to delivering enough power to successfully site and operate these machines, as well as concerns about energy usage while running large simulations. Since obtaining accurate power measurements can be challenging, it may be tempting to use the processor thermal design power (TDP) as a surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advanced technology systems at Lawrence Livermore and Sandia National Labs, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale Lawrence Livermore simulation codes are significantly more efficient than a simple processor TDP model might suggest.

HPC↗

Understanding Power and Energy Utilization in Large Scale Production Physics Simulation Codes

Power is an often-cited reason for moving to advanced architectures on the path to Exascale computing. This is due to the practical concern of delivering enough power to successfully site and operate these machines, as well as concerns over energy usage while running large simulations. Since accurate power measurements can be difficult to obtain, processor thermal design power (TDP) is a possible surrogate due to its simplicity and availability. However, TDP is not indicative of typical power usage while running simulations. Using commodity and advance technology systems at Lawrence Livermore National Laboratory (LLNL) and Sandia National Laboratory, we performed a series of experiments to measure power and energy usage in running simulation codes. These experiments indicate that large scale LLNL simulation codes are significantly more efficient than a simple processor TDP model might suggest.

97 MATHEMATICS AND COMPUTING↗

Secondary gamma-ray production in a coded aperture mask

The application of the coded aperture mask to high energy gamma-ray astronomy will provide the capability of locating a cosmic gamma-ray point source with a precision of a few arc-minutes above 20 MeV. Recent tests using a mask in conjunction with drift chamber detectors have shown that the expected point spread function is achieved over an acceptance cone of 25 deg. A telescope employing this technique differs from a conventional telescope only in that the presence of the mask modifies the radiation field in the vicinity of the detection plane. In addition to reducing the primary photon flux incident on the detector by absorption in the mask elements, the mask will also be a secondary radiator of gamma-rays. The various background components in a CAMTRAC (Coded Aperture Mask Track Chamber) telescope are considered. Monte-Carlo calculations are compared with recent measurements obtained using a prototype instrument in a tagged photon beam line.

Owens, A.↗

NASA Technologies for Product Identification

Since 1975 bar codes on products at the retail counter have been accepted as the standard for entering product identity for price determination. Since the beginning of the 21st century, the Data Matrix symbol has become accepted as the bar code format that is marked directly on a part, assembly or product that is durable enough to identify that item for its lifetime. NASA began the studies for direct part marking Data Matrix symbols on parts during the Return to Flight activities after the Challenger Accident. Over the 20 year period that has elapsed since Challenger, a mountain of studies, analyses and focused problem solutions developed by and for NASA have brought about world changing results. NASA Technical Standard 6002 and NASA Handbook 6003 for Direct Part Marking Data Matrix Symbols on Aerospace Parts have formed the basis for most other standards on part marking internationally. NASA and its commercial partners have developed numerous products and methods that addressed the difficulties of collecting part identification in aerospace operations. These products enabled the marking of Data Matrix symbols in virtually every situation and the reading of symbols at great distances, severe angles, under paint and in the dark without a light. Even unmarkable delicate parts now have a process to apply a chemical mixture called NanocodesTM that can be converted to a Data Matrix. The accompanying intellectual property is protected by 10 patents, several of which are licensed. Direct marking Data Matrix on NASA parts virtually eliminates data entry errors and the number of parts that go through their life cycle unmarked, two major threats to sound configuration management and flight safety. NASA is said to only have people and stuff with information connecting them. Data Matrix is one of the most significant improvements since Challenger to the safety and reliability of that connection. This presentation highlights the accomplishments of NASA in its efforts to develop technologies for automatic identification, its efforts to implement them and its vision on their role in space.

Schramm, Fred, Jr.↗

Development of a model and computer code to describe solar grade silicon production processes

Two computer codes were developed for describing flow reactors in which high purity, solar grade silicon is produced via reduction of gaseous silicon halides. The first is the CHEMPART code, an axisymmetric, marching code which treats two phase flows with models describing detailed gas-phase chemical kinetics, particle formation, and particle growth. It can be used to described flow reactors in which reactants, mix, react, and form a particulate phase. Detailed radial gas-phase composition, temperature, velocity, and particle size distribution profiles are computed. Also, deposition of heat, momentum, and mass (either particulate or vapor) on reactor walls is described. The second code is a modified version of the GENMIX boundary layer code which is used to compute rates of heat, momentum, and mass transfer to the reactor walls. This code lacks the detailed chemical kinetics and particle handling features of the CHEMPART code but has the virtue of running much more rapidly than CHEMPART, while treating the phenomena occurring in the boundary layer in more detail.

Gould, R. K.↗

Performance of the Dot Product Function in Radiative Transfer Code SORD

The successive orders of scattering radiative transfer (RT) codes frequently call the scalar (dot) product function. In this paper, we study performance of some implementations of the dot product in the RT code SORD using 50 scenarios for light scattering in the atmosphere-surface system. In the dot product function, we use the unrolled loops technique with different unrolling factor. We also considered the intrinsic Fortran functions. We show results for two machines: ifort compiler under Windows, and pgf90 under Linux. Intrinsic DOT_PRODUCT function showed best performance for the ifort. For the pgf90, the dot product implemented with unrolling factor 4 was the fastest. The RT code SORD together with the interface that runs all the mentioned tests are publicly available from ftp:maiac.gsfc.nasa.govpubskorkinSORD_IP_16B (current release) or by email request from the corresponding (first) author.

polarized radiative transfer↗

An Automatic Code Synthesizer

This paper describes the design and architecture fo an automatic code sysnthesizer we call the ACG. The input to the ACG is a list of facts that must hold for the generated code, along with domain-specific knowledge (design rules and patterns).

programming programming code programming productiv↗