Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43

Genarris 2.0: A Random Structure Generator for Molecular Crystals

Genarris is an open source Python package for generating random molecular crystal structures with physical constraints for seeding crystal structure prediction algorithms and training machine learning models. Here we present a new version of the code, containing several major improvements. A MPI-based parallelization scheme has been implemented, which facilitates the seamless sequential execution of user-defined workflows. A new method for estimating the unit cell volume based on the single molecule structure has been developed using a machine-learned model trained on experimental structures. A new algorithm has been implemented for generating crystal structures with molecules occupying special Wyckoff positions. A new hierarchical structure check procedure has been developed to detect unphysical close contacts efficiently and accurately. New intermolecular distance settings have been implemented for strong hydrogen bonds. To demonstrate these new features, we study two specific cases: benzene and glycine. Genarris finds the experimental structures of the two polymorphs of benzene and the three polymorphs of glycine. Program summary Program Title: Genarris 2.0 Program Files doi: http://dx.doi.org/10.17632/grx6mz4pjn.1 Licensing provisions: BSD-3 Clause Programming language: Python, C External routines/libraries: Spglib, ASE, pymatgen, SciPy, mpi4py, scikit-learn, PyTorch, FHI-aims. Nature of problem: Molecular crystal structure prediction. Solution method: Genarris 2.0 generates molecular crystal structures over the 230 space groups, on general and special Wyckoff positions, using physical constraints. Down-sampling of the generated structures may be performed subsequently, based on molecular crystal packing descriptors and an unsupervised machine learning algorithm. Lastly, ab initio structure relaxation may be performed for the final pool. Depending on the user-defined workflow implemented, Genarris may be used to generate diverse molecular crystal datasets to seed evolutionary algorithms or to train machine learning algorithms or as a standalone crystal structure prediction method. Restrictions: For crystal structure generation, the molecule of interest must be semi-rigid with no bond rotational degrees of freedom. Unusual features: Genarris 2.0 is a highly distributed program, making use of MPI for Python parallelization. The user has the ability to design and implement workflows by executing a user-defined list of procedures. Genarris 2.0 offers new features including a machine learning model for estimating the molecular volume in the solid state from the single molecule structure, structure generation in special Wyckoff positions of space groups, hierarchical structure checks including rigorous treatment of non-orthogonal structures, and clustering and down-selection workflows combining first principles simulations with machine learning. (C) 2020 Elsevier B.V. All rights reserved.

Crystal structure prediction↗

Near-continuum, hypersonic oxygen flow over a double cone simulated by direct simulation Monte Carlo informed from quantum chemistry

A large-scale, fully resolved direct simulation Monte Carlo (DSMC) computation of a non-equilibrium, reactive flow of pure oxygen over a double cone is presented. Under the simulated near-continuum conditions, the computational demands are shown to be significant because of the wide range of length scales that must be resolved. Therefore, robust grid adaption capabilities and efficient parallelization of the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code that is utilized in this work are essential. The thermochemical and transport collision models were selected for efficiency and simplicity. First-principles data, obtained from the highly accurate direct molecular simulation method, were used to inform the collision models’ parameters. Importantly, because SPARTA implements molecular collision models using collision-specific energies, the resulting macroscopic relaxation rates were evaluated a posteriori via zero-dimensional heat bath simulations. The comparisons of surface properties, namely heat flux and pressure, show very close agreement with previous computational fluid dynamics (CFD) results. Differences with the measurements were found to be similar to the CFD simulations. The unresolved discrepancy with the measurements could be due to inconsistent free stream conditions with the actual experimental data or missing physical phenomena altogether, for example atomic and molecular oxygen electronically excited states, three-dimensional effects, or more complex gas–surface interactions. As shown in this work, the advantages of obtaining a DSMC particle solution for these flows reside in the method's ability to be directly informed from first principles and to seamlessly describe internal energy non-equilibrium for all modes. With the advent of exascale computing and beyond, particle methods will be an increasingly important tool to verify the validity of physical assumptions in reduced-order models via fully resolved, experimental-scale simulations, down to the level of molecular-level distributions.

Mechanics↗

PETSc TSAdjoint: A Discrete Adjoint ODE Solver for First-Order and Second-Order Sensitivity Analysis

Here, we present a new software system PETSc TSAdjoint for first-order and second order adjoint sensitivity analysis of time-dependent nonlinear differential equations. The derivative calculation in PETSc TSAdjoint is essentially a high-level algorithmic differentiation process. The adjoint models are derived by differentiating the timestepping algorithms and implementing them based on the parallel infrastructure in PETSc. Full differentiation of the library code, including MPI routines, is avoided, and users do not need to derive their own adjoint models for their specific applications. PETSc TSAdjoint can compute the first-order derivative, that is, the gradient of a scalar functional, and the Hessian-vector product, which carries second-order derivative information, while requiring minimal input (a few callbacks) from the users. The adjoint model employs optimal checkpointing schemes in a manner that is transparent to users. Finally, usability, efficiency, and scalability are demonstrated through examples from a variety of applications.

79 ASTRONOMY AND ASTROPHYSICS↗

GPU Profiling and Optimizing xRAGE (Final Report)

Our project’s objective is to increase the efficiency of GPU-enabled kernels in xRAGE. To do so, we conduct GPU profiling with NSight Systems on xRAGE tests unsplit_sod_1d and unsplit_sedov_2d to identify bottlenecks and understand the behavior of the GPU during code execution. Next, we analyze these generated GPU profiles to locate the lines of code whose optimization have the most potential for improving runtime. We replicate the structure of the code in smaller test problems that are easier to understand, edit, and run quickly. Within these test problems, we implement two different methods of improving performance: transformation of nested loops into a single MDRangePolicy and hierarchical parallelization using teams of threads. Both methods show speedups in the test code, and after transferring them to xRAGE, they both show up to 30x speedups on various computing platforms. Profiling the edited versions of xRAGE reveals that the GPU successfully executed the bottlenecks with greater efficiency

97 MATHEMATICS AND COMPUTING↗

Substructure analysis using NICE/SPAR and applications of force to linear and nonlinear structures

Parallel computing studies are presented for a variety of structural analysis problems. Included are the substructure planar analysis of rectangular panels with and without a hole, the static analysis of space mast, using NICE/SPAR and FORCE, and substructure analysis of plane rigid-jointed frames using FORCE. The computations are carried out on the Flex/32 MultiComputer using one to eighteen processors. The NICE/SPAR runstream samples are documented for the panel problem. For the substructure analysis of plane frames, a computer program is developed to demonstrate the effectiveness of a substructuring technique when FORCE is enforced. Ongoing research activities for an elasto-plastic stability analysis problem using FORCE, and stability analysis of the focus problem using NICE/SPAR are briefly summarized. Speedup curves for the panel, the mast, and the frame problems provide a basic understanding of the effectiveness of parallel computing procedures utilized or developed, within the domain of the parameters considered. Although the speedup curves obtained exhibit various levels of computational efficiency, they clearly demonstrate the excellent promise which parallel computing holds for the structural analysis problem. Source code is given for the elasto-plastic stability problem and the FORCE program.

Razzaq, Zia↗

Implementation and analysis of a Navier-Stokes algorithm on parallel computers

The results of the implementation of a Navier-Stokes algorithm on three parallel/vector computers are presented. The object of this research is to determine how well, or poorly, a single numerical algorithm would map onto three different architectures. The algorithm is a compact difference scheme for the solution of the incompressible, two-dimensional, time-dependent Navier-Stokes equations. The computers were chosen so as to encompass a variety of architectures. They are the following: the MPP, an SIMD machine with 16K bit serial processors; Flex/32, an MIMD machine with 20 processors; and Cray/2. The implementation of the algorithm is discussed in relation to these architectures and measures of the performance on each machine are given. The basic comparison is among SIMD instruction parallelism on the MPP, MIMD process parallelism on the Flex/32, and vectorization of a serial code on the Cray/2. Simple performance models are used to describe the performance. These models highlight the bottlenecks and limiting factors for this algorithm on these architectures. Finally, conclusions are presented.

Fatoohi, Raad A.↗

Parallel PAB3D: Experiences with a Prototype in MPI

PAB3D is a three-dimensional Navier Stokes solver that has gained acceptance in the research and industrial communities. It takes as computational domain, a set disjoint blocks covering the physical domain. This is the first report on the implementation of PAB3D using the Message Passing Interface (MPI), a standard for parallel processing. We discuss briefly the characteristics of tile code and define a prototype for testing. The principal data structure used for communication is derived from preprocessing "patching". We describe a simple interface (COMMSYS) for MPI communication, and some general techniques likely to be encountered when working on problems of this nature. Last, we identify levels of improvement from the current version and outline future work.

Guerinoni, Fabio↗

Analysis of Physical Properties of Dust Suspended in the Mars Atmosphere

Methods for iteratively determining the infrared optical constants for dust suspended in the Mars atmosphere are described. High quality spectra for wavenumbers from 200 to 2000 1/cm were obtained over a wide range of view angles by the Mariner 9 spacecraft, when it observed a global Martian dust storm in 1971-2. In this research, theoretical spectra of the emergent intensity from Martian dust clouds are generated using a 2-stream source-function radiative transfer code. The code computes the radiation field in a plane-parallel, vertically homogeneous, multiply scattering atmosphere. Calculated intensity spectra are compared with the actual spacecraft data to iteratively retrieve the optical properties and opacity of the dust, as well as the surface temperature of Mars at the time and location of each measurement. Many different particle size distributions a-re investigated to determine the best fit to the data. The particles are assumed spherical and the temperature profile was obtained from the CO2 band shape. Given a reasonable initial guess for the indices of refraction, the searches converge in a well-behaved fashion, producing a fit with error of less than 1.2 K (rms) to the observed brightness spectra. The particle size distribution corresponding to the best fit was a lognormal distribution with a mean particle radius, r(sub m) 0.66 pm, and variance, omega(sup 2) = 0.412 (r(sub eff) = 1.85 microns, v(sub eff) =.51), in close agreement with the size distribution found to be the best fit in the visible wavelengths in recent studies. The optical properties and the associated single scattering properties are shown to be a significant improvement over those used in existing models by demonstrating the effects of the new properties both on heating rates of the Mars atmosphere and in example spectral retrieval of surface characteristics from emission spectra.

Snook, Kelly↗

Modeling of High Speed Reacting Flows: Established Practices and Future Challenges

Computational fluid dynamics (CFD) has proven to be an invaluable tool for the design and analysis of high- speed propulsion devices. Massively parallel computing, together with the maturation of robust CFD codes, has made it possible to perform simulations of complete engine flowpaths. Steady-state Reynolds-Averaged Navier-Stokes simulations are now routinely used in the scramjet engine development cycle to determine optimal fuel injector arrangements, investigate trends noted during testing, and extract various measures of engine efficiency. Unfortunately, the turbulence and combustion models used in these codes have not changed significantly over the past decade. Hence, the CFD practitioner must often rely heavily on existing measurements (at similar flow conditions) to calibrate model coefficients on a case- by-case basis. This paper provides an overview of the modeled equations typically employed by commercial- quality CFD codes for high-speed combustion applications. Careful attention is given to the approximations employed for each of the unclosed terms in the averaged equation set. The salient features (and shortcomings) of common models used to close these terms are covered in detail, and several academic efforts aimed at addressing these shortcomings are discussed.

Baurle, R. A.↗

An Analog Processor for Image Compression

This paper describes a novel analog Vector Array Processor (VAP) that was designed for use in real-time and ultra-low power image compression applications. This custom CMOS processor is based architectually on the Vector Quantization (VQ) algorithm in image coding, and the hardware implementation fully exploits the inherent parallelism built-in the VQ algorithm.

Vector Array Processor VAP analog processor image ↗

Using EIGER for Antenna Design and Analysis

EIGER (Electromagnetic Interactions GenERalized) is a frequency-domain electromagnetics software package that is built upon a flexible framework, designed using object-oriented techniques. The analysis methods used include moment method solutions of integral equations, finite element solutions of partial differential equations, and combinations thereof. The framework design permits new analysis techniques (boundary conditions, Green#s functions, etc.) to be added to the software suite with a sensible effort. The code has been designed to execute (in serial or parallel) on a wide variety of platforms from Intel-based PCs and Unix-based workstations. Recently, new potential integration scheme s that avoid singularity extraction techniques have been added for integral equation analysis. These new integration schemes are required for facilitating the use of higher-order elements and basis functions. Higher-order elements are better able to model geometrical curvature using fewer elements than when using linear elements. Higher-order basis functions are beneficial for simulating structures with rapidly varying fields or currents. Results presented here will demonstrate curren t and future capabilities of EIGER with respect to analysis of installed antenna system performance in support of NASA#s mission of exploration. Examples include antenna coupling within an enclosed environment and antenna analysis on electrically large manned space vehicles.

Champagne, Nathan J.↗

NASA Tech Briefs, April 2009

Topics covered include: Direct-Solve Image-Based Wavefront Sensing; Use of UV Sources for Detection and Identification of Explosives; Using Fluorescent Viruses for Detecting Bacteria in Water; Gradiometer Using Middle Loops as Sensing Elements in a Low-Field SQUID MRI System; Volcano Monitor: Autonomous Triggering of In-Situ Sensors; Wireless Fluid-Level Sensors for Harsh Environments; Interference-Detection Module in a Digital Radar Receiver; Modal Vibration Analysis of Large Castings; Structural/Radiation-Shielding Epoxies; Integrated Multilayer Insulation; Apparatus for Screening Multiple Oxygen-Reduction Catalysts; Determining Aliasing in Isolated Signal Conditioning Modules; Composite Bipolar Plate for Unitized Fuel Cell/Electrolyzer Systems; Spectrum Analyzers Incorporating Tunable WGM Resonators; Quantum-Well Thermophotovoltaic Cells; Bounded-Angle Iterative Decoding of LDPC Codes; Conversion from Tree to Graph Representation of Requirements; Parallel Hybrid Vehicle Optimal Storage System; and Anaerobic Digestion in a Flooded Densified Leachbed.

Source record↗

Limitations of Phased Array Beamforming in Open Rotor Noise Source Imaging

Phased array beamforming results of the F31/A31 historical baseline counter-rotating open rotor blade set were investigated for measurement data taken on the NASA Counter-Rotating Open Rotor Propulsion Rig in the 9- by 15-Foot Low-Speed Wind Tunnel of NASA Glenn Research Center as well as data produced using the LINPROP open rotor tone noise code. The planar microphone array was positioned broadside and parallel to the axis of the open rotor, roughly 2.3 rotor diameters away. The results provide insight as to why the apparent noise sources of the blade passing frequency tones and interaction tones appear at their nominal Mach radii instead of at the actual noise sources, even if those locations are not on the blades. Contour maps corresponding to the sound fields produced by the radiating sound waves, taken from the simulations, are used to illustrate how the interaction patterns of circumferential spinning modes of rotating coherent noise sources interact with the phased array, often giving misleading results, as the apparent sources do not always show where the actual noise sources are located. This suggests that a more sophisticated source model would be required to accurately locate the sources of each tone. The results of this study also have implications with regard to the shielding of open rotor sources by airframe empennages.

Horvath, Csaba↗

A Single Thread to Fortran Coarray Transition Process for the Control Algorithm in the Space Radiation Code HZETRN

Exa-scale computing is the direction by industry and government are going to generate solutions to problems they deem necessary. Computing hardware is being developed to achieve the transition from Peta-scale to Exa-scale with more CPUs (Central Processing Units) that have more cores per CPU and more accelerators (GPGPUs (General Purpose Graphics Processing Units) and MICs (Many Integrated Cores)) per node. To fully utilize the hardware available now and in the future, algorithms must become multi-threaded. There are a few methods to generate multi-threaded software such as MPI (Message Passing Interface) and OpenMP (Multi-Processing) / OpenACC (ACCelerator). This paper concentrates on using Coarray Fortran to convert the Fortran 95 based HZETRN (High Z and Energy TRaNsport) code's control algorithm from a single threaded code to a multithreaded code. The resultant Coarray code was 32.5 times faster (with a theoretical speed-up of 74.5 times) than the single threaded version on the hardware tested, as reliable as the Fortran 95 version, and, as it uses native Fortran, was as maintainable as the Fortran 95 version. The Coarray code can be maintained by the same project engineers and scientists who created the original single threaded code. This transition process can be utilized on a C language based code with a compiler that has the UPC (Universal Parallel C) extensions to C.

Singleterry, Robert C., Jr.↗

Implementation of active sites to capture pitting of oxidizing carbon materials in DSMC.

In this work we demonstrate a newly developed capability to capture pitting of carbon fibers in DSMC simulations, specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code [1]. State-of-the-art reactive surface models in DSMC compute collision dependent carbon consumption rates (usually through desorption of CO) based on a set of surface reactions that has been derived from molecular beam experiments [2]. The reactivity on each carbon surface element is constant in those models, such that the carbon surface recedes uniformly as a result of ablation. However, it is well known that in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale [3]. These defective sites have a much higher reactivity than the average sites (2-3 orders of magnitude) and are first to react during ablation leading to its removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these ”active” sites (as shown in Fig. 1). In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various reaction mechanisms such as adsorption, desorption, Eley-Rideal (ER) and Langmuir-Hinshelwood (LH) mechanisms [4]. Within this framework, we have implemented the capability of a single surface having multiple site sets with different reactivities. Using this feature, we can simulate the presence of active sites on carbon surfaces, whose reactivity is much greater than an average site as a result of defects. We have implemented the active site fraction as a property of surface elements within SPARTA, which is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface, and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we are able to capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗

Implementation of Active Sites in DSMC to Capture Pitting of Oxidizing Carbon Materials

In this work we demonstrate a newly developed capability to capture pitting of carbon fibers in DSMC simulations, specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code. State-of-the-art reactive surface models in DSMC compute collision dependent carbon consumption rates (usually through desorption of CO) based on a set of surface reactions that has been derived from molecular beam experiments. The reactivity on each carbon surface element is constant in those models, such that the carbon surface recedes uniformly as a result of ablation. However, it is well known that in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale. These defective sites have a much higher reactivity than the average sites (2-3 orders of magnitude) and are first to react during ablation leading to its removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these ”active” sites. In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various reaction mechanisms such as adsorption, desorption, Eley-Rideal (ER) and Langmuir-Hinshelwood (LH) mechanisms. Within this framework, we have implemented the capability of a single surface having multiple site sets with different reactivities. Using this feature, we can simulate the presence of active sites on carbon surfaces, whose reactivity is much greater than an average site as a result of defects. We have implemented the active site fraction as a property of surface elements within SPARTA, which is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface, and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we are able to capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗

Simulation of etch pit formation through active sites in carbon fiber micro-structures

Erosion of carbon due to oxidation does not occur uniformly but through localized etch pit formation as a result of active surface sites. In this work we demonstrate a newly developed capability to capture pitting of carbon fiber microstructures such as FiberForm, which is commonly used as the base material for NASA’s spacecraft ablative thermal protection systems (TPS). The simulations are performed at the meso-scale in order to capture the pit formation and growth using direct simulation Monte Carlo (DSMC), specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code. Legacy and latest models both assume uniform reactivity of carbon surface sites with oxygen even at the meso-scale level. However, in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale. The defective nature of these sites enhances their reactivity with atmospheric gases compared to the non-defective sites (2-3 orders of magnitude) and are termed as “active sites”. Thus, these sites tend to be the first to react and eventually get removed through the formation of gasses such as CO, CO2, and CN. Their removal results in all the neighboring atoms becoming defective, thus leading to chain reaction of localized carbon removal and formation of etch pits. Capturing the formation of pits during the ablation simulation of carbon micro-structures is critical to predicting their structural failure. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various reaction mechanisms such as adsorption, desorption, Eley-Rideal (ER) and Langmuir-Hinshelwood (LH) mechanisms. We have implemented the capability of a single surface having multiple site sets with different reactivities within this framework. We have used this feature to model the presence of active sites on carbon surfaces, whose reactivity is orders of magnitude higher than that of the passive sites due to the presence of defects. The active site fraction is a property of surface elements within SPARTA and is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we can capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗

Simulation of Etch Pit Formation in DSMC Through Active Sites in Carbon Fiber Micro-Structures

Erosion of carbon due to oxidation does not occur uniformly but through localized etch pit formation because of active surface sites. In this work we demonstrate a newly developed capability to capture pitting of carbon fiber microstructures such as FiberForm, which is commonly used as the base material for NASA’s spacecraft ablative thermal protection systems (TPS). The simulations are performed at the meso-scale in order to capture the pit formation and growth using direct simulation Monte Carlo (DSMC), specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code. Legacy and latest models both assume uniform reactivity of carbon surface sites with oxygen even at the meso-scale level. However, in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale. The defective nature of these sites enhances their reactivity with atmospheric gases compared to the non-defective sites (2-3 orders of magnitude) and are termed as “active sites”. Thus, these sites tend to be the first to react and eventually get removed through the formation of gases such as CO, CO2, and CN. Their removal results in all the neighboring atoms becoming defective, thus leading to chain reaction of localized carbon removal and formation of etch pits. Capturing the formation of pits during the ablation simulation of carbon micro-structures is critical to predicting their structural failure. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various detailed surface reaction mechanisms. We have implemented the capability of a single surface having multiple site sets with different reactivities within this framework. We have used this feature to model the presence of active sites on carbon surfaces, whose reactivity is orders of magnitude higher than that of the passive sites due to the presence of defects. The active site fraction is a property of surface elements within SPARTA and is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we can capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗