Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30

Integrating machine learning interatomic potentials with hybrid reverse Monte Carlo structure refinements in RMCProfile

Structure refinement with reverse Monte Carlo (RMC) is a powerful tool for interpreting experimental diffraction data. To ensure that the under-constrained RMC algorithm yields reasonable results, the hybrid RMC approach applies interatomic potentials to obtain solutions that are both physically sensible and in agreement with experiment. To expand the range of materials that can be studied with hybrid RMC, we have implemented a new interatomic potential constraint in RMCProfile that grants flexibility to apply potentials supported by the Large-scale Atomic/Molecular Massively Parallel Simulator ( LAMMPS ) molecular dynamics code. This includes machine learning interatomic potentials, which provide a pathway to applying hybrid RMC to materials without currently available interatomic potentials. To this end, we present a methodology to use RMC to train machine learning interatomic potentials for hybrid RMC applications.

Cuillier, Paul↗

A GPU-Accelerated Population Generation, Sorting, and Mutation Kernel for an Optimization-Based Causal Inference Model

We develop a GPU-accelerated machine learning generative adversarial network model that can be used with observational data for the purpose of constructing causal inferences. The theoretical basis of our machine learning model is novel and is conceptualized to be operable and scalable for high performance computing platforms. Our GPU-accelerated code enables large-scale parallelization of the computation within a common and accessible computing environment. This will expand the reach of our model and empower research in new substantive domains while maintaining the underlying theoretical properties.

Cho, Wendy K. Tam↗

Concurrent Relaxation through Accelerated Deep Learning

CRADL captures performance metrics of machine learning algorithms operating on mesh data from multiphysics codes This proxy application is a tool to explore scalability of inference on HPC platforms, and also gather performance metrics for inference on new machine learning specific hardware. CRADL is designed to give users as fine a control as possible over an inference simulation. Users may select the number of cycles, amount of data, and batch size to pass to the accelerator of choice. Additionally the user may select a number of performance optimization libraries and flags. CRADL comes packaged with a repository of anonymized multi-physics simulation data, as well as a pretrained model for inference. The code allows a user to load their own pre-trained model and data if they wish. The code can operate in multiple parallelization schemes, with performance enhancing options such as half-precision libraries, PyTorch benchmarking, and pinned memory with non-blocking data transfers.

Zieb, KristoferJ.↗

Tacho

SAND2022-7470 O Tacho is a performance portable sparse direct solver for use with various sparse matrix problems which typically arise from finite element methods. The code relies on the Kokkos parallel programming model porting to both CPUs and GPUs. It is also part of the Trilinos library. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kim, Kyungjoo↗

Coupling MOOSE-Wrapped MPACT to BISON

As part of the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, an effort is being made to leverage codes like Virtual Environment for Reactor Analysis (VERA) and Michigan Parallel Characteristics Transport (MPACT) by coupling them with other NEAMS codes. To facilitate coupling with other Multiphysics Objected Oriented Simulation Environment (MOOSE) applications, a MOOSE-wrapped MPACT app is created. The capability of this app, named Trogdor, is demonstrated by coupling it with another MOOSE app, BISON. Several single pin problems were run to test the Trogdor and BISON coupling.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Viterbi algorithm on a hypercube: Concurrent formulation

The similarity between the Fast Fourier Transform and the Viterbi algorithm is exploited to develop a Concurrent Viterbi Algorithm suitable for a multiprocessor system interconnected as a hypercube. The proposed algorithm can efficiently decode large constraint length convolutional codes, using different degrees of parallelism, and is attractive for VLSI implementation.

Pllara, F.↗

The nondeterministic divide

The nondeterministic divide partitions a vector into two non-empty slices by allowing the point of division to be chosen nondeterministically. Support for high-level divide-and-conquer programming provided by the nondeterministic divide is investigated. A diva algorithm is a recursive divide-and-conquer sequential algorithm on one or more vectors of the same range, whose division point for a new pair of recursive calls is chosen nondeterministically before any computation is performed and whose recursive calls are made immediately after the choice of division point; also, access to vector components is only permitted during activations in which the vector parameters have unit length. The notion of diva algorithm is formulated precisely as a diva call, a restricted call on a sequential procedure. Diva calls are proven to be intimately related to associativity. Numerous applications of diva calls are given and strategies are described for translating a diva call into code for a variety of parallel computers. Thus diva algorithms separate logical correctness concerns from implementation concerns.

Charlesworth, Arthur↗

Implementation of the Lanczos eigen-solver for the CSI code on high performance computers

The focus of this research is to implement a Lanczos algorithm for the Control-Structure Integration (CSI) code which can exploit both parallel and vector capabilities provided by modern, high performance computers. A partial restoring orthogonality scheme is also developed and incorporated into the basic Lanczos algorithm. The numerical performance of the proposed parallel-vector Lanczos algorithm is demonstrated by solving for the frequencies and mode shapes of the Phase Zero CSI model. The superior performance of the Lanczos algorithm is illustrated in tabular form.

Nguyen, Duc T.↗

The design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a three-dimensional unstructured grid Euler-solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data structure has been adopted, and a number of optimizations have been devised and implemented in order to accelerate the parallel communication rates. The implementation is carried out by creating a set of software tools, which provide an interface between the parallelization issues and the sequential code, while providing a basis for future automatic run-time compilation support. Large practical unstructured grid problems are solved on the Intel iPSC/860 hypercube and Intel Touchstone Delta machine. The quantitative effect of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of three performance improvement. The overall solution efficiency is compared with that obtained on the CRAY-YMP vector supercomputer.

Das, R.↗

A hydrodynamic approach to cosmology: The mixed dark matter cosmological scenario

We compute the evolution of spatially flat, mixed cold and hot dark matter models containing both baryonic matter and two kinds of dark matter. Hydrodynamics is treated with a highly developed Eulerian hydrodynamic code (see Cen 1992). A standard particle-mesh (PM) code is also used in parallel to calculate the motion of the dark matter components. We adopt the following parameters: h equivalent to (sub 0)/100 km/s Mpc(exp -1) = 0.5, OMEGA(sub C) = 0.3, and OMEGA(sub B) = 0.06, with amplitude of the perturbation spectrum fixed by the Cosmic Background Explorer Satellite (COBE) Dark Matter Radiation (DMR) measurements (Smoot et al. 1992) being sigma (sub 8) = 0.67. Four different boxes are simulated with box sizes of L = (64, 16, 4, 1) h(exp -1) Mpc, respectively, the two small boxes providing good resolution but little valid information due to absence of large-scale power. We use 128(exp 3) approximate 10(exp 6.3) baryonic cells, 128(exp .3) cold dark matter particles, and 2 x 128(exp 3) hot dark matter particles. In addition to the dark matter we follow separately six baryonic species (H, H(+), He, He(+), He(++), e(-)) with allowance for both (nonequilibrium) collisional and radiative ionization in every cell. The background radiation field is also followed in detail with allowance made for both continuum and line processes, to allow nonequilibrium heating and cooling processes to be followed in detail. The mean final Zeldovich-Sunyaev y parameter is estimated to be y Bar = (5.4 + or - 2.7) x 10(exp -7) below currently attainable observations, with a rms fluctuation of approximately delta bar y = (0.6 + or - 3.0) x 10(exp -7) on arcminute scales. The rate of galaxy formation peaks at an even later epoch (z approximate 0.3) than in the standard (OMEGA = 1, sigma sub 8 = 0.67) cold dark matter (CDM) model (z approximate 0.5) and, at a redshift of z = 4, is nearly a factor of 100 lower than for the CDM model with the same value of sigma sub 8. With regard to mass function, the smallest objects are stabilized against collapse by thermal energy: the mass-weighted mass spectrum has a broad peak in the vicinity of M(sub B) = 10(exp 9.5) solar mass with a reasonable fit to the Schechter luminosity function if the ratio of baryon mass to blue light is approximately 4. In addition, one very large PM simulation was made in a box with size (320 h(exp - 1) Mpc) containing 3 x 200(exp 3) = 10(exp 7.4) particles. Utilizing this simulation we find that the model yields a cluster mass function which is about a factor of 4 higher than observed, but a cluster-cluster correlation length marginally lower than observed, but that both are closer to observations than in the (COBE) normalized CDM model. The one-dimensional pairwise velocity dispersion is 605 + or - 8 km/s at 1/h separation, lower than that of the DCM model normalized to COBE, but still significant higher than observations (Davis & Peebles 1983). A plausible velocity bias b(sub v) = 0.8 + or - 0.1 on this scale will reduce but not remove the discrepancy. The velocity auto-correlat ion function has a coherence length of 40/h Mpc, which is somewhat lower than the observed counterpart. In all these respects the model would be improved by decreasing the cold fraction of the dark OMEGA(sub CDM)/ (OMEGA(sub CDM) + OMEGA(sub HDB). But formation of galaxies and clusters of galaxies is much later in this model than in COBE-normalized CDM, perhaps too late. To improve on these constraints a larger ratio of OMEGA(sub CDM)/ (OMEGA(sub CDM) + OMEGA(sub HDM)) is required than the value of 0.67 adopted here. It does not seem possible to find a value for this ratio which would satisfy all tests. Overall, the model is similar both on large and intermediate scales to the standard CDM model normalized to the same value of sigma(sub B), but the problem with regard to late formation of galaxies is more severe in this model than in that CDM model. Adding hot dark matter, significantly improves the ability of the COBE-normalized CDM scenario to fit existing observations, but the model is in fact not as good as the CDM model with the same sigma(sub 8) and is still probably unsatisfactory with regard to several critical tests.

Cen, Renyue↗

Design and implementation of a parallel unstructured Euler solver using software primitives

This paper is concerned with the implementation of a three-dimensional unstructured-grid Euler solver on massively parallel distributed-memory computer architectures. The goal is to minimize solution time by achieving high computational rates with a numerically efficient algorithm. An unstructured multigrid algorithm with an edge-based data structure has been adopted, and a number of optimizations have been devised and implemented to accelerate the parallel computational rates. The implementation is carried out by creating a set of software tools, which provide an interface between the parallelization issues and the sequential code, while providing a basis for future automatic run-time compilation support. Large practical unstructured grid problems are solved on the Intel iPSC/860 hypercube and Intel Touchstone Delta machine. The quantitative effects of the various optimizations are demonstrated, and we show that the combined effect of these optimizations leads to roughly a factor of 3 performance improvement. The overall solution efficiency is compared with that obtained on the Cray Y-MP vector supercomputer.

Das, R.↗

Implementing Multidisciplinary and Multi-Zonal Applications Using MPI

Multidisciplinary and multi-zonal applications are an important class of applications in the area of Computational Aerosciences. In these codes, two or more distinct parallel programs or copies of a single program are utilized to model a single problem. To support such applications, it is common to use a programming model where a program is divided into several single program multiple data stream (SPMD) applications, each of which solves the equations for a single physical discipline or grid zone. These SPMD applications are then bound together to form a single multidisciplinary or multi-zonal program in which the constituent parts communicate via point-to-point message passing routines. Unfortunately, simple message passing models, like Intel's NX library, only allow point-to-point and global communication within a single system-defined partition. This makes implementation of these applications quite difficult, if not impossible. In this report it is shown that the new Message Passing Interface (MPI) standard is a viable portable library for implementing the message passing portion of multidisciplinary applications. Further, with the extension of a portable loader, fully portable multidisciplinary application programs can be developed. Finally, the performance of MPI is compared to that of some native message passing libraries. This comparison shows that MPI can be implemented to deliver performance commensurate with native message libraries.

Fineberg, Samuel A.↗

Full 3D Analysis of the GE90 Turbofan Primary Flowpath

The multistage simulations of the GE90 turbofan primary flowpath components have been performed. The multistage CFD code, APNASA, has been used to analyze the fan, fan OGV and booster, the 10-stage high-pressure compressor and the entire turbine system of the GE90 turbofan engine. The code has two levels of parallel, and for the 18 blade row full turbine simulation has 87.3 percent parallel efficiency with 121 processors on an SGI ORIGIN. Grid generation is accomplished with the multistage Average Passage Grid Generator, APG. Results for each component are shown which compare favorably with test data.

Turner, Mark G.↗

3D Simulations of the Richtmyer-Meshkov Instability with Re-Shock

We present results of inviscid simulations, in three dimensions, of Richtmyer-Meshkov instability for high incident shock Mach number. The growth rate of a single harmonic perturbation is quantified and compared with the results of a 2D calculation. Upon re-shock, the perturbation amplitude undergoes a phase reversal while the mean velocity of the interface is zero. Before re-shock the normalized growth rate of a 2D and 3D interface are nearly the same, but the growth rate after re-shock is significantly larger for the 3D than the 2D case. We also examine the evolution of multiple harmonic perturbations. Computational and parallelization issues of the simulation code will also be briefly discussed. The computations were done on the T3E at Pittsburgh Supercomputing Center.

Meiron, Daniel I.↗

Low Density Parity Check Codes: Bandwidth Efficient Channel Coding

Low Density Parity Check (LDPC) Codes provide near-Shannon Capacity performance for NASA Missions. These codes have high coding rates R=0.82 and 0.875 with moderate code lengths, n=4096 and 8176. Their decoders have inherently parallel structures which allows for high-speed implementation. Two codes based on Euclidean Geometry (EG) were selected for flight ASIC implementation. These codes are cyclic and quasi-cyclic in nature and therefore have a simple encoder structure. This results in power and size benefits. These codes also have a large minimum distance as much as d,,, = 65 giving them powerful error correcting capabilities and error floors less than lo- BER. This paper will present development of the LDPC flight encoder and decoder, its applications and status.

Fong, Wai↗

Using Tabulated Experimental Data to Drive an Orthotropic Elasto-Plastic Three-Dimensional Model for Impact Analysis

An orthotropic elasto-plastic-damage three-dimensional model with tabulated input has been developed to analyze the impact response of composite materials. The theory has been implemented as MAT 213 into a tailored version of LS-DYNA being developed under a joint effort of the FAA and NASA and has the following features: (a) the theory addresses any composite architecture that can be experimentally characterized as an orthotropic material and includes rate and temperature sensitivities, (b) the formulation is applicable for solid as well as shell element implementations and utilizes input data in a tabulated form directly from processed experimental data, (c) deformation and damage mechanics are both accounted for within the material model, (d) failure criteria are established that are functions of strain and damage parameters, and mesh size dependence is included, and (e) the theory can be efficiently implemented into a commercial code for both sequential and parallel executions. The salient features of the theory as implemented in LS-DYNA are illustrated using a widely used composite - the T800S/3900-2B[P2352W-19] BMS8-276 Rev-H-Unitape fiber/resin unidirectional composite. First, the experimental tests to characterize the deformation, damage and failure parameters in the material behavior are discussed. Second, the MAT213 input model and implementation details are presented with particular attention given to procedures that have been incorporated to ensure that the yield surfaces in the rate and temperature dependent plasticity model are convex. Finally, the paper concludes with a validation test designed to test the stability, accuracy and efficiency of the implemented model.

Polymer Matrix Composites↗

InSight's Reconstructed Aerothermal Environments

The InSight Mars Lander successfully landed on the surface on November 26, 2018. This poster will describe the methodologies and margins used in developing the aerothermal environments for design of the thermal protection systems (TPS), as well as a prediction of as-flown environments based on the best estimated trajectory. The InSight mission spacecraft design approach included the effects of radiant heat flux to the aft body from the wake for the first time on a US Mars Mission, due to overwhelming evidence in ground testing for the European ExoMars mission (2009/2010) [1] and 2010 tests in the Electric Arc Shock Tube (EAST) facility [2]. The radiant energy on an aftbody was also recently confirmed via measurement on the Schiaparelli mission [3]. In addition, the InSight mission expected to enter the Mars atmosphere during the dust storm season, so the heatshield TPS was designed to accommodate the extra recession due to the potential dust impact. This poster will compare the predicted aerothermal environments using the reconstructed best estimated trajectory to the design environments. Design Approach: The InSight spacecraft was planned to be a near-design-to-print copy of the Phoenix spacecraft. The determination of the heatshield TPS requirements was approached as if it was a new design due to the new requirement of flying through a dust storm. The baseline for aftbody was build-to-print, and all analyses focused on ensuring adequate margin. This proved to be a challenge because the Phoenix aftbody was designed to withstand only convective heating and the InSight aftbody was evaluated for both convective and radiative heating. Aerothermal environments were predicted using the Langley Aerothermodynamic Upwind Relaxation Algorithm (LAURA) and the Data Parallel Line Relaxation (DPLR) CFD codes, and the Nonequilibrium Radiative Transport and Spectra Program (NEQAIR) utilizing bounding design trajectories derived from Monte Carlo analyses from the Program to Optimize Simulated Trajectories II (POST2). In all cases, super-catalytic flowfields were assigned to ensure the most conservative heating results. Two trajectories were evaluated: 1) the trajectory with the maximum heat flux was utilized to determine the flowfield characteristics and the viability of the selection of TPS materials; and 2) the trajectory with the maximum heat load was used to determine the required thicknesses of the TPS materials. Evaluation of the MEDLI data [4], along with ground test data [5] led to the determination of whether or not the flow would transition from laminar to turbulent on the heatshield, which also determined the TPS sizing location for the heatshield. Aerothermal margins were added for the convective heating and developed for the radiative heating. TPS material sizing was determined with the Reaction Kinetic Ablation Program (REKAP) and the Fully Implicit Ablation and Thermal Analysis program (FIAT) using a three-branched approach to account for aerothermal, material response, and material properties uncertainties. In addition, the heatshield recession was augmented by an analysis of the effect of entry through a potential dusty atmosphere using a methodology developed in References [6] and [7]. These analyses resulted in an increase to the Phoenix heatshield TPS thickness. Reconstruction Efforts: Once the best estimated trajectory is reconstructed by the team, the LAURA/HARA (High-Temperature Aerothermo-dynamic Radiation model) and DPLR/NEQAIR code pairs will be used to predict the as-flown aerothermal conditions. In these runs, fully-catalytic flowfields will be assigned because it is a more physically accurate description of the chemistry in the flow. Once again, determination of the onset of turbulence on the heatshield will be evaluated. The as-flown aerothermal environments will then be compared to the design environments.

Beck, R. A.↗

NEQAIR v15.0 Release Notes: Nonequilibrium and Equilibrium Radiative Transport and Spectra Program

NEQAIR v15.0 provides the first steps to improved coupling between NEQAIR and the DPLR CFD code, which will be fully realized in v15.1. The plan is to release NEQAIR v15.1 and DPLR 4.05 at the same time. The improvements implemented in NEQAIR v15.0 have focused on improving stability, solution robustness, usability and providing different options for running the code. It is also the first version of the code to have a new input file and line of sight format since 2009. Backward compatibility with previous formats of the input files (neqair.inp and LOS.dat) has also been provided. NEQAIR v15.0 supersedes the prerelease of this version, as well as NEQAIR v14.0, v13.2, v13.1 and the suite of NEQAIR2009 versions. These updates have predominantly been performed by Brett Cruden and Aaron Brandis from AMA Inc at NASA Ames Research Center between 2016 and 2018. NEQAIR v15.0 is a standalone software tool for line-by-line spectral computation of radiative intensities and/or radiative heat flux, with one-dimensional transport of radiation. In order to accomplish this, NEQAIR v15.0, as in previous versions, requires the specification of distances (in cm), temperatures (in K) and number densities (in parts/cc) of constituent species along lines of sight. Therefore, it is assumed that flow quantities have been extracted from flow fields computed using other tools, such as CFD codes like DPLR or LAURA, and that lines of sight have been constructed and written out in the format required by NEQAIR v15.0. There are two principal modes for running NEQAIR v15.0. In the first mode NEQAIR v15.0 is used as a tool for creating synthetic spectra of any desired resolution (including convolution with a specified instrument/slit function). The first mode is typically exercised in simulating/interpreting spectroscopic measurements of different sources (e.g. shock tube data, plasma torches, etc.). In the second mode, NEQAIR v15.0 is used as a radiative heat flux prediction tool for flight projects. Correspondingly, NEQAIR has also been used to simulate the radiance measured on previous flight missions. This report summarizes the database updates, corrections that have been made to the code, changes to input files, parallelization, the current usage recommendations, including test cases, and an indication of the performance enhancements achieved.

Brandis, Aaron M.↗