Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules

Abstract Materials simulations involving strongly correlated electrons pose fundamental challenges to state‐of‐the‐art electronic structure methods but are hypothesized to be the ideal use case for quantum computing algorithms. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real‐world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near‐term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using the VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies using it have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near‐term feasibility of modeling a molecule with d‐orbitals on real quantum hardware. We demonstrate that the inclusion of d‐orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE + UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Investigation of Cycle-to-Cycle Variations in Internal Combustion Engine Using Proper Orthogonal Decomposition

The understanding, modeling and control of the cycle-to-cycle variation (CCV) in the modern internal combustion engine (ICE) is a key scientific challenge to achieve stable engine operation. High CCV in the engine combustion chamber may contribute to partial burn, misfire and knock, which adversely affects the engine performance and may potentially damage the engine. The objective of the current study is to leverage high-fidelity numerical simulations to improve the understanding of the causes of CCV. Using the massively parallel code, Nek5000, multi-cycle, wall-resolved large-eddy simulations (LES) were performed for the General Motors (GM), Transparent Combustion Chamber (TCC-III) optical engine under motored operating conditions. Further, the large-scale structures of the in-cylinder flow were investigated using a triple proper orthogonal decomposition (POD) technique to explore the characteristics of different parts of the flow and their contributions to CCV. The kinetic energy of the subset of flow structures were determined and correlated between the intake and compression strokes. The insights from the analysis of the large-scale flow structures will be used to assist the development of improved engine designs with reduced CCV and enhance the engine performance.

33 ADVANCED PROPULSION SYSTEMS↗

Kinetic Monte Carlo simulations of structural evolution during anneal of additively manufactured materials

Our experiments indicated that upon a post-processing anneal, an additively manufactured 316L stainless steel exhibits cubic grains rather than the conventional equiaxed grains. In this work, we have used kinetic Monte Carlo simulations to explore the origin of these cubic grains. First, we implemented a new kinetic Monte Carlo model in parallel code SPPARKS to simulate grain growth and recrystallization under a residual energy distribution. Our model incorporates physical properties and real-time, as opposed to generic properties and relative time. We further validated that our SPPARKS simulations reproduced the expected kinetic behavior of single-grain evolution. We then used the validated approach to simulate the anneal of an additively manufactured material under the same conditions used in our experiments. We found that the cubic grains can origin from a periodically varying residual energy that may be present in additively manufactured materials.

36 MATERIALS SCIENCE↗

Permutationally Invariant Polynomial Expansions with Unrestricted Complexity

A general strategy is presented for constructing and validating permutationally invariant polynomial (PIP) expansions for chemical systems of any stoichiometry. Demonstrations are made for three categories of gas-phase dynamics and kinetics: collisional energy-transfer trajectories for predicting pressure-dependent kinetics, three-body collisions for describing transient van der Waals adducts relevant to atmospheric chemistry, and nonthermal reactivity via quasiclassical trajectories. In total, 30 systems are considered with up to 15 atoms and 39 degrees of freedom. Permutational invariance is enforced in PIP expansions with as many as 13 million terms and 13 permutationally distinct atom types by taking advantage of petascale computational resources. The quality of the PIP expansions is demonstrated through the systematic convergence of in-sample and out-of-sample errors with respect to both the number of training data and the order of the expansion, and these errors are shown to predict errors in the dynamics for both reactive and nonreactive applications. Here, the parallelized code distributed as part of this work enables the automation of PIP generation for complex systems with multiple channels and flexible user-defined symmetry constraints and for automatically removing unphysical unconnected terms from the basis set expansions, all of which are required for simulating complex reactive systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

tih_vqe [SWR-23-32]

This software supports the paper, "Exploring the scaling limitations of the variational quantum eigensolver with the bond dissociation of hydride diatomic molecules," published in the International Journal of Quantum Chemistry, whose abstract is as follows: Materials simulations involving strongly correlated electrons pose fundamental challenges to state-of-the-art electronic structure methods but are hypothesized to be the ideal use case for quantum computing. To date, no quantum computer has simulated a molecule of a size and complexity relevant to real-world applications, despite the fact that the variational quantum eigensolver (VQE) algorithm can predict chemically accurate total energies. Nevertheless, because of the many applications of moderately-sized, strongly correlated systems, such as molecular catalysts, the successful use of the VQE stands as an important waypoint in the advancement toward useful chemical modeling on near-term quantum processors. In this paper, we take a significant step in this direction. We lay out the steps, write, and run parallel code for an (emulated) quantum computer to compute the bond dissociation curves of the TiH, LiH, NaH, and KH diatomic hydride molecules using VQE. TiH was chosen as a relatively simple chemical system that incorporates d orbitals and strong electron correlation. Because current VQE implementations on existing quantum hardware are limited by qubit error rates, the number of qubits available, and the allowable gate depth, recent studies have focused on chemical systems involving s and p block elements. Through VQE + UCCSD calculations of TiH, we evaluate the near-term feasibility of modeling a molecule with d-orbitals on real quantum hardware. We demonstrate that the inclusion of d-orbitals and the use of the UCCSD ansatz, which are both necessary to capture the correct TiH physics, dramatically increase the cost of this problem. We estimate the approximate error rates necessary to model TiH on current quantum computing hardware using VQE+UCCSD and show them to likely be prohibitive until significant improvements in hardware and error correction algorithms are available.

Graf, Peter↗

Efficient phase-space generation for hadron collider event simulation

We present a simple yet efficient algorithm for phase-space integration at hadron colliders. Individual mappings consist of a single t-channel combined with any number of s-channel decays, and are constructed using diagrammatic information. The factorial growth in the number of channels is tamed by providing an option to limit the number of s-channel topologies. We provide a publicly available, parallelized code in C++ and test its performance in typical LHC scenarios.

47 OTHER INSTRUMENTATION↗

An efficient formulation of the coupled finite element-integral equation technique for solving large 3D scattering problems

It is often desirable to calculate the electromagnetic fields inside and about a complicated system of scattering bodies, as well as in their far-field region. The finite element method (FE) is well suited to solving the interior problem, but the domain has to be limited to a manageable size. At the truncation of the FE mesh one can either impose approximate (absorbing) boundary conditions or set up an integral equation (IE) for the fields scattered from the bodies. The latter approach is preferable since it results in higher accuracy. Hence, the two techniques can be successfully combined by introducing a surface that encloses the scatterers, applying a FE model to the inner volume and setting up an IE for the tangential fields components on the surface. Here the continuity of the tangential fields is used bo obtain a consistent solution. A few coupled FE-IE methods have recently appeared in the literature. The approach presented here has the advantage of using edge-based finite elements, a type of finite elements with degrees of freedom associated with edges of the mesh. Because of their properties, they are better suited than the conventional node based elements to represent electromagnetic fields, particularly when inhomogeneous regions are modeled, since the node based elements impose an unnatural continuity of all field components across boundaries of mesh elements. Additionally, our approach is well suited to handle large size problems and lends itself to code parallelization. We will discuss the salient features that make our approach very efficient from the standpoint of numerical computation, and the fields and RCS of a few objects are illustrated as examples.

Cwik, T.↗

Portable parallel stochastic optimization for the design of aeropropulsion components

This report presents the results of Phase 1 research to develop a methodology for performing large-scale Multi-disciplinary Stochastic Optimization (MSO) for the design of aerospace systems ranging from aeropropulsion components to complete aircraft configurations. The current research recognizes that such design optimization problems are computationally expensive, and require the use of either massively parallel or multiple-processor computers. The methodology also recognizes that many operational and performance parameters are uncertain, and that uncertainty must be considered explicitly to achieve optimum performance and cost. The objective of this Phase 1 research was to initialize the development of an MSO methodology that is portable to a wide variety of hardware platforms, while achieving efficient, large-scale parallelism when multiple processors are available. The first effort in the project was a literature review of available computer hardware, as well as review of portable, parallel programming environments. The first effort was to implement the MSO methodology for a problem using the portable parallel programming language, Parallel Virtual Machine (PVM). The third and final effort was to demonstrate the example on a variety of computers, including a distributed-memory multiprocessor, a distributed-memory network of workstations, and a single-processor workstation. Results indicate the MSO methodology can be well-applied towards large-scale aerospace design problems. Nearly perfect linear speedup was demonstrated for computation of optimization sensitivity coefficients on both a 128-node distributed-memory multiprocessor (the Intel iPSC/860) and a network of workstations (speedups of almost 19 times achieved for 20 workstations). Very high parallel efficiencies (75 percent for 31 processors and 60 percent for 50 processors) were also achieved for computation of aerodynamic influence coefficients on the Intel. Finally, the multi-level parallelization strategy that will be needed for large-scale MSO problems was demonstrated to be highly efficient. The same parallel code instructions were used on both platforms, demonstrating portability. There are many applications for which MSO can be applied, including NASA's High-Speed-Civil Transport, and advanced propulsion systems. The use of MSO will reduce design and development time and testing costs dramatically.

Sues, Robert H.↗

Ropes: Support for collective opertions among distributed threads

Lightweight threads are becoming increasingly useful in supporting parallelism and asynchronous control structures in applications and language implementations. Recently, systems have been designed and implemented to support interprocessor communication between lightweight threads so that threads can be exploited in a distributed memory system. Their use, in this setting, has been largely restricted to supporting latency hiding techniques and functional parallelism within a single application. However, to execute data parallel codes independent of other threads in the system, collective operations and relative indexing among threads are required. This paper describes the design of ropes: a scoping mechanism for collective operations and relative indexing among threads. We present the design of ropes in the context of the Chant system, and provide performance results evaluating our initial design decisions.

Haines, Matthew↗

Physics of Boundaries and their Interactions in Space Plasmas

This report describes the work done by SciberNet, Inc. during the month of August. We have resolved the issues associated with the implementation of the dipole field in our large scale hybrid simulations of the magnetopause. We have setup several runs and will spend the next several months analyzing the data. The results will be presented at the Fall AGU. We are also continuing our analysis of the 3-D simulations of thin current sheets at the magnetopause, paying special attention to the conditions under which Kelvin-Helmholtz would lead to sizable perturbations of the magnetopause. In a related study, we are in the process of developing a new kinetic linear code that would for the first time enable us to examine the linear properties of the Kelvin-Helmholtz instability in the fully kinetic regime. Finally, we are continuing our code development to include inflow-outflow boundary conditions in our 2-D and 3-D hybrid codes. We are also comparing the different methods of code parallelization in order to extend the limits of our calculations.

Omidi, Nojan↗

An Efficient Objective Analysis System for Parallel Computers

A new atmospheric objective analysis system designed for parallel computers will be described. The system can produce a global analysis (on a 1 X 1 lat-lon grid with 18 levels of heights and winds and 10 levels of moisture) using 120,000 observations in 17 minutes on 32 CPUs (SGI Origin 2000). No special parallel code is needed (e.g. MPI or multitasking) and the 32 CPUs do not have to be on the same platform. The system is totally portable and can run on several different architectures at once. In addition, the system can easily scale up to 100 or more CPUS. This will allow for much higher resolution and significant increases in input data. The system scales linearly as the number of observations and the number of grid points. The cost overhead in going from 1 to 32 CPUs is 18%. In addition, the analysis results are identical regardless of the number of processors used. This system has all the characteristics of optimal interpolation, combining detailed instrument and first guess error statistics to produce the best estimate of the atmospheric state. Static tests with a 2 X 2.5 resolution version of this system showed it's analysis increments are comparable to the latest NASA operational system including maintenance of mass-wind balance. Results from several months of cycling test in the Goddard EOS Data Assimilation System (GEOS DAS) show this new analysis retains the same level of agreement between the first guess and observations (O-F statistics) as the current operational system.

Stobie, J.↗

An Efficient Objective Analysis System for Parallel Computers

A new objective analysis system designed for parallel computers will be described. The system can produce a global analysis (on a 2 x 2.5 lat-lon grid with 20 levels of heights and winds and 10 levels of moisture) using 120,000 observations in less than 3 minutes on 32 CPUs (SGI Origin 2000). No special parallel code is needed (e.g. MPI or multitasking) and the 32 CPUs do not have to be on the same platform. The system Ls totally portable and can run on -several different architectures at once. In addition, the system can easily scale up to 100 or more CPUS. This will allow for much higher resolution and significant increases in input data. The system scales linearly as the number of observations and the number of grid points. The cost overhead in going from I to 32 CPus is 18%. in addition, the analysis results are identical regardless of the number of processors used. T'his system has all the characteristics of optimal interpolation, combining detailed instrument and first guess error statistics to produce the best estimate of the atmospheric state. It also includes a new quality control (buddy check) system. Static tests with the system showed it's analysis increments are comparable to the latest NASA operational system including maintenance of mass-wind balance. Results from a 2-month cycling test in the Goddard EOS Data Assimilation System (GEOS DAS) show this new analysis retains the same level of agreement between the first guess and observations (0-F statistics) throughout the entire two months.

Stobie, James G.↗

Time-Dependent Simulations of Incompressible Flow in a Turbopump Using Overset Grid Approach

This viewgraph presentation provides information on mathematical modelling of the SSME (space shuttle main engine). The unsteady SSME-rig1 start-up procedure from the pump at rest has been initiated by using 34.3 million grid points. The computational model for the SSME-rig1 has been completed. Moving boundary capability is obtained by using DCF module in OVERFLOW-D. MPI (Message Passing Interface)/OpenMP hybrid parallel code has been benchmarked.

Kiris, Cetin↗

Combustor Simulation

The goal was to perform 3D simulation of GE90 combustor, as part of full turbofan engine simulation. Requirements of high fidelity as well as fast turn-around time require massively parallel code. National Combustion Code (NCC) was chosen for this task as supports up to 999 processors and includes state-of-the-art combustion models. Also required is ability to take inlet conditions from compressor code and give exit conditions to turbine code.

Norris, Andrew↗

Simulation of Foam Impact Effects on Components of the Space Shuttle Thermal Protection System

A series of three dimensional simulations has been performed to investigate analytically the effect of insulating foam impacts on ceramic tile and reinforced carbon-carbon components of the Space Shuttle thermal protection system. The simulations employed a hybrid particle-finite element method and a parallel code developed for use in spacecraft design applications. The conclusions suggested by the numerical study are in general consistent with experiment. The results emphasize the need for additional material testing work on the dynamic mechanical response of thermal protection system materials, and additional impact experiments for use in validating computational models of impact effects.

Fahrenthold, Eric P.↗

Simulation of Aerosols and Chemistry with a Unified Global Model

This project is to continue the development of the global simulation capabilities of tropospheric and stratospheric chemistry and aerosols in a unified global model. This is a part of our overall investigation of aerosol-chemistry-climate interaction. In the past year, we have enabled the tropospheric chemistry simulations based on the GEOS-CHEM model, and added stratospheric chemical reactions into the GEOS-CHEM such that a globally unified troposphere-stratosphere chemistry and transport can be simulated consistently without any simplifications. The tropospheric chemical mechanism in the GEOS-CHEM includes 80 species and 150 reactions. 24 tracers are transported, including O3, NOx, total nitrogen (NOy), H2O2, CO, and several types of hydrocarbon. The chemical solver used in the GEOS-CHEM model is a highly accurate sparse-matrix vectorized Gear solver (SMVGEAR). The stratospheric chemical mechanism includes an additional approximately 100 reactions and photolysis processes. Because of the large number of total chemical reactions and photolysis processes and very different photochemical regimes involved in the unified simulation, the model demands significant computer resources that are currently not practical. Therefore, several improvements will be taken, such as massive parallelization, code optimization, or selecting a faster solver. We have also continued aerosol simulation (including sulfate, dust, black carbon, organic carbon, and sea-salt) in the global model to cover most of year 2002. These results have been made available to many groups worldwide and accessible from the website http://code916.gsfc.nasa.gov/People/Chin/aot.html.

Chin, Mian↗

Turbopump Performance Improved by Evolutionary Algorithms

The development of design optimization technology for turbomachinery has been initiated using the multiobjective evolutionary algorithm under NASA's Intelligent Synthesis Environment and Revolutionary Aeropropulsion Concepts programs. As an alternative to the traditional gradient-based methods, evolutionary algorithms (EA's) are emergent design-optimization algorithms modeled after the mechanisms found in natural evolution. EA's search from multiple points, instead of moving from a single point. In addition, they require no derivatives or gradients of the objective function, leading to robustness and simplicity in coupling any evaluation codes. Parallel efficiency also becomes very high by using a simple master-slave concept for function evaluations, since such evaluations often consume the most CPU time, such as computational fluid dynamics. Application of EA's to multiobjective design problems is also straightforward because EA's maintain a population of design candidates in parallel. Because of these advantages, EA's are a unique and attractive approach to real-world design optimization problems.

Oyama, Akira↗