Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

kokkosSZ (kSZ) A Portable Cellerator Implementation of SZ using Kokkos Programming Model

kSZ is a Kokkos-based implementation of the world-widely used SZ lossy compressor. We use Kokkos because it provides abstractions for both parallel execution of code and data management, which can be used to support portable implementation across different accelerator technologies. Kokkos can support OpenMP/OpenMPTarget, oneAPI, Pthreads, and CUDA as backend programming models.

ECP↗

ML-PSA

The computer code uses a parallel simulated annealing framework with embedded machine learning components to solve multi-constrained optimization problems. The software automatically balances the execution of low and high fidelity physics models within the optimization procedure. The low fidelity model is used to rapidly explore the design space while the high fidelity physics model is executed sparingly to account for complex design constraints that are not resolved by the quickly executing low fidelity model.

Gurecky, William↗

gismo-cloud-deployment

Tools for executing time-consuming tasks with developer-defined custom code blocks in parallel on the AWS EKS platform.

Leu, Jimmy↗

A New Capability of E4D For 3D Parallel Joint Inversion of DC Resistivity And Traveltime Data on Unstructured Mesh

A major challenge in interpreting geophysical data is how to derive consistent three-dimensional (3D) earth models of different physical properties from spatially and temporally limited measurements. Joint inversion with cross-gradient constraints is an approach to find such models by imposing structural similarities between different physical parameters. We have developed a parallel distributed-memory joint inversion code for direct-current (DC) resistivity and traveltime data using the cross-gradient constraint on unstructured mesh. The code utilizes existing E4D framework for parallel forward simulation, distributed storage and computation of the Jacobian matrix of forward operator, and parallel execution of matrix-vector multiplication during inversion. Besides, the joint inversion is solved by nonlinear conjugate gradient algorithm parallelized for DC resistivity and traveltime data. The joint inversion capability of E4D was tested using synthetic data from cross-borehole DC resistivity and traveltime data. The results indicate that the shape and size of the anomalies from the joint inversion are more reliable than those from separate inversions.

58 GEOSCIENCES↗

Beam Dynamics Simulations of Transient Beam Loading Effects in the 10GeV EIC Electron Storage Ring

We report on beam dynamics studies of transient beam loading effects in the 10 GeV EIC electron storage ring [1]. The studies are carried out with time-dependent Vlasov-Fokker-Planck simulations performed with the parallel, particle tracking code SPACE [2], which allows to follow self-consistently the dynamics of h bunches, where h in the number of RF buckets, in arbitrary multi-bunch configurations. The specific goal of the numerical simulations is to determine stable RF cavity settings under heavy beam loading. We also study the option to operate with a passive, third-harmonic cavity (3HC) system for bunch lengthening, addressing both stability and the performance limitation due to a gap in the uniform filling pattern for ion clearing.

43 PARTICLE ACCELERATORS↗

A class of hybrid finite element methods for electromagnetics: A review

Integral equation methods have generally been the workhorse for antenna and scattering computations. In the case of antennas, they continue to be the prominent computational approach, but for scattering applications the requirement for large-scale computations has turned researchers' attention to near neighbor methods such as the finite element method, which has low O(N) storage requirements and is readily adaptable in modeling complex geometrical features and material inhomogeneities. In this paper, we review three hybrid finite element methods for simulating composite scatterers, conformal microstrip antennas, and finite periodic arrays. Specifically, we discuss the finite element method and its application to electromagnetic problems when combined with the boundary integral, absorbing boundary conditions, and artificial absorbers for terminating the mesh. Particular attention is given to large-scale simulations, methods, and solvers for achieving low memory requirements and code performance on parallel computing architectures.

Volakis, J. L.↗

Parallel computing for probabilistic fatigue analysis

This paper presents the results of Phase I research to investigate the most effective parallel processing software strategies and hardware configurations for probabilistic structural analysis. We investigate the efficiency of both shared and distributed-memory architectures via a probabilistic fatigue life analysis problem. We also present a parallel programming approach, the virtual shared-memory paradigm, that is applicable across both types of hardware. Using this approach, problems can be solved on a variety of parallel configurations, including networks of single or multiprocessor workstations. We conclude that it is possible to effectively parallelize probabilistic fatigue analysis codes; however, special strategies will be needed to achieve large-scale parallelism to keep large number of processors busy and to treat problems with the large memory requirements encountered in practice. We also conclude that distributed-memory architecture is preferable to shared-memory for achieving large scale parallelism; however, in the future, the currently emerging hybrid-memory architectures will likely be optimal.

Sues, Robert H.↗

High Performance FORTRAN

High performance FORTRAN is a set of extensions for FORTRAN 90 designed to allow specification of data parallel algorithms. The programmer annotates the program with distribution directives to specify the desired layout of data. The underlying programming model provides a global name space and a single thread of control. Explicitly parallel constructs allow the expression of fairly controlled forms of parallelism in particular data parallelism. Thus the code is specified in a high level portable manner with no explicit tasking or communication statements. The goal is to allow architecture specific compilers to generate efficient code for a wide variety of architectures including SIMD, MIMD shared and distributed memory machines.

Mehrotra, Piyush↗

Using PVM to host CLIPS in distributed environments

It is relatively easy to enhance CLIPS (C Language Integrated Production System) to support multiple expert systems running in a distributed environment with heterogeneous machines. The task is minimized by using the PVM (Parallel Virtual Machine) code from Oak Ridge Labs to provide the distributed utility. PVM is a library of C and FORTRAN subprograms that supports distributive computing on many different UNIX platforms. A PVM deamon is easily installed on each CPU that enters the virtual machine environment. Any user with rsh or rexec access to a machine can use the one PVM deamon to obtain a generous set of distributed facilities. The ready availability of both CLIPS and PVM makes the combination of software particularly attractive for budget conscious experimentation of heterogeneous distributive computing with multiple CLIPS executables. This paper presents a design that is sufficient to provide essential message passing functions in CLIPS and enable the full range of PVM facilities.

Myers, Leonard↗

Preliminary large-eddy simulations of flow around a NACA 4412 airfoil using unstructured grids

Large-eddy simulation (LES) has matured to the point where application to complex flows is desirable. The extension to higher Reynolds numbers leads to an impractical number of grid points with existing structured-grid methods. Furthermore, most real world flows are rather difficult to represent geometrically with structured grids. Unstructured-grid methods offer a release from both of these constraints. However, just as it took many years for structured-grid methods to be well understood and reliable tools for LES, unstructured-grid methods must be carefully studied before we can expect them to attain their full potential. In the past two years, important building blocks have been put into place making possible a careful study of LES on unstructured grids. The first building block was an efficient mesh generator which allowed the placement of points according to smooth variation of physical length scales. This variation of length scales is in all three directions independently, which allows a large reduction in points when compared to structured-grid methods, which can only vary length scales in one direction at a time. The second building block was the development of a dynamic model appropriate for unstructured grids. The principle obstacle was the development of an unstructured-grid filtering operator. In the past year, some of the new filters developed by Jansen have been implemented into a highly parallelized finite element code based on the Galerkin/least-squares finite element method. We have chosen the NACA 4412 airfoil at maximum lift as the first simulation for a variety of reasons. First, it is a problem of significant interest since it would be the first LES of an aircraft component. Second, this flow has been the subject of three experimental studies. The third reason for considering this flow is the variety of flow features which provide an important test of the dynamic model. Only the dynamic model can be expected to perform satisfactorily in this variety of situations: from the laminar regions where it must not modify the flow at all to the turbulent boundary layers and wake where it must represent a wide variety of subgrid-scale structures. The flow configuration we have chosen is that of Wadcock (1987) at Reynolds number based on chord Re(sub c) = u(sub infinity)c/v = 1.64 x 10(exp 6), Mach number M = 0.2, and 12 deg angle of attack.

Jansen, Kenneth↗

Computations of Boiling in Microgravity

The absence (or reduction) of gravity, can lead to major changes in boiling heat transfer. On Earth, convection has a major effect on the heat distribution ahead of an evaporation front, and buoyancy determines the motion of the growing bubbles. In microgravity, convection and buoyancy are absent or greatly reduced and the dynamics of the growing vapor bubbles can change in a fundamental way. In particular, the lack of redistribution of heat can lead to a large superheat and explosive growth of bubbles once they form. While considerable efforts have been devoted to examining boiling experimentally, including the effect of microgravity, theoretical and computational work is limited to very simple models. In this project, the growth of boiling bubbles is studied by direct numerical simulations where the flow field is fully resolved and the effects of inertia, viscosity, surface deformation, heat conduction and convection, as well as the phase change, are fully accounted for. The proposed work is based on previously funded NASA work that allowed us to develop a two-dimensional numerical method for boiling flows and to demonstrate the ability of the method to simulate film boiling. While numerical simulations of multi-fluid flows have been advanced in a major way during the last five years, or so, similar capability for flows with phase change are still in their infancy. Although the feasibility of the proposed approach has been demonstrated, it has yet to be extended and applied to fully three-dimensional simulations. Here, a fully three-dimensional, parallel, grid adaptive code will be developed. The numerical method will be used to study nucleate boiling in microgravity, with particular emphasis on two aspects of the problem: 1) Examination of the growth of bubbles at a wall nucleation site and the instabilities of rapidly growing bubbles. Particular emphasis will be put on accurately capturing the thin wall layer left behind as a bubble expands along a wall, on computing instabilities on bubble surfaces as bubbles grow, and on quantifying the effects of both these phenomena on heat transfer; and 2) Examination of the effect of shear flow on bubble growth and heat transfer.

Tryggvason, Gretar↗

Dendritic Growth with Fluid Flow for Pure Materials

We have developed a three-dimensional, adaptive, parallel finite element code to examine solidification of pure materials under conditions of forced flow. We have examined the effect of undercooling, surface tension anisotropy and imposed flow velocity on the growth. The flow significantly alters the growth process, producing dendrites that grow faster, and with greater tip curvature, into the flow. The selection constant decreases slightly with flow velocity in our calculations. The results of the calculations agree well with the transport solution of Saville and Beaghton at high undercooling and high anisotropy. At low undercooling, significant deviations are found. We attribute this difference to the influence of other parts of the dendrite, removed from the tip, on the flow field.

Jeong, Jun-Ho↗

Computation of Sensitivity Derivatives of Navier-Stokes Equations using Complex Variables

Accurate computation of sensitivity derivatives is becoming an important item in Computational Fluid Dynamics (CFD) because of recent emphasis on using nonlinear CFD methods in aerodynamic design, optimization, stability and control related problems. Several techniques are available to compute gradients or sensitivity derivatives of desired flow quantities or cost functions with respect to selected independent (design) variables. Perhaps the most common and oldest method is to use straightforward finite-differences for the evaluation of sensitivity derivatives. Although very simple, this method is prone to errors associated with choice of step sizes and can be cumbersome for geometric variables. The cost per design variable for computing sensitivity derivatives with central differencing is at least equal to the cost of three full analyses, but is usually much larger in practice due to difficulty in choosing step sizes. Another approach gaining popularity is the use of Automatic Differentiation software (such as ADIFOR) to process the source code, which in turn can be used to evaluate the sensitivity derivatives of preselected functions with respect to chosen design variables. In principle, this approach is also very straightforward and quite promising. The main drawback is the large memory requirement because memory use increases linearly with the number of design variables. ADIFOR software can also be cumber-some for large CFD codes and has not yet reached a full maturity level for production codes, especially in parallel computing environments.

Vatsa, Veer N.↗

Large-Eddy Simulation Code Developed for Propulsion Applications

A large-eddy simulation (LES) code was developed at the NASA Glenn Research Center to provide more accurate and detailed computational analyses of propulsion flow fields. The accuracy of current computational fluid dynamics (CFD) methods is limited primarily by their inability to properly account for the turbulent motion present in virtually all propulsion flows. Because the efficiency and performance of a propulsion system are highly dependent on the details of this turbulent motion, it is critical for CFD to accurately model it. The LES code promises to give new CFD simulations an advantage over older methods by directly computing the large turbulent eddies, to correctly predict their effect on a propulsion system. Turbulent motion is a random, unsteady process whose behavior is difficult to predict through computer simulations. Current methods are based on Reynolds-Averaged Navier- Stokes (RANS) analyses that rely on models to represent the effect of turbulence within a flow field. The quality of the results depends on the quality of the model and its applicability to the type of flow field being studied. LES promises to be more accurate because it drastically reduces the amount of modeling necessary. It is the logical step toward improving turbulent flow predictions. In LES, the large-scale dominant turbulent motion is computed directly, leaving only the less significant small turbulent scales to be modeled. As part of the prediction, the LES method generates detailed information on the turbulence itself, providing important information for other applications, such as aeroacoustics. The LES code developed at Glenn for propulsion flow fields is being used to both analyze propulsion system components and test improved LES algorithms (subgrid-scale models, filters, and numerical schemes). The code solves the compressible Favre-filtered Navier- Stokes equations using an explicit fourth-order accurate numerical scheme, it incorporates a compressible form of Smagorinsky s model for the subgrid-scale turbulence, and it uses generalized curvilinear coordinates to allow analysis of a wide range of geometries. The code runs in parallel on shared memory multiprocessor computers and is written in Fortran 90 with dynamic memory allocation. A sample result for a Mach-1.4 round jet is presented in the figure. Instantaneous Mach number contours in several cross-planes downstream of the nozzle exit are shown, illustrating how an LES captures the large unsteady three-dimensional turbulent structures present in the jet.

DeBonis, James R.↗

Simulations of the Aerosol Index and the Absorption Aerosol Optical Depth and Comparisons with OMI Retrievals During ARCTAS-2008 Campaign

We have computed the Aerosol Index (AI) at 354 nm, useful for observing the presence of absorbing aerosols in the atmosphere, from aerosol simulations conducted with the Goddard Chemistry, Aerosol, Radiation, and Transport (GOCART) module running online the GEOS-5 Atmospheric GCM. The model simulates five aerosol types: dust, sea salt, black carbon, organic carbon and sulfate aerosol and can be run in replay or data assimilation modes. In the assimilation mode, information's provided by the space-based MODIS and MISR sensors constrains the model aerosol state. Aerosol optical properties are then derived from the simulated mass concentration and the Al is determined at the OMI footprint using the radiative transfer code VLIDORT. In parallel, model derived Absorption Aerosol Optical Depth (AAOD) is compared with OMI retrievals. We have focused our study during ARCTAS (June - July 2008), a period with a good sampling of dust and biomass burning events. Our ultimate goal is to use OMI measurements as independent validation for our MODIS/MISR assimilation. Towards this goal we document the limitation of OMI aerosol absorption measurements on a global scale, in particular sensitivity to aerosol vertical profile and cloud contamination effects, deriving the appropriate averaging kernels. More specifically, model simulated (full) column integrated AAOD is compared with model derived Al, this way identifying those regions and conditions under which OMI cannot detect absorbing aerosols. Making use of ATrain cloud measurements from MODIS, C1oudSat and CALIPSO we also investigate the global impact on clouds on OMI derived Al, and the extent to which GEOS-5 clouds can offer a first order representation of these effects.

Source record↗

Computational Study of the CC3 Impeller and Vaneless Diffuser Experiment

Centrifugal compressors are compatible with the low exit corrected flows found in the high pressure compressor of turboshaft engines and may play an increasing role in turbofan engines as engine overall pressure ratios increase. Centrifugal compressor stages are difficult to model accurately with RANS CFD solvers. A computational study of the CC3 centrifugal impeller in its vaneless diffuser configuration was undertaken as part of an effort to understand potential causes of RANS CFD mis-prediction in these types of geometries. Three steady, periodic cases of the impeller and diffuser were modeled using the TURBO Parallel Version 4 code: 1) a k-epsilon turbulence model computation on a 6.8 million point grid using wall functions, 2) a k-epsilon turbulence model computation on a 14 million point grid integrating to the wall, and 3) a k-omega turbulence model computation on the 14 million point grid integrating to the wall. It was found that all three cases compared favorably to data from inlet to impeller trailing edge, but the k-epsilon and k-omega computations had disparate results beyond the trailing edge and into the vaneless diffuser. A large region of reversed flow was observed in the k-epsilon computations which extended from 70% to 100% span at the exit rating plane, whereas the k-omega computation had reversed flow from 95% to 100% span. Compared to experimental data at near-peak-efficiency, the reversed flow region in the k-epsilon case resulted in an under-prediction in adiabatic efficiency of 8.3 points, whereas the k-omega case was 1.2 points lower in efficiency.

Kulkarni, Sameer↗

Computational Study of the CC3 Impeller and Vaneless Diffuser Experiment

Centrifugal compressors are compatible with the low exit corrected flows found in the high pressure compressor of turboshaft engines and may play an increasing role in turbofan engines as engine overall pressure ratios increase. Centrifugal compressor stages are difficult to model accurately with RANS CFD solvers. A computational study of the CC3 centrifugal impeller in its vaneless diffuser configuration was undertaken as part of an effort to understand potential causes of RANS CFD mis-prediction in these types of geometries. Three steady, periodic cases of the impeller and diffuser were modeled using the TURBO Parallel Version 4 code: (1) a k-ε turbulence model computation on a 6.8 million point grid using wall functions, (2) a k-ε turbulence model computation on a 14 million point grid integrating to the wall, and (3) a k-ω turbulence model computation on the 14 million point grid integrating to the wall. It was found that all three cases compared favorably to data from inlet to impeller trailing edge, but the k-ε and k-ω computations had disparate results beyond the trailing edge and into the vaneless diffuser. A large region of reversed flow was observed in the k-ε computations which extended from 70 to 100 percent span at the exit rating plane, whereas the k-ω computation had reversed flow from 95 to 100 percent span. Compared to experimental data at near-peak-efficiency, the reversed flow region in the k-ε case resulted in an underprediction in adiabatic efficiency of 8.3 points, whereas the k-ω case was 1.2 points lower in efficiency.

Kulkarni, Sameer↗