Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35

2nd-Order CESE Results For C1.4: Vortex Transport by Uniform Flow

The Conservation Element and Solution Element (CESE) method was used as implemented in the NASA research code ez4d. The CESE method is a time accurate formulation with flux-conservation in both space and time. The method treats the discretized derivatives of space and time identically and while the 2nd-order accurate version was used, high-order versions exist, the 2nd-order accurate version was used. In regards to the ez4d code, it is an unstructured Navier-Stokes solver coded in C++ with serial and parallel versions available. As part of its architecture, ez4d has the capability to utilize multi-thread and Messaging Passage Interface (MPI) for parallel runs.

Space-time CE/SE method↗

Transition Analysis for the CRM-NLF Wind Tunnel Configuration

This paper reports the results of a comprehensive linear stability analysis of the boundary layer flow over the common research model with natural laminar flow (CRM-NLF) aircraft configuration. The flow conditions match selected test conditions from a recent wind tunnel experiment in the National Transonic Facility at the NASA Langley Research Center. Previous work has shown that the measured onset of laminar-turbulent transition during the experiments can be correlated with the linear amplification of Tollmien-Schlichting (TS) and stationary crossflow (CF) instabilities in the swept wing boundary layer. However, a significant scatter ( N ∈ (4,9)) was observed in the values of the logarithmic amplification factor along the measured transition front. This previous analysis was based on an approximate basic state (based on a boundary layer code with conical flow approximation) and parallel stability computations without surface curvature effects. Here, we examine the effects of these various approximations with the goal of quantifying the resulting changes in the N-factor correlations. Specifically, both linear stability theory (LST) and the parabolized stability equations (PSE)are used in conjunction with an accurate definition of the laminar boundary layer flow as computed with a Navier-Stokes solver with a Reynolds-Averaged-Navier-Stokes (RANS) based turbulence model within the turbulent parts of the flow. Furthermore, the effects of instability wave propagation within a fully three-dimensional boundary layer are also evaluated by integrating the disturbance growth rates along suitably chosen, curvilinear (i.e., nonplanar) propagation trajectories. The results of this analysis are also used in an accompanying paper by Venkatachari et al. to develop improved, physics based transition predictions for the same CRM-NLF configuration.

Boundary layer transition↗

Geologic Disposal Safety Assessment (GDSA) Biosphere Model Development

The Spent Fuel and Waste Science and Technology Campaign of the U.S. Department of Energy Office of Nuclear Energy, Office of Spent Fuel and Waste Disposition is conducting research and development on geologic disposal of spent nuclear fuel and high-level nuclear waste. This work includes the Geologic Disposal Safety Assessment (GDSA) program which is charged with development of generic deep geologic repository concepts and system performance assessment models. One part of the GDSA framework is the development of a biosphere model capable of assessing doses to potential receptors exposed to radionuclides released from geologic disposal sites. As part of the GDSA framework, a biosphere model compatible with the PFLOTRAN massively parallel subsurface flow and reactive transport code is under development. The PFLOTRAN model provides the radionuclide source term for the biosphere model. The GDSA Biosphere model then assesses the potential movement of radionuclides through the surface biosphere and the subsequent exposure to a human receptor living in the biosphere. The biosphere model includes pathways originating from the groundwater as well as pathways originating from surface water bodies that have a water exchange with a contaminated groundwater body. The pathways for human exposure include consumption of drinking water, irrigated crops, meat animals, aquatic vegetation, and animals, etc.; external exposure from irrigated ground surfaces, surface water bodies, recreational activities, etc.; and inadvertent exposures such as ingestion of contaminated soils or shower water, etc. The GDSA Biosphere model was designed to be flexible and generic in order to accommodate a variety of different sites and climate states. This presentation will present the on the purpose, design, and development progress of the GDSA Biosphere Model.

GDSA, biosphere, repository↗

Developing Ultrahigh-Resolution E3SM Land Model for GPU Systems

Designing and refactoring complex scientific code, such as the E3SM land model (ELM), for new computing architectures is challenging. This paper presents design strategies and technical approaches to develop a data-oriented, GPU-ready ELM model using compiler directives (OpenACC/OpenMP). We first analyze the datatypes and processes in the original ELM code. Then we present design considerations for ultrahigh-resolution ELM (uELM) development for massive GPU systems. These techniques include the global data-oriented simulation workflow, domain partition, code porting and data copy, memory reduction, parallel loop restructure and flattening, and race condition detection. We implemented the first version of uELM using OpenACC targeting the NVidia GPUs in the Summit supercomputer at Oak Ridge National Laboratory. During the implementation, we developed a software tool (named SPEL) to facilitate code generation, verification, and performance tuning using these techniques. The first uELM implementation for Nvidia GPUs on Summit delivered promising results: 1) over 98% of the ELM code was automatically generated and tuned by scripts. Most ELM modules had better computational performances than the original ELM code for CPUs. The GPU-ready uELM is more scalable than the CPU code on fully-loaded Summit nodes. Example profiling results from several modules are also presented to illustrate the performance improvements and race condition detection. The lessons learned and toolkit developed in the study are also suitable for further uELM deployment using OpenMP on the first US exascale computer, Frontier, equipped with AMD CPUs and GPUs.

Schwartz, Peter↗

Parallel computation of 3-D Navier-Stokes flowfields for supersonic vehicles

Multidisciplinary design optimization of aircraft will require unprecedented capabilities of both analysis software and computer hardware. The speed and accuracy of the analysis will depend heavily on the computational fluid dynamics (CFD) module which is used. A new CFD module has been developed to combine the robust accuracy of conventional codes with the ability to run on parallel architectures. This is achieved by parallelizing the ARC3D algorithm, a central-differenced Navier-Stokes method, on the Intel iPSC/860. The computed solutions are identical to those from conventional machines. Computational speed on 64 processors is comparable to the rate on one Cray Y-MP processor and will increase as new generations of parallel computers become available.

Ryan, James S.↗

Dynamics of face and annular seals with two-phase flow

A detailed study was made of face and annular seals under conditions where boiling, i.e., phase change of the leaking fluid, occurs within the seal. Many seals operate in this mode because of flashing due to pressure drop and/or heat input from frictional heating. Some of the distinctive behavior characteristics of two phase seals are discussed, particularly their axial stability. The main conclusions are that seals with two phase flow may be unstable if improperly balanced. Detailed theoretical analyses of low (laminar) and high (turbulent) leakage seals are presented along with computer codes, parametric studies, and in particular a simplified PC based code that allows for rapid performance prediction: calculations of stiffness coefficients, temperature and pressure distributions, and leakage rates for parallel and coned face seals. A simplified combined computer code for the performance prediction over the laminar and turbulent ranges of a two phase flow is described and documented. The analyses, results, and computer codes are summarized.

Hughes, William F.↗

The JOREK non-linear extended MHD code and applications to large-scale instabilities and their control in magnetically confined fusion plasmas

JOREK is a massively parallel fully implicit non-linear extended magneto-hydrodynamic (MHD) code for realistic tokamak X-point plasmas. It has become a widely used versatile simulation code for studying large-scale plasma instabilities and their control and is continuously developed in an international community with strong involvements in the European fusion research programme and ITER organization. This article gives a comprehensive overview of the physics models implemented, numerical methods applied for solving the equations and physics studies performed with the code. A dedicated section highlights some of the verification work done for the code. A hierarchy of different physics models is available including a free boundary and resistive wall extension and hybrid kinetic-fluid models. The code allows for flux-surface aligned iso-parametric finite element grids in single and double X-point plasmas which can be extended to the true physical walls and uses a robust fully implicit time stepping. Particular focus is laid on plasma edge and scrape-off layer (SOL) physics as well as disruption related phenomena. Among the key results obtained with JOREK regarding plasma edge and SOL, are deep insights into the dynamics of edge localized modes (ELMs), ELM cycles, and ELM control by resonant magnetic perturbations, pellet injection, as well as by vertical magnetic kicks. Also ELM free regimes, detachment physics, the generation and transport of impurities during an ELM, and electrostatic turbulence in the pedestal region are investigated. Regarding disruptions, the focus is on the dynamics of the thermal quench (TQ) and current quench triggered by massive gas injection and shattered pellet injection, runaway electron (RE) dynamics as well as the RE interaction with MHD modes, and vertical displacement events. Also the seeding and suppression of tearing modes (TMs), the dynamics of naturally occurring TQs triggered by locked modes, and radiative collapses are being studied.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Address tracing of parallel systems via TRAPEDS

Trace-driven simulation is an important aid in performance analysis of computer systems. Capturing address traces to use in these simulations, however, is a difficult problem for parallel processor architectures. A technique termed TRAPEDS modifies executable code (at the assembly language level) to dynamically collect the address trace from executing code. TRAPEDS has recently been implemented on both a hypercube multicomputer and a shared-memory multiprocessor. Particular attention is focused on strategies for efficiently and accurately collecting traces from both classes of parallel machines. The iPSC/2 hypercube multicomputer implementation traces both user and system code, and performs simulation on-the-fly to avoid large storage costs. Strategies are detailed for mitigating address trace distortion when collecting operating system traces. The Encore Multimax multiprocessor implementation uses a timer-based approach to reflect the interleaving of the processor traces and stores the traces to disc. Time and space overhead results are presented for both TRAPEDS implementations. Experimental cache simulation results derived from iPSC/2 address traces are presented to illustrate the importance of tracing operating system references.

Stunkel, Craig B.↗

Application of the Loci-Based CFD Code Chem at MSFC: Preliminary Results

Contents include the following: 1. Objectives. Concentrate on determining the qualitative accuracy, performance and robustness of the Chen code. 2. What is the Loci-Chem CFD code? Density-based, finite volume, generalized unstructured grid, Navier-Stokes solver. The algorithm was implemented using the Loci framework, which allows implementation issues such as parallel processing to be handled transparently to the coding of the CFD algorithm. 3. Application to Bifurcating Duct problem. Flow splits from single duct to two ducts. 4. Application to single element injector. 5. Application to PSU RBCC rig. 6. 90 degree elbow benchmark problem. 7. Future work.

West, Jeff S.↗

TOUGH3-FLAC3D: a modeling approach for parallel computing of fluid flow and geomechanics

The recent development of the TOUGH3 code allows for a faster and more reliable fluid flow simulator. At the same time, new versions of FLAC3D are released periodically, allowing for new features and faster execution. In this paper, we present the first implementation of the coupling between TOUGH3 and FLAC3Dv6/7, maintaining parallel computing capabilities for the coupled fluid flow and geomechanical codes. We compare the newly developed version with analytical solutions and with the previous approach, and provide some performance analysis on different meshes and varying the number of running processors. Finally, we present two case studies related to fault reactivation during CO 2 sequestration and nuclear waste disposal. The use of parallel computing allows for meshes with a larger number of elements, and hence more detailed understanding of thermo-hydro-mechanical processes occurring at depth.

58 GEOSCIENCES↗

Coset Codes Viewed as Terminated Convolutional Codes

In this paper, coset codes are considered as terminated convolutional codes. Based on this approach, three new general results are presented. First, it is shown that the iterative squaring construction can equivalently be defined from a convolutional code whose trellis terminates. This convolutional code determines a simple encoder for the coset code considered, and the state and branch labelings of the associated trellis diagram become straightforward. Also, from the generator matrix of the code in its convolutional code form, much information about the trade-off between the state connectivity and complexity at each section, and the parallel structure of the trellis, is directly available. Based on this generator matrix, it is shown that the parallel branches in the trellis diagram of the convolutional code represent the same coset code C(sub 1), of smaller dimension and shorter length. Utilizing this fact, a two-stage optimum trellis decoding method is devised. The first stage decodes C(sub 1), while the second stage decodes the associated convolutional code, using the branch metrics delivered by stage 1. Finally, a bidirectional decoding of each received block starting at both ends is presented. If about the same number of computations is required, this approach remains very attractive from a practical point of view as it roughly doubles the decoding speed. This fact is particularly interesting whenever the second half of the trellis is the mirror image of the first half, since the same decoder can be implemented for both parts.

Fossorier, Marc P. C.↗

Optimization of Particle-in-Cell Codes on RISC Processors

General strategies are developed to optimize particle-cell-codes written in Fortran for RISC processors which are commonly used on massively parallel computers. These strategies include data reorganization to improve cache utilization and code reorganization to improve efficiency of arithmetic pipelines.

particle-cell-codes cache utilization code reorgan↗

The inviscid incompressible limit of Kelvin–Helmholtz instability for plasmas

The Kelvin–Helmholtz Instability (KHI) is an interface instability that develops between two fluids or plasmas flowing with a common shear layer. KHI occurs in astrophysical jets, solar atmosphere, solar flows, cometary tails, planetary magnetospheres. Two applications of interest, encompassing both space and fusion applications, drive this study: KHI formation at the outer flanks of the Earth’s magnetosphere and KHI growth from non-uniform laser heating in magnetized direct-drive implosion experiments. Here, we study 2D KHI with or without a magnetic field parallel to the flow. We use both the GAMERA code, which solves the compressible Euler equations, and the STRATOSPEC code, which solves the Navier-Stokes equations under the Boussinesq approximation, coupled with the magnetic field dynamics. GAMERA is a global three-dimensional MHD code with high-order reconstruction in arbitrary nonorthogonal curvilinear coordinates, which is developed for a large range of astrophysical applications. STRATOSPEC is a three-dimensional pseudo-spectral code with an accuracy of infinite order (no numerical diffusion). Magnetized KHI is a canonical case for benchmarking hydrocode simulations with extended MHD options. An objective is to assess whether or not, and under which conditions, the incompressibility hypothesis allows to describe a dynamic compressible system. For comparing both codes, we reach the inviscid incompressible regime, by decreasing the Mach number in GAMERA, and viscosity and diffusion in STRATOSPEC. Here, we specifically investigate both single-mode and multi-mode initial perturbations, either with or without magnetic field parallel to the flow. The method relies on comparisons of the density fields, 1D profiles of physical quantities averaged along the flow direction, and scale-by-scale spectral densities. We also address the triggering, formation and damping of filamentary structures under varying Mach number or Atwood number, with or without a parallel magnetic field. Comparisons show very satisfactory results between the two codes. The vortices dynamics is well reproduced, along with the breaking or damping of small-scale structures. We end with the extraction of growth rates of magnetized KHI from the compressible regime to the incompressible limit in the linear regime assessing the effects of compressibility under increasing magnetic field. The observed differences between the two codes are explained either from diffusion or non-Boussinesq effects.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Wing design for a civil tiltrotor transport aircraft

The goal of this research is the proper tailoring of the civil tiltrotor's composite wing-box structure leading to a minimum-weight wing design. With focus on the structural design, the wing's aerodynamic shape and the rotor-pylon system are held fixed. The initial design requirement on drag reduction set the airfoil maximum thickness-to-chord ratio to 18 percent. The airfoil section is the scaled down version of the 23 percent-thick airfoil used in V-22's wing. With the project goal in mind, the research activities began with an investigation of the structural dynamic and aeroelastic characteristics of the tiltrotor configuration, and the identification of proper procedures to analyze and account for these characteristics in the wing design. This investigation led to a collection of more than thirty technical papers on the subject, some of which have been referenced here. The review of literature on the tiltrotor revealed the complexity of the system in terms of wing-rotor-pylon interactions. The aeroelastic instability or whirl flutter stemming from wing-rotor-pylon interactions is found to be the most critical mode of instability demanding careful consideration in the preliminary wing design. The placement of wing fundamental natural frequencies in bending and torsion relative to each other and relative to the rotor 1/rev frequencies is found to have a strong influence on the whirl flutter. The frequency placement guide based on a Bell Helicopter Textron study is used in the formulation of frequency constraints. The analysis and design studies are based on two different finite-element computer codes: (1) MSC/NASATRAN and (2) WIDOWAC. These programs are used in parallel with the motivation to eventually, upon necessary modifications and validation, use the simpler WIDOWAC code in the structural tailoring of the tiltrotor wing. Several test cases were studied for the preliminary comparison of the two codes. The results obtained so far indicate a good overall agreement between the two codes.

Rais-Rohani, Masoud↗

Development of MGMC: A proxy Multi-Group Monte Carlo Particle Transport Application

This document details the development and current state of MGMC: a proxy Multi-Group Monte Carlo transport application. The goal of this work was to develop a light-weight Monte Carlo transport solver that could be easily reconfigured to test parallelization strategies targeting advanced architectures and heterogeneous computing environments. Additionally, MGMC is able to act as a proxy-app for testing the development of algorithms and the integration of other libraries and into Monte Carlo transport codes. MGMC is currently able to run in parallel using OpenMP for CPU threads or CUDA for NVIDIA GPUs. MGMC was also used to investigate the using SYCL for portable parallelization targeting either CPU threads, vender-agnostic GPUs, and ARM chips, detailed in Section II. In the process of developing MGMC, several other C++ libraries have been created with the intent for re-use in future research projects and potential inclusion in production-level codes, detailed in Section III. The physics capabilities in MGMC are detailed in Section IV, code verification is detailed in Section V, and performance results are detailed in Section VI.

97 MATHEMATICS AND COMPUTING↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Automatic Multilevel Parallelization Using OpenMP

In this paper we describe the extension of the CAPO (CAPtools (Computer Aided Parallelization Toolkit) OpenMP) parallelization support tool to support multilevel parallelism based on OpenMP directives. CAPO generates OpenMP directives with extensions supported by the NanosCompiler to allow for directive nesting and definition of thread groups. We report some results for several benchmark codes and one full application that have been parallelized using our system.

Jin, Hao-Qiang↗