Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28

PyOMP: Multithreaded Parallel Programming in Python

We know that Python is a widely used language in scientific computing. When the goal is high performance, however, Python lags far behind low-level languages such as C and Fortran. To support applications that stress performance, Python needs to access the full capabilities of modern CPUs. That means support for parallel multithreading. In this paper, we describe PyOMP, a system that enables OpenMP in Python. Programmers write code in Python with OpenMP, Numba generates code that compiles to LLVM, and the resulting programs run with performance that approaches that from code written with C and OpenMP. In this paper we provide an update on the PyOMP project and explain how to install it and use it to write parallel multithreaded code in Python.

97 MATHEMATICS AND COMPUTING↗

Multitasking for flows about multiple body configurations using the chimera grid scheme

The multitasking of a finite-difference scheme using multiple overset meshes is described. In this chimera, or multiple overset mesh approach, a multiple body configuration is mapped using a major grid about the main component of the configuration, with minor overset meshes used to map each additional component. This type of code is well suited to multitasking. Both steady and unsteady two dimensional computations are run on parallel processors on a CRAY-X/MP 48, usually with one mesh per processor. Flow field results are compared with single processor results to demonstrate the feasibility of running multiple mesh codes on parallel processors and to show the increase in efficiency.

Dougherty, F. C.↗

Parallel ALLSPD-3D: Speeding Up Combustor Analysis Via Parallel Processing

The ALLSPD-3D Computational Fluid Dynamics code for reacting flow simulation was run on a set of benchmark test cases to determine its parallel efficiency. These test cases included non-reacting and reacting flow simulations with varying numbers of processors. Also, the tests explored the effects of scaling the simulation with the number of processors in addition to distributing a constant size problem over an increasing number of processors. The test cases were run on a cluster of IBM RS/6000 Model 590 workstations with ethernet and ATM networking plus a shared memory SGI Power Challenge L workstation. The results indicate that the network capabilities significantly influence the parallel efficiency, i.e., a shared memory machine is fastest and ATM networking provides acceptable performance. The limitations of ethernet greatly hamper the rapid calculation of flows using ALLSPD-3D.

Fricker, David M.↗

Simulating coupled surface–subsurface flows with ParFlow v3.5.0: capabilities, applications, and ongoing development of an open-source, massively parallel, integrated hydrologic model

Surface flow and subsurface flow constitute a naturally linked hydrologic continuum that has not traditionally been simulated in an integrated fashion. Recognizing the interactions between these systems has encouraged the development of integrated hydrologic models (IHMs) capable of treating surface and subsurface systems as a single integrated resource. IHMs are dynamically evolving with improvements in technology, and the extent of their current capabilities are often only known to the developers and not general users. This article provides an overview of the core functionality, capability, applications, and ongoing development of one open-source IHM, ParFlow. ParFlow is a parallel, integrated, hydrologic model that simulates surface and subsurface flows. ParFlow solves the Richards equation for three-dimensional variably saturated groundwater flow and the two-dimensional kinematic wave approximation of the shallow water equations for overland flow. The model employs a conservative centered finite-difference scheme and a conservative finite-volume method for subsurface flow and transport, respectively. ParFlow uses multigrid-preconditioned Krylov and Newton–Krylov methods to solve the linear and nonlinear systems within each time step of the flow simulations. The code has demonstrated very efficient parallel solution capabilities. ParFlow has been coupled to geochemical reaction, land surface (e.g., the Common Land Model), and atmospheric models to study the interactions among the subsurface, land surface, and atmosphere systems across different spatial scales. This overview focuses on the current capabilities of the code, the core simulation engine, and the primary couplings of the subsurface model to other codes, taking a high-level perspective.

58 GEOSCIENCES↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

Implementation of a Parallel Kalman Filter for Stratospheric Chemical Tracer Assimilation

A Kalman filter for the assimilation of long-lived atmospheric chemical constituents has been developed for two-dimensional transport models on isentropic surfaces over the globe. An important attribute of the Kalman filter is that it calculates error covariances of the constituent fields using the tracer dynamics. Consequently, the current Kalman-filter assimilation is a five-dimensional problem (coordinates of two points and time), and it can only be handled on computers with large memory and high floating point speed. In this paper, an implementation of the Kalman filter for distributed-memory, message-passing parallel computers is discussed. Two approaches were studied: an operator decomposition and a covariance decomposition. The latter was found to be more scalable than the former, and it possesses the property that the dynamical model does not need to be parallelized, which is of considerable practical advantage. This code is currently used to assimilate constituent data retrieved by limb sounders on the Upper Atmosphere Research Satellite. Tests of the code examined the variance transport and observability properties. Aspects of the parallel implementation, some timing results, and a brief discussion of the physical results will be presented.

Chang, Lang-Ping↗

Fully-Implicit Navier-Stokes (FIN-S)

FIN-S is a SUPG finite element code for flow problems under active development at NASA Lyndon B. Johnson Space Center and within PECOS: a) The code is built on top of the libMesh parallel, adaptive finite element library. b) The initial implementation of the code targeted supersonic/hypersonic laminar calorically perfect gas flows & conjugate heat transfer. c) Initial extension to thermochemical nonequilibrium about 9 months ago. d) The technologies in FIN-S have been enhanced through a strongly collaborative research effort with Sandia National Labs.

Kirk, Benjamin S.↗

Micro- and meso-scale simulations of magnetospheric processes related to the aurora and substorm morphology

The primary methodology during the grant period has been the use of micro or meso-scale simulations to address specific questions concerning magnetospheric processes related to the aurora and substorm morphology. This approach, while useful in providing some answers, has its limitations. Many of the problems relating to the magnetosphere are inherently global and kinetic. Effort during the last year of the grant period has increasingly focused on development of a global-scale hybrid code to model the entire, coupled magnetosheath - magnetosphere - ionosphere system. In particular, numerical procedures for curvilinear coordinate generation and exactly conservative differencing schemes for hybrid codes in curvilinear coordinates have been developed. The new computer algorithms and the massively parallel computer architectures now make this global code a feasible proposition. Support provided by this project has played an important role in laying the groundwork for the eventual development or a global-scale code to model and forecast magnetospheric weather.

Swift, Daniel W.↗

Task 7: ADPAC User's Manual

The overall objective of this study was to develop a 3-D numerical analysis for compressor casing treatment flowfields. The current version of the computer code resulting from this study is referred to as ADPAC (Advanced Ducted Propfan Analysis Codes-Version 7). This report is intended to serve as a computer program user's manual for the ADPAC code developed under Tasks 6 and 7 of the NASA Contract. The ADPAC program is based on a flexible multiple- block grid discretization scheme permitting coupled 2-D/3-D mesh block solutions with application to a wide variety of geometries. Aerodynamic calculations are based on a four-stage Runge-Kutta time-marching finite volume solution technique with added numerical dissipation. Steady flow predictions are accelerated by a multigrid procedure. An iterative implicit algorithm is available for rapid time-dependent flow calculations, and an advanced two equation turbulence model is incorporated to predict complex turbulent flows. The consolidated code generated during this study is capable of executing in either a serial or parallel computing mode from a single source code. Numerous examples are given in the form of test cases to demonstrate the utility of this approach for predicting the aerodynamics of modem turbomachinery configurations.

Hall, E. J.↗

A Framework for the Analysis of Compiler Optimizations

Compilers transform program source code to machine executable code. During this transformation, they perform a number of compiler optimizations to improve the performance of the generated executable code. Importantly, applying those optimizations depends on the source code structure, such as the parallel programming model used to parallelize an algorithm. Often, implementations of the same algorithm with different programming models have vastly different performance because the compiler optimized them differently. We create FAROS, a framework to structure and automate the analysis of compiler optimizations on programs. FAROS automates the building process, execution profiling, and analysis of compiler optimization of programs, through a configuration interface. It outputs compiler optimization reports to show which optimizations applied to which line of source code, leveraging compilation remarks output by the compiler. Also, FAROS supports benchmarking performance of different program versions by collecting execution time results. In this first release of FAROS, we provide a configuration file to analyze compiler optimization differences for sequential vs OpenMP compilation, including 38 programs consisting of HPC proxy/mini/large applications, and NAS and Rodina kernels for analysis.

Georgakoudis, Giorgis↗

All-Atom Simulation of 3D Hot Spot Formation in Shocked TATB Explosive

TATB is an insensitive high explosive (IHE) critical to the stockpile that is challenging to model at the continuum scale. Advanced detonation models in the Cheetah high explosive chemistry code require validation though subscale simulations. High explosive initiation is determined by micron-scale physics of hot spots formed a shock-collapsed pores. Pore sizes between 100 nm and 1 μm are believed to be the most important for determining the shock sensitivity of TATB. This range of pore sizes is difficult to access at the atomic scale through allatom molecular dynamics (MD) simulations, even with Sierra-class computers. Quasi-2D simulations are widely used and allow much larger pore sizes (up to 400 nm) to be studied, but the applicability of 2D simulations to the actual 3D pore response is not understood. Resolving these uncertainties through “full physics” MD modeling is key for generalizing, parameterizing, and validating the kinds of continuum models used to inform design, safety, and performance. This work was a continuation of FY20 efforts pushing simulations to full 3D with the largest-ever all-atom simulations of an explosive. These were the first all-atom full-3D simulations of large hot spots thought to govern explosive detonation and required over a billion atoms. Simulations were performed using LAMMPS, an open SNL science code. MD explosive models present unique challenges, even for established codes such as LAMMPS. Their model forms are more complex than typical models for metals, while simulating high temperature-pressure conditions is demanding and increases computational cost. Scaling problems in GPU-enabled MD algorithms initially limited simulations to <100 million atoms but were resolved through collaboration with SNL. An overall 24x speedup was obtained relative to CPU machines. Specialized analysis of these simulations required a bottom-up refactoring and algorithm parallelization of in-house codes and application of computer vision algorithms to extract meaningful information.

36 MATERIALS SCIENCE↗

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Recent Advancements in the PATO Material Response Code

Introduction: Predicting the complicated multiphysics phenomena during atmospheric entry requires high-fidelity modeling tools to refine estimates of mission risks during entry. To this end, new capabilities are being added to the Porous-material Analysis Toolbox based on OpenFOAM (PATO). PATO is an open-source software for Computational Material Response (CMR) of reactive porous materials submitted to high-temperature environments. The objective of this work is to highlight current efforts to add to and improve upon the modeling capabilities of PATO. These include efforts to loosely couple PATO with other discipline specialized codes including hypersonic Computational Fluid Dynamics (CFD), to assess the interaction effects between pyrolysis gas blowing and the boundary layer, and Computational Solid Mechanics (CSM), to address modeling of mechanical erosion. Other refinements include surface phenomena modeling capabilities to address the effects of silicone-based coatings applied to the TPS during flight preparation, and a unified multiphase solver for a mixed porous-material and plain-fluid domain. Coupling CMR with CFD (CMR/CFD): A loose coupling between PATO and the Data Parallel Line Relaxation (DPLR) CFD code has been achieved by making use of a blowing boundary condition at the heatshield surface available in DPLR. Starting with heat flux estimates with no pyrolysis gas blowing at the surface, blowing gases are computed by the CMR and passed to the CFD such that aerothermal properties of the environment can be recomputed for a new CMR computation. This leads to an iterative process which is supplemented with an estimate of the radiative heat flux using the Nonequilibrium air radiation (NEQAIR) program. The entire iterative process is illustrated in Figure 1. This coupling strategy has been utilized in computing the MSL material response. The goal is to compare the coupled CMR/CFD results with material response results obtained using traditional blowing corrections. Coupling CMS with CMR: A mechanical erosion model is currently being implemented in PATO to account for the additional mass removal induced by high shear conditions. The modeling process at each timestep consists of updating the mechanical properties as a function of temperature and computing the stress tensor and displacement fields of the material. Then, a failure criteria model determines the regions in which the stress exceeds the ultimate strength values resulting in mesh movement to account for mass removal. This model allows the material response simulation to compute the recession due to both oxidation and shear-induced erosion. The model is demonstrated by computing material response of sphere-cone arc jet samples. Surface Modeling Capabilities: NuSil, a silicone-based coating, was sprayed onto the MSL and Mars 2020 heatshields to mitigate shedding of phenolic dust. To better understand the effects of the NuSil coating on the material response, a novel model has been implemented in PATO. In this model, the equilibrium of the charred NuSil surface is modeled as pure silica, and a constant offset, inspired by the classical spallation model, is added to the the char blowing rate and wall enthalpy to reproduce HyMETS experimental results. The model has also been used to estimate the 3D material response of the MSL heatshield. Unified Solver: In addition to the iterative loose coupling approach mentioned above, a multiphase unified solver is being developed to couple the environment (plain-fluid phase) and the porous-material phase. The solver is based on the volume averaged conservation of mass, momentum, and energy for the macroscale with closure models which include microscale effects through effective physicochemical properties. The unified solver has been used to compute flow through a porous plug and solve the Beavers and Joseph problem. Since the strong coupling between phases is inherent to this solver, modeling assumptions present in other coupling methods of material response are mitigated. This strategy also makes it feasible to capture the competition between surface and volume ablation in the same computational domain, which is usually not possible with other coupling approaches.

Material Response↗

MPIRUN: A Portable Loader for Multidisciplinary and Multi-Zonal Applications

Multidisciplinary and multi-zonal applications are an important class of applications in the area of Computational Aerosciences. In these codes, two or more distinct parallel programs or copies of a single program are utilized to model a single problem. To support such applications, it is common to use a programming model where a program is divided into several single program multiple data stream (SPMD) applications, each of which solves the equations for a single physical discipline or grid zone. These SPMD applications are then bound together to form a single multidisciplinary or multi-zonal program in which the constituent parts communicate via point-to-point message passing routines. One method for implementing the message passing portion of these codes is with the new Message Passing Interface (MPI) standard. Unfortunately, this standard only specifies the message passing portion of an application, but does not specify any portable mechanisms for loading an application. MPIRUN was developed to provide a portable means for loading MPI programs, and was specifically targeted at multidisciplinary and multi-zonal applications. Programs using MPIRUN for loading and MPI for message passing are then portable between all machines supported by MPIRUN. MPIRUN is currently implemented for the Intel iPSC/860, TMC CM5, IBM SP-1 and SP-2, Intel Paragon, and workstation clusters. Further, MPIRUN is designed to be simple enough to port easily to any system supporting MPI.

Fineberg, Samuel A.↗

Line-drawing algorithms for parallel machines

The fact that conventional line-drawing algorithms, when applied directly on parallel machines, can lead to very inefficient codes is addressed. It is suggested that instead of modifying an existing algorithm for a parallel machine, a more efficient implementation can be produced by going back to the invariants in the definition. Popular line-drawing algorithms are compared with two alternatives; distance to a line (a point is on the line if sufficiently close to it) and intersection with a line (a point on the line if an intersection point). For massively parallel single-instruction-multiple-data (SIMD) machines (with thousands of processors and up), the alternatives provide viable line-drawing algorithms. Because of the pixel-per-processor mapping, their performance is independent of the line length and orientation.

Pang, Alex T.↗

Rapid Prediction of Unsteady Three-Dimensional Viscous Flows in Turbopump Geometries

A program is underway to improve the efficiency of a three-dimensional Navier-Stokes code and generalize it for nozzle and turbopump geometries. Code modifications have included the implementation of parallel processing software, incorporation of new physical models and generalization of the multiblock capability. The final report contains details of code modifications, numerical results for several nozzle and turbopump geometries, and the implementation of the parallelization software.

Dorney, Daniel J.↗

NWChem: Past, present, and future

Specialized computational chemistry packages have permanently reshaped the landscape of chemical sciences by providing tools to support and guide the experimental effort and for prediction of chemical and materials properties. In this regard, a special role has been played by electronic structure packages where complex chemical and materials processes can be modeled using first-principle-driven methodologies. Over the last few decades, the rapid development of computing technologies and a tremendous increase in computational power has offered a unique chance to study complex chemical transformations using sophisticated and predictive many-body techniques to describe correlated behavior of electrons in molecular and condensed phase systems at different levels of theory. In enabling these simulations, a critical role has been played by novel parallel algorithms capable of taking advantage of computational resources to address polynomial scaling of electronic structure methods. NWChem was among the first electronic structure codes that focused on delivering scalable performance for electronic structure simulations. Herein, we briefly review the NWChem suite of computational codes including its history, design principles, parallel tools, current capabilities, outreach and outlook.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Initial Development of Fusion Magnet Simulation Capabilities for Performance and Safety Evaluation Using the MOOSE Framework

Fusion energy holds the promise of being a transformative technology as a carbon-neutral, sustainable source of energy. Whole device modeling and the development of fusion digital twins will be increasingly important for emerging fusion device concepts at both national laboratories and within the commercial fusion industry. However, meeting the challenge of whole device modeling of fusion energy devices requires robust, multiphysics, multiscale modeling and simulation technologies capable of running on large-scale supercomputers. Detailed analysis of individual systems at-scale is also required to ensure safe and efficient operation as well as provide the safety basis for future device designs and licensing activities. In a tokamak, toroidal and poloidal magnets confine and shape the fusion plasma to promote the fusion reaction. High plasma temperatures and high magnetic field requirements in modern design concepts (leading to high amounts of energy stored within each magnet) impose electrical, thermal, and mechanical loads on the magnet components, which in turn impacts the safety considerations of the magnet and their supporting systems. Idaho National Laboratory (INL) has a history of working in this space, including development and benchmarking of the Magnetic System Circuitry Analysis Program (MSCAP) and Magnet Arcing (MAGARC) codes to study magnet quench events; notably, MAGARC was used to study quenching during the ITER Engineering Design Activity. However, these legacy codes and capabilities are not parallel and scalable, and new tools are required for future advances in this area, which leads to the INL-developed Multiphysics Object-Oriented Simulation Environment (MOOSE) framework. Developed originally for fission reactor systems under United States Department of Energy, Office of Nuclear Energy modeling and simulation programs, the MOOSE framework has been well-suited to multiscale, multiphysics modeling and simulation needs for nuclear systems. The framework is open-source, well-tested, under continuous development and deployment, and developed to a Nuclear Quality Assurance, Level 1 software quality standard. MOOSE has also been used in the fusion space previously in several projects: INL’s Tritium Migration Analysis Program, Version 8 (TMAP8) for tritium migration, UK Atomic Energy Authority’s A Unified Resource for OpenMC (fusion) Reactor Applications (AURORA) code for fusion thermo-mechanical and neutronics analysis, and Argonne National Laboratory’s Cardinal for high-fidelity computational fluid dynamics and neutronics. However, to model superconducting magnets, several MOOSE enhancements are required: additions to the current MOOSE electromagnetic capabilities, new material libraries for superconductors of interest (such as YBCO), as well as fusion-specific models for thermo-mechanics. This talk will discuss initial development activities to build these capabilities in MOOSE, focusing on initial validation and benchmarking activities. Proposed coupling workflows and future work to support the simulation of fusion magnets and magnet structural assemblies for performance and safety evaluation in MOOSE will also be discussed.

70 - PLASMA PHYSICS AND FUSION TECHNOLOGY↗