Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Event Generators for Simulating Heavy Ion Interactions of Interest in Evaluating Risks in Human Spaceflight

Simulating the Space Radiation environment with Monte Carlo Codes, such as FLUKA, requires the ability to model the interactions of heavy ions as they penetrate spacecraft and crew member's bodies. Monte-Carlo-type transport codes use total interaction cross sections to determine probabilistically when a particular type of interaction has occurred. Then, at that point, a distinct event generator is employed to determine separately the results of that interaction. The space radiation environment contains a full spectrum of radiation types, including relativistic nuclei, which are the most important component for the evaluation of crew doses. Interactions between incident protons with target nuclei in the spacecraft materials and crew member's bodies are well understood. However, the situation is substantially less comfortable for incident heavier nuclei (heavy ions). We have been engaged in developing several related heavy ion interaction models based on a Quantum Molecular Dynamics-type approach for energies up through about 5 GeV per nucleon (GeV/A) as part of a NASA Consortium that includes a parallel program of cross section measurements to guide and verify this code development.

Wilson, Thomas L.↗

PHASM: A Toolkit for Creating AI Surrogate Models within Legacy Codebases

PHASM (“Parallel Hardware viA Surrogate Models”) is a software toolkit for creating AI-based surrogate models of scientific code. AI-based surrogate models are widely used for creating fast and inverse simulations. PHASM anticipates an additional future use case: adapting legacy code to modern hardware. While data centers are investing in heterogeneous hardware such as GPUs and FPGAs, many established scientific codebases remain unable to take advantage of the hardware’s higher parallelism without undergoing a costly rewrite. An alternative is to train a AI-based surrogate model to mimic computationally intensive functions in the code, and run the surrogate instead. PHASM formalizes a development lifecycle for such surrogate models, including discovering functions amenable to replacement with a surrogate model, predicting the resulting performance, identifying the function’s space of inputs and outputs, binding the model to the code, and managing model versions. A suite of software tools for facilitating these steps was written and validated against a set of model problems.

97 MATHEMATICS AND COMPUTING↗

Radiative interaction of atmosphere and surface: write up with elements of code

In passive satellite remote sensing of the Earth, separation of the path radiance (atmosphere-only contribution) from the surface reflection remains a “significant challenge”. Recent literature names it among the gaps in radiative transfer (RT) topics that “require continued research in the near future”. The challenge comes from multiple reflections (bouncing) between the atmosphere and surface – radiative interaction. In this paper we use a known RT technique, the matrix-operator method (MOM), and a new modification of the monochromatic vector RT (vRT) code IPOL (Intensity and POLarization) to simulate the interaction of a plane-parallel atmosphere and a few widely used surface reflection models. Following the idea of the Green’s function method, IPOL no longer takes the surface model parameters on input. Instead, it provides the path radiance, and the atmospheric reflection and transmission matrices as output. Despite many RT codes use the MOM formalism, this output does not seem common. The surface reflection matrix is computed externally. Therefore, this paper extends the Green’s function atmospheric correction technique to the case of polarized light. Aiming clarity rather than performance, we explain in Python the structure of the surface matrices for the isotropic (Lambertian), directional unpolarized, and polarized ocean reflection models. We then combine these surface matrices and the precomputed IPOL output to get numerically accurate signal at the top of atmosphere (TOA) and test it vs. published benchmarks. Then, for each benchmark scenario we show how to get the surface from the TOA signal, i.e. perform the RT-based atmospheric correction.

radiative transfer↗

Enhancing Application Performance Using Mini-Apps: Comparison of Hybrid Parallel Programming Paradigms

In many fields, real-world applications for High Performance Computing have already been developed. For these applications to stay up-to-date, new parallel strategies must be explored to yield the best performance; however, restructuring or modifying a real-world application may be daunting depending on the size of the code. In this case, a mini-app may be employed to quickly explore such options without modifying the entire code. In this work, several mini-apps have been created to enhance a real-world application performance, namely the VULCAN code for complex flow analysis developed at the NASA Langley Research Center. These mini-apps explore hybrid parallel programming paradigms with Message Passing Interface (MPI) for distributed memory access and either Shared MPI (SMPI) or OpenMP for shared memory accesses. Performance testing shows that MPI+SMPI yields the best execution performance, while requiring the largest number of code changes. A maximum speedup of 23 was measured for MPI+SMPI, but only 11 was measured for MPI+OpenMP.

Lawson, Gary↗

A weighted state redistribution algorithm for embedded boundary grids

State redistribution is an algorithm that stabilizes cut cells for embedded boundary grid methods. This work extends the earlier algorithm in several important ways. First, state redistribution is extended to three spatial dimensions. Second, we discuss several algorithmic changes and improvements motivated by the more complicated cut cell geometries that can occur in higher dimensions. In particular, we introduce a weighted version with less dissipation in an easily generalizable framework. Third, we demonstrate that state redistribution can also stabilize a solution update that includes both advective and diffusive contributions. Notably, the stabilization algorithm is shown to be effective for incompressible as well as compressible reacting flows. Finally, we discuss the implementation of the algorithm for several exascale-ready simulation codes based on AMReX, demonstrating ease of use in combination with domain decomposition, hybrid parallelism and complex physics.

97 MATHEMATICS AND COMPUTING↗

Structured Adaptive Mesh Refinement Adaptations to Retain Performance Portability With Increasing Heterogeneity

Adaptive mesh refinement (AMR) is an important method that enables many mesh-based applications to run at effectively higher resolution within limited computing resources by allowing high resolution only where really needed. This advantage comes at a cost, however: greater complexity in the mesh management machinery and challenges with load distribution. With the current trend of increasing heterogeneity in hardware architecture, AMR presents an orthogonal axis of complexity. Additionally, the usual techniques, such as asynchronous communication and hierarchy management for parallelism and memory that are necessary to obtain reasonable performance are very challenging to reason about with AMR. Different groups working with AMR are bringing different approaches to this challenge. Here, we examine the design choices of several AMR codes and also the degree to which demands placed on them by their users influence these choices.

42 ENGINEERING↗

Sierra/SolidMechanics 5.6 User's Manual

Sierra/SolidMechanics (Sierra/SM) is a three-dimensional solid mechanics code with a versatile element library, nonlinear material models, large deformation capabilities, and contact. It is built on the SIERRA Framework. SIERRA provides a data management framework in a parallel computing environment that allows the addition of capabilities in a modular fashion. Contact capabilities are parallel and scalable. This document provides information about the functionality in Sierra/SM and the command structure required to access this functionality in a user input file. This document is divided into chapters based primarily on functionality. For example, the command structure related to the use of various element types is grouped in one chapter; descriptions of material models are grouped in another chapter. Sierra/SM provides both explicit transient dynamics and implicit quasistatics and dynamics capabilities. Both the explicit and implicit modules are highly scalable in a parallel computing environment. In the past, the explicit and implicit capabilities were provided by two separate codes, known as Presto and Adagio, respectively. These capabilities have been consolidated into a single code. The executable is named Adagio, but it provides the full suite of solid mechanics capabilities, for both implicit and explicit. The Presto executable has been disabled as a consequence of this consolidation.

97 MATHEMATICS AND COMPUTING↗

Introduction of the ASGARD Code

ASGARD stands for 'Automated Selection and Grouping of events in AIA Regional Data'. The code is a refinement of the event detection method in Ugarte-Urra & Warren (2014). It is intended to automatically detect and group brightenings ('events') in the AIA EUV channels, to record event parameters, and to find related events over multiple channels. Ultimately, the goal is to automatically determine heating and cooling timescales in the corona and to significantly increase statistics in this respect. The code is written in IDL and requires the SolarSoft library. It is parallelized and can run with multiple CPUs. Input files are regions of interest (ROIs) in time series of AIA images from the JSOC cutout service (http://jsoc.stanford.edu/ajax/exportdata.html). The ROIs need to be tracked, co-registered, and limited in time (typically 12 hours).

Code development↗

SPEL: Software tool for Porting E3SM Land Model with OpenACC in a Function Unit Test Framework

Most high-end computers adopt hybrid architecture, porting a large-scale scientific code onto accelerators is necessary. The paper presents a generic method for porting large-scale scientific code onto accelerators using compiler directives within a modularized function unit test platform. We have implemented the method and designed a software tool (SPEL) to port the E3SM Land Model (ELM) onto the GPUs in the Summit computer. SPEL automatically generates GPU-ready test modules for all ELM functions, such as CanopyFlux, SoilTemperature, and EcosystemDynamics. SPEL breaks the ELM into a collection of standalone unit test programs for easy code verification and further performance improvement. We further optimize several ELM test modules with advanced techniques, including memory reduction, reconstructed parallel loops, and asynchronous GPU kernel launch. We hope our study will inspire new toolkit developments that expedite large-scale scientific code porting with compiler directives.

Schwartz, Peter↗

MEUMAPPS (Microstructure Evolution Using Massively Parallel Phase-field Simulations)

The software “MEUMAPPS” is a h high-performance computing code used to simulate microstructure evolution associated with diffusional solid-state transformations in structural alloys. Understanding microstructure evolution during thermo-mechanical processing of structural alloys is the first step towards designing and processing alloys for specific applications by meeting property requirements demanded by the application. In this instance, the code is an important component in predicting processing-structure linkages during additive manufacturing of structural alloys as a function of processing parameters and alloy composition used in an additive manufacturing process. The code provides a detailed three-dimensional distribution of different types of phases / constituents and their morphologies, as well as the compositions of these phases that make up the microstructure. The microstructure is simulated by solving the governing partial differential equations using a Fourier Spectral Method.

Radhakrishnan, Balasubram↗

RX-PSA

This code is built on top of the ML-PSA utility and implements the ability to implement a multi-fidelity parallel simulated annealing optimization for a range of engineering problems. RX-PSA provides specializations to interact with modern nuclear reactor codes such as VERA to perform assembly and core optimization in a robust way. RX-PSA also provides LWR specific objective functions and constraints to provide a straightforward user interface that reactor designers are familiar with.

Collins, BenjaminS.↗

Simulation of electrostatic turbulence due to sheared flows parallel and transverse to the magnetic field

A spatially two-dimensional electrostatic particle simulation code is used to examine the stability of a plasma equilibrium characterized by a localized transverse dc electric field and a magnetic-field-aligned electron drift for L much less than Lx, where Lx is the simulation length in the x direction and L is the scale length associated with the dc electric field. It is found that the dc electric field and the field-aligned current can together play a synergistic role to enable the excitation of electrostatic waves even when the threshold values of the field-aligned drift and the E x B drift are individually subcritical. The simulation results indicate that a broadband turbulence is associated with such an equilibrium.

Nishikawa, K.-I.↗

Release Factor Inhibiting Antimicrobial Peptides Improve Nonstandard Amino Acid Incorporation in Wild-type Bacterial Cells

We report a tunable chemical genetics approach for enhancing genetic code expansion in different wild-type bacterial strains that employs apidaecin-like, anti-microbial peptides observed to temporarily sequester and thereby inhibit Release Factor 1 (RF1). In a concentration-dependent matter, these peptides granted a conditional lambda phage resistance to a recoded Escherichia coli strain with non-essential RF1 activity and promoted multi-site non-standard amino acid (nsAA) incorporation at in-frame amber stop codons in vivo and in vitro. When exogenously added, the peptides stimulated specific nsAA incorporation in a variety of sensitive, wild-type (RF1+) strains including Agrobacterium tumefaciens, a species in which nsAA incorporation has not been previously reported. Improvement in nsAA incorporation was typically 2–15-fold in E. coli BL21, MG1655, DH10B strains and A. tumefaciens with the >20-fold improvement observed in probiotic E. coli Nissle 1917. In-cell expression of these peptides promoted multi-site nsAA incorporation in transcripts with up to 6 amber codons, with a >35-fold increase in BL21 showing moderate toxicity. Leveraging this RF1 sensitivity allowed multiplexed partial recoding of MG1655 and DH10B that rapidly resulted in resistant strains that showed an additional ~2 fold boost to nsAA incorporation independent of the peptide. Lastly, in-cell expression of an apidaecin-like peptide library allowed the discovery of new peptide variants with reduced toxicity that still improved multi-site nsAA incorporation >25-fold. In parallel to genetic reprogramming efforts, these new approaches can facilitate genetic code expansion technologies in a variety of wild-type bacterial strains.

59 BASIC BIOLOGICAL SCIENCES↗

High-Performance, Low-Complexity Codes Researched for Communication Channels

NASA Lewis Research Center s Communications Technology Division has an ongoing program in the development of efficient channel coding schemes for satellite communications applications. Through a university grant, as a part of this research, the University of Toledo is investigating the performance of turbocodes, which use parallel concatenation of non-systematic convolutional encoders with an interleaver. The error correcting capacity of these codes is close to the Shannon limit. The research emphasis is on the development of low-complexity, but higher rate (greater than one half), turbocodes and on the iterative decoding of block codes.

Kwatra, Subhash C.↗

SCEPTRE 2.2 Quick Start Guide

This report provides a summary of notes for building and running the Sandia Computational Engine for Particle Transport for Radiation Effects (SCEPTRE) code. SCEPTRE is a general- purpose C++ code for solving the li near Boltzmann transport equation in serial or parallel using unstructured spatial finite elements, multigroup energy treatment, and a variety of angular treatments including discrete ordinates and spherical harmonics. Either the first-order form of the Boltzmann equation or one of the second-order forms may be solved. SCEPTRE requires a small number of open-source Third Part y Libraries (TPL) to be available, and example scripts for building these TPLs are provided. The TPLs needed by SCEPTRE are Trilinos, boost, and netcdf. SCEPTRE uses an autotools build system , and a sample configure script is provided. Running the SCEPTRE code requires that the user provide a spatial finite-elements mesh in Exodus format and a cross section library in a format that will be described. SCEPTRE uses an xml-based input, and several examples will be provided.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

GeantV: Results from the Prototype of Concurrent Vector Particle Transport Simulation in HEP

Full detector simulation was among the largest CPU consumers in all CERN experiment software stacks for the first two runs of the Large Hadron Collider. In the early 2010s, it was projected that simulation demands would scale linearly with increasing luminosity, with only partial compensation from increasing computing resources. The extension of fast simulation approaches to cover more use cases that represent a larger fraction of the simulation budget is only part of the solution, because of intrinsic precision limitations. The remainder corresponds to speeding up the simulation software by several factors, which is not achievable by just applying simple optimizations to the current code base. In this context, the GeantV R&D project was launched, aiming to redesign the legacy particle transport code in order to benefit from features of fine-grained parallelism, including vectorization and increased locality of both instruction and data. This paper provides an extensive presentation of the results and achievements of this R&D project, as well as the conclusions and lessons learned from the beta version prototype.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

FLARE: field line analysis and reconstruction for 3D boundary plasma modeling

The FLARE code is a magnetic mesh generator that is integrated within a suite of tools for the analysis of the magnetic geometry in toroidal fusion devices. A magnetic mesh is constructed from field line segments and permits fast reconstruction of field lines in 3D boundary plasma codes such as EMC3-EIRENE. Both intrinsically non-axisymmetric configurations (stellarators) and those with symmetry breaking perturbations of an axisymmetric equilibrium (tokamaks) are supported. The code itself is written in Modern Fortran with MPI support for parallel computing, and it incorporates object-oriented programming for the definition of the magnetic field and the material surface geometry. Extended derived types for a number of different magnetohydrodynamic equilibrium and plasma response models are implemented. The core element of FLARE is a field line tracer with adaptive step-size control, and this is integrated into tools for the construction of Poincaré maps and invariant manifolds of X-points. A collection of high-level procedures that generate output files for visualization is build on top of that. The analysis modules are build with Python frontends that facilitate customization of tasks and/or scripting of parameter scans.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗