Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Computational Astrophysics in the Era of Technological Heterogeneity [Slides]

The HPC landscape is changing and we’re headed toward an era where compute specialization will be prevalent. There are opportunities for co-design that can influence this future. Technological heterogeneity will be a major challenge unless we shift our approach to developing the computational tools for astrophysics. Parthenon provides convenient functionality and a bright future for block structured AMR applications. Phoebus is a new (soon-to-be) open source code for relativistic astro that promises excellent performance, portability, and unique physics capabilities.

79 ASTRONOMY AND ASTROPHYSICS↗

Spiner

Spiner is a library for storing, indexing, and interpolating multidimensional data in a performance-portable way. It’s intended to run on CPUs, GPUs and everything in-between. You can create a table on a CPU, copy it to a GPU, and interpolate on it in a GPU kernel, for example. Spiner also defines (via hdf5) a file format that bundles data together with instructions for interpolating it. This means you don’t have to specify anything to start interpolating, simple load the file and evaluate where you want.

97 MATHEMATICS AND COMPUTING↗

FY2022 Q4: Demonstrate multi-turbine simulation with hybrid-structured/unstructured-moving-grid software stack running primarily on GPUs and propose improvements for successful KPP-2 [Poster]

Milestone accomplishments staged the ExaWind team for successful completion of KPP-2 challenge problem in FY23, which requires the simulation on Frontier of at least four MW-scale turbines in an atmospheric boundary layer with at least 20B gridpoints. The ExaWind project and software stack is many faceted, with team members working on multiple areas, including linear-system solvers (Trilinos, hypre, AMReX), overset meshes, turbulence modeling, and in situ visualization, all with an aim for high fidelity predictions and performance portability. This milestone marks significant improvements on many fronts and provides the team with a pathway to exascale wind farm simulations in FY23.

17 WIND ENERGY↗

Extending PETSc's Composable Hierarchical Solvers (Final Technical Report)

This report documents research activities conducted at CU Boulder as part of Extending PETSc’s Composable Hierarchical Solvers, which has been part of a collaboration with Argonne National Laboratory (separate award). Our work has focused on performance-portable end-to-end GPU solvers demonstrated via exemplary applications in nonlinear fluid and structural mechanics. We describe advances in algorithmic composition and analysis in the context of these applications, but the implementations are fully documented and decoupled, and in use by other projects. We believe the vertical integration achieved through collaboration with ECP’s CEED and the PSAAP center at CU was necessary to take risks with data structures and algorithms.

42 ENGINEERING↗

To Interoperability And Beyond: Interoperable Types Through the Promises of C and C++ and ABI Abuse [Slides]

This presentation presents a technique that allows passing of Fortran nested-type hierarchies interoperably to C++. The technique is motivated by the Eulerian Application Project’s need to port code from Fortran to C++ to utilize the Kokkos performance portability library. The resulting method allows hierarchies of types to become interoperable while at the same time transforming Fortran array members of the original Fortran type into Kokkos::Views in the resulting C++ type.

97 MATHEMATICS AND COMPUTING↗

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

CabanaPD

CabanaPD is a meshfree peridynamics application built with Cabana and Kokkos. Kokkos enables performance portability across hardware architectures and Cabana provides particle capabilities including multi-node MPI support. The main components of CabanaPD are particle initialization, neighbor list generation, force and energy computation, time integration, and multi-node particle communication. In addition, options for creating pre-cracked regions and particle boundary conditions are available. CabanaPD currently enables two common force models: prototype microelastic brittle (PMB) and linear peridynamic solid (LPS). For both of these models, versions with and without fracture as well as linearized model options are available. CabanaPD is designed to be extensible for addition of more complex force models, boundary conditions, etc.

Reeve, Sam↗

Developing Information Power Grid Based Algorithms and Software

This exploratory study initiated our effort to understand performance modeling on parallel systems. The basic goal of performance modeling is to understand and predict the performance of a computer program or set of programs on a computer system. Performance modeling has numerous applications, including evaluation of algorithms, optimization of code implementations, parallel library development, comparison of system architectures, parallel system design, and procurement of new systems. Our work lays the basis for the construction of parallel libraries that allow for the reconstruction of application codes on several distinct architectures so as to assure performance portability. Following our strategy, once the requirements of applications are well understood, one can then construct a library in a layered fashion. The top level of this library will consist of architecture-independent geometric, numerical, and symbolic algorithms that are needed by the sample of applications. These routines should be written in a language that is portable across the targeted architectures.

Dongarra, Jack↗

Comparison of Knee and Ankle Dynamometry between NASA's X1 Exoskeleton and Biodex System 4

Pre- and post-flight dynamometry is performed on International Space Station crewmembers to characterize microgravity-induced strength changes. Strength is not assessed in flight due to hardware limitations and there is poor understanding of the time course of in-flight changes. PURPOSE: To assess the reliability of a prototype dynamometer, the X1 Exoskeleton (EXO) and its agreement with a Biodex System 4 (BIO). METHODS: Eight subjects (4 M/4 F) completed 2 counterbalanced testing sessions of knee extension/flexion (KE/KF), 1 with BIO and 1 with EXO, with repeated measures within each session in normal gravity. Test-retest reliability (test 1 and 2) and device agreement (BIO vs. EXO) were evaluated. Later, to assess device agreement for ankle plantarflexion (PF), 10 subjects (4 M/6 F) completed 3 test conditions (BIO, EXO, and BIOEXO); BIOEXO was a hybrid condition comprised of the Biodex dynamometer motor and the X1 footplate and ankle frame. Ankle comparisons were: BIO vs. BIOEXO (footplate differences), BIOEXO vs. EXO (motor differences), and BIO vs. EXO (all differences). Reliability for KE/KF was determined by intraclass correlation (ICC). Device agreement was assessed with: 1) repeated measures ANOVA, 2) a measure of concordance (rho), and 3) average difference. RESULTS: ICCs for KE/KF were 0.99 for BIO and 0.96 to 0.99 for EXO. Agreement was high for KE (concordance: 0.86 to 0.95; average differences: -7 to +9 Nm) and low to moderate for KF (concordance: 0.64 to 0.78; average differences: -4 to -29 Nm, P<0.05). BIO vs. BIOEXO PF concordance ranged from 0.89 to 0.92 and mean differences ranged from -9 to +3 Nm (BIO < BIOEXO). BIOEXO vs. EXO PF concordance ranged from 0.73 to 0.80 while mean differences were -18 to -36 Nm (BIOEXO < EXO, P<0.05). PF concordance for BIO vs. EXO was slightly lower (0.61 to 0.84) and mean differences were greater (-27 to -33 Nm; BIO < EXO, P<0.05). CONCLUSION: BIO and EXO were similarly reliable for KE and KF. KE measures produced high agreement between devices; KF did not. For ankle PF, torque differences due to the two footplates were small. However, the X1 motor reports greater torques than the Biodex motor during PF. This first prototype provides proof of concept for a reliable, robotic-based exoskeleton to perform portable dynamometry for large muscle groups of the lower body.

English, K. L.↗

Towards Generic Parallel Programming in Computer Science Education with Kokkos

Parallel patterns, views, and spaces are promising abstractions to capture the programmer's intent as well as the contextual information that can be used by an underlying runtime to efficiently map software to parallel hardware. These abstractions can be valuable in cases where an algorithm must accommodate requirements of code and performance portability across hardware architectures and vendor programming models. Kokkos is a parallel programming model for host- and accelerator architectures that relies on these abstractions and targets these requirements. It consists of a pure C++ interface, a specification, and a programming library. The programming library exposes patterns and types and maps them to an underlying abstract machine model. The abstract machine model offers a generic view of parallel hardware. While Kokkos is gaining popularity in large-scale HPC applications at some DOE laboratories, we believe that the implemented concepts are of interest to a broader audience including academia as they may contribute to a generic, vendor, and architecture-independent education of parallel programming. In this work, we give an insight into the design considerations of this programming model and list important abstractions. Further, we document best practices obtained from giving virtual classes on Kokkos and give pointers to resources that the reader may consider valuable for a lecture on generic parallel programming for students with preexisting knowledge on this matter.

Ciesko, Jan↗

autoGEMM: Pushing the Limits of Irregular Matrix Multiplication on Arm Architectures

This paper presents an open-source library that pushes the limits of performance portability for irregular General Matrix Multiplication (GEMM) on the widely-used Arm architectures. Our library, autoGEMM, is designed to support a wide range of Arm processors: from edge devices to HPC-grade CPUs. autoGEMM generates optimized kernels for various hardware configurations by auto-combining fragments of autogenerated micro-kernels that employ hand-written optimizations to maximize computational efficiency. We optimize the kernel pipeline by tuning the register reuse and the data load/store overlapping. In addition, we use a dynamic tiling scheme to generate balanced tile shapes. Finally, we position autoGEMM on top of the TVM framework where our dynamic tiling scheme prunes the search space for TVM to identify the optimal combination of parameters for code optimization. Evaluations on five different classes of Arm chips demonstrate the advantages of autoGEMM. For small matrices, autoGEMM achieves 98% of peak and up to 2.0x speedup over state-of-the-art libraries such as LIBXSMM and LibShalom. For irregular matrices (i.e. tall skinny and long rectangles), autoGEMM is 1.3-2.0x faster than widely-used libraries such as OpenBLAS and Eigen. autoGEMM is publicly available at: https://github.com/wudu98/autoGEMM.

Wu, Du↗

QMCPACK v3.16.0

QMCPACK is an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support.

Kent, Paul R. C. [Oak Ridge National Laboratory] (↗

QMCPACK v3.17.0

QMCPACK is an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support.

Kent, Paul R. C. [Oak Ridge National Laboratory] (↗

QMCPACK v3.17.1

QMCPACK is an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support.

Kent, Paul R. C. [Oak Ridge National Laboratory] (↗

QMCPACK v3.14.0

QMCPACK is an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support.

Kent, Paul R. C.↗

Portable microwave test packages for beam-waveguide antenna performance evaluations

Portable microwave test packages used to evaluate a new 34-m-diameter beam-waveguide (BWG) antenna are described. The experimental methodology involved transporting test packages to different focal points of the BWG system and making noise temperature, antenna efficiency, and holography measurements. Comparisons of data measured at the different focal points enabled determinations of performance degradations caused by various mirrors in the BWG system. It is shown that, due to remarkable stabilities and accuracies of radiometric data obtained through the use of the microwave test packages, degradations caused by the BWG system were successfully determined.

Otoshi, Tom Y.↗

Neutral Buoyancy Portable Life Support System performance study

The Neutral Buoyancy Portable Life Support System (NBPSS) has been designed to support astronaut underwater training activities associated with EVA operations. The performance of competing NBPSS configurations has been analyzed on the basis of a modified 'Metabolic Man' program. NBPSS success is dependent on the development of novel cryogen supply tank and liquid-cooling garment vaporizer. Attention is given to mass and thermal balances and the evaluation results for the vent-loop ejector and heat-exchanger designs.

Chang, Chi-Min↗

Portability and Cross-Platform Performance of an MPI-Based Parallel Polygon Renderer

Visualizing the results of computations performed on large-scale parallel computers is a challenging problem, due to the size of the datasets involved. One approach is to perform the visualization and graphics operations in place, exploiting the available parallelism to obtain the necessary rendering performance. Over the past several years, we have been developing algorithms and software to support visualization applications on NASA's parallel supercomputers. Our results have been incorporated into a parallel polygon rendering system called PGL. PGL was initially developed on tightly-coupled distributed-memory message-passing systems, including Intel's iPSC/860 and Paragon, and IBM's SP2. Over the past year, we have ported it to a variety of additional platforms, including the HP Exemplar, SGI Origin2OOO, Cray T3E, and clusters of Sun workstations. In implementing PGL, we have had two primary goals: cross-platform portability and high performance. Portability is important because (1) our manpower resources are limited, making it difficult to develop and maintain multiple versions of the code, and (2) NASA's complement of parallel computing platforms is diverse and subject to frequent change. Performance is important in delivering adequate rendering rates for complex scenes and ensuring that parallel computing resources are used effectively. Unfortunately, these two goals are often at odds. In this paper we report on our experiences with portability and performance of the PGL polygon renderer across a range of parallel computing platforms.

Crockett, Thomas W.↗