Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance Portability”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Tacho

SAND2022-7470 O Tacho is a performance portable sparse direct solver for use with various sparse matrix problems which typically arise from finite element methods. The code relies on the Kokkos parallel programming model porting to both CPUs and GPUs. It is also part of the Trilinos library. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kim, Kyungjoo↗

QMCPACK v3.15.0

QMCPACK is an open-source production level many-body ab initio Quantum Monte Carlo code for computing the electronic structure of atoms, molecules, and solids with full performance portable GPU support.

Kent, Paul R. C. [Oak Ridge National Laboratory] (↗

MAM4xx

SAND2023-05377O The Modal Aerosol Model with 4 fixed modes (MAM4xx) is a performance-portable, C++ implementation for modeling the chemical reactions of aerosols in the atmosphere. MAM4xx is part of the Energy Exascale Earth System Model (E3SM). The E3SM project is an ongoing Earth system modeling, simulation, and prediction project that optimizes the use of DOE laboratory resources to meet the science needs of the nation and the mission needs of DOE. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Overfelt, James↗

Flash-X

Flash-X is a highly composable multiphysics software system that can be used to simulate physical phenomena in several scientific domains. It is derived from FLASH, which has been a community code for several communities over the last 20 years. The Flash-X architecture has been redesigned to be compatible with increasingly heterogeneous hardware platforms. A part of the redesign is a utilizing a newly designed performance portability layer that is language agnostic.

Dubey, Anshu [Argonne National Laboratory (ANL), A↗

eagles-project/haero

A toolbox for constructing performance portable aerosol packages

Johnson, Jeffrey N. [Cohere Consulting LLC]↗

LAPIS: Linear Algebra Performance for Intermediate Subprograms

SAND2025-11594O LAPIS (Linear Algebra Performance for Intermediate Subprograms) is a compiler infrastructure for linear algebra that targets both high productivity and performance portability. It is based on the open-source MLIR package from the LLVM project. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kelley, Brian↗

Flash-X Recipe Tools

SF-24-102 A code generation tool for Flash-X to support their performance portability using a domain-specific runtime library.

Lee, Youngjun↗

hymera

Hymera is a performance portable multiscale simulation framework based on parthenon and Kokkos, that designed to push relativistic particles, governed by slowly varying, quasi static background fields

Edelmann, Philipp [@LANL]↗

High-performance finite elements with MFEM

The MFEM (Modular Finite Element Methods) library is a high-performance C++ library for finite element discretizations. MFEM supports numerous types of finite element methods and is the discretization engine powering many computational physics and engineering applications across a number of domains. Furthermore, this paper describes some of the recent research and development in MFEM, focusing on performance portability across leadership-class supercomputing facilities, including exascale supercomputers, as well as new capabilities and functionality, enabling a wider range of applications. Much of this work was undertaken as part of the Department of Energy’s Exascale Computing Project (ECP) in collaboration with the Center for Efficient Exascale Discretizations (CEED).

97 MATHEMATICS AND COMPUTING↗

Time-temperature history and input files for ExaCA v2.0 scaling, performance, and demonstration simulations

The files in this data repository are used in various sections of the manuscript "ExaCA v2.0: A versatile, scalable, and performance portable cellular automata application for additive manufacturing solidification" by Rolchigo et al. (DOI: 10.1016/j.commatsci.2025.113734). The README file references dataset numbers as given in the manuscript's Table 2, as well as the manuscript's relevant subsections.

36 MATERIALS SCIENCE↗

Twelve turbine wind farm simulation with AMR-Wind

This dataset contains simulation data of a twelve turbine wind farm simulation performed with AMR-Wind. The simulation is documented in Kuhn, M. B., Henry de Frahan, M. T., Mohan, P., Deskos, G., Churchfield, M., Cheung, L., ... & Sprague, M. (2025). AMR‐Wind: A Performance‐Portable, High‐Fidelity Flow Solver for Wind Farm Simulations. Wind Energy, 28(5), e70010 (https://doi.org/10.1002/we.70010).

17 WIND ENERGY↗

Nyx: A Massively Parallel AMR Code for Computational Cosmology

Nyx is a highly parallel, adaptive mesh, finite-volume N-body compressible hydrodynamics solver for cosmological simulations. It has been used to simulate different cosmological scenarios with a recent focus on the intergalactic medium and Lyman alpha forest. Together, Nyx, the compressible astrophysical simulation code, Castro, and the low Mach number code MAESTROeX, make up the AMReX-Astrophysics Suite of open-source, adaptive mesh, performance-portable astrophysical simulation codes. Other examples of cosmological simulation research codes include Enzo, Enzo-P/Cello, RAMSES, ART, FLASH, Cholla, as well as Gadget, Gasoline, Arepo, Gizmo, and SWIFT.

79 ASTRONOMY AND ASTROPHYSICS↗

ERF: Energy Research and Forecasting

The Energy Research and Forecasting (ERF) code is a new model that simulates the mesoscale and microscale dynamics of the atmosphere using the latest high-performance computing architectures. It employs hierarchical parallelism using an MPI+X model, where X may be OpenMP on multicore CPU-only systems, or CUDA, HIP, or SYCL on GPU-accelerated systems. ERF is built on AMReX (Zhang et al., 2019, 2021), a block-structured adaptive mesh refinement (AMR) software framework that provides the underlying performance-portable software infrastructure for block-structured mesh operations. The "energy" aspect of ERF indicates that the software has been developed with renewable energy applications in mind. In addition to being a numerical weather prediction model, ERF is designed to provide a flexible computational framework for the exploration and investigation of different physics parameterizations and numerical strategies, and to characterize the flow field that impacts the ability of wind turbines to extract wind energy. The ERF development is part of a broader effort led by the US Department of Energy's Wind Energy Technologies Office.

17 WIND ENERGY↗

Fiats: Functional inference and training for surrogates

Fiats provides a platform for research on the training and deployment of neural-network surrogate models for computational science. Fiats also supports exploring, advancing, and combining functional, object-oriented, and parallel programming patterns in Fortran 2023. As such, the Fiats name has dual expansions: “Functional Inference And Training for Surrogates” or “Fortran Inference And Training for Science.” Fiats inference and training procedures are pure and therefore satisfy a language constraint imposed on procedure invocations inside Fortran’s parallel loop construct: do concurrent. Furthermore, the Fiats training procedures are built around a do concurrent parallel reduction. Several compilers can automatically parallelize do concurrent on Central Processing Units (CPUs) or Graphics Processing Units (GPUs). Fiats thus aims to achieve performance portability through standard language mechanisms.

Rouson, Damian [Lawrence Berkeley National Laborat↗

Contra: A New Language for Task- and Data-Parallellism [Slides]

A new language is beneficial to take advantage of emerging architectures, and existing and new software technologies. Contra is a new language to provide performance portable code and consists of two innovations, which are described. A current status and illustrations of the code are shown.

97 MATHEMATICS AND COMPUTING↗

Integration of Kokkos into MonteRay [Slides]

Monte Carlo Neutron Transport codes are C++-based and simulate interaction of nuclear particles with materials. MonteRay is a library for accelerating Monte Carlo ray-casting tallies with GPUs. Kokkos is commonly used to write performance portable applications for HPC platforms. A goal is to implement Kokkos into the Expected Path Length file that uses Cuda as a backend. The methodology and results are summarized.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The Multiphysics on Advanced Platforms Project

In 2015, the Lawrence Livermore National Laboratory started development of next-generation multiphysics simulation capabilities for the National Nuclear Security Administration under the Advanced Technologies Development and Mitigation (ATDM) element of the Advanced Simulation and Computing program in collaboration with the Exascale Computing Project (ECP). A key driver for this effort across the NNSA tri-lab was the emergence of advanced high performance computing (HPC) architectures based on heterogeneous compute capabilities, including GPU based systems, as part of the national drive toward exascale computing platforms at multiple Department of Energy (DOE) facilities. Developing a multiphysics code capable of meeting the various simulation needs of the NNSA as defined by the current generation of integrated codes (or ICs), initially developed as part of the Accelerated Strategic Computing Initiative (ASCI) program beginning in 1996, and able to scale to the current 100 petaflop class pre-exascale systems, as well the forthcoming exaflop class computers, is a daunting challenge. To accomplish this ambitious goal, LLNL has embraced two key themes: use of high-order numerical methods and a modular approach to code development. The LLNL next generation effort is organized under the Multi-Physics on Advanced Platforms Project (MAPP). A foundational component of MAPP is the Axom computer science (CS) toolkit which provides infrastructure for the development of modular, performance portable, multi-physics application codes. MARBL is a next-generation application code built on the Axom base to address the modeling needs of the high energy density physics (HEDP) community for simulating high-explosive, magnetic or laser driven experiments such as inertial confinement fusion (ICF), pulsed-power magneto-hydrodynamics (MHD), equation of state (EOS) and material strength studies as part of the NNSA’s stockpile stewardship program (SSP).

97 MATHEMATICS AND COMPUTING↗

CodeFlow: A Code Generation System for Flash-X Orchestration Runtime

We propose the CodeFlow toolchain for Flash-X that realizes the “recipe-to-source” code transformation for Flash-X simulations and that is necessary to achive performance portability. We design a high-level language to express operations of simulations in so-called recipes, which are given as input to the toolchain. The tools of the CodeFlow pipeline include code transformation with tree-based source code representation techniques and code orchestration and generation based on control flow graphs. The generated source code utilizes a new runtime, developed for Flash-X, that orchestrates dynamic and asynchronous data movement and task execution. The functionality of CodeFlow is demonstrated using a hydrodynamic problem with a strong shock.

97 MATHEMATICS AND COMPUTING↗