Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MPI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

PPO And Friends

PPO and Friends (PPOAF) is a pytorch implementation of proximal policy optimization for single- and multi-agent reinforcement learning (the PPO), along with several optimizations and add-ons (the Friends) to enable efficient MPI-parallelized model training on HPC clusters.

Maguire, AlisterO↗

VerifyIO: Verifying Adherence to Parallel I/O Consistency Semantics

VerifyIO is a tool designed for verifying I/O consistency semantics in High-Performance Computing (HPC) applications. It addresses the challenges of ensuring correctness and portability across different I/O consistency models, such as POSIX, Commit, Session, and MPI-IO. By analyzing execution traces, detecting conflicts, and verifying synchronization adherence, VerifyIO provides actionable insights for both application developers and I/O library designers.

Wang, Chen [Lawrence Livermore National Laboratory↗

BULKI-Store v0.3.2

BULKI-Store is a distributed object storage system optimized for high-performance computing environments. Built with a Rust core and Python bindings, it efficiently manages scientific and machine learning datasets across HPC clusters. The system employs a client-server architecture with MPI integration, enabling seamless scaling on supercomputers like Perlmutter. BULKI-Store's object-oriented approach provides intuitive data organization with rich metadata support, contrasting with traditional file-based solutions. Key optimizations include selective checkpoint loading, unified checkpoint files, and object chunking for large data transfers. For machine learning workloads, BULKI-Store offers advantages through fine-grained access patterns, dynamic data sharing between training instances, and reduced memory pressure. Memory management features include strategic Python GC calls, minimized data copies, and batch processing capabilities. The system leverages Rayon's thread pool for asynchronous data prefetching and supports multiple CPU architectures (ARM64, x86, AMD, RISC-V). By combining performance optimizations with developer-friendly APIs, BULKI-Store addresses the complex data management challenges of modern HPC applications while maintaining compatibility across heterogeneous computing environments.

Zhang, Wei [Lawrence Berkeley National Laboratory ↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

MultifidelityOpt- bohydra

Multifidelity Bayesian optimization with serial and MPI-enabled (parallel, asynchronous) workflows.

Grosskopf, Mike [Los Alamos National Laboratory]↗

Photochemically Induced Acousto-optics Fluid Simulations

PIAFS is a finite-difference code to solve the compressible Navier-Stokes equations with chemical heating on Cartesian grids. It models chemical reactions of air (oxygen and carbon dioxide) with ozone subject to radiation. It uses a high-order WENO spatial discretization and explicit Runge-Kutta time integration. It is capable of parallel simulations using MPI. The code is written in C/C++.

Oudin, AlbertineN [Lawrence Livermore National Lab↗

PRISMA: PARALLEL REFINEMENT AND INTEGRATION SYSTEM FOR MULTI-AZIMUTHAL ANALYSIS

The Parallel Refinement and Integration System for Multi-azimuthal Analysis (PRISMA, version 1.1.0) is a Python application for processing X-ray diffraction (XRD) image data. PRISMA wraps GSAS-II to perform azimuthally-binned peak refinement, computes per-frame strain and d-spacing from those fits, and provides three PyQt5 graphical interfaces: (1) a Recipe Builder for selecting GSAS-II control (.imctrl) files, optional mask (.immask) files or threshold-ased masking, reference and experiment image sets, peaks, zimuthal range and bin size, and an optional ceria-based auto-calibration; (2) a Batch Processor that uses Dask on local workstations and pure MPI (mpi4py.futures.MPICommExecutor) on HPC to distribute GSAS-II refinement across cores or compute nodes and write results to a 4-dimensional (peaks x frames x azimuths x measurements) Zarr dataset; and (3) a Data Analyzer that renders heatmaps of fit parameters, strain, frame-to-frame deltas, and percent-change-vs-reference, and exports user-defined subsections to CSV or Excel. The peak-refinement algorithm is deterministic. Benchmark on ALCF Crux: a 20,000-image set, single-peak fit in frame mode with 44 azimuthal bins on 128 nodes x 128 workers, 48 seconds total wall time.

Lorenzo Martin, Maria De La Cinta [Argonne Nationa↗

Scientific Core Library Stack (SCLS) v2026

SCLS (Scientific Core Library Stack) is an opinionated build and packaging system for scientific computing libraries developed at Lawrence Berkeley National Laboratory. It produces a coherent, reproducible stack of numerical libraries — including BLAS/LAPACK, MPI, sparse direct and iterative solvers, graph partitioners, and parallel I/O libraries (e.g., PETSc, SLEPc, HDF5, NetCDF, MUMPS, OpenBLAS) — that work together without manual repair by downstream scientific software. From a single recipe-and-flavor model, SCLS produces native RPM packages for RHEL-family Linux, DEB packages for Debian/Ubuntu, direct Unix-style prefix installs for HPC and locked-down environments, and native macOS builds. Multiple build "flavors" (e.g., GCC+OpenBLAS, GCC+MKL, Intel+MKL, debug) coexist in distinct prefixes on the same host. Compared to general-purpose meta-build frameworks, SCLS is deliberately curated rather than infinitely configurable. It enforces deterministic, audit-friendly behavior: explicit build dependencies, no silent feature autodetection, a clear open-source license policy, and rpath-based runtime linkage so installs integrate cleanly with standard package-manager workflows.

Messe, Christian [Lawrence Berkeley National Labor↗

West African Monsoon System’s Responses to Global Ocean–Regional Atmosphere Coupling

This study explores the added value (AV) of a regional Earth system model (ESM) compared to an atmosphere-only regional climate model (RCM) in simulating West African monsoon (WAM) rainfall. The primary goals are to foster discussions on the suitability of coupled RCMs for WAM projections and deepen our understanding of ocean–atmosphere coupling’s influence on the WAM system. The study employs results from dynamical downscaling of the ERA-Interim reanalysis and Max Plank Institute ESM, low resolution (MPI-ESM-LR), by two RCMs, atmosphere only (REMO) and REMO coupled with Max Planck Institute Ocean Model (MPIOM) (ROM), at ~25-km horizontal resolution. Results show that in regions distant from coupling domain boundaries such as West Africa (WA), constraint conditions from ERA-Interim are more beneficial than coupling effects. REMO, reliant on oceanic sea surface temperatures (SSTs) from observations and influenced by ERA-Interim, is biased under coupling conditions, although coupling offers potential advantages in representing heat and mass fluxes. Contrastingly, as intended, coupling improves SSTs and monsoon fluxes’ relationships under ESM-forced conditions. In this latter case, the coupling features a dipole-like spatial structure of AV, improving precipitation over the Guinea Coast but degrading precipitation over half of the Sahel. Our extensive examination of physical processes and mechanisms underpinning the WAM system supports the plausibility of AV. Additionally, we found that the monsoonal dynamics over the ocean respond to convective activity, with the Sahara–Sahel surface temperature gradient serving as the maintenance mechanism. While further efforts are needed to enhance the coupled RCM, we advocate for its use in the context of WAM rainfall forecasts and projections.

54 ENVIRONMENTAL SCIENCES↗

How Well Can CMIP6 Models Represent the Observed Influence of the Pacific and Indian Oceans on the Indian Summer Monsoon Rainfall?

This study evaluates the ability of CMIP6 climate models to simulate the observed effects of tropical Pacific and Indian Ocean sea surface temperature anomalies (SSTAs) on Indian summer monsoon rainfall (ISMR) variability. Using observational data and the large ensemble historical simulations of seven CMIP6 models from 1950 to 2014, we applied a cyclostationary linear inverse model (CS-LIM) to isolate the impacts of tropical Pacific SSTAs, Indian Ocean SSTAs and their interaction on the interannual variability of ISMR. Overall, CMIP6 models well reproduced the observed enhanced (reduced) ISMR variability from Pacific SSTAs (Indian Ocean SSTAs and the Indo-Pacific interaction), but with varying spatial patterns and magnitudes. While CESM2 and E3SM-2-0 showed the best agreement with observations for the effects of Pacific SSTAs and the Indo-Pacific interaction, respectively, CMIP6 models showed mixed results for the impacts from Indian Ocean SSTAs. Composite analysis of ISMR anomalies during the developing phases of pure and co-occurring El Niño-Southern Oscillation (ENSO) and Indian Ocean dipole (IOD) events revealed that the impacts from Pacific SSTAs were captured reasonably well by E3SM-2-0, CESM2, MIROC6, and MPI-ESM1-2-LR, while E3SM-2-0 also showed the best agreement with observations for the effects from the Indo-Pacific interaction. However, all models showed substantial biases in simulating the Indian Ocean SSTA impacts on ISMR, especially for pure El Niño events. Overall, this study provides new insights into how individual CMIP6 models simulate the isolated impacts from the tropical Pacific and Indian Oceans, which has important applications for improving ISMR predictions and interpreting ISMR future projections.

monsoon↗

Refining HPCToolkit for application performance analysis at exascale

As part of the US Department of Energy’s Exascale Computing Project (ECP), Rice University has been refining its HPCToolkit performance tools to better support measurement and analysis of applications executing on exascale supercomputers. To efficiently collect performance measurements of GPU-accelerated applications, HPCToolkit employs novel non-blocking data structures to communicate performance measurements between tool threads and application threads. To attribute performance information in detail to source lines, loop nests, and inlined call chains, HPCToolkit performs parallel analysis of large CPU and GPU binaries involved in the execution of an exascale application to rapidly recover mappings between machine instructions and source code. To analyze terabytes of performance measurements gathered during executions at exascale, HPCToolkit employs distributed-memory parallelism, multithreading, sparse data structures, and out-of-core streaming analysis algorithms. To support interactive exploration of profiles up to terabytes in size, HPCToolkit’s hpcviewer graphical user interface uses out-of-core methods to visualize performance data. The result of these efforts is that HPCToolkit now supports collection, analysis, and presentation of profiles and traces of GPU-accelerated applications at exascale. These improvements have enabled HPCToolkit to efficiently measure, analyze and explore terabytes of performance data for executions using as many as 64K MPI ranks and 64K GPU tiles on ORNL’s Frontier supercomputer. HPCToolkit’s support for measurement and analysis of GPU-accelerated applications has been employed to study a collection of open-science applications developed as part of ECP. This paper reports on these experiences, which provided insight into opportunities for tuning applications, strengths and weaknesses of HPCToolkit itself, as well as unexpected behaviors in executions at exascale.

Adhianto, Laksono↗

Advances in ArborX to support exascale applications

ArborX is a performance portable geometric search library developed as part of the Exascale Computing Project (ECP). In this paper, we explore a collaboration between ArborX and a cosmological simulation code HACC. Large cosmological simulations on exascale platforms encounter a bottleneck due to the in-situ analysis requirements of halo finding, a problem of identifying dense clusters of dark matter (halos). This problem is solved by using a density-based DBSCAN clustering algorithm. With each MPI rank handling hundreds of millions of particles, it is imperative for the DBSCAN implementation to be efficient. In addition, the requirement to support exascale supercomputers from different vendors necessitates performance portability of the algorithm. We describe how this challenge problem guided ArborX development, and enhanced the performance and the scope of the library. We explore the improvements in the basic algorithms for the underlying search index to improve the performance, and describe several implementations of DBSCAN in ArborX. Further, we report the history of the changes in ArborX and their effect on the time to solve a representative benchmark problem, as well as demonstrate the real world impact on production end-to-end cosmology simulations.

97 MATHEMATICS AND COMPUTING↗

Observed and Imputed Volumetric Soil Water Content Timeseries for the New Mexico Elevation Gradient

Reliable soil water content (SWC) data are essential for understanding dryland ecosystem dynamics, but high-frequency SWC sensors often fail, creating gaps in critical datasets. To address this, we developed a Bayesian mixture model that imputes missing SWC using both linear interpolation and an ecosystem water balance model (SOILWAT2), tested across six AmeriFlux eddy covariance tower sites in the New Mexico Elevation Gradient, demonstrating its effectiveness in reconstructing SWC patterns while providing insights into the factors driving SWC variability. Daily volumetric soil water content (SWC) data are provided as csv-formatted spreadsheets for the six AmeriFlux sites (US-Seg, US-Ses, US-Wjs, US-Mpi, US-Vcp, and US-Vcs). For each site there is an observed SWC file (site_SWC_gapfill.csv) and a file that contains imputed SWC (imputed_SWC_site.csv). The observed SWC files contain temperature corrected sensor values, tower precipitation data, as well as outputs from SOILWAT2 simulations that were used to impute SWC. The imputed files contain the original observed SWC values and the imputed missing SWC values. When SWC was missing from the original data, the missing value was imputed based on the Bayesian imputation mixture model. The posterior mean of all imputed values is reported as "mean_X". When the observed SWC was NOT missing, mean_X = observed SWC value (original data). The standard deviation, 2.5th percentile and the 97.5th percentile for the imputed values are also reported in the imputed files. There are readme text files for each file type explaining the contents of each column.

54 ENVIRONMENTAL SCIENCES↗

Progressive Hedging Decomposition for Solutions of Large-Scale Process Family Design Problems

In previous work, we have introduced a mathematical model for solving a discretized version of the process family design problem. This involves two sets of decision variables. One set selects which unit module designs are included in the process platform out of a candidate set of options; the other set determines which of these unit module designs are assigned to each variant. In this work, we exploit a parallelized Progressive Hedging (PH) algorithm to solve even larger scale design problems. PH is a well-known algorithm traditionally used to solve stochastic programming problems. While our problem is not a two-stage stochastic programming problem, the structure is similar, and it can be directly mapped to the PH approach, which we employ here to solve this deterministic optimization problem. We decompose our problem by process variant. We treat the platform unit module design variables as first-stage and the assignment of unit module designs to variants as second-stage, solving the problem using mpi-sppy. We demonstrate this approach on case studies of CC, water desalination, and refrigeration.

Stinchfield, Georgia↗

PETSc/TAO Users Manual Revision 3.22

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗

Decomposition and Algorithmic Approaches for Solving Large-Scale Process Family Design Problems

Our most recent work expands the water desalination case study from 76 variants to 10,897 variants using the equation-oriented model built in Pyomo as part of the PARETO project. Using the discretization formulation presented in Stinchfield (2024a), rather than solving for all 10,897 variants simultaneously, we decompose the formulation into subproblems containing subsets of variants from the process family. We solve the overall problem with Progressive Hedging (PH) deployed in parallel on a distributed HPC cluster using the open-source Python package mpi-sppy (Knueven et al., 2023). This approach allowed us to solve this process family design problem to ~1.5% relative optimality gap in about 5 hours; in comparison, Gurobi reached ~50% relative optimality gap in about 6 hours (Stinchfield et al., 2024b). However, this approach still requires discretization of the common unit module design ranges; additionally, PH acts as a heuristic for MILP’s with gap-closing capabilities. Ideally, we would not have to use ML surrogates or discretization to solve this problem, instead solving the process family design problem with the equation-oriented model directly to achieve the most accurate results. However, recall that we did not consider solving the MINLP directly due to complexity and size. In this work, we aim to decompose and solve this large-scale MINLP using a Structured Nonlinear Global Optimization algorithm presented by Cao and Zavala (2019).

Stinchfield, Georgia↗

Selection of Global Climate Model Data for Downscaling With Generative Machine Learning and Use in the Power Planning for Alignment of Climate and Energy Systems Project

The range of results from climate models and scenarios is important to the understanding of uncertainty in power planning analysis. A U.S. Department of Energy-funded analytic project called Power Planning for Alignment of Climate and Energy Systems is developing data and analytic methods to reflect the effects of climate change on key variables for power system planning, as part of the Grid Modernization Lab Consortium. This project will select and prepare global climate model results for use in power system planning models. A related report (Evaluation of Global Climate Models for Use in Energy Analysis) assesses the performance of various global climate models from the Coupled Model Intercomparison Project Phase 6 data archive for their historical skill with respect to energy system performance and for their future projections under multiple climate change scenarios. Building from that report, we describe the selection of a climate scenario (Shared Socioeconomic Pathway [SSP] 2-4.5) and five climate models: TaiESM1, EC-Earth3-CC, GFDL-CM4, EC-Earth3-Veg, and MPI-ESM1-2-HR. We describe the model selection criteria, which were based on the quality of the match between model results under historical conditions and on the representation of the range of future values for several variables. These results will be downscaled via an open-source generative machine learning method called Super-Resolution for Renewable Energy Resource Data with Climate Change Impacts.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

PETSc/TAO Users Manual Revision 3.23

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations (PDEs) and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for implementing large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication.

97 MATHEMATICS AND COMPUTING↗