Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel systems”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Parallel-in-Time Solution of Hyperbolic PDE Systems via Characteristic-Variable Block Preconditioning

We consider the parallel-in-time solution of both linear and nonlinear hyperbolic partial differential equation (PDE) systems in one spatial dimension. In the nonlinear setting, the discretized equations are solved with a preconditioned residual iteration based on a global linearization. The linear(ized) equation systems are approximately solved parallel-in-time using a block preconditioner applied in the characteristic variables of the underlying linear(ized) hyperbolic PDE. This change of variables is motivated by the observation that intervariable coupling between characteristic variables is weak, at least locally where spatio-temporal variations in the eigenvectors of the associated flux Jacobian are sufficiently small, while that between the original variables is not. For an ℓ-dimensional system of PDEs, applying the preconditioner consists of solving a sequence of ℓ scalar linear(ized)-advection-like problems, each associated with a different characteristic wave-speed in the underlying linear(ized) PDE. Furthermore, we approximately solve these linear advection problems using multigrid reduction-in-time (MGRIT); however, any other suitable parallel-in-time method could be used. Numerical examples are shown for the (linear) acoustics equations in heterogeneous media and for the (nonlinear) shallow water equations and Euler equations of gas dynamics with shocks and rarefactions. For many test problems, the solver converges in just a handful of iterations and with mesh-independent convergence rates.

97 MATHEMATICS AND COMPUTING↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Status report on HFIR irradiation of optimized alumina forming alloys

Properties of FeCrAl alloys under neutron irradiation are of interest because of these materials’ potential application as accident-tolerant fuel cladding in nuclear systems. In parallel, alumina-forming austenitic (AFA) alloys are of interest for use as structural materials in advanced nuclear systems for their potential higher resistance to embrittlement and high-temperature steam oxidation resistance. An irradiation campaign for fiscal year 2024 has been developed under the Advanced Fuels Campaign to perform irradiation testing of various FeCrAl and AFA alloys in Oak Ridge National Laboratory’s High Flux Isotope Reactor (HFIR). The goals of this irradiation campaign are to (1) study the impact of minor alloying elements on the neutron-irradiated mechanical properties of FeCrAl alloys and (2) collect neutron-irradiated mechanical properties on AFA alloys for comparison with those of FeCrAl alloys. This campaign will include both tensile and fracture toughness specimens tested following HFIR irradiation at temperatures representative of normal operating conditions in light-water reactors. The pre-irradiation characterization to date, the irradiation plan for the FeCrAl and AFA specimens, and the subsequent post-irradiation experimental test plan are presented in this report, along with the status of HFIR builds and scheduled insertion dates.

36 MATERIALS SCIENCE↗

Exploiting user activeness for data retention in HPC systems

HPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology.

Zhang, Wei↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

TuckerMPI: A Parallel C++/MPI Software Package for Large-scale Data Compression via the Tucker Tensor Decomposition

With this study, our goal is compression of massive-scale grid-structured data, such as the multi-terabyte output of a high-fidelity computational simulation. For such data sets, we have developed a new software package called TuckerMPI, a parallel C++/MPI software package for compressing distributed data. The approach is based on treating the data as a tensor, i.e., a multidimensional array, and computing its truncated Tucker decomposition, a higher-order analogue to the truncated singular value decomposition of a matrix. The result is a low-rank approximation of the original tensor-structured data. Compression efficiency is achieved by detecting latent global structure within the data, which we contrast to most compression methods that are focused on local structure. In this work, we describe TuckerMPI, our implementation of the truncated Tucker decomposition, including details of the data distribution and in-memory layouts, the parallel and serial implementations of the key kernels, and analysis of the storage, communication, and computational costs. We test the software on 4.5 and 6.7 terabyte data sets distributed across 100 s of nodes (1,000 s of MPI processes), achieving compression ratios between 100 and 200,000×, which equates to 99--99.999% compression (depending on the desired accuracy) in substantially less time than it would take to even read the same dataset from a parallel file system. Moreover, we show that our method also allows for reconstruction of partial or down-sampled data on a single node, without a parallel computer so long as the reconstructed portion is small enough to fit on a single machine, e.g., in the instance of reconstructing/visualizing a single down-sampled time step or computing summary statistics. The code is available at https://gitlab.com/tensors/TuckerMPI.

97 MATHEMATICS AND COMPUTING↗

Case Study of Using Kokkos and SYCLs Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.

Dufek, Amanda S↗

Point-by-point inscribed sapphire parallel fiber Bragg gratings in a fully multimode system for multiplexed high-temperature sensing

In this work, we study the point-by-point inscription of sapphire parallel fiber Bragg gratings (sapphire pFBGs) in a fully multimode system. A parallel FBG is shown to be critical in enabling detectable and reliable high-order grating signals. The impacts of modal volume, spatial coherence, and grating location on reflectivity are examined. Three cascaded seventh-order pFBGs are fabricated in one sapphire fiber for wavelength multiplexed temperature sensing. Using a low-cost, fully multimode 850-nm interrogator, reliable measurement up to 1500°C is demonstrated.

42 ENGINEERING↗

Designing a parallel Feel-the-Way clustering algorithm on HPC systems

This paper introduces a new parallel clustering algorithm, named Feel-the-Way clustering algorithm, that provides better or equivalent convergence rate than the traditional clustering methods by optimizing the synchronization and communication costs. Our algorithm design centers on how to optimize three factors simultaneously: reduced synchronizations, improved convergence rate, and retained same or comparable optimization cost. To compare the optimization cost, we use the Sum of Square Error (SSE) cost as the metric, which is the sum of the square distance between each data point and its assigned clusters. Compared with the traditional MPI k-means algorithm, the new Feel-the-Way algorithm requires less communications among participating processes. As for the convergence rate, the new algorithm requires fewer number of iterations to converge. As for the optimization cost, it obtains the SSE costs that are close to the k-means algorithm. In the paper, we first design the full-step Feel-the-Way k-means clustering algorithm that can significantly reduce the number of iterations that are required by the original k-means clustering method. Next, we improve the performance of the full-step algorithm by adopting an optimized sampling-based approach, named reassignment-history-aware sampling. Our experimental results show that the optimized sampling-based Feel-the-Way method is significantly faster than the widely used k-means clustering method, and can provide comparable optimization costs. More extensive experiments with several synthetic datasets and real-world datasets (e.g., MNIST, CIFAR-10, ENRON, and PLACES-2) show that the new parallel algorithm can outperform the open source MPI k-means library by up to 110% on a high-performance computing system using 4,096 CPU cores. In addition, the new algorithm can take up to 51% fewer iterations to converge than the k-means clustering algorithm.

97 MATHEMATICS AND COMPUTING↗

DC Microgrid Reliability Enhancement with Adaptive Converter Thermal Management

Due to the different device selections, aging levels, and thermal dissipation performance, some converters may take additional thermal stress on switching devices than others in paralleled converter systems, which will reduce system reliability. To address this problem, this paper proposes a power-sharing strategy with adaptive thermal management. First, the temperature-based power loss model and electrical-thermal model are established. Based on that, a high-accuracy IGBT junction temperature estimate considering the power loss-temperature coupling can be achieved. Further, the thermal-sharing for all the switching devices in paralleled converters can be achieved with the proposed adaptive thermal management strategy. The proposed strategy can change the power-sharing ratio adaptively according to the system operation conditions, which will contribute to the system reliability enhancement. The effectiveness of the proposed strategy is verified through PLECS thermal simulation and joint real-time simulation with Dspace and RT-box.

DC microgrid↗

Random fields from quenched disorder in an archetype for correlated electrons: The parallel spin stripe phase of La 1.6 – x Nd 0.4 Sr x CuO 4 at the 1/8 anomaly

The parallel stripe phase is remarkable both in its own right, and in relation to the other phases with which it coexists. Its inhomogeneous nature makes such states susceptible to random fields from quenched magnetic vacancies. Here we argue this is the case by introducing low concentrations of nonmagnetic Zn impurities (0%–10%) into La 1.6–x ⁢Nd 0.4⁢ Sr x ⁢CuO 4 (Nd-LSCO) with x=0.125 in single-crystal form, well below the percolation threshold of ~41% for a two-dimensional square lattice. Elastic neutron scattering measurements on these crystals show clear magnetic quasi-Bragg peaks at all Zn dopings. While all the Zn-doped crystals display order parameters that merge into each other and the background at ~68 K, the temperature dependence of the order parameter as a function of Zn concentration is drastically different. This result is consistent with meandering charge stripes within the parallel stripe phase, which are pinned in the presence of quenched magnetic vacancies. In turn it implies vacancies that preferentially occupy sites within the charge stripes, and hence that can be very effective at disrupting superconductivity in Nd-LSCO (x=0.125), and, by extension, in all systems exhibiting parallel stripes.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Level-2 Milestone 9009: Flux and Rabbit Capabilities on El Capitan

This document is the milestone delivery report for the ASC 2025 L2 milestone (See Table 1) for advanced I/O capabilities for El Capitan via Flux Workload Manager support and the new I/O hardware designed for El Capitan, the Rabbit Storage System. In this document we describe the design of the Rabbit Storage System and how it is managed by Flux. We evaluate the performance and usability of Rabbit using ARES, IOR, and an AI inference workload. Overall, we find that Rabbit shows good scalability, especially in node-local storage configurations, and is more scalable than the global Lustre parallel file system.

97 MATHEMATICS AND COMPUTING↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Accelerating shared file checkpoint with local burst buffers

A data management system and method for accelerating shared file checkpointing. Written application data is aggregated in an application data file created in a local burst buffer memory at a compute node, and an associated data mapping built index to maintain information related to the offsets into a shared file at which segments of the application data is to be stored in a parallel file system, and where in the buffer those segments are located. The node asynchronously transfers a data file containing the application data and the associated data mapping index to a file server for shared file storage. The data management system and method further accelerates shared file checkpointing in which a shared file, together with a map file that specifies how the shared file is to be distributed, is asynchronously transferred to local burst buffer memories at the nodes to accelerate reading of the shared file.

Gooding, Thomas↗

Accelerating Application Bulk Synchronous Writes in HPC Environments

High-bandwidth storage tiers are becoming more common for their capability to absorb high-rate, bursty I/Os. Notably, the designs of these fast storage tiers differ from system to system. The variation of these layers and non-uniform methods of access can pose chal- lenges for applications seeking to run at multiple HPC facilities. Therefore, in this work, we present Spectral, a rapid-output ab- straction library to accelerate application, bulk-synchronous writes on HPC systems. We design Spectral to enable applications to use high-bandwidth storage, such as node-local storage and dis- tributed, write-caches (e.g., burst buffers) transparently without requiring modifications to the application or file system source code. The key idea is to allow applications to spend most of the time performing productive work and to not require any source code changes for maximum portability on different HPC archi- tectures. Spectral internally re-routes write-only files through available, high-performance I/O resources before ultimately mi- grating them to the shared global parallel file system. For instance, on Summit, Spectral transparently places application outputs on node-local storage and then utilizes asynchronous migration to the center-wide GPFS file system. We evaluate Spectral on the Summit HPC system (1024 nodes) using the IOR benchmark and real scientific applications. Spectral shows linear performance scaling, improving application write performance by over an order of magnitude when compared to GPFS.

Khan, Awais↗

Frontier (HPE Cray EX) Exascale Supercomputer at the Oak Ridge Leadership Computing Facility

Frontier is the HPE Cray EX exascale supercomputer deployed and operated by the Oak Ridge Leadership Computing Facility (OLCF) at Oak Ridge National Laboratory (ORNL). Frontier is designed for large-scale modeling, simulation, and AI workloads and is built from HPE Cray EX system architecture with AMD CPUs and AMD Instinct GPU accelerators connected by the HPE Slingshot interconnect. System composition (representative production configuration): Frontier is composed of approximately 74 cabinets with 128 compute nodes per cabinet (~9,400 compute nodes total). Each compute node contains one 64-core AMD EPYC CPU and four AMD Instinct MI250X GPUs. Nodes are connected using HPE Slingshot (Slingshot-200 class) networking with multiple NIC ports per node providing high injection bandwidth. Frontier is connected to the Orion parallel file system (multi-tier Lustre) providing a large, center-wide high-performance storage namespace. Operational context: Frontier entered public prominence as the first system to reach No. 1 on the TOP500 list in May 2022 (HPL benchmark), establishing the first widely recognized exascale-era performance milestone. The system supports DOE Office of Science mission workloads and enables leadership-class computational science and AI for open science users.

AMD EPYC↗

PRO-X Fuel Cycle Transportation and Crosscutting Progress Report

The PRO-X program is actively supporting the design of nuclear systems by developing a framework to both optimize the fuel cycle infrastructure for advanced reactors (ARs) and minimize the potential for production of weapons-usable nuclear material. Three study topics are currently being investigated by Sandia National Laboratories (SNL) with support from Argonne National Laboratories (ANL). This multi-lab collaboration is focused on three study topics which may offer proliferation resistance opportunities or advantages in the nuclear fuel cycle. These topics are: 1) Transportation Global Landscape, 2) Transportation Avoidability, and 3) Parallel Modular Systems vs Single Large System (Crosscutting Activity).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

kiloAMPS (Final Report)

This project aims to reduce the testing time of various electrochemical techniques for neural interface electrodes by developing an automated PCB design which would be capable of testing multiple interface electrodes in parallel. The system should be able to perform the full battery of electrochemical tests, though electrochemical impedance spectroscopy (EIS) will be of primary interest. Previous versions of this project have already improved upon traditional interface testing by automating many parts of the process, reducing the test time to around a day and requiring the technician to check up on the process once every six hours. A key bottleneck is that the electrode testing is still done sequentially. Testing time could be heavily reduced if electrodes could be tested in parallel.

42 ENGINEERING↗