Engineering PapersSearch

SEARCH · Engineering Papers

Results for “MPI”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Methodology and Application of HPC I/O Characterization with MPIProf and IOT

Combining the strengths of MPIProf and IOT, an efficient and systematic method is devised for I/O characterization at the per-job, per-rank, per-file and per-call levels of HPC programs running on the NASA Advanced Supercomputing Center. This method is applied to answer four I/O questions in this paper. A total of 13 MPI programs and 15 cases, ranging from 24 to 5968 ranks, are analyzed to establish the I/O landscape from answers to the four questions. Four of the 13 programs use MPI I/O and the behavior of their collective writes depends on the specific implementation of the MPI library used. The SGI MPT library, the prevailing MPI library for our systems, was found to gather small writes from a large number of ranks to perform larger writes by a small subset of collective buffering ranks. The number of collective buffering ranks invoked by MPT depends on the Lustre stripe count and the number of nodes used for the run. A demonstration of varying the stripe count to achieve double-digit speedup of one program's I/O was presented. Another program, which concurrently opens private files by all ranks and could potentially create a heavy load on the Lustre servers, was identified. The ability to systematically characterize I/O for a large number of programs running on a supercomputer, seek I/O optimization opportunity and identify programs that could cause a high load and instability on the filesystems is important for pursuing exascale in a real production environment.

Characterization

Multiplicity dependent 𝐽/𝜓 and 𝜓⁡(2⁢𝑆) production at forward and backward rapidity in 𝑝 + 𝑝 collisions at $\sqrt{𝑠}$ = 200 GeV

Recent measurements of 𝐽/𝜓 production as a function of event charged-particle multiplicity at the collision energies of both the Large Hadron Collider (LHC) and the Relativistic Heavy Ion Collider (RHIC) show enhanced 𝐽/𝜓 production yields with increasing multiplicity. One potential explanation for this type of dependence is multiparton interactions (MPI). We present the first study of potential autocorrelations at RHIC energies and forward and backward rapidity of self-normalized 𝐽/𝜓 yields and 𝜓⁡(2⁢𝑆) to 𝐽/𝜓 ratio, as a function of self-normalized multiplicity in 𝑝 + 𝑝 collisions. In addition, detailed pythia studies tuned to RHIC energies were performed to investigate the MPI impacts. We find that the PHENIX data at RHIC are consistent with recent LHC measurements and can only be described by pythia calculations that include MPI effects. The forward and backward 𝜓⁡(2⁢𝑆) to 𝐽/𝜓 ratio is found to be less dependent on the charged-particle multiplicity.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Phloem

Phloem is a Message Passing Interface (MPI) micro-benchmarking suite featuring sub-communicator collectives, methods for finding slow links on MPI interconnects, and point-to-point MPI benchmarks, including a messaging rate benchmark. All of the benchmarks except for ones related exclusively to finding slow links are GPU-aware via the Umpire resource management library.

Moody, AdamT [Lawrence Livermore National Laborato

Evolution of the SLATE linear algebra library

SLATE (Software for Linear Algebra Targeting Exascale) is a distributed, dense linear algebra library targeting both CPU-only and GPU-accelerated systems, developed over the course of the Exascale Computing Project (ECP). While it began with several documents setting out its initial design, significant design changes occurred throughout its development. In some cases, these were anticipated: an early version used a simple consistency flag that was later replaced with a full-featured consistency protocol. In other cases, performance limitations and software and hardware changes prompted a redesign. Sequential communication tasks were parallelized; host-to-host MPI calls were replaced with GPU device-to-device MPI calls; more advanced algorithms such as Communication Avoiding LU and the Random Butterfly Transform (RBT) were introduced. Early choices that turned out to be cumbersome, error prone, or inflexible have been replaced with simpler, more intuitive, or more flexible designs. Applications have been a driving force, prompting a lighter weight queue class, nonuniform tile sizes, and more flexible MPI process grids. Of paramount importance has been building a portable library that works across several different GPU architectures – AMD, Intel, and NVIDIA – while keeping a clean and maintainable codebase. Here we explore the evolving design choices and their effects, both in terms of performance and software sustainability.

Gates, Mark

Preparing MPICH for exascale

The advent of exascale supercomputers heralds a new era of scientific discovery, yet it introduces significant architectural challenges that must be overcome for MPI applications to fully exploit its potential. Among these challenges is the adoption of heterogeneous architectures, particularly the integration of GPUs to accelerate computation. Additionally, the complexity of multithreaded programming models has also become a critical factor in achieving performance at scale. The efficient utilization of hardware acceleration for communication, provided by modern NICs, is also essential for achieving low latency and high throughput communication in such complex systems. In response to these challenges, the MPICH library, a high-performance and widely used Message Passing Interface (MPI) implementation, has undergone significant enhancements. Here, this paper presents four major contributions that prepare MPICH for the exascale transition. First, we describe a lightweight communication stack that leverages the advanced features of modern NICs to maximize hardware acceleration. Second, our work showcases a highly scalable multithreaded communication model that addresses the complexities of concurrent environments. Third, we introduce GPU-aware communication capabilities that optimize data movement in GPU-integrated systems. Finally, we present a new datatype engine aimed at accelerating the use of MPI derived datatypes on GPUs. These improvements in the MPICH library not only address the immediate needs of exascale computing architectures but also set a foundation for exploiting future innovations in high-performance computing. By embracing these new designs and approaches, MPICH-derived libraries from HPE Cray and Intel were able to achieve real exascale performance on OLCF Frontier and ALCF Aurora respectively.

Guo, Yanfei [Argonne National Laboratory (ANL), Ar

Historical and Future Windstorms in the Northeastern United States

Large-scale windstorms represent an important atmospheric hazard in the Northeastern US (NE) and are associated with substantial socioeconomic losses. Regional simulations performed with the Weather Research and Forecasting (WRF) model using lateral boundary conditions from three Earth System Models (ESMs: Geophysical Fluid Dynamics Laboratory (GFDL), Hadley Centre Global Environment Model (HadGEM) and Max Planck Institute (MPI)) are used to quantify possible future changes in windstorm characteristics and/or changes in the parent cyclone types responsible for windstorms. WRF nested within MPI ESM best represents important aspects of historical windstorms and the cyclone types responsible for generating windstorms compared with a reference simulation performed with the ERA-Interim reanalysis for the historical climate. The spatial scale and frequency of the largest windstorms in each simulation defined using the greatest extent of exceedance of local 99.9th percentile wind speeds (U > U999) plus 50-year return period wind speeds (U50,RP) do not exhibit secular trends. Projections of extreme wind speeds and windstorm intensity/frequency/geolocation and dominant parent cyclone type associated with windstorms vary markedly across the simulations. Only the MPI nested simulations indicate statistically significant differences in windstorm spatial scale, frequency and intensity over the NE in the future and historical periods. This model chain, which also exhibits the highest fidelity in the historical climate, yields evidence of future increases in 99.9th percentile 10 m height wind speeds, the frequency of simultaneous U > U999 over a substantial fraction (5–25%) of the NE and the frequency of maximum wind speeds above 22.5 ms−1. These geophysical changes, coupled with a projected doubling of population, leads to a projected tripling of a socioeconomic loss index, and hence risk to human systems, from future windstorms.

Pryor, Sara C. (ORCID:0000000348473440)

RANS-MP: A Portable Parallel Navier-Stokes Solver

RANS-MP, a new implementation of a single-grid Navier-Stokes solver using the diagonalized Beam-Warming approximate-factorization scheme, is presented. This first release of the completely rewritten solver employs the following optimizations: (1) Bi-directional multi-partition method for the ADI solver part; this improves granularity and load balance; (2) Improved cache usage through elimination of non-unit-stride array access (possible in part due to multi-partitioning); (3) Preprocessing of communicating boundary conditions to streamline logic during time stepping; (4) Truly parallel, high-performance I/O using the newly-developed MPI-IO library; (5) Elimination of large amounts of redundant operations through efficient use of workspace. Results of some realistic wing computations on the IBM SP2 computer will be presented. We will demonstrate that excellent absolute performance and scalability are obtained with RANS-MP, even for relatively small grid sizes. Besides high performance, an outstanding feature of RANS-MP is its true portability, due to the use of the portable message passing and I/O libraries MPI and MPI-IO.

VanderWijngaart, Rob F.

Analysis of 2D Torus and Hub Topologies of 100Mb/s Ethernet for the Whitney Commodity Computing Testbed

A variety of different network technologies and topologies are currently being evaluated as part of the Whitney Project. This paper reports on the implementation and performance of a Fast Ethernet network configured in a 4x4 2D torus topology in a testbed cluster of 'commodity' Pentium Pro PCs. Several benchmarks were used for performance evaluation: an MPI point to point message passing benchmark, an MPI collective communication benchmark, and the NAS Parallel Benchmarks version 2.2 (NPB2). Our results show that for point to point communication on an unloaded network, the hub and 1 hop routes on the torus have about the same bandwidth and latency. However, the bandwidth decreases and the latency increases on the torus for each additional route hop. Collective communication benchmarks show that the torus provides roughly four times more aggregate bandwidth and eight times faster MPI barrier synchronizations than a hub based network for 16 processor systems. Finally, the SOAPBOX benchmarks, which simulate real-world CFD applications, generally demonstrated substantially better performance on the torus than on the hub. In the few cases the hub was faster, the difference was negligible. In total, our experimental results lead to the conclusion that for Fast Ethernet networks, the torus topology has better performance and scales better than a hub based network.

Pedretti, Kevin T.

Comparison of 250 MHz R10K Origin 2000 and 400 MHz Origin 2000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on Steger, a 250 MHz Origin 2000 system with R10K processors, currently installed at the NASA Ames National Advanced Supercomputing (NAS) facility. For comparison purposes, the tests were also run on Lomax, a 400 MHz Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to measure system performance. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versions used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.

Performance Evaluation of Remote Memory Access (RMA) Programming on Shared Memory Parallel Computers

The purpose of this study is to evaluate the feasibility of remote memory access (RMA) programming on shared memory parallel computers. We discuss different RMA based implementations of selected CFD application benchmark kernels and compare them to corresponding message passing based codes. For the message-passing implementation we use MPI point-to-point and global communication routines. For the RMA based approach we consider two different libraries supporting this programming model. One is a shared memory parallelization library (SMPlib) developed at NASA Ames, the other is the MPI-2 extensions to the MPI Standard. We give timing comparisons for the different implementation strategies and discuss the performance.

Jin, Hao-Qiang

Space-Spurred Metallized Materials

Spurred R&D toward improved vacuum metallizing techniques led to an extensive line of commercial products, from insulated outdoor garments to packaging for foods, from wall coverings to window shades, from life rafts to candy wrappings, reflective blankets to photographic reflectors. Metallized Products, Inc. (MPI) was one of the companies that worked with NASA in development of the original space materials. MPI markets its own metallized products and supplies materials to other manufacturers. One of the most widely used MPI products is TXG laminate. An example is a reflective kite, the S.O.S. Signal Kite that can be flown as high as 200 feet to enhance radar and visual detectability. It offers a boon to campers, hikers, mountain climbers and boaters. It is produced by Solar Reflections, Inc. The company also markets a solar reflective hat. Another example is by Pro-Tektion, Inc. to provide protection for expensive musical equipment that have sensitive electronic components subject to damage from the heat of stage lights, dust, or rain at outdoor concerts. MP supplied the material and acceptance of the covers by the sound industry has been excellent.

Source record

Comparison of Origin 2000 and Origin 3000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on the Origin 3000 system currently being installed at the NASA Ames National Advanced Supercomputing facility. This machine will ultimately contain 1024 R14K processors. The first part of the system, installed in November, 2000 and named mendel, is an Origin 3000 with 128 R12K processors. For comparison purposes, the tests were also run on lomax, an Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to determine system performance and measure the impact of changes on the machine as it evolves. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versiqns used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN 3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.

Performance Evaluation of Supercomputers using HPCC and IMB Benchmarks

The HPC Challenge (HPCC) benchmark suite and the Intel MPI Benchmark (IMB) are used to compare and evaluate the combined performance of processor, memory subsystem and interconnect fabric of five leading supercomputers - SGI Altix BX2, Cray XI, Cray Opteron Cluster, Dell Xeon cluster, and NEC SX-8. These five systems use five different networks (SGI NUMALINK4, Cray network, Myrinet, InfiniBand, and NEC IXS). The complete set of HPCC benchmarks are run on each of these systems. Additionally, we present Intel MPI Benchmarks (IMB) results to study the performance of 11 MPI communication functions on these systems.

Saini, Subhash

High-frequency Electrocardiogram Analysis in the Ability to Predict Reversible Perfusion Defects during Adenosine Myocardial Perfusion Imaging

Background: A previous study has shown that analysis of high-frequency QRS components (HF-QRS) is highly sensitive and reasonably specific for detecting reversible perfusion defects on myocardial perfusion imaging (MPI) scans during adenosine. The purpose of the present study was to try to reproduce those findings. Methods: 12-lead high-resolution electrocardiogram recordings were obtained from 100 patients before (baseline) and during adenosine Tc-99m-tetrofosmin MPI tests. HF-QRS were analyzed regarding morphology and changes in root mean square (RMS) voltages from before the adenosine infusion to peak infusion. Results: The best area under the curve (AUC) was found in supine patients (AUC=0.736) in a combination of morphology and RMS changes. None of the measurements, however, were statistically better than tossing a coin (AUC=0.5). Conclusion: Analysis of HF-QRS was not significantly better than tossing a coin for determining reversible perfusion defects on MPI scans.

Tragardh, Elin

High-Performance Data Analysis Tools for Sun-Earth Connection Missions

The data analysis tool of choice for many Sun-Earth Connection missions is the Interactive Data Language (IDL) by ITT VIS. The increasing amount of data produced by these missions and the increasing complexity of image processing algorithms requires access to higher computing power. Parallel computing is a cost-effective way to increase the speed of computation, but algorithms oftentimes have to be modified to take advantage of parallel systems. Enhancing IDL to work on clusters gives scientists access to increased performance in a familiar programming environment. The goal of this project was to enable IDL applications to benefit from both computing clusters as well as graphics processing units (GPUs) for accelerating data analysis tasks. The tool suite developed in this project enables scientists now to solve demanding data analysis problems in IDL that previously required specialized software, and it allows them to be solved orders of magnitude faster than on conventional PCs. The tool suite consists of three components: (1) TaskDL, a software tool that simplifies the creation and management of task farms, collections of tasks that can be processed independently and require only small amounts of data communication; (2) mpiDL, a tool that allows IDL developers to use the Message Passing Interface (MPI) inside IDL for problems that require large amounts of data to be exchanged among multiple processors; and (3) GPULib, a tool that simplifies the use of GPUs as mathematical coprocessors from within IDL. mpiDL is unique in its support for the full MPI standard and its support of a broad range of MPI implementations. GPULib is unique in enabling users to take advantage of an inexpensive piece of hardware, possibly already installed in their computer, and achieve orders of magnitude faster execution time for numerically complex algorithms. TaskDL enables the simple setup and management of task farms on compute clusters. The products developed in this project have the potential to interact, so one can build a cluster of PCs, each equipped with a GPU, and use mpiDL to communicate between the nodes and GPULib to accelerate the computations on each node.

Messmer, Peter

pFUnit 3.0 Tutorial Advanced

This tutorial will introduce Fortran developers to unit-testing and test-driven development (TDD) using pFUnit. As with other unit-testing frameworks, pFUnit, simplifies the process of writing, collecting, and executing tests while providing clear diagnostic messages for failing tests. pFUnit specifically targets the development of scientific-technical software written in Fortran and includes customized features such as: assertions for multi-dimensional arrays, distributed (MPI) and thread-based (OpenMP) parallellism, and flexible parameterized tests.These sessions will include numerous examples and hands-on exercises that gradually build in complexity. Attendees are expected to have working knowledge of F90, but familiarity with object-oriented syntax in F2003 and MPI will be of benefit for the more advanced examples. By the end of the tutorial the audience should feel comfortable in applying pFUnit within their own development environment.

Test Driven using pFUnit

Exploration of Nirmatrelvir Derivatives as Optimized SARS‐CoV‐2 Antivirals

Nirmatrelvir (NMV) is a SARS‐CoV‐2 antiviral component of the approved COVID‐19 therapeutic Paxlovid. It is a reversible covalent inhibitor of SARS‐CoV‐2 main protease (M Pro ) that is effluxed from human cells by P‐glycoprotein (P‐gp). To identify NMV analogs with improved potency and reduced P‐gp efflux, a structure–activity relationship campaign was conducted. Warheads alternative to nitrile for engaging the active site cysteine were tested showing aldehyde and dichloroacetamide with better enzyme inhibition potency. Crystal structure of MPI‐136−M Pro shows its aldehyde warhead forming a thiohemiacetal with active Cys145 of M Pro . Several S4 binders were explored revealing that an O‐to‐S shift at the N ‐terminal amide leads to better enzyme inhibition. By exploring different combinations of S2, S3, and S4 binders, two inhibitors with better enzyme inhibition potency than NMV were found. Crystal structure of MPI‐148, with ( S )‐2‐azaspiro[4,5]decane‐3‐carboxylate as an alternative S2 binder, shows extensive hydrogen‐bond networks for locking the inhibitor in active site, explaining high affinity of NMV analogs. Further characterization of cellular M Pro engagement and antiviral potency against SARS‐CoV‐2 revealed four inhibitors with greater potency than NMV in P‐gp‐expressing cells. Studies with the P‐gp inhibitor CP‐100356 showed that these compounds were less sensitive to P‐gp inhibition than NMV, consistent with reduced P‐gp‐mediated efflux.

Alugubelli, Yugendar R. [Texas A&M Drug Discovery

Lessons Learned and Scalability Achieved When Porting Uintah to DOE Exascale Systems

A key challenge faced when preparing codes for Department of Energy (DOE) exascale systems was designing scalable applications for systems featuring hardware and software not yet available at leadership-class scale. With such systems now available, it is important to evaluate scalability of the resulting software solutions on these target systems. One such code designed with the exascale DOE Aurora and DOE Frontier systems in mind is the Uintah Computational Framework, an open-source asynchronous many-task (AMT) runtime system. To prepare for exascale, Uintah adopted a portable MPI+X hybrid parallelism approach using the Kokkos performance portability library (i.e., MPI+Kokkos). This paper complements recent work with additional details and an evaluation of the resulting approach on Aurora and Frontier. Results are shown for a challenging benchmark demonstrating interoperability of 3 portable codes essential to Uintah-related combustion research. These results demonstrate single-source portability across Aurora and Frontier with scaling characteristics shown to 3,072 Aurora nodes and 9,216 Frontier nodes. In addition to showing results run to new scales on new systems, this paper also discusses lessons learned through efforts preparing Uintah for exascale systems.

Holmen, John [ORNL] (ORCID:0000000259342641)