Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “MPI characterization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

Evaluating MPI resource usage summary statistics

The Message Passing Interface (MPI) remains the dominant programming model for scientific applications running on today’s high-performance computing (HPC) systems. This dominance stems from MPI’s powerful semantics for inter-process communication that has enabled scientists to write applications for simulating important physical phenomena. MPI does not, however, specify how messages and synchronization should be carried out. Those details are typically dependent on low-level architecture details and the message characteristics of the application. Therefore, analyzing an application’s MPI resource usage is critical to tuning MPI’s performance on a particular platform. The result of this analysis is typically a discussion of the mean message sizes, queue search lengths and message arrival times for a workload or set of workloads. While a discussion of the arithmetic mean in MPI resource usage might be the most intuitive summary statistic, it is not always the most accurate in terms of representing the underlying data. In this paper, we analyze MPI resource usage for a number of key MPI workloads using an existing MPI trace collector and discrete-event simulator. Our analysis demonstrates that the average, while easy and efficient to calculate, is a useful metric for characterizing latency and bandwidth measurements, but may not be a good representation of application message sizes, match list search depths, or MPI inter-operation times. Additionally, we show that the median and mode are superior choices in many cases. We also observe that the arithmetic mean is not the best representation of central tendency for data that are drawn from distributions that are multi-modal or have heavy tails. Furthermore, the results and analysis of our work provide valuable guidance on how we, as a community, should discuss and analyze MPI resource usage data for scientific applications.

97 MATHEMATICS AND COMPUTING↗

Multihead Attention U‐Net for Magnetic Particle Imaging–Computed Tomography Image Segmentation

Magnetic particle imaging (MPI) is an emerging noninvasive molecular imaging modality with high sensitivity and specificity, exceptional linear quantitative ability, and potential for successful applications in clinical settings. Computed tomography (CT) is typically combined with the MPI image to obtain more anatomical information. Herein, a deep learning‐based approach for MPI‐CT image segmentation is presented. The dataset utilized in training the proposed deep learning model is obtained from a transgenic mouse model of breast cancer following administration of indocyanine green (ICG)‐conjugated superparamagnetic iron oxide nanoworms (NWs‐ICG) as the tracer. The NWs‐ICG particles progressively accumulate in tumors due to the enhanced permeability and retention (EPR) effect. The proposed deep learning model exploits the advantages of the multihead attention mechanism and the U‐Net model to perform segmentation on the MPI‐CT images, showing superb results. In addition, the model is characterized with a different number of attention heads to explore the optimal number for our custom MPI‐CT dataset.

Juhong, Aniwat↗

Parallel Adjective High-Order CFD Simulations Characterizing SOFIA Cavity Acoustics

This paper presents large-scale MPI-parallel computational uid dynamics simulations for the Stratospheric Observatory for Infrared Astronomy (SOFIA). SOFIA is an airborne, 2.5-meter infrared telescope mounted in an open cavity in the aft fuselage of a Boeing 747SP. These simulations focus on how the unsteady ow eld inside and over the cavity interferes with the optical path and mounting structure of the telescope. A temporally fourth-order accurate Runge-Kutta, and spatially fth-order accurate WENO- 5Z scheme was used to perform implicit large eddy simulations. An immersed boundary method provides automated gridding for complex geometries and natural coupling to a block-structured Cartesian adaptive mesh re nement framework. Strong scaling studies using NASA's Pleiades supercomputer with up to 32k CPU cores and 4 billion compu- tational cells shows excellent scaling. Dynamic load balancing based on execution time on individual AMR blocks addresses irregular numerical cost associated with blocks con- taining boundaries. Limits to scaling beyond 32k cores are identi ed, and targeted code optimizations are discussed.

CFD↗

Parallel Adaptive High-Order CFD Simulations Characterizing SOFIA Cavitiy Acoustics

This paper presents large-scale MPI-parallel computational uid dynamics simulations for the Stratospheric Observatory for Infrared Astronomy (SOFIA). SOFIA is an airborne, 2.5-meter infrared telescope mounted in an open cavity in the aft fuselage of a Boeing 747SP. These simulations focus on how the unsteady ow eld inside and over the cavity interferes with the optical path and mounting structure of the telescope. A tempo- rally fourth-order accurate Runge-Kutta, and a spatially fth-order accurate WENO-5Z scheme were used to perform implicit large eddy simulations. An immersed boundary method provides automated gridding for complex geometries and natural coupling to a block-structured Cartesian adaptive mesh re nement framework. Strong scaling studies using NASA's Pleiades supercomputer with up to 32k CPU cores and 4 billion compu- tational cells shows excellent scaling. Dynamic load balancing based on execution time on individual AMR blocks addresses irregular numerical cost associated with blocks con- taining boundaries. Limits to scaling beyond 32k cores are identi ed, and targeted code optimizations are discussed.

Parallell↗

Characterizing the performance of node-aware strategies for irregular point-to-point communication on heterogeneous architectures

Supercomputer architectures are trending toward higher computational throughput due to the inclusion of heterogeneous compute nodes. These multi-GPU nodes increase on-node computational efficiency, while also increasing the amount of data to be communicated and the number of potential data flow paths. In this work, we characterize the performance of irregular point-to-point communication with MPI on heterogeneous compute environments through performance modeling, demonstrating the limitations of standard communication strategies for both device-aware and staging-through-host communication techniques. Presented models suggest staging communicated data through host processes then using node-aware communication strategies for high inter-node message counts. Notably, the models also predict that node-aware communication utilizing all available CPU cores to communicate inter-node data leads to the most performant strategy when communicating with a high number of nodes. Furthermore, model validation is provided via a case study of irregular point-to-point communication patterns in distributed sparse matrix–vector products. Importantly, we include a discussion on the implications model predictions have on communication strategy design for emerging supercomputer architectures.

97 MATHEMATICS AND COMPUTING↗

Parallel Adaptive High-Order CFD Simulations Characterizing Cavity Acoustics for the Complete SOFIA Aircraft

This paper presents one-of-a-kind MPI-parallel computational fluid dynamics simulations for the Stratospheric Observatory for Infrared Astronomy (SOFIA). SOFIA is an airborne, 2.5-meter infrared telescope mounted in an open cavity in the aft of a Boeing 747SP. These simulations focus on how the unsteady flow field inside and over the cavity interferes with the optical path and mounting of the telescope. A temporally fourth-order Runge-Kutta, and spatially fifth-order WENO-5Z scheme was used to perform implicit large eddy simulations. An immersed boundary method provides automated gridding for complex geometries and natural coupling to a block-structured Cartesian adaptive mesh refinement framework. Strong scaling studies using NASA's Pleiades supercomputer with up to 32,000 cores and 4 billion cells shows excellent scaling. Dynamic load balancing based on execution time on individual AMR blocks addresses irregularities caused by the highly complex geometry. Limits to scaling beyond 32K cores are identified, and targeted code optimizations are discussed.

Acoustics↗

Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads

Scientific workflows increasingly involve both HPC and machine-learning tasks, combining MPI-based simulations, training, and inference in a single execution. Launchers such as Slurm’s srun constrain concurrency and throughput, making them unsuitable for dynamic and heterogeneous workloads. We present a performance study of RADICAL-Pilot (RP) integrated with Flux and Dragon, two complementary runtime systems that enable hierarchical resource management and high-throughput function execution. Using synthetic and production-scale workloads on Frontier, we characterize the task execution properties of RP across runtime configurations. RP+Flux sustains up to 930 tasks/s, and RP+Flux+Dragon exceeds 1,500 tasks/s with over 99.6% utilization. In contrast, srun peaks at 152 tasks/s and degrades with scale, with utilization below 50%. For IMPECCABLE.v2 drug discovery campaign, RP+Flux reduces makespan by 30–60% relative to srun/Slurm and increases throughput more than four times on up to 1,024. These results demonstrate hybrid runtime integration in RP as a scalable approach for hybrid AI-HPC workloads.

HPC-AI↗

Reflective Insulation

NRC-2 Superinsulation, manufactured by Metallized Products, Inc. (MPI), is designed for superconducting magnets used in MRI systems and particle accelerators. It is a thin, polyester film characterized by a unique crinkled surface that provides surface stand-off between layers and minimizes heat transfer in multilayer applications. NRC-2/Two is a two-sided metallized film. The material, originally developed as a skin for balloon-like satellites, was later used by NASA as a thermal barrier. MPI worked with NASA on the development of the original material and now supplies it for both space and consumer applications.

Source record↗

Methodology and Application of HPC I/O Characterization with MPIProf and IOT

Combining the strengths of MPIProf and IOT, an efficient and systematic method is devised for I/O characterization at the per-job, per-rank, per-file and per-call levels of HPC programs running on the NASA Advanced Supercomputing Center. This method is applied to answer four I/O questions in this paper. A total of 13 MPI programs and 15 cases, ranging from 24 to 5968 ranks, are analyzed to establish the I/O landscape from answers to the four questions. Four of the 13 programs use MPI I/O and the behavior of their collective writes depends on the specific implementation of the MPI library used. The SGI MPT library, the prevailing MPI library for our systems, was found to gather small writes from a large number of ranks to perform larger writes by a small subset of collective buffering ranks. The number of collective buffering ranks invoked by MPT depends on the Lustre stripe count and the number of nodes used for the run. A demonstration of varying the stripe count to achieve double-digit speedup of one program's I/O was presented. Another program, which concurrently opens private files by all ranks and could potentially create a heavy load on the Lustre servers, was identified. The ability to systematically characterize I/O for a large number of programs running on a supercomputer, seek I/O optimization opportunity and identify programs that could cause a high load and instability on the filesystems is important for pursuing exascale in a real production environment.

Characterization↗

A Reference Implementation for a Quantum Message Passing Interface

Practical applications of quantum computing are currently limited by the number of qubits that can be set with reasonable fidelities for each system. Therefore, a distributed quantum computing system with multiple quantum computers coherently connected is highly demanding. To realize the internode communication of quantum information, the software interface, Quantum Message Passing Interface (QMPI), leveraging the framework built for classical MPI but taking advantage of quantum teleportation to communicate between different quantum nodes was proposed. In this project, we develop the QMPI with point-to-point and collective operations in Qiskit and characterize its performance by demonstrating the application implementations. Moreover, we developed a new technique for optimizing collective communication of the distributed quantum programs with Multi-Controlled Toffoli gates. This technique beats the state-of-the-art in terms of fidelity and the number of remote EPR pairs consumed in both simulations and experiments.

Shi, Yue↗

Exploration of Nirmatrelvir Derivatives as Optimized SARS‐CoV‐2 Antivirals

Nirmatrelvir (NMV) is a SARS‐CoV‐2 antiviral component of the approved COVID‐19 therapeutic Paxlovid. It is a reversible covalent inhibitor of SARS‐CoV‐2 main protease (M Pro ) that is effluxed from human cells by P‐glycoprotein (P‐gp). To identify NMV analogs with improved potency and reduced P‐gp efflux, a structure–activity relationship campaign was conducted. Warheads alternative to nitrile for engaging the active site cysteine were tested showing aldehyde and dichloroacetamide with better enzyme inhibition potency. Crystal structure of MPI‐136−M Pro shows its aldehyde warhead forming a thiohemiacetal with active Cys145 of M Pro . Several S4 binders were explored revealing that an O‐to‐S shift at the N ‐terminal amide leads to better enzyme inhibition. By exploring different combinations of S2, S3, and S4 binders, two inhibitors with better enzyme inhibition potency than NMV were found. Crystal structure of MPI‐148, with ( S )‐2‐azaspiro[4,5]decane‐3‐carboxylate as an alternative S2 binder, shows extensive hydrogen‐bond networks for locking the inhibitor in active site, explaining high affinity of NMV analogs. Further characterization of cellular M Pro engagement and antiviral potency against SARS‐CoV‐2 revealed four inhibitors with greater potency than NMV in P‐gp‐expressing cells. Studies with the P‐gp inhibitor CP‐100356 showed that these compounds were less sensitive to P‐gp inhibition than NMV, consistent with reduced P‐gp‐mediated efflux.

Alugubelli, Yugendar R. [Texas A&M Drug Discovery ↗

Design and Performance Characterization of RADICAL-Pilot on Leadership-Class Platforms

Many extreme scale scientific applications have workloads comprised of a large number of individual highperformance tasks. The Pilot abstraction decouples workload specification, resource management, and task execution via job placeholders and late-binding. As such, suitable implementations of the Pilot abstraction can support the collective execution of large number of tasks on supercomputers. We introduce RADICAL-Pilot (RP) as a portable, modular and extensible Pilot enabled runtime system. We describe RP's design, architecture and implementation. We characterize its performance and show its ability to scalably execute workloads comprised of tens of thousands heterogeneous tasks on DOE and NSF leadership-class HPC platforms. Specifically, we investigate RP's weak/strong scaling with CPU/GPU, single/multi core, (non)MPI tasks and python functions when using most of ORNL Summit and TACC Frontera. RADICAL-Pilot can be used stand-alone, as well as the runtime for third-party workflow systems.

97 MATHEMATICS AND COMPUTING↗

High Resolution Aerospace Applications using the NASA Columbia Supercomputer

This paper focuses on the parallel performance of two high-performance aerodynamic simulation packages on the newly installed NASA Columbia supercomputer. These packages include both a high-fidelity, unstructured, Reynolds-averaged Navier-Stokes solver, and a fully-automated inviscid flow package for cut-cell Cartesian grids. The complementary combination of these two simulation codes enables high-fidelity characterization of aerospace vehicle design performance over the entire flight envelope through extensive parametric analysis and detailed simulation of critical regions of the flight envelope. Both packages. are industrial-level codes designed for complex geometry and incorpor.ats. CuStomized multigrid solution algorithms. The performance of these codes on Columbia is examined using both MPI and OpenMP and using both the NUMAlink and InfiniBand interconnect fabrics. Numerical results demonstrate good scalability on up to 2016 CPUs using the NUMAIink4 interconnect, with measured computational rates in the vicinity of 3 TFLOP/s, while InfiniBand showed some performance degradation at high CPU counts, particularly with multigrid. Nonetheless, the results are encouraging enough to indicate that larger test cases using combined MPI/OpenMP communication should scale well on even more processors.

Mavriplis, Dimitri J.↗

Workshop on Extraterrestrial Materials from Cold and Hot Deserts

The workshop was held July 6-8, 1999 before the Meteoritical Society meeting in Johannesburg, South Africa. The venue was Kwa Maritane Resort in the Pilanesburg Game Reserve. Conveners were Ludolf Schultz (Chair, MPI fur Chemie), Ian Franchi (Open University), Arch Reid (University of Houston), and Mike Zolensky (NASA JSC). Extended abstracts will be published as an LPI Technical Report. In the first session, Marvin discussed three African iron meteorites: Cape of Good Hope, Gibeon, and Hoba. Grady presented a statistical analysis of meteorites from hot and cold deserts. Wasson discussed types of Antarctic iron meteorites. Several presentations characterized populations of meteorites from individual desert areas: Libyan Desert (Weber et al.), Nullarbor Region (Bevan et al.), and Mojave Desert (Kring et al. and Verish et al). Pairing among EET87503-group howardites was discussed by Buchanan et al. Based on 14C terrestrial ages of Allan Hills ordinary chondrites, Bland et al. suggested that ice flow may be the principal sink for Antarctic meteorites. The effects of preterrestrial and terrestrial alteration were considered in the second session. Nakamura et al. and Lipschutz discussed asteroidal metamorphism of carbonaceous chondrites. Zolensky presented evidence for preterrestrial halide and sulfide in meteorites. Crozaz and Wadhwa described terrestrial alteration of Dar al Gani 476. Welten and Nishiizumi discussed terrestrial weathering of chondrites from Frontier Mountain, Antarctica. Most of the third session dealt with terrestrial meteorite ages. Based on 14C-10Be ages, Jull et al. discussed the exponential decay in numbers of meteorites with increased age. Nishiizumi et al. concluded that some Allan Hills meteorites have much older terrestrial ages than any meteorites from Lewis Cliffs. Welten et al. discussed terrestrial ages determined by 41Ca/36CI of metal separates from hot desert meteorites. Based on a comparison with large IDPs, Flynn et al. suggested that polar micrometeorites, lost a water-soluble sulfate phase by terrestrial alteration. Nyquist suggested that Type I cosmic spheres from deep sea sediments and polar ice were derived from carbonaceous chondrite-like asteroidal sources. The final session considered noble gases, cosmic ray effects, and thermoluminescence. Calculations presented by Reedy indicate that cosmic ray-produced nuclides are more likely to be preserved in small objects than in larger objects. Murty and Mohapatra. reported that trapped gas in Dar al Gani 476 includes a Martian atmospheric component. Wieler et al. and Scherer et al. reported noble gas abundances in different types of desert meteorites. Patzer and Schultz discussed the influence of terrestrial weathering on cosmic ray exposure ages of enstatite chondrites. Franchi et al. suggested that it is difficult to discriminate whether differences in gas release profiles of lunar meteorites from hot deserts and returned lunar samples are the result of terrestrial weathering or shock metamorphism. Benoit and Sears surveyed natural and induced thermoluminescence of Antarctic ordinary chondrites. Merchel et al. analyzed Saharan meteorites with short or complex exposure histories.

Buchanan, Paul C.↗

Uncovering grain and subgrain microstructure at the scale of additive manufacturing melt tracks with a scalable cellular automaton solidification model

Metal additive manufacturing, characterized by rapid solidification, yields refined grains with a distinctive cellular subgrain microstructure that plays a pivotal role in determining material properties. Due to the significant computational expense demanded to simulate the required physics with submicron spatial resolution, their numerical simulations have been limited to proof-of-concept studies to either 2D or small subregions of a melt pool. In this study, an open-source, scalable, solidification code, muMatScale, based on the cellular automaton method, has been developed to predict the grain and the underlying subgrain microstructure over an entire melt pool. The model incorporates flexible parallelization schemes, utilizing MPI and OpenMP GPU Offloading, in addition to appropriate multi-physics specific to non-equilibrium rapid solidification in AM. The impact of nucleation parameters on grain microstructures was investigated with a focus on grain size variations and morphology transitions. With selected nucleation parameters, the simulation predicted the grain size, subgrain morphology, crystallographic orientation, and microsegregation aligned with experimental measurements. The model demonstrates that epitaxial grain growth is a dominant factor at the melt pool boundary, influencing grain size variation under different grain sizes in the build plate while maintaining consistent primary dendrite arm spacing under identical thermal conditions. Here, the highly efficient numerical model enables large-scale simulations with a spatial resolution of 100 nm or less, unveiling unprecedented insights into thermal and solutal diffusion driven grain growth, and the subgrains with microsegregation within grains in 3D across scales. muMatScale will enable the linking of submicron length-scale microstructure to part-level material behavior by investigating fundamental solidification problems at the intercellular scale in many-track and many-layer builds.

36 MATERIALS SCIENCE↗

ERF: Energy Research and Forecasting

The Energy Research and Forecasting (ERF) code is a new model that simulates the mesoscale and microscale dynamics of the atmosphere using the latest high-performance computing architectures. It employs hierarchical parallelism using an MPI+X model, where X may be OpenMP on multicore CPU-only systems, or CUDA, HIP, or SYCL on GPU-accelerated systems. ERF is built on AMReX (Zhang et al., 2019, 2021), a block-structured adaptive mesh refinement (AMR) software framework that provides the underlying performance-portable software infrastructure for block-structured mesh operations. The "energy" aspect of ERF indicates that the software has been developed with renewable energy applications in mind. In addition to being a numerical weather prediction model, ERF is designed to provide a flexible computational framework for the exploration and investigation of different physics parameterizations and numerical strategies, and to characterize the flow field that impacts the ability of wind turbines to extract wind energy. The ERF development is part of a broader effort led by the US Department of Energy's Wind Energy Technologies Office.

17 WIND ENERGY↗

Three practical workflow schedulers for easy maximum parallelism

Runtime scheduling and workflow systems are an increasingly popular algorithmic component in HPC because they allow full system utilization with relaxed synchronization requirements. There are so many special-purpose tools for task scheduling, one might wonder why more are needed. Use cases seen on the Summit supercomputer needed better integration with MPI and greater flexibility in job launch configurations. Preparation, execution, and analysis of computational chemistry simulations at the scale of tens of thousands of processors revealed three distinct workflow patterns. A separate job scheduler was implemented for each one using extremely simple and robust designs: file-based, task-list based, and bulk-synchronous. Comparing to existing methods shows unique benefits of this work, including simplicity of design, suitability for HPC centers, short startup time, and well-understood per-task overhead. All three new tools have been shown to scale to full utilization of Summit, and have been made publicly available with tests and documentation. This work presents a complete characterization of the minimum effective task granularity for efficient scheduler usage scenarios. Here, these schedulers have the same bottlenecks, and hence similar task granularities as those reported for existing tools following comparable paradigms.

97 MATHEMATICS AND COMPUTING↗

Transient Triplet Metallopnictinidenes M–Pn (M = Pd II , Pt II ; Pn = P, As, Sb): Characterization and Dimerization

Nitrenes (R–N) have been subject to a large body of experimental and theoretical studies. The fundamental reactivity of this important class of transient intermediates has been attributed to their electronic structures, particularly the accessibility of triplet vs singlet states. In contrast, electronic structure trends along the heavier pnictinidene analogues (R–Pn; Pn = P–Bi) are much less systematically explored. We here report the synthesis of a series of metallodipnictenes, {M–Pn=Pn–M} (M = Pd II , Pt II ; Pn = P, As, Sb, Bi) and the characterization of the transient metallopnictinidene intermediates, {M–Pn} for Pn = P, As, Sb. Structural, spectroscopic, and computational analysis revealed spin triplet ground states for the metallopnictinidenes with characteristic electronic structure trends along the series. In comparison to the nitrene, the heavier pnictinidenes exhibit lower-lying ground state SOMOs and singlet excited states, thus suggesting increased electrophilic reactivity. Furthermore, the splitting of the triplet magnetic microstates is beyond the phosphinidenes {M–P} dominated by heavy pnictogen atom induced spin–orbit coupling.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗