Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel communication”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Employing artificial intelligence to steer exascale workflows with colmena

Computational workflows are a common class of application on supercomputers, yet the loosely coupled and heterogeneous nature of workflows often fails to take full advantage of their capabilities. We created Colmena to leverage the massive parallelism of a supercomputer by using Artificial Intelligence (AI) to learn from and adapt a workflow as it executes. Colmena allows scientists to define how their application should respond to events (e.g., task completion) as a series of cooperative agents. In this paper, we describe the design of Colmena, the challenges we overcame while deploying applications on exascale systems, and the science workflows we have enhanced through interweaving AI. The scaling challenges we discuss include developing steering strategies that maximize node utilization, introducing data fabrics that reduce communication overhead of data-intensive tasks, and implementing workflow tasks that cache costly operations between invocations. These innovations coupled with a variety of application patterns accessible through our agent-based steering model have enabled science advances in chemistry, biophysics, and materials science using different types of AI. In conclusion, our vision is that Colmena will spur creative solutions that harness AI across many domains of scientific computing.

Workflows↗

Decentralized Carrier Phase Shifting for Optimal Harmonic Minimization in Asymmetric Parallel-Connected Inverters

This paper presents a carrier phase shifting technique for minimizing the aggregate harmonics in networks of asymmetric parallel-connected inverters for distributed power generation system applications. The proposed technique is: 1) implemented in a decentralized manner, relying only on local voltage and current measurements, and 2) optimal in the sense that it minimizes a cost function representing the carrier-frequency current harmonics. The analysis indicates that the proposed optimal carrier phase shifting technique can enable order-of-magnitude reductions in harmonic power, and also universal improvements compared to symmetric carrier interleaving for asymmetric inverter networks. Moreover, compared to existing methods that require either centralized communication or information exchange between inverters to coordinate carriers, the proposed technique is completely decentralized, which provides important practical benefits for implementation, including improved robustness and reduced cost. The technique is experimentally validated on a network of three single-phase 2-kW inverters and demonstrates a 36.5% reduction in the weighted total harmonic distortion factor of the aggregate inverter current, and the ability to converge to the optimal carrier phase spacing dynamically in less than one line frequency cycle (16.7 ms) in steady state and transient operating conditions.

42 ENGINEERING↗

Characterization and identification of HPC applications at leadership computing facility

High Performance Computing (HPC) is an important method for scientific discovery via large-scale simulation, data analysis, or artificial intelligence. Leadership-class supercomputers are expensive, but essential to run large HPC applications. The Petascale era of supercomputers began in 2008, with the first machines achieving performance in excess of one petaflops, and with the advent of new supercomputers in 2021 (e.g., Aurora, Frontier), the Exascale era will soon begin. However, the high theoretical computing capability (i.e., peak FLOPS) of a machine is not the only meaningful target when designing a supercomputer, as the resources demand of applications varies. A deep understanding of the characterization of applications that run on a leadership supercomputer is one of the most important ways for planning its design, development and operation. In order to improve our understanding of HPC applications, user demands and resource usage characteristics, we perform correlative analysis of various logs for different subsystems of a leadership supercomputer. This analysis reveals surprising, sometimes counter-intuitive patterns, which, in some cases, conflicts with existing assumptions, and have important implications for future system designs as well as supercomputer operations. For example, our analysis shows that while the applications spend significant time on MPI, most applications spend very little time on file I/O. Combined analysis of hardware event logs and task failure logs show that the probability of a hardware FATAL event causing task failure is low. Combined analysis of control system logs and file I/O logs reveals that pure POSIX I/O is used more widely than higher level parallel I/O. Based on holistic insights of the application gained through combined and co-analysis of multiple logs from different perspectives and general intuition, we engineer features to "fingerprint" HPC applications. We use t-SNE (a machine learning technique for dimensionality reduction) to validate the explainability of our features and finally train machine learning models to identify HPC applications or group those with similar characteristic. To the best of our knowledge, this is the first work that combines logs on file I/O, computing, and inter-node communication for insightful analysis of HPC applications in production.

Liu, Zhengchun↗

A parallel strategy for density functional theory computations on accelerated nodes

Using the Löwdin orthonormalization of tall-skinny matrices as a proxy-app for wavefunction-based Density Functional Theory solvers, we investigate a distributed memory parallel strategy focusing on Graphics Processing Unit (GPU)-accelerated nodes as available on some of the top ranked supercomputers at the present time. Here we present numerical results in the strong limit regime, as it is particularly relevant for First-Principles Molecular Dynamics. We also examine how matrix product-based iterative solvers provide a competitive alternative to dense eigensolvers on GPUs, allowing to push the strong scaling limit of these computations to a larger number of distributed tasks. Our strategy, which relies on replicated Gram matrices and efficient collective communications using the NCCL library, leads to a time-to-solution under 0.5 s for the Löwdin orthonormalization of a tall-skinny matrix of 3000 columns on Summit at Oak Ridge Leadership Facility (OLCF). Given the similarity in computational operations between one iteration of a DFT solver and this proxy-app, this shows the possibility of solving accurately the DFT equations well under a minute for 3000 electronic wave functions, and thus perform First-Principles molecular dynamics of physical systems much larger than traditionally solved on CPU systems.

97 MATHEMATICS AND COMPUTING↗

Evaluating asynchronous Schwarz solvers on GPUs

With the commencement of the exascale computing era, we realize that the majority of the leadership supercomputers are heterogeneous and massively parallel. Even a single node can contain multiple co-processors such as GPUs and multiple CPU cores. For example, ORNL’s Summit accumulates six NVIDIA Tesla V100 GPUs and 42 IBM Power9 cores on each node. Synchronizing across compute resources of multiple nodes can be prohibitively expensive. Hence, it is necessary to develop and study asynchronous algorithms that circumvent this issue of bulk-synchronous computing. In this study, we examine the asynchronous version of the abstract Restricted Additive Schwarz method as a solver. We do not explicitly synchronize, but allow the communication between the sub-domains to be completely asynchronous, thereby removing the bulk synchronous nature of the algorithm. We accomplish this by using the one-sided Remote Memory Access (RMA) functions of the MPI standard. We study the benefits of using such an asynchronous solver over its synchronous counterpart. We also study the communication patterns governed by the partitioning and the overlap between the sub-domains on the global solver. Finally, we show that this concept can render attractive performance benefits over the synchronous counterparts even for a well-balanced problem.

Nayak, Pratik↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)↗

Implementing Software Resiliency in HPX for Extreme Scale Computing

The DOE Office of Science Exascale Computing Project (ECP) outlines the next milestones in the supercomputing domain. The target computing systems under the project will deliver 10x performance while keeping the power budget under 30 megawatts. With such large machines, the need to make applications resilient has become paramount. The benefits of adding resiliency to mission critical and scientific applications, includes the reduced cost of restarting the failed simulation both in terms of time and power. Most of the current implementation of resiliency at the software level makes use of a Coordinated Checkpoint and Restart (C/R). This technique of resiliency generates a consistent global snapshot, also called a checkpoint. Generating snapshots involves global communication and coordination and is achieved by synchronizing all running processes. The generated checkpoint is then stored in some form of persistent storage. On failure detection, the runtime initiates a global rollback to the most recent previously saved checkpoint. This involves aborting all running processes, rolling them back to the previous state and restarting them.

97 MATHEMATICS AND COMPUTING↗

Intern Poster Session 08/13: Autonomous Nuclear Robotics: Applications in nuclear waste inspection and hot cell experiments

The nuclear industry is experiencing renewed interest in autonomous robotics, yet most deployed systems remain teleoperated with limited autonomy. This work presents two contributions toward fully autonomous nuclear robotic systems: autonomous waste inspection at the Hanford Site and an autonomous hot cell laboratory framework. Inspections of Hanford's underground waste storage tanks are performed manually at significant cost and personnel exposure. We developed a reinforcement-learning (RL) training pipeline for a custom-built inspection arm. In parallel, we are designing an autonomous laboratory framework for post-irradiation examination in hot cells at the Specimen Preparation Laboratory (SPL) that integrates computer vision, task and motion planning, hardware execution, and operator-in-the-loop control. These systems demonstrate a path toward safer, more efficient nuclear operations by reducing human exposure while maintaining rigorous human oversight at critical decision points.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A fast and large bandwidth superconducting variable coupler

Variable microwave-frequency couplers are highly useful components in classical communication systems and likely will play an important role in quantum communication applications. Conventional semiconductor-based microwave couplers have been used with superconducting quantum circuits, enabling, for example, the in situ measurements of multiple devices via a common readout chain. However, the semiconducting elements are lossy and furthermore dissipate energy when switched, making them unsuitable for cryogenic applications requiring rapid, repeated switching. Superconducting Josephson junction-based couplers can be designed for dissipation-free operation with fast switching and are easily integrated with superconducting quantum circuits. These enable on-chip, quantum-coherent routing of microwave photons, providing an appealing alternative to semiconductor switches. Here, we present and characterize a chip-based broadband microwave variable coupler, tunable over 4-8GHz with over 1.5GHz instantaneous bandwidth, based on the superconducting quantum interference device with two parallel Josephson junctions. The coupler is dissipation-free and features large on-off ratios in excess of 40dB, and the coupling can be changed in about 10ns. The simple design presented here can be readily integrated with superconducting qubit circuits and can be easily generalized to realize a four- or more port device.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Evaluating the Effectiveness of an Ultrasonic Acoustic Deterrent in Reducing Bat Fatalities at Wind Energy Facilities

This project was designed to use thermal video cameras and fatality monitoring to evaluate the effectiveness of an ultrasonic acoustic deterrent (UAD) on bat activity and mortality, respectively. Our goals were to redesign the UAD device and installation infrastructure, determine the placement on wind turbines to optimize safety, compatibility and functionality, and to compare the mortality among the following conditions: Control (deterrents off and turbines feathered up to the manufacturer’s cut-in speed of 3.5 m/s), Deterrent (deterrents on and turbines feathered up to the manufacturer’s cut-in speed of 3.5 m/s), Curtailment (deterrents off and turbines feathered up to 5/ m/s), and combination (deterrents on and turbines feathered up to 5 m/s). The project was divided into a Feasibility Study and Comparative Study. The objectives for the Feasibility Study were to develop an installation strategy, redesign a previous iteration of a UAD to improve performance and weatherization, and test the effectiveness of the deterrents on bat activity. The Feasibility Study was intended to work out potential issues using a relatively small number of devices prior to manufacturing and installing numerous devices for a larger-scale comparative study. The division of the project into these separate studies was based on previous experience and the challenges of assessing the capabilities of an untested UAD. During the initial development and manufacturing of the UAD, NRG Systems decided to use a piezoelectric transducer rather than an electrostatic transducer, which was used by a previous device (i.e. Deaton UAD). NRG Systems conducted lab testing (i.e., IP67 or Ingress Protection) to ensure no water or dust ingress. In addition, shock/drop trials and variations in temperature exposure were conducted as part of the reliability testing. NRG Systems also developed a communications system to allow for continuous performance monitoring of the UADs. For the Feasibility Study, we installed 6 UADs on each of 2 Gamesa G90 2-MW wind turbines (Turbine 14 and 8) at the South Chestnut Wind Energy Facility, Pennsylvania and monitored activity under control and treatment (i.e. Deterrent) conditions using thermal video monitoring. There are two major sources of variation in bat activity (beyond the anticipated treatment effect): 1) environment around the turbines might inherently favor more activity at one than the other; 2) weather conditions on any given night or within season difference (e.g., migration later during the study period) might favor more activity on some nights than on other nights. Because we could only monitor two turbines on any night, we sought to control the potential influences of these two sources by alternating the turbine on which deterrents were activated each night. If there were no loss of data due to technical failures, this design would result in an equal number of deterrent and control nights at each turbine through the monitoring period, balancing the effects of both sources of variation. We compared the time bats spent and the number of events that occurred in overlapping cameras FOV as an indicator of risk, since 80% of the overlapping FOV of the cameras was in the RSA. Equipment failures, majority due to lightning, resulted in only 17 nights with useable data, with unbalanced treatment assignment within turbines and uneven distribution of treatment assignment throughout the observational period. This resulted in a confounding of treatment assignment and seasonal change. Deterrent treatment was measured at Turbine 14 on only 2 of the first 8 usable nights (spanning the period from 8/18-9/17), whereas 6 times on Turbine 8. From 9/18-927, deterrent was on at Turbine 14 on 6 of the remaining 10 nights, and 4 on Turbine 8. We recorded a total of 1,057 bats and observed a reduction in number of events and duration of events at the UAD-activated turbine when it was Turbine 14. When Turbine 8 had the UAD activated, we observed no difference in number or duration of events. Variation between turbines is not unusual and can cause issues when study designs have no true replication, i.e., multiple turbines per treatment. Within-turbine differences suggested a trend for reduced activity when UADs were activated, particularly for Turbine 14. These results may be caused by the overall higher bat activity at Turbine 8 and confounding of treatment assignment and seasonal trends. We mapped 58 bat events in 3-dimensional space (3D), 30 and 28 events during control and treatment conditions, respectively. We observed bats crossing the rotor plane under both control and treatment conditions and observed a total of 40 confirmed or near-collisions (i.e., target close to blade but no visual confirmation of a strike) out of a total of 1,491 medium and high confidence bat observations (880 control, 611 at treatment). Twice as many collisions/possible collisions were observed during control conditions. Given the challenges with the equipment and potential confounding of the data (i.e., different activity levels at the two wind turbines), we were unable to determine whether this initial turbine placement and orientation was optimal. Given no new information on how best to install the devices, we elected to use the same placement and orientation for the comparative study. For the Comparative Study, the objectives were to investigate the relative mortality rates among 4 treatments. We searched the area within 90 m of each turbine daily to recover the highest number of fresh fatalities possible. We were unable to detect a clear reduction in mortality from deterrents alone for any individual species. Surprisingly, mortality rate of the eastern red bat (Lasiurus borealis) was estimated to be 1.3–4.2 times as much when turbines were operating normally and UADs were on than when UADs were off. Reduction in mortality of all bat species combined due to curtailment of turbines was estimated to be between 0%–38%. This effect was nullified when, in addition to curtailment, UADs were on, with 95% confidence interval ranging from a 45% reduction to a 36% increase in mortality. This was likely due to the large proportion of eastern red bats in the total carcass population. Mortality of all low-frequency echolocating bats combined (i.e. hoary bat [L. cinereus], big brown bat [Eptesicus fuscus], silver-haired bat [Lasionycteris noctivagans]) relative to control was lower when curtailed (95% CI: 0%–74%), but the addition of UADs had no detectable effect (95%CI: 13%–79%). The combined treatment reduced mortality in silver-haired bats relative to control by 11%–99%, compared to curtailment (81% reduction–67% increase) or deterrent (82% reduction–67% increase) alone. Because silver-haired bats comprised a large proportion of low-frequency calling bats found during this study, a similar effect was seen for that group. The higher mortality observed for eastern red bats at UAD compared to control could have been caused by several factors, such as the effective range of the UAD, particularly at higher frequencies, behavior, positioning of the devices on the nacelle, or a combination of these. We used 3D thermal videography to compare control and UAD bat behavior from two turbines using a total of 203 3D bat-tracks across 34 nights. We recorded a similar number of bat-tracks between treatment groups, with 51% and 49% for control and UAD, respectively. Due to potential differences in bat behavior around spinning vs stationary turbine blades, we examined UAD effectiveness separately for non-operating turbines (feathered below cut-in speed of 3.5 m/s) and operating (normal operation above wind speed of 3.5 m/s). We found a higher proportion of bat-tracks at operating turbines (82%) compared to non-operating turbines (18%), although this does not account for overall time turbines were operating versus not. At non-operating turbines the UAD appears to be effective at reducing the amount of time, flight length, and number of passes through the rotor plane, compared to control. In addition, we found bats approached turbines similarly between control and treatment turbines, with 61% of control and 63% of UAD bat-tracks originating leeward of the hub. In contrast, at operating turbines, we saw little change in bat behavior in response to UADs. For example, we found an increase in the average duration of bat-tracks between non-operating and operating turbines for UAD but at control turbines we found average duration decreased once turbines became operational. Both control and UAD had a high proportion of bat-tracks that crossed the rotor plane (i.e., collision risk) originate from the windward side when turbines were operational 65% to 92%, respectively. Given that the UAD devices closest to the blades were orientated parallel to the blades, its possible bats were not exposed to the signal until they were close to the turbine blades, as suggested by the slightly higher mean duration within 5 meters of the blades for UAD turbines. Future research should consider concentrating UAD intensity on the areas of risk (i.e. blades) with enough buffer to allow bats to react to the sound before entering the rotor-swept area (RSA). In addition, investigating the potential of installing UAD units windward of the turbine blades (e.g. a hub-mounted UAD), particularly since even under control conditions, a relatively high proportion (65%) of crosses through the blade plane originated windward. 3D thermal videography provided valuable information on future testing strategies (e.g. device placement, UAD orientation) to improve UAD effectiveness when bats are at risk (i.e., operating wind turbines). Across the entire project we experienced issues with the operation and communication with the UADs. Most of the issues occurred during the Feasibility Study and were resolved prior to the Comparability Study. Additional challenges surfaced during the Comparability Study but were remedied immediately and are thought to have little impact on the results. We had logistical constraints at the project that limited our ability use traditional methods in our camera calibration. Several calibrations showed inaccurate scales, which may have been related to inadequate spatial coverage of “points” in the camera calibration volume. Because we were using actual video recordings of bats at a wind turbine, the behavior of bats could have concentrated “points” in specific areas of the turbine (i.e. leeward of nacelle), and limited “points” in other areas, resulting in camera calibration issues. We have plans to address these inconsistencies and improving the software and related methodologies by early 2020.

17 WIND ENERGY↗

Multilevel analysis between Physcomitrium patens and Mortierellaceae endophytes explores potential long‐standing interaction among land plants and fungi

SUMMARY The model moss species Physcomitrium patens has long been used for studying divergence of land plants spanning from bryophytes to angiosperms. In addition to its phylogenetic relationships, the limited number of differential tissues, and comparable morphology to the earliest embryophytes provide a system to represent basic plant architecture. Based on plant–fungal interactions today, it is hypothesized these kingdoms have a long‐standing relationship, predating plant terrestrialization. Mortierellaceae have origins diverging from other land fungi paralleling bryophyte divergence, are related to arbuscular mycorrhizal fungi but are free‐living, observed to interact with plants, and can be found in moss microbiomes globally. Due to their parallel origins, we assess here how two Mortierellaceae species, Linnemannia elongata and Benniella erionia , interact with P. patens in coculture. We also assess how Mollicute ‐related or Burkholderia ‐related endobacterial symbionts (MRE or BRE) of these fungi impact plant response. Coculture interactions are investigated through high‐throughput phenomics, microscopy, RNA‐sequencing, differential expression profiling, gene ontology enrichment, and comparisons among 99 other P. patens transcriptomic studies. Here we present new high‐throughput approaches for measuring P. patens growth, identify novel expression of over 800 genes that are not expressed on traditional agar media, identify subtle interactions between P. patens and Mortierellaceae, and observe changes to plant–fungal interactions dependent on whether MRE or BRE are present. Our study provides insights into how plants and fungal partners may have interacted based on their communications observed today as well as identifying L. elongata and B. erionia as modern fungal endophytes with P. patens .

59 BASIC BIOLOGICAL SCIENCES↗

DIMPLES: Distributed Influence Maximization for Pandemic pLanning on Exascale Systems

We study exascale parallel algorithms for the selection of intervention or monitoring strategies in massive realistic socio-technical networks through scalable Influence Maximization (InfMax) algorithms. We employ novel techniques to enable efficient scaling on up to 8k nodes of OLCF Frontier, with 65k AMD GPUs and 458k AMD CPU cores. Current state-of-the-art InfMax tools are limited to networks with only a few million actors (vertices) and a few hundred million interactions (edges). By overcoming these limitations, we show that our approach is capable of processing a realistic social contact network of the United States with 285 million nodes and about 8 billion edges. This two orders-of-magnitude improvement over the previous state-of-the-art is obtained by leveraging algorithmic advancements for the InfMax problem and designing several problem-specific approaches to overlap communication with computation, improve GPU efficiency, and lower the application’s memory requirements. We evaluate strong scaling for computing 10k most influential seeds using up to 8k nodes of an exascale system, and weak scaling from 128 to 8k system nodes for seed sets ranging from 625 to 40k seeds. We achieve the fastest-known runtime of 25 minutes while performing 48 million diffusion simulations totaling 2.31 petabytes to identify 40k influential seeds using 8k nodes, and take 5.75 minutes to identify 10k seeds while using 4k nodes.

Minutoli, Marco [Pacific Northwest National Labora↗

Distributed out-of-memory NMF on CPU/GPU architectures

We propose an efficient distributed out-of-memory implementation of the non-negative matrix factorization (NMF) algorithm for heterogeneous high-performance-computing systems. The proposed implementation is based on prior work on NMFk, which can perform automatic model selection and extract latent variables and patterns from data. In this work, we extend NMFk by adding support for dense and sparse matrix operation on multi-node, multi-GPU systems. The resulting algorithm is optimized for out-of-memory problems where the memory required to factorize a given matrix is greater than the available GPU memory. Memory complexity is reduced by batching/tiling strategies, and sparse and dense matrix operations are significantly accelerated with GPU cores (or tensor cores when available). Input/output latency associated with batch copies between host and device is hidden using CUDA streams to overlap data transfers and compute asynchronously, and latency associated with collective communications (both intra-node and inter-node) is reduced using optimized NVIDIA Collective Communication Library (NCCL) based communicators. Benchmark results show significant improvement, from 32X to 76x speedup, with the new implementation using GPUs over the CPU-based NMFk. Good weak scaling was demonstrated on up to 4096 multi-GPU cluster nodes with approximately 25,000 GPUs when decomposing a dense 340 Terabyte-size matrix and an 11 Exabyte-size sparse matrix of density 10 -6 .

97 MATHEMATICS AND COMPUTING↗

Communication-Avoiding and Memory-Constrained Sparse Matrix-Matrix Multiplication at Extreme Scale

Sparse matrix-matrix multiplication (SpGEMM) is a widely used kernel in various graph, scientific computing and machine learning algorithms. In this paper, we consider SpGEMMs performed on hundreds of thousands of processors generating trillions of nonzeros in the output matrix. Distributed SpGEMM at this extreme scale faces two key challenges: (1) high communication cost and (2) inadequate memory to generate the output. Furthermore, we address these challenges with an integrated communication-avoiding and memory-constrained SpGEMM algorithm that scales to 262,144 cores (more than 1 million hardware threads) and can multiply sparse matrices of any size as long as inputs and a fraction of output fit in the aggregated memory. As we go from 16,384 cores to 262,144 cores on a Cray XC40 supercomputer, the new SpGEMM algorithm runs 10x faster when multiplying large-scale protein-similarity matrices.

97 MATHEMATICS AND COMPUTING↗

An MPMD approach coupling electromagnetic continuum mechanics approximations in ALEGRA

In this work, two complementary approximations for describing aspects of continuum electromagnetics in moving media are discussed: electroquasistatic and magnetoquasistatic. Each has been implemented in the finite element shock code ALEGRA for modeling dynamic electromechanical phenomena on typical engineering time scales, with fully integrated circuit coupling. The approximations can be obtained by consistent asymptotic balancing of Maxwell’s equations relative to timescales associated with magnetic diffusion, charge relaxation, and electromagnetic wave propagation. In ALEGRA, the electroquasistatic approximation is used for ferroelectric (FE) modeling, while the magnetoquasistatic approximation is used for magnetohydrodynamic (MHD) modeling. In this paper we introduce for the first time a detailed derivation of a useful quasi-steady “low-R m ” variant of the MHD approximation applicable for cases, such as with detonators, where the thermodynamic pressure arising from Joule heating dominates over magnetic forces. An additional purpose of this paper is to present a coupling mode using Multiple Program-Multiple Data (MPMD) message passing communication that allows the user to run 3D FE problems together with 2D and/or 3D MHD problems with the respective simulation domains coupled through a common circuit equation. The MPMD coupling capability is used here to model the dynamic coupling of a notional ferroelectric generator with an RP-87 exploding bridgewire detonator. The simulated bridgewire heats up and bursts under current generated by simulated depoling of the ferroelectric generator, as a demonstration of the MPMD capability.

42 ENGINEERING↗