Engineering PapersSearch

Engineering topics

Ghosh, Sayan

Publications and source records attributed to Ghosh, Sayan.

MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs

Graph Neural Networks (GNN) are indispensable in learning from graph-structured data, yet their rising computational costs, especially on massively connected graphs, pose significant challenges in terms of execution performance. To tackle this, distributed-memory solutions such as partitioning the graph to concurrently train multiple replicas of GNNs are in practice. However, approaches requiring a partitioned graph usually suffer from communication overhead and load imbalance, even under optimal partitioning and communication strategies due to irregularities in the neighborhood minibatch sampling. This paper proposes practical trade-offs for improving the sampling and communication overheads for representation learn- ing on distributed graphs (using popular GraphSAGE architecture) by developing a parameterized prefetch and eviction scheme on top of the state-of-the-art Amazon DistDGL distributed GNN framework, demonstrating about 15–40% improvement in end-to-end training performance on the NERSC Perlmutter supercomputer for various OGB datasets.

Machine Leanring, high performance comptuing, grap

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning

Analyzing inference workloads for spatiotemporal modeling

Ensuring power grid resiliency, forecasting climate conditions, and optimization of transportation infrastructure are some of the many application areas where data is collected in both space and time. Spatiotemporal modeling is about modeling those patterns for forecasting future trends and carrying out critical decision-making by leveraging machine learning/deep learning. Once trained offline, field deployment of trained models for near real-time inference could be challenging because performance can vary significantly depending on the environment, available compute resources and tolerance to ambiguity in results. Users deploying spatiotemporal models for solving complex problems can benefit from analytical studies considering a plethora of system adaptations to understand the associated performance-quality trade-offs. To facilitate the co-design of next-generation hardware architectures for field deployment of trained models, it is critical to characterize the workloads of these deep learning (DL) applications during inference and assess their computational patterns at different levels of the execution stack. In this paper, we develop several variants of deep learning applications that use spatiotemporal data from dynamical systems. We study the associated computational patterns for inference workloads at different levels, considering relevant models (Long short-term Memory, Convolutional Neural Network and Spatio-Temporal Graph Convolution Network), DL frameworks (Tensorflow and PyTorch), precision (FP16, FP32, AMP, INT16 and INT8), inference runtime (ONNX and AI Template), post-training quantization (TensorRT) and platforms (Nvidia DGX A100 and Sambanova SN10 RDU). Overall, our findings indicate that although there is potential in mixed-precision models and post-training quantization for spatiotemporal modeling, extracting efficiency from contemporary GPU systems might be challenging. Instead, co-designing custom accelerators by leveraging optimized High Level Synthesis frameworks (such as SODA High-Level Synthesizer for customized FPGA/ASIC targets) can make workload-specific adjustments to enhance the efficiency.

97 MATHEMATICS AND COMPUTING

Graph Analytics on Jellyfish topology

Because large unstructured datasets is important for many science domains, distributed graph analytics is critical to many scientists. Unfortunately, obtaining scaling and performance for irregular communication is challenging because contemporary network interconnects are primarily designed to maximize bandwidths of fixed-neighborhoods large-message exchanges (e.g., stencils). Although there is no consensus on the “best” network topologies for irregular communication, unstructured graph-based interconnects can be more suitable. We analyze three popular graph workloads – clustering, pattern enumeration, and traversal — on comparable networks (in terms of resources and costs) constructed from Jellyfish Random Regular, Dragonfly and Fat tree topologies, varying the routing algorithms. Using packet-level simulations, we demonstrate up to 60% improvement in communication time with Jellyfish due to diversity of the short paths between arbitrary endpoints, which can reduce overall network stalls and congestion.

Graph Analytics, network topology, interconnect, H

Origin of spin-driven ferroelectricity and effect of external pressure on the complex magnetism of the 6⁢ H perovskite Ba 3 ⁢Ho⁢Ru 2⁢ O 9

The compound Ba 3 HoRu 2 O 9 magnetically orders at 50 K (T N1 ), followed by another complex magnetic ordering at 10.2 K (T N2 ). The second magnetic phase transition was characterized by the coexistence of two competing magnetic ground states associated with two different magnetic wave vectors (K 1 = 0.5 0 0 and K 2 = 0.25 0.25 0). The multiferroicity and magnetoelectric coupling were predicted below T N2 in these 4d-based materials. Here, in this work, we have discussed the origin of spin-driven ferroelectricity, which is not known yet, and the nature of magnetoelectric domains. We have investigated the compound through time-of-flight neutron diffraction, synchrotron x-ray diffraction (XRD), ac susceptibility, frequency-dependent complex dielectric spectroscopy, and dc magnetization under external pressure. We have demonstrated that the noncollinear structure involving two different magnetic ions, Ru (4d) and Ho (4f), breaks the spatial inversion symmetry via inverse Dzyaloshinskii-Moriya (DM) interaction through strong 4d-4f magnetic correlation, which shifts the oxygen atoms and results in nonzero polarization. Such an observation of inverse DM interaction from two different magnetic ions which cause ferroelectricity is rarely observed. The stronger spin-orbit coupling of 4d orbital might play a major role in creating DM interaction of noncollinear spins. We have systematically studied the spin and dipolar dynamics, which exhibit intriguing behavior with shorter coherence lengths of second magnetic phase associated with the k 2 wave vector. The results manifest the development of finite-size magnetoelectric domains instead of true long-range ordering, which justifies the experimentally obtained low value of ferroelectric polarization. The lattice parameters and volume show a sharp anomaly at T N2 obtained by analyzing the temperature-dependence XRD, which is consistent with the ferroelectric transition, predicting a noncentrosymmetric space group, $P\bar{6}2c$, for this compound. Furthermore, we have investigated the effect of external pressure on this complex magnetism. The result reveals an enhancement of ordering temperature by the application of external pressure (~1.6 K/GPa). The external pressure might favor stabilizing the magnetic ground state associated with second magnetic phase. Our study shows an unconventional mechanism of spin-driven ferroelectricity involving inverse DM interaction between Ru (4d) and Ho (4f) magnetic ions due to strong 4d-4f cross coupling.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND