Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph partitioning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

129 records · Page 8

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Oxidation States of Grim Glasses in EET79001 Based on Vanadium Valence

Gas-rich impact-melt (GRIM) glasses in SNC meteorites are very rich in Martian atmospheric noble gases and sulfur suggesting a possible occurrence of regolith-derived secondary mineral assemblages in these samples. Previously, we have studied two GRIM glasses, 506 and 507, from EET79001 Lith A and Lith B, respectively, for elemental abundances and spatial distribution of sulfur using EMPA (WDS) and FE-SEM (EDS) techniques and for sulfur-speciation using K-edge XANES techniques. These elemental and FE-SEM micro-graph data at several locations in the GRIM glasses from Shergotty (DBS), Zagami 994 and EET79001, Lith B showed that FeO and SO3 are positively correlated (SO3 represents a mixture of sulfide and sulfate). FE-SEM (EDS) study revealed that the sulfur-rich pockets in these glasses contain numerous micron-sized iron-sulfide (Fe-S) globules sequestered throughout the volume. However, in some areas (though less frequently), we detected significant Fe-S-O signals suggesting the occurrence of iron sulfate. These GRIM glasses were studied by K-edge microXANES techniques for sulfur speciation in association with iron in sulfur-rich areas. In both samples, we found the sulfur speciation dominated by sulfide with minor oxidized sulfur mixed in with various proportions. The abundance of oxidized sulfur was greater in 506 than in 507. Based on these results, we hypothesize that sulfur initially existed as sulfate in the glass precursor materials and, on shock-impact melting of the precursor materials producing these glasses, the oxidized sulfur was reduced to predominately sulfide. In order to further test this hypothesis, we have used microXANES to measure the valence states of vanadium in GRIM glasses from Lith A and Lith B to complement and compare with previous analogous measurements on Lith C (note: 506 and 507 contain the largest amounts of martian atmospheric gases but the gas-contents in Lith C measured by are unknown). Vanadium is ideal for addressing this re-dox issue because it has multiple valence states and is a well-studied element. Ferrous-dominated iron valences determined by microXANES on the Lith A and Lith B glasses provide little redox sensitivity. Vanadium valence measurements for impact glass in Lith C at three different locations yielded valence values of 3.1, 3.2 and 3.4 with inferred fO2 values of IW-0.7, IW-0.1 and IW+0.7, respectively. This range of oxygen-fugacity values is understandable because the glasses are shock-molten impact glasses which are heterogeneous in nature. Oxygen fugacity values obtained from the analysis of Fe-Ti oxides and Eu partitioning in pyroxenes from EET79001 Lith A and Lith B (host lithologies) were in the range of IW+0.3 to IW+1.9 suggesting that V in the Lith C impact glass was reduced in the impact process. Here, we examine whether the 506 from Lith A and 507 from Lith B GRIM glasses yield similar or different fO2 values from those of Lith C using the vanadium K-edge microXANES technique.

Sutton, S. R.↗