Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Graph partitioning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

99 records · Page 6

GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism

Graph neural networks (GNNs), an emerging class of machine learning models for graphs, have gained popularity for their superior performance in various graph analytical tasks. Mini-batch training is commonly used to train GNNs on large graphs, and data parallelism is the standard approach to scale mini-batch training across multiple GPUs. Data parallel approaches contain redundant work as subgraphs sampled by different GPUs contain significant overlap. To address this issue, we introduce a hybrid parallel mini-batch training paradigm called Split parallelism. Split parallelism avoids redundant work by splitting the sampling, loading, and training of each mini-batch across multiple GPUs. Split parallelism, however, introduces communication overheads that can be more than the savings from removing redundant work. We further present a lightweight partitioning algorithm that probabilistically minimizes these overheads. We implement spllit parllelism in GSplit and show that it outperforms state-of-the-art mini-batch training systems like DGL, Quiver, and P3.

Lim, Seung-Hwan [ORNL] (ORCID:0000000194616866)↗

Leveraging Natural Language Processing and Generative Models in Molecular Chemistry: Property Prediction and Novel Compound Generation

The accurate prediction of molecular properties is important for the rational design and the advancement of green chemistry and sustainable materials research. However, the predictive power of traditional computational chemistry methods is limited due to computational restrictions. Here, in this study, we examine an alternative approach to the accurate prediction of properties of organic compounds: natural language processing (NLP)-based molecular embedding. Using viscosity, partition coefficient (log P), and enthalpy of vaporization as test properties through a survey of comprehensive datasets comprising 5695 data points for viscosity, 25 870 data points for log P, and 2296 data points for enthalpy of vaporization. These are important properties for the design of greener, safer, and sustainable chemical processes. Models were trained using NLP methods such as Mol2vec and fine-tuned ChemBERTa, and results were compared with traditional input featurization techniques such as Morgan fingerprints and quantum chemistry derived sigma profiles and DFT features. Among the various machine learning models, Mol2vec demonstrated superior predictive capabilities, achieving the highest correlation coefficient (R 2 = 0.945) and lowest RMSE (0.106 mPa s) for viscosity, as well as high accuracy for log P and enthalpy of vaporization predictions. These findings establish the Mol2vec featurization technique, graph-convolutional neural networks (GCNN), and fine-tuned ChemBERTa model as powerful tools for predictive modeling of organic compounds properties, offering a significant improvement over previously used featurization techniques and opening up strategies for very-high-throughput computational screening. Finally, we integrated ML models with hybrid language-model-based generative adversarial networks (LM-GAN) to generate novel molecular sequences with desirable properties for different research applications. The ability to computationally design solvents with lower viscosity, lower log P, and lower enthalpy of vaporization offers a data-driven route to accelerating the discovery of sustainable alternatives to traditionally toxic solvents.

ChemBERTa↗

Fault-Tolerant Self-Stabilizing Distributed Clock Synchronization Protocol for Arbitrary Digraphs

A self-stabilizing network in the form of an arbitrary, non-partitioned digraph includes K nodes having a synchronizer executing a protocol. K-1 monitors of each node may receive a Sync message transmitted from a directly connected node. When the Sync message is received, the logical clock value for the receiving node is set to between 0 and a communication latency value (gamma) if the clock value is less than a minimum event-response delay (D). A new Sync message is also transmitted to any directly connected nodes if the clock value is greater than or equal to both D and a graph threshold (T(sub S)). When the Sync message is not received the synchronizer increments the clock value if the clock value is less than a resynchronization period (P), and resets the clock value and transmits a new Sync message to all directly connected nodes when the clock value equals or exceeds P.

Malekpour, Mahyar R.↗

'Virtual triple Schmidt' - Wide field two-stage optics

The design concept of an unobscured-wide-field two-stage optical system based on a virtual triple Schmidt (VTS) configuration is presented. It is pointed out that the single large aperture and field-partitioning capability of two-stage systems can lower material and fabrication costs, making the VTS optics suitable for ground-based and space telescopes. The VTS design combines a Schmidt-camera first stage and a second stage comprising two back-to-back Schmidt systems as a 1:1 relay. Aspheric Schmidt correction is achieved at the relayed pupil location for all three systems. The effects of the separation between the error-producing surface and the aperture stop are discussed; the performance of the wavefront-correction system is analyzed; and extensive diagrams, drawings, and graphs of projected performance data are provided.

Manhart, Paul K.↗

Correlations in cosmic density fields

A method is proposed to place constraints on the functional form of the high-order correlation functions zeta(sub n) that arise in cosmic density fields at large scales. This technique is based on a mass-in-cell statistic and a difference of mass in partitions of a cell. The relationship between these measures is sensitive to the formal structure of the zeta(sub n) as well as their amplitudes. This relationship is quantified in several theoretical models of structure, based on the hierarchical clustering paradigm. The results lead to a test for specific types of hierarchical clustering that is sensitive to correlations of all orders. The method is applied to examples of simulated large-scaled structure dominated by cold dark matter. In the preliminary study, the hierarchical paradigm appears to be a realistic approximation over a broad range of the scales. Furthermore, there is evidence that graphs of low-order vertices are dominant. On the basis of simulated data a phenomological model is specified that gives a good representation of clustering from linear scales to the strongly clustered regime (zeta(sub 2) approximately 500).

Bromley, B. C.↗

High-throughput predictions of metal–organic framework electronic properties: theoretical challenges, graph neural networks, and data exploration

Abstract With the goal of accelerating the design and discovery of metal–organic frameworks (MOFs) for electronic, optoelectronic, and energy storage applications, we present a dataset of predicted electronic structure properties for thousands of MOFs carried out using multiple density functional approximations. Compared to more accurate hybrid functionals, we find that the widely used PBE generalized gradient approximation (GGA) functional severely underpredicts MOF band gaps in a largely systematic manner for semi-conductors and insulators without magnetic character. However, an even larger and less predictable disparity in the band gap prediction is present for MOFs with open-shell 3 d transition metal cations. With regards to partial atomic charges, we find that different density functional approximations predict similar charges overall, although hybrid functionals tend to shift electron density away from the metal centers and onto the ligand environments compared to the GGA point of reference. Much more significant differences in partial atomic charges are observed when comparing different charge partitioning schemes. We conclude by using the dataset of computed MOF properties to train machine-learning models that can rapidly predict MOF band gaps for all four density functional approximations considered in this work, paving the way for future high-throughput screening studies. To encourage exploration and reuse of the theoretical calculations presented in this work, the curated data is made publicly available via an interactive and user-friendly web application on the Materials Project.

36 MATERIALS SCIENCE↗

Distributed-Memory Sparse Deep Neural Network Inference Using Global Arrays

Partitioned Global Address Space (PGAS) models exhibit tremendous promise in developing efficient and productive distributed-memory parallel applications. They have been used extensively in scientific computations due to conveniently offering a ``shared-memory''-like model and convenient interfaces that separate communication with synchronization. Traditionally, PGAS communication models have been applied to dense/contiguously distributed data, but most modern applications depict varied levels of sparsity. Existing PGAS models require certain adaptations to support distributed sparse computations, since associated computations often require matrix arithmetic, in addition to data movement. The Global Arrays toolkit from Pacific Northwest National Laboratory (PNNL) is one of the earliest PGAS models to combine one-sided data communication and distributed matrix operations and is still used in the popular NWChem quantum chemistry suite. Recently, we have expanded the Global Arrays toolkit to support common sparse operations, like sparse matrix-dense matrix multiplies (SpMM), sparse matrix-sparse matrix multiplication (SpGEMM) and Sampled Dense-Dense Matrix Multiplication (SDDMM). As it turns out, these operations are the bedrock of sparse Deep Learning (DL); sparse deep neural networks and Graph Neural Networks (GNNs) have gained increasing attention recently in achieving speedups on training and inference with reduced memory footprints. Unlike scientific applications in High Performance Computing (HPC), modern (distributed-memory capable) DL toolkits often rely on non-standardized and closed-source vendor software optimizations, creating challenges in software-hardware co-design at scale. Our goal is to support a variety of distributed-memory sparse matrix operations and helper functions in the newly created Sparse Global Arrays (SGA), such that it is possible to build portable and productive Machine Learning scenarios for algorithm/software and hardware codesign purposes. Contemporary data-parallel schemes for training/inference are undergoing a major overhaul since model replication limits scalability and causes resource inefficiencies. As such, we have adopted tensor parallelism in decomposing the model and inputs, to mitigate memory issues. Current implementation is built on top of MPI and uses CPUs to maximize the portability across the platforms.

Distributed computing, machine learning↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

Oxidation States of Grim Glasses in EET79001 Based on Vanadium Valence

Gas-rich impact-melt (GRIM) glasses in SNC meteorites are very rich in Martian atmospheric noble gases and sulfur suggesting a possible occurrence of regolith-derived secondary mineral assemblages in these samples. Previously, we have studied two GRIM glasses, 506 and 507, from EET79001 Lith A and Lith B, respectively, for elemental abundances and spatial distribution of sulfur using EMPA (WDS) and FE-SEM (EDS) techniques and for sulfur-speciation using K-edge XANES techniques. These elemental and FE-SEM micro-graph data at several locations in the GRIM glasses from Shergotty (DBS), Zagami 994 and EET79001, Lith B showed that FeO and SO3 are positively correlated (SO3 represents a mixture of sulfide and sulfate). FE-SEM (EDS) study revealed that the sulfur-rich pockets in these glasses contain numerous micron-sized iron-sulfide (Fe-S) globules sequestered throughout the volume. However, in some areas (though less frequently), we detected significant Fe-S-O signals suggesting the occurrence of iron sulfate. These GRIM glasses were studied by K-edge microXANES techniques for sulfur speciation in association with iron in sulfur-rich areas. In both samples, we found the sulfur speciation dominated by sulfide with minor oxidized sulfur mixed in with various proportions. The abundance of oxidized sulfur was greater in 506 than in 507. Based on these results, we hypothesize that sulfur initially existed as sulfate in the glass precursor materials and, on shock-impact melting of the precursor materials producing these glasses, the oxidized sulfur was reduced to predominately sulfide. In order to further test this hypothesis, we have used microXANES to measure the valence states of vanadium in GRIM glasses from Lith A and Lith B to complement and compare with previous analogous measurements on Lith C (note: 506 and 507 contain the largest amounts of martian atmospheric gases but the gas-contents in Lith C measured by are unknown). Vanadium is ideal for addressing this re-dox issue because it has multiple valence states and is a well-studied element. Ferrous-dominated iron valences determined by microXANES on the Lith A and Lith B glasses provide little redox sensitivity. Vanadium valence measurements for impact glass in Lith C at three different locations yielded valence values of 3.1, 3.2 and 3.4 with inferred fO2 values of IW-0.7, IW-0.1 and IW+0.7, respectively. This range of oxygen-fugacity values is understandable because the glasses are shock-molten impact glasses which are heterogeneous in nature. Oxygen fugacity values obtained from the analysis of Fe-Ti oxides and Eu partitioning in pyroxenes from EET79001 Lith A and Lith B (host lithologies) were in the range of IW+0.3 to IW+1.9 suggesting that V in the Lith C impact glass was reduced in the impact process. Here, we examine whether the 506 from Lith A and 507 from Lith B GRIM glasses yield similar or different fO2 values from those of Lith C using the vanadium K-edge microXANES technique.

Sutton, S. R.↗