Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Graph algorithms”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

GraMeR: Gra ph Me ta R einforcement learning for multi-objective influence maximization

Influence maximization (IM) is a combinatorial problem of identifying a subset of seed nodes in a network (graph), which when activated, provide a maximal spread of influence in the network for a given diffusion model and a budget for seed set size. IM has numerous applications such as viral marketing, epidemic control, sensor placement and other network-related tasks. However, its practical uses are limited due to the computational complexity of current algorithms. Recently, deep reinforcement learning has been leveraged to solve IM in order to ease the computational burden. However, there are serious limitations in current approaches, including narrow IM formulation that only consider influence via spread and ignore self-activation, low scalability to large graphs, and lack of generalizability across graph families leading to a large running time for every test network. In this work, we address these limitations through a unique approach that involves: (1) Formulating a generic IM problem as a Markov decision process that handles both intrinsic and influence activations; (2)incorporating generalizability via meta-learning across graph families. There are previous works that combine deep reinforcement learning with graph neural network, but this work solves a more realistic IM problem and incorporates generalizability across graphs via meta reinforcement learning. Extensive experiments are carried out in various standard networks to validate performance of the proposed Graph Meta Reinforcement learning (GraMeR) framework. Finally, the results indicate that GraMeR is multiple orders faster and generic than conventional approaches when applied on small to medium scale graphs.

97 MATHEMATICS AND COMPUTING↗

Decomposition Algorithms for Solving NP-hard Problems on a Quantum Annealer

NP-hard problems such as the maximum clique or minimum vertex cover problems, two of Karp’s 21 NP-hard problems, have several applications in computational chemistry, biochemistry and computer network security. Adiabatic quantum annealers can search for the optimum value of such NP-hard optimization problems, given the problem can be embedded on their hardware. However, this is often not possible due to certain limitations of the hardware connectivity structure of the annealer. This paper studies a general framework for a decomposition algorithm for NP-hard graph problems aiming to identify an optimal set of vertices. Our generic algorithm allows us to recursively divide an instance until the generated subproblems can be embedded on the quantum annealer hardware and subsequently solved. Furthermore, the framework is applied to the maximum clique and minimum vertex cover problems, and we propose several pruning and reduction techniques to speed up the recursive decomposition. The performance of both algorithms is assessed in a detailed simulation study.

97 MATHEMATICS AND COMPUTING↗

Heterogeneous Graph Neural Network for identifying hadronically decayed tau leptons at the High Luminosity LHC

Here, we present a new algorithm that identifies reconstructed jets originating from hadronic decays of tau leptons against those from quarks or gluons. No tau lepton reconstruction algorithm is used. Instead, the algorithm represents jets as heterogeneous graphs with tracks and energy clusters as nodes and trains a Graph Neural Network to identify tau jets from other jets. Different attributed graph representations and different GNN architectures are explored. We propose to use differential track and energy cluster information as node features and a heterogeneous sequentially-biased encoding for the inputs to final graph-level classification.

47 OTHER INSTRUMENTATION↗

Integrated Land Suitability Assessment for Depots Siting in a Sustainable Biomass Supply Chain

A sustainable biomass supply chain would require not only an effective and fluid transportation system with a reduced carbon footprint and costs, but also good soil characteristics ensuring durable biomass feedstock presence. Unlike existing approaches that fail to account for ecological factors, this work integrates ecological as well as economic factors for developing sustainable supply chain development. For feedstock to be sustainably supplied, it necessitates adequate environmental conditions, which need to be captured in supply chain analysis. Using geospatial data and heuristics, we present an integrated framework that models biomass production suitability, capturing the economic aspect via transportation network analysis and the environmental aspect via ecological indicators. Production suitability is estimated using scores, considering both ecological factors and road transportation networks. These factors include land cover/crop rotation, slope, soil properties (productivity, soil texture, and erodibility factor) and water availability. This scoring determines the spatial distribution of depots with priority to fields scoring the highest. Two methods for depot selection are presented using graph theory and a clustering algorithm to benefit from contextualized insights from both and potentially gain a more comprehensive understanding of biomass supply chain designs. Graph theory, via the clustering coefficient, helps determine dense areas in the network and indicate the most appropriate location for a depot. Clustering algorithm, via K-means, helps form clusters and determine the depot location at the center of these clusters. An application of this innovative concept is performed on a case study in the US South Atlantic, in the Piedmont region, determining distance traveled and depot locations, with implications on supply chain design. The findings from this study show that a more decentralized depot-based supply chain design with 3depots, obtained using the graph theory method, can be more economical and environmentally friendly compared to a design obtained from the clustering algorithm method with 2 depots. In the former, the distance from fields to depots totals 801,031,476 miles, while in the latter, it adds up to 1,037,606,072 miles, which represents about 30% more distance covered for feedstock transportation.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Examination of Haines Jump in Microfluidic Experiments via Evolution Graphs and Interface Tracking

This work examines a type of rapid pore-filling event in multiphase flow through permeable media that is better known as Haines Jump. While existing microfluidic experiments on Haines Jump mostly seek to maintain quasi-steady states through very low bulk flow rates over long periods of time, this work explores the combined use of a highly structured microscale transport network, high-speed fluorescent microscopy, displacement front segmentation algorithms, and a tracking algorithm to build evolution graphs that track displacement fronts as they evolve through high-speed video recording. The resulting evolution graph allows the segmentation of a high-speed recording in both space and time, potentially facilitating topology-cognitive computation on the transport network. Occurrences of Haines Jump are identified in the microfluidic displacement experiments and their significance in bulk flow rates is qualitatively analyzed. The bulk flow rate has little effect on the significance of Haines Jump during merging and splitting, but large bulk flow rates may obscure small bursts at the narrowest part of the throat.

permeable media↗

ScaWL: Scaling k-WL (Weisfeiler-Lehman) Algorithms in Memory and Performance on Shared and Distributed-Memory Systems

The k-dimensional Weisfeiler-Lehman (k-WL) algorithm—developed as an efficient heuristic for testing if two graphs are isomorphic—is a fundamental kernel for node embedding in the emerging field of graph neural networks. Unfortunately, the k-WL algorithm has exponential storage requirements, limiting the size of graphs that can be handled. This work presents a novel k-WL scheme with a storage requirement orders of magnitude lower while maintaining the same accuracy as the original k-WL algorithm. Due to the reduced storage requirement, our scheme allows for processing much bigger graphs than previously possible on a single compute node. For even bigger graphs, we provide the first distributed-memory implementation. Our k-WL scheme also has significantly reduced communication volume and offers high scalability. Our experimental results demonstrate that our approach is significantly faster and has superior scalability compared to five other implementations employing state-of-the-art techniques.

algorithims↗

Quantum Algorithm for Approximating Maximum Independent Sets

We present a quantum algorithm for approximating maximum independent sets of a graph based on quantum non-Abelian adiabatic mixing in the sub-Hilbert space of degenerate ground states, which generates quantum annealing in a secondary Hamiltonian. For both sparse and dense random graphs G , numerical simulation suggests that our algorithm on average finds an independent set of size close to the maximum size α ( G ) in low polynomial time. The best classical algorithms, by contrast, produce independent sets of size about half of α ( G ) in polynomial time.

Physics↗

New graph-neural-network flavor tagger for Belle II and measurement of sin 2⁢𝜙 1 in 𝐵 0 → 𝐽/𝜓⁢𝐾$^0_ S$ decays

We present GFlaT, a new algorithm that uses a graph-neural-network to determine the flavor of neutral 𝐵 mesons produced in ϒ⁡(4⁢𝑆) decays. It improves previous algorithms by using the information from all charged final-state particles and the relations between them. We evaluate its performance using 𝐵 decays to flavor-specific hadronic final states reconstructed in a 362 fb −1 sample of electron-positron collisions collected at the ϒ⁡(4⁢𝑆) resonance with the Belle II detector at the SuperKEKB collider. We achieve an effective tagging efficiency of (37.40 ± 0.43 ± 0.36%), where the first uncertainty is statistical and the second systematic, which is 18% better than the previous Belle II algorithm. Demonstrating the algorithm, we use 𝐵 0 →𝐽/𝜓⁢𝐾$^0_ S$ decays to measure the mixing-induced and direct 𝐶⁢𝑃 violation parameters, 𝑆 = (0.724 ± 0.035 ± 0.009) and 𝐶 = (−0.035 ± 0.026 ± 0.029).

CP violation↗

Accelerated Charged Particle Tracking with Graph Neural Networks on FPGAs

We develop and study FPGA implementations of algorithms for charged particle tracking based on graph neural networks. The two complementary FPGA designs are based on OpenCL, a framework for writing programs that execute across heterogeneous platforms, and hls4ml, a high-level-synthesis-based compiler for neural network to firmware conversion. We evaluate and compare the resource usage, latency, and tracking performance of our implementations based on a benchmark dataset. We find a considerable speedup over CPU-based execution is possible, potentially enabling such algorithms to be used effectively in future computing workflows and the FPGA-based Level-1 trigger at the CERN Large Hadron Collider.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Graph Neural Networks and Applied Linear Algebra v.1.0

SAND2024-01365O The Graph Neural Networks and Applied Linear Algebra is companion software for the educational article with the same title. The software provides illustrative examples of graph neural networks in Matlab and Python. These stand-alone algorithms are for educational purposes. The software also includes graph neural network-based algorithms for a trainable Jacobi iteration as well as diffusion coefficient estimation. The software provides human-interpretable implementations of Graph Neural Networks in Matlab and Python. These implementations are not optimized for performance and instead emphasize readability. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

0BGRaman: Graph Network based Simulator for Forecasting Molecular Polarizability

This report presents the work performed under the GRaman project, sponsored by the PCSD LDRD Seed program. The project aimed at accelerating ab initio molecular dynamics simulation using Graph Networks. The Graph Network framework is a ML framework that has been successfully employed to simulate the dynamics of several physical systems: including water splashing in a container and flags moving with the wind. In this effort, we performed a data collection campaign for 3 different molecules of interest. We have built tools for preprocessing the trajectories obtained by simulating Raman Spectroscopy with NWChem and translating them into a suitable format for training. We have developed a training algorithm to train the Graph Network based simulators based on our data and developed a simulator that produces trajectories in the same NWChem format. While the tool has improved with each iteration of development and subsequent experiments, the current state of the tool does not allow to directly incorporate the technology within the NWChem framework because the trajectories produced by the tool are not yet accurate enough. However, the technology has proved to have good potential and it is certainly worth further research and development.

97 MATHEMATICS AND COMPUTING↗

Photon Reconstruction in the Belle II Calorimeter Using Graph Neural Networks

We present the study of a fuzzy clustering algorithm for the Belle II electromagnetic calorimeter using Graph Neural Networks. We use a realistic detector simulation including simulated beam backgrounds and focus on the reconstruction of both isolated and overlapping photons. We find significant improvements of the energy resolution compared to the currently used reconstruction algorithm for both isolated and overlapping photons of more than 30% for photons with energies E γ < 0.5 GeV and high levels of beam backgrounds. Overall, the GNN reconstruction improves the resolution and reduces the tails of the reconstructed energy distribution and therefore is a promising option for the upcoming high luminosity running of Belle II.

Calorimeter↗

Theoretically and practically efficient parallel nucleus decomposition

This paper studies the nucleus decomposition problem, which has been shown to be useful in finding dense substructures in graphs. We present a novel parallel algorithm that is efficient both in theory and in practice. Our algorithm achieves a work complexity matching the best sequential algorithm while also having low depth (parallel running time), which significantly improves upon the only existing parallel nucleus decomposition algorithm (Sariyüce et al. , PVLDB 2018). The key to the theoretical efficiency of our algorithm is a new lemma that bounds the amount of work done when peeling cliques from the graph, combined with the use of a theoretically-efficient parallel algorithms for clique listing and bucketing. We introduce several new practical optimizations, including a new multi-level hash table structure to store information on cliques space-efficiently and a technique for traversing this structure cache-efficiently. On a 30-core machine with two-way hyper-threading on real-world graphs, we achieve up to a 55x speedup over the state-of-the-art parallel nucleus decomposition algorithm by Sariyüce et al. , and up to a 40x self-relative parallel speedup. We are able to efficiently compute larger nucleus decompositions than prior work on several million-scale graphs for the first time.

Computer Science↗

Picasso: Memory-Efficient Graph Coloring Using Palettes With Applications in Quantum Computing

A coloring of a graph is an assignment of colors to vertices such that no two neighboring vertices have the same color. The need for memory-efficient coloring algorithms is motivated by their application in computing clique partitions of graphs arising in quantum computations where the objective is to map a large set of Pauli strings into a compact set of unitaries. We present Picasso, a randomized memory-efficient iterative parallel graph coloring algorithm with theoretical sublinear space guarantees under practical assumptions. The parameters of our algorithm provide a trade-off between coloring quality and resource consumption. To assist the user, we also propose a machine learning model to predict the coloring algorithm’s parameters considering these trade-offs. We provide a sequential and a parallel implementation of the proposed algorithm. We perform an experimental evaluation on a 64-core AMD CPU equipped with 512 GB of memory and an Nvidia A100 GPU with 40GB of memory. For a small dataset where existing coloring algorithms can be executed within the 512 GB memory budget, we show up to 68× memory savings. On massive datasets we demonstrate that GPU-accelerated Picasso can process inputs with 49.5× more Pauli strings (vertex set in our graph) and 2,478× more edges than state-of-the-art parallel approaches.

artificial intelligence, quantum computing↗

Scalable All-pairs Shortest Paths for Huge Graphs on Multi-GPU Clusters

We present an optimized Floyd-Warshall (Floyd-Warshall) algorithm that computes the All-pairs shortest path (APSP) for GPU accelerated clusters. The Floyd-Warshall algorithm due to its structural similarities to matrix-multiplication is well suited for highly parallel GPU architectures. To achieve high parallel efficiency, we address two key algorithmic challenges: reducing high communication overhead and addressing limited GPU memory. To reduce high communication costs, we redesign the parallel (a) to expose more parallelism, (b) aggressively overlap communication and computation with pipelined and asynchronous scheduling of operations, and (c) tailored MPI-collective. To cope with limited GPU memory, we employ an offload model, where the data resides on the host and is transferred to GPU on-demand. The proposed optimizations are supported with detailed performance models for tuning. Our optimized parallel Floyd-Warshall implementation is up to 5x faster than a strong baseline and achieves 8.1 PetaFLOPS/sec on 256~nodes of the Summit supercomputer at Oak Ridge National Laboratory. This performance represents 70% of the theoretical peak and 80% parallel efficiency. The offload algorithm can handle 2.5x larger graphs with a 20% increase in overall running time.

Sao, Piyush↗

Grid Topology Discovery Algorithm Evaluation of Suitability for Utility Deployment (CRADA 606 Final Report)

This work presents the results of a field-informed demonstration aimed at evaluating the practical suitability of a topology discovery algorithm for utility environments. We demonstrated an algorithm that uses a graph-theory-informed state estimation approach for model selection. In collaboration with Survalent and Peninsula Light Co., the algorithm was applied to real feeder models and field measurements from supervisory control and data acquisition (SCADA) and advanced metering infrastructure (AMI) systems to identify the operational topology of a power distribution system. The demonstration assessed the algorithm’s performance under realistic data conditions, including sparse and noisy measurements, and examined its ability to identify the most likely network configurations. The results confirmed that the approach can effectively narrow down feasible topologies, providing operators with improved situational awareness of network status. Key lessons learned emphasize the need for systematic data validation and strategic sensor placement to enhance observability. These insights inform future deployment strategies and guide refinements for broader adoption in utility operations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

Graphical Gaussian Process Regression Model for Aqueous Solvation Free Energy Prediction of Organic Molecules in Redox Flow Battery

The solvation free energy of organic molecules is a critical parameter in determining emergent properties such as solubility, liquid-phase equilibrium constants, and pKa and redox potentials in an organic redox flow battery. In this work, we present a machine learning (ML) model that can learn and predict the aqueous solvation free energy of an organic molecule using Gaussian process regression method based on a new molecular graph kernel. To investigate the performance of the ML model on electrostatic interaction, the nonpolar interaction contribution of solvent and the conformational entropy of solute in solvation free energy, three data sets with implicit or explicit water solvent models, and contribution of conformational entropy of solute are tested. We demonstrate that our ML model can predict the solvation free energy of molecules at chemical accuracy with a mean absolute error of less than 1 kcal/mol for subsets of the QM9 dataset and the Freesolv database. To solve the general data scarcity problem for a graph-based ML model, we propose a dimension reduction algorithm based on the distance between molecular graphs, which can be used to examine the diversity of the molecular data set. It provides a promising way to build a minimum training set to improve prediction for certain test sets where the space of molecular structures is predetermined.

25 ENERGY STORAGE↗