Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph partitioning”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Modern chemical graph theory

Abstract Graph theory has a long history in chemistry. Yet as the breadth and variety of chemical data is rapidly changing, so too do graph encoding methods and analyses that yield qualitative and quantitative insights. Using illustrative cases within a basic mathematical framework, we showcase modern chemical graph theory's utility in Chemists' analysis and model development toolkit. The encoding of both experimental and simulation data is discussed at various levels of granularity of information. This is followed by a discussion of the two major classes of graph theoretical analyses: identifying connectivity patterns and partitioning methods. Measures, metrics, descriptors, and topological indices are then introduced with an emphasis upon enhancing interpretability and incorporation into physical models. Challenging data cases are described that include strategies for studying time dependence. Throughout, we incorporate recent advancements in computer science and applied mathematics that are propelling chemical graph theory into new domains of chemical study. This article is categorized under: Molecular and Statistical Mechanics > Molecular Dynamics and Monte‐Carlo Methods Structure and Mechanism > Computational Materials Science Structure and Mechanism > Molecular Structures

Leite, Leonardo S. G.↗

Mechanism of Rad26-assisted rescue of stalled RNA polymerase II in transcription-coupled repair

Transcription-coupled repair is essential for the removal of DNA lesions from the transcribed genome. The pathway is initiated by CSB protein binding to stalled RNA polymerase II. Mutations impairing CSB function cause severe genetic disease. Yet, the ATP-dependent mechanism by which CSB powers RNA polymerase to bypass certain lesions while triggering excision of others is incompletely understood. Here we build structural models of RNA polymerase II bound to the yeast CSB ortholog Rad26 in nucleotide-free and bound states. This enables simulations and graph-theoretical analyses to define partitioning of this complex into dynamic communities and delineate how its structural elements function together to remodel DNA. We identify an allosteric pathway coupling motions of the Rad26 ATPase modules to changes in RNA polymerase and DNA to unveil a structural mechanism for CSB-assisted progression past less bulky lesions. Our models allow functional interpretation of the effects of Cockayne syndrome disease mutations.

59 BASIC BIOLOGICAL SCIENCES↗

Spectral Clustering-Based Partitioning of Large-Scale Power Electronics-Based Power Systems for Small-Signal Stability Analysis

The nodal admittance matrix (NAM)-based approach is well-suited for small-signal stability analysis of large-scale power electronics-based power systems (PEPSs), as it preserves the system structure through its admittance matrix. Previous studies have explored partitioning such systems into subareas and interconnections to reduce computational burden; however, they lacked a formal algorithmic procedure for determining feasible partitions. While several grid partitioning methods, such as those based on graph theory or machine learning, exist in the literature, they cannot be directly applied to NAM-based analysis due to differing objectives and constraints. Here, this paper addresses this gap by presenting a systematic, step-by-step procedure for applying a spectral partitioning algorithm that yields a division of the system into subareas suitable for NAM-based analysis. The computational complexity of the proposed method is also derived to demonstrate its efficiency and justify the practicality of the resulting subarea decomposition. The performance of the partitioning method is evaluated by applying the spectral clustering-derived subareas and interconnections to the NAM-based partitioning approach on a 140-bus system. Computational times for the full-system and partitioned NAM analyses are compared using MATLAB. Additionally, PSCAD simulations of the complete system and partitioned subareas are carried out to verify the effectiveness of the proposed method.

Nupur [Univ. of Tennessee, Knoxville, TN (United S↗

Stream Temperature Prediction in a Shifting Environment: Explaining the Influence of Deep Learning Architecture

Stream temperature is a fundamental control on ecosystem health. Recent efforts incorporating process guidance into deep learning models for predicting stream temperature have been shown to outperform existing statistical and physical models. This performance is in part because deep learning architectures can actively learn spatiotemporal relationships that govern how water and energy propagate through a river network. However, exploration of how spatiotemporal awareness and process guidance influence a model's generalizability under shifting environmental conditions such as climate change is limited. Here, we use Explainable Artificial Intelligence (XAI) to interrogate how differing deep learning architectures affect a model's learned spatial and temporal dependencies, and how those learned dependencies affect a model's ability to maintain high accuracy when applied to unseen environmental conditions. Using the Delaware River Basin in the northeastern United States as a test case, we compare two spatiotemporally aware process–guided deep learning models for predicting stream temperature (a recurrent graph convolution network—RGCN, and a temporal convolution graph model—Graph WaveNet). Both models achieve equally high predictive performance when testing data are well represented in the training data (test root mean squared errors of 1.64°C and 1.65°C); however, Graph WaveNet significantly outperforms RGCN in 4 out of 5 experiments where test partitions represent different types of unseen environmental conditions. XAI results show that the architecture of Graph WaveNet leads to learned spatial relationships with greater fidelity to physical processes, and that this fidelity improves the generalizability of the model when applied to shifting and/or unseen environmental conditions.

54 ENVIRONMENTAL SCIENCES↗

Wave function analysis with a maximum flow algorithm

An efficient algorithm for computing the maximum-flow path in a network is applied to the identification of the dominant configuration state functions (CSFs) in a graphically contracted function (GCF), configuration interaction, wave function. The flow network is a space of spin-adapted CSFs represented by a Shavitt graph, wherein the nodes correspond to orbital occupations and spin quantum numbers. The graph nodes are connected by arcs, and an arc density is defined as sums of the associated squared CSF coefficients. A max-min approach determines an upper bound to the maximum possible incoming flow for each graph node. A backtracking step generates a candidate walk and is followed by a limited search of alternative branching paths for the dominant CSF. The arc density contributions are removed from the graph, and the algorithm is reapplied to the updated graph. This list of generated walks can be partitioned in order to guarantee that the dominant CSFs have been identified. All of the steps in this algorithm are computationally efficient and do not depend on the potentially large dimension of the underlying linear CSF expansion space. An analysis of low-lying valence states of C-2 illustrates the method.

74 ATOMIC AND MOLECULAR PHYSICS↗

Deep learning for characterizing the self-assembly of three-dimensional colloidal systems

Creating a systematic framework to characterize the structural states of colloidal self-assembly systems is crucial for unraveling the fundamental understanding of these systems’ stochastic and non-linear behavior. The most accurate characterization methods create high-dimensional neighborhood graphs that may not provide useful information about structures unless these are well-defined reference crystalline structures. Dimensionality reduction methods are thus required to translate the neighborhood graphs into a low-dimensional space that can be easily interpreted and used to characterize non-reference structures. We investigate a framework for colloidal system state characterization that employs deep learning methods to reduce the dimensionality of neighborhood graphs. The framework next uses agglomerative hierarchical clustering techniques to partition the low-dimensional space and assign physically meaningful classifications to the resulting partitions. In this work, we first demonstrate the proposed colloidal self-assembly state characterization framework on a three-dimensional in silico system of 500 multi-flavored colloids that self-assemble under isothermal conditions. We next investigate the generalizability of the characterization framework by applying the framework to several independent self-assembly trajectories, including a three-dimensional in silico system of 2052 colloidal particles that undergo evaporation-induced self-assembly.

36 MATERIALS SCIENCE↗

CORE-BFS: Communication-Optimized REctangular-partitioned BFS Achieving 160.845 TeraTEPS on Frontier Supercomputer

Distributed Breadth-First Search (BFS) is fundamental to many large-scale graph applications, but its performance on parallel systems is often limited by high communication overhead. This paper presents CORE-BFS, an extremely scalable GPU-based BFS implementation that introduces a unique rectangular 2D partitioning-based design for Frontier supercomputer. To further improve performance, we propose four key optimizations: (1) Rectangular 2D-partition specific data formats that use two compressed row and one compressed column status array bitmaps combined with a Double Compressed Sparse Row (DCSR) format per partition, reducing memory footprint and inter-rank traffic; (2) Adaptive frontier & communication strategy that unifies top-down and bottom-up traversal on the rectangular layout, uses lazy synchronization in top-down levels, and switches variants based on frontier size to minimize communication overhead; (3) Frontier-split degree-aware update that maps frontier vertices to thread-centric, wavefront-centric, and block-centric kernels based on their degree to improve GPU utilization and memory coalescing; (4) Row-reduction pipeline that overlaps bottom-up adjacency list processing with row-wise bitmap reduction to hide inter-rank latency. Together, these techniques increase parallelism while reducing memory and communication overhead. On the Graph500 benchmark, CORE - BFS scales up to 9,248 Frontier nodes with scale-42 graphs and reaches 160.845 TTEPS, delivering a 5.42 × speedup over our previous Frontier implementation.

Yang, Haoshen [Rutgers University]↗

Scaling neural simulations in STACS

Abstract As modern neuroscience tools acquire more details about the brain, the need to move towards biological-scale neural simulations continues to grow. However, effective simulations at scale remain a challenge. Beyond just the tooling required to enable parallel execution, there is also the unique structure of the synaptic interconnectivity, which is globally sparse but has relatively high connection density and non-local interactions per neuron. There are also various practicalities to consider in high performance computing applications, such as the need for serializing neural networks to support potentially long-running simulations that require checkpoint-restart. Although acceleration on neuromorphic hardware is also a possibility, development in this space can be difficult as hardware support tends to vary between platforms and software support for larger scale models also tends to be limited. In this paper, we focus our attention on Simulation Tool for Asynchronous Cortical Streams (STACS), a spiking neural network simulator that leverages the Charm++ parallel programming framework, with the goal of supporting biological-scale simulations as well as interoperability between platforms. Central to these goals is the implementation of scalable data structures suitable for efficiently distributing a network across parallel partitions. Here, we discuss a straightforward extension of a parallel data format with a history of use in graph partitioners, which also serves as a portable intermediate representation for different neuromorphic backends. We perform scaling studies on the Summit supercomputer, examining the capabilities of STACS in terms of network build and storage, partitioning, and execution. We highlight how a suitably partitioned, spatially dependent synaptic structure introduces a communication workload well-suited to the multicast communication supported by Charm++. We evaluate the strong and weak scaling behavior for networks on the order of millions of neurons and billions of synapses, and show that STACS achieves competitive levels of parallel efficiency.

59 BASIC BIOLOGICAL SCIENCES↗

Molecular architecture and functional dynamics of the pre-incision complex in nucleotide excision repair

Nucleotide excision repair (NER) is vital for genome integrity. Yet, our understanding of the complex NER protein machinery remains incomplete. Combining cryo-EM and XL-MS data with AlphaFold2 predictions, we build an integrative model of the NER pre-incision complex(PInC). Here TFIIH serves as a molecular ruler, defining the DNA bubble size and precisely positioning the XPG and XPF nucleases for incision. Using simulations and graph theoretical analyses, we unveil PInC’s assembly, global motions, and partitioning into dynamic communities. Remarkably, XPG caps XPD’s DNA-binding groove and bridges both junctions of the DNA bubble, suggesting a novel coordination mechanism of PInC’s dual incision. XPA rigging interlaces XPF/ERCC1 with RPA, XPD, XPB, and 5' ssDNA, exposing XPA’s crucial role in licensing the XPF/ERCC1 incision. Mapping disease mutations onto our models reveals clustering into distinct mechanistic classes, elucidating xeroderma pigmentosum and Cockayne syndrome disease etiology.

60 APPLIED LIFE SCIENCES↗

Minimizing Optimal Transport for Functions with Fixed-Size Nodal Sets

Consider the class of zero-mean functions with fixed L ∞ and L 1 norms and exactly N ϵ N nodal points. Which functions f minimize W p (f + ,f – ), the Wasserstein distance between the measures whose densities are the positive and negative parts? We provide a complete solution to this minimization problem on the line and the circle, which provides sharp constants for previously proven “uncertainty principle”-type inequalities, i.e., lower bounds on N • W p (f + ,f – ). We further show that, while such inequalities hold in many metric measure spaces, they are no longer sharp when the non-branching assumption is violated; indeed, for metric star-graphs, the optimal lower bound on W p (f + ,f – ) is not inversely proportional to the size of the nodal set, N. Here, based on similar reductions, we make connections between the analogous problem of minimizing W p (f + ,f – ) for f defined on Ω C R d with an equivalent optimal domain partition problem.

97 MATHEMATICS AND COMPUTING↗

Dynamic conformational switching underlies TFIIH function in transcription and DNA repair and impacts genetic diseases

Transcription factor IIH (TFIIH) is a protein assembly essential for transcription initiation and nucleotide excision repair (NER). Yet, understanding of the conformational switching underpinning these diverse TFIIH functions remains fragmentary. TFIIH mechanisms critically depend on two translocase subunits, XPB and XPD. To unravel their functions and regulation, we build cryo-EM based TFIIH models in transcription- and NER-competent states. Using simulations and graph-theoretical analysis methods, we reveal TFIIH’s global motions, define TFIIH partitioning into dynamic communities and show how TFIIH reshapes itself and self-regulates depending on functional context. Our study uncovers an internal regulatory mechanism that switches XPB and XPD activities making them mutually exclusive between NER and transcription initiation. By sequentially coordinating the XPB and XPD DNA-unwinding activities, the switch ensures precise DNA incision in NER. Mapping TFIIH disease mutations onto network models reveals clustering into distinct mechanistic classes, affecting translocase functions, protein interactions and interface dynamics.

59 BASIC BIOLOGICAL SCIENCES↗

Feynman path integrals for discrete-variable systems: Walks on Hamiltonian graphs

We propose a natural, parameter-free, discrete-variable formulation of Feynman path integrals. We show that for discrete-variable quantum systems, Feynman path integrals take the form of walks on the graph whose weighted adjacency matrix is the Hamiltonian. By working out expressions for the partition function and transition amplitudes of discretized versions of continuous-variable quantum systems, and then taking the continuum limit, we explicitly recover Feynman's continuous-variable path integrals. We also discuss the implications of our result.

Feynman diagrams↗

Multitask graph neural networks for elastoplastic response prediction in dual-phase polycrystals

Microstructure-sensitive prediction of elastoplastic response remains a recurring bottleneck in multiscale damage and fatigue modeling, where large ensembles of statistically distinct polycrystals are required to quantify variability and extreme-value behavior. In this work, we develop a multitask graph neural network (GNN) surrogate that maps dual-phase ferrite–martensite polycrystal microstructures to Statistical Volume Element (SVE)-level elastoplastic Quantities of Interest (QoIs). Each SVE is represented as a grain-adjacency graph, with node features encoding phase, geometry, and crystallographic orientation, and edge features encoding relative misorientation. A message-passing graph convolution generates node embeddings, which are pooled into a graph representation and passed to a multitask regression head that jointly predicts 10 scalar QoIs and vector-valued stress–strain responses in orthogonal loading directions across multiple martensite volume fractions and SVE sizes. Results show high accuracy for scalar QoIs and strong agreement for full stress–strain trajectories, with population envelopes reproducing both median behavior and finite-SVE variability across compositions and partition scales. A unified model trained on pooled volume-fraction data preserves most within-regime accuracy relative to regime-specific models while also capturing the broader cross-regime variation reflected in the pooled test set. Distributional comparisons further demonstrate that the surrogate preserves heterogeneity under SVE partitioning, enabling statistically consistent block-wise random-field construction for mesoscale analyses. Overall, the proposed grain-graph surrogate provides a practical pathway to accelerate ensemble-based studies of SVE-level constitutive variability in dual-phase polycrystals.

Crystal plasticity↗

Intelligent Experiments through Real-Time AI: Fast Data Processing and Autonomous Detector Control for High-Energy Nuclear Experiments

The aim of this project is to develop software and hardware for fast real-time data processing and autonomous detector control and calibration for the sPHENIX and the future EIC experiments. Below summarizes Georgia Tech team efforts in the past year: 1. We developed a real-time clustering algorithm and FPGA-based pipeline architecture for processing fired pixel data from ALPIDE sensors in sPHENIX experiments. Our Columnar Clustering Co-Design introduces a hardware-aware, stream-friendly approach that segments pixel data by column pairs using a Column Pair Clustering (CPC) strategy, followed by Cluster Stitching to merge adjacent subclusters. Implemented in Vitis HLS, the pipeline comprises five stages—read-in, subclustering, stitching, analysis, and write-out—connected by tagged HLS streams with custom end-of-event signaling for robust synchronization. We designed a pipelined dataflow model optimized for throughput, low latency, and minimal buffering, enabling scalable clustering across events of arbitrary size. Our system maintains spatial precision via center-of-mass and shape key extraction and efficiently handles edge cases such as fragmented or nested clusters. Compared against DBSCAN in both software and hardware, our approach demonstrates competitive performance under FPGA constraints. 2. We also conducted a comprehensive algorithm-to-hardware co-design of connected component analysis tailored for sPHENIX experiments, focusing on real-time, low-latency processing using FPGAs and High-Level Synthesis (HLS). Starting from a Python-based particle tracking pipeline, the team translated the core logic—graph traversal via DFS and Union-Find—into an HLS-compatible C++ model, replacing dynamic memory and recursion with static arrays and pipelined control flow. The final design includes a fully streamed and dataflow-compatible Union-Find kernel optimized across five iterations, incorporating loop pipelining, array partitioning, AXI/FIFO interface tuning, and function flattening. Experimental results show up to 14.8× speedup over the CPU baseline, reducing per-graph latency to 1.58 μs and demonstrating strong resource efficiency with only ~7k LUTs and zero BRAM usage. The design maintains functional correctness against the Python reference using a Python-based C-simulation framework and Mean Squared Error metrics. This work validates the potential of HLS-driven FPGA designs for edge-level HEP data acquisition, laying a scalable foundation for future integration with real-time detector pipelines and multi-graph processing systems.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

Triangle Counting with Cyclic Distributions

Triangles are the simplest non-trivial subgraphs and triangle counting is used in a number of different applications. The order in which vertices are processed in triangle counting strongly effects the amount of work that needs to be done (and thus the overall performance). Ordering vertices by degree has been shown to be one particularly effective ordering approach. However, for graphs with skewed degree distributions (such as power-law graphs), ordering by degree effects the distribution of work; parallelization must account for this distribution in order to balance work among workers. In this paper we provide an in- depth analysis of the ramifications of degree-based ordering on parallel triangle counting. We present approach for partitioning work in triangle counting, based on cyclic distribution and some surprisingly simple C++ implementations. Experimental results demonstrate the effectiveness of our approach, particularly for power-law (and social network) graphs.

Graph algorithms, parallel algorithms↗

DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware

Spiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches.

Computer Science↗