Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “difference graphs”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Network-Level Traffic Signal Cooperation: A Higher-Order Conflict Graph Approach

Traffic signal control and cooperation are extremely important to alleviate traffic congestion in a large traffic network. This study develops a higher-order conflict graph approach for network-wide traffic signal control and cooperation. A conflict graph is applied to model the traffic signal configurations, which identifies the conflict and unconflicted movements for each intersection. In conflict graph, the node represents each movement. The weight of each node can be defined as traffic volume, queue length, fuel consumption, or any weighted combinations of these measurements. The calculation of the optimal green light duration and green light sequence (for different movements) is equivalent to sequentially finding the maximum weight independent set (MWIS) in the conflict graph. The conflict graph also provides a uniform and efficient way to connect traffic signal operations among nearby intersections spatially. Then, we introduced the concept of the k -th order neighborhood to model the degree of connectivity between each movement to the movements at upstream or downstream intersections. The weight of each node in the higher-order conflict graph not only represents its own congestion level, but also relates to the traffic conditions of nearby intersections. Through this approach, the cooperation of multiple intersections can be realized by incorporating their spatial connectivity into conflict graph and solving the MWIS problem. A simulation network is built in SUMO to test the effectiveness of the proposed method. Results suggested that the proposed model outperformed other state-of-the-art signal control methods. Also, the scheme maintains good performance under varying traffic demands.

42 ENGINEERING↗

Graph Neural Networks for Particle Reconstruction in High Energy Physics detectors

Pattern recognition problems in high energy physics are notably different from traditional machine learning applications in computer vision. Reconstruction algorithms identify and measure the kinematic properties of particles produced in high energy collisions and recorded with complex detector systems. Two critical applications are the reconstruction of charged particle trajectories in tracking detectors and the reconstruction of particle showers in calorimeters. These two problems have unique challenges and characteristics, but both have high dimensionality, high degree of sparsity, and complex geometric layouts. Graph Neural Networks (GNNs) are a relatively new class of deep learning architectures which can deal with such data effectively, allowing scientists to incorporate domain knowledge in a graph structure and learn powerful representations leveraging that structure to identify patterns of interest. In this work we demonstrate the applicability of GNNs to these two diverse particle reconstruction problems.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

User Manual - HydraGNN: Distributed PyTorch Implementation of Multi-Headed Graph Convolutional Neural Networks

This document serves as user manual for HydraGNN, a scalable graph neural network (GNN) architecture that allows for a simultaneous prediction of multiple target properties using multi-task learning (MTL). The HydraGNN architecture is constructed by successive superposition of three different sets of layers. The first set is made of message-passing layers to exchange information across nodes in the graph and use this to update the nodal features. The second set is made of global pooling layers that aggregate information from all the nodes in the graph and map it into a scalar, and is needed only for global target properties that are related to the entire graph. The third set of layers is dedicated to the implementation of MTL, which is enabled by forking of the architecture into separate heads, each one of them dedicated to the predictive task of one specific target property. Through an object-oriented programming paradigm, HydraGNN is templated over different message-passing policies, which allows for a user-friendly hyperparameter study to assess the sensitivity of the predictive performance of the HydraGNN architecture on a specific dataset with respect to the choice of the message-passing policy. The object-oriented paradigm used by HydraGNN also allows for a user-friendly inclusion of newly developed message passing policies within the existing framework. HydraGNN supports distributed computing capabilities for scalable data reading and scalable training on leadership-class supercomputers.

97 MATHEMATICS AND COMPUTING↗

An ontology-based knowledge graph for representing interactions involving RNA molecules

The "RNA world" represents a novel frontier for the study of fundamental biological processes and human diseases and is paving the way for the development of new drugs tailored to each patient's biomolecular characteristics. Although scientific data about coding and non-coding RNA molecules are constantly produced and available from public repositories, they are scattered across different databases and a centralized, uniform, and semantically consistent representation of the "RNA world" is still lacking. We propose RNA-KG, a knowledge graph (KG) encompassing biological knowledge about RNAs gathered from more than 60 public databases, integrating functional relationships with genes, proteins, and chemicals and ontologically grounded biomedical concepts. To develop RNA-KG, we first identified, pre-processed, and characterized each data source; next, we built a meta-graph that provides an ontological description of the KG by representing all the bio-molecular entities and medical concepts of interest in this domain, as well as the types of interactions connecting them. Finally, we leveraged an instance-based semantically abstracted knowledge model to specify the ontological alignment according to which RNA-KG was generated. RNA-KG can be downloaded in different formats and also queried by a SPARQL endpoint. A thorough topological analysis of the resulting heterogeneous graph provides further insights into the characteristics of the "RNA world". RNA-KG can be both directly explored and visualized, and/or analyzed by applying computational methods to infer bio-medical knowledge from its heterogeneous nodes and edges. The resource can be easily updated with new experimental data, and specific views of the overall KG can be extracted according to the bio-medical problem to be studied.

59 BASIC BIOLOGICAL SCIENCES↗

Multiscale characterization and representation of variability in ceramic matrix composites

Low density, high strength, and high creep and oxidation resistance properties of ceramic matrix composites (CMCs) make them an ideal choice for use in extreme environments in space and military applications. This paper presents a detailed characterization study of structural and manufacturing flaws in Carbon fiber Silicon-Carbide-Nitride matrix (C/SiNC) CMCs at different length-scales. Energy-dispersive spectroscopy (EDS) is used for the chemical characterization of the material’s elemental constituents. High-resolution multiscale graphs obtained from scanning electron microscope (SEM) and confocal laser scanning microscope (LSM) are used to characterize the distribution and morphology of defects at different length scales. This is followed by the classification and quantification of the common manufacturing defects. An image processing algorithm based on the image segmentation process is developed to quantify the variability of various scale-dependent architectural parameters. Finally, a three-dimensional stochastic representative volume element (SRVE) generation algorithm is developed to provide precise representations of material textures at multiple length scales. The developed algorithm accurately accounts for material features and flaws based on a range of multiscale structural and defects characterization results.

36 MATERIALS SCIENCE↗

ConnectIt: a framework for static and incremental parallel graph connectivity algorithms

Connected components is a fundamental kernel in graph applications. The fastest existing multicore algorithms for solving graph connectivity are based on some form of edge sampling and/or linking and compressing trees. However, many combinations of these design choices have been left unexplored. In this paper, we design the ConnectIt framework, which provides different sampling strategies as well as various tree linking and compression schemes. ConnectIt enables us to obtain several hundred new variants of connectivity algorithms, most of which extend to computing spanning forest. In addition to static graphs, we also extend ConnectIt to support mixes of insertions and connectivity queries in the concurrent setting. We present an experimental evaluation of ConnectIt on a 72-core machine, which we believe is the most comprehensive evaluation of parallel connectivity algorithms to date. Compared to a collection of state-of-the-art static multicore algorithms, we obtain an average speedup of 12.4x (2.36x average speedup over the fastest existing implementation for each graph). Using ConnectIt, we are able to compute connectivity on the largest publicly-available graph (with over 3.5 billion vertices and 128 billion edges) in under 10 seconds using a 72-core machine, providing a 3.1x speedup over the fastest existing connectivity result for this graph, in any computational setting. For our incremental algorithms, we show that our algorithms can ingest graph updates at up to several billion edges per second. To guide the user in selecting the best variants in ConnectIt for different situations, we provide a detailed analysis of the different strategies. Finally, we show how the techniques in ConnectIt can be used to speed up two important graph applications: approximate minimum spanning forest and SCAN clustering.

Computer Science↗

Leveraging graph clustering techniques for cyber‐physical system analysis to enhance disturbance characterisation

Abstract Cyber‐physical systems have behaviour that crosses domain boundaries during events such as planned operational changes and malicious disturbances. Traditionally, the cyber and physical systems are monitored separately and use very different toolsets and analysis paradigms. The security and privacy of these cyber‐physical systems requires improved understanding of the combined cyber‐physical system behaviour and methods for holistic analysis. Therefore, the authors propose leveraging clustering techniques on cyber‐physical data from smart grid systems to analyse differences and similarities in behaviour during cyber‐, physical‐, and cyber‐physical disturbances. Since clustering methods are commonly used in data science to examine statistical similarities in order to sort large datasets, these algorithms can assist in identifying useful relationships in cyber‐physical systems. Through this analysis, deeper insights can be shared with decision‐makers on what cyber and physical components are strongly or weakly linked, what cyber‐physical pathways are most traversed, and the criticality of certain cyber‐physical nodes or edges. This paper presents several types of clustering methods for cyber‐physical graphs of smart grid systems and their application in assessing different types of disturbances for informing cyber‐physical situational awareness. The collection of these clustering techniques provide a foundational basis for cyber‐physical graph interdependency analysis.

97 MATHEMATICS AND COMPUTING↗

Generalist multimodal AI: A review of architectures, challenges and opportunities

Multimodal models are expected to be a critical component to future advances in artificial intelligence. Here, this field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural language processing (NLP) and vision. It is widely hoped that further extending the foundation models to multiple modalities (e.g., text, image, video, sensor, time series, graph, etc.) will ultimately lead to generalist multimodal models, i.e. one model across different data modalities and tasks. However, there is little research that systematically analyzes recent multimodal models (particularly the ones that work beyond text and vision) with respect to the underling architecture proposed. Therefore, this work provides a fresh perspective on generalist multimodal models (GMMs) via a novel architecture and training configuration specific taxonomy. This includes factors such as Unifiability, Modularity, and Adaptability that are pertinent and essential to the wide adoption and application of GMMs. The review further highlights key challenges and prospects for the field and guide the researchers into the new advancements.

Artificial intelligence (AI)↗

Quantile-dependent expressivity of postprandial lipemia

“Quantile-dependent expressivity” describes an effect of the genotype that depends upon the level of the phenotype (e.g., whether a subject’s triglycerides are high or low relative to its population distribution). Prior analyses suggest that the effect of a genetic risk score (GRS) on fasting plasma triglyceride levels increases with the percentile of the triglyceride distribution. Postprandial lipemia is well suited for testing quantile-dependent expressivity because it exposes each individual’s genotype to substantial increases in their plasma triglyceride concentrations. Ninety-seven published papers were identified that plotted mean triglyceride response vs. time and genotype, which were converted into quantitative data. Separately, for each published graph, standard least-squares regression analysis was used to compare the genotype differences at time t (dependent variable) to average triglyceride concentrations at time t (independent variable) to assess whether the genetic effect size increased in association with higher triglyceride concentrations and whether the phenomenon could explain purported genetic interactions with sex, diet, disease, BMI, and drugs.

59 BASIC BIOLOGICAL SCIENCES↗

Statistical Learning for Nonlinear Model Reduction from Local Simulations of Stochastic and Particle- and Agent-Based Systems

Stochastic physical systems across the sciences that have very high-dimensional state spaces, with a large number of fast degrees of freedom that force direct simulators to proceed by integration steps that are orders of magnitude smaller than events of interests (e.g., particle collisions). Examples range from molecular motion to dynamics of large populations of cells. A grand challenge in the simulation and understanding of such systems is the systematic construction of accurate, interpretable, reduced models, enabling faster simulations, revealing fundamental properties of the dynamics, and predicting phenomena of interest that the original simulator could not reached with sufficient accuracy or within a given computational budget. In this projected we developed novel statistical estimation/machine learning techniques for analyzing and building empirical reduced models for important families of high-dimensional stochastic systems, in particular: - we developed techniques for estimating interaction kernels in interacting particle- and agent-based systems, which are ubiquitous in Physics, Biology and many other sciences, given observed trajectories of the system; - we developed techniques for nonlinear model reduction for high-dimensional stochastic systems that have a small number of unknown, nonlinear slow variables, and a large number of fast modes, that are possibly of large magnitude, given observed short trajectories of the system in the form of bursts of trajectories from different initial conditions; - we developed novel techniques for estimating linear dynamical systems on graphs when both the dynamics and the underlying graph are unknown, and we have a sparse set of space-time observations; - we considered the problem of estimating an unknown nonlinear observation function of a standard process (e.g. Brownian motion), so that we can recognized if an observed dynamics is "just" a nonlinear version of a known dynamics; we also developed benchmarks for learning algorithms aimed at learning and classifying diffusion processes.

97 MATHEMATICS AND COMPUTING↗

Illuminating the Material World: Autonomous Microscopy to Understand Order, Disorder, and Everything In Between

Artificial intelligence (AI) holds immense promise for revolutionizing microscopy, yet its widespread adoption has been hindered by challenges ranging from user inexperience to limited model transferability and difficulties in operationalizing machine learning. This presentation showcases our approach to developing practical autonomy for materials discovery, aiming to accelerate the integration of AI into everyday microscopy workflows. As shown in Fig. 1, I will focus on three key areas: understanding order-disorder transitions, quantifying point defects, and achieving truly device-scale microscopy. First, I will demonstrate the power of multi-modal knowledge graphs for integrating diverse microscopy data. By combining imaging, spectroscopy, and diffraction data, these graphs provide a holistic view of material behavior, capturing the intricate relationships between different modalities [1,2]. I will present a case study on how these models illuminate the structural and chemical changes associated with irradiation in oxide thin films, revealing critical insights for designing materials for extreme environments like spaceflight and nuclear energy. Specifically, I will show how multi-modal analysis clarifies the evolution of order-disorder transitions under irradiation, a key factor influencing material performance in these applications. Next, I will address the challenge of quantifying point defects in 2D materials. We demonstrate the application of computer vision and transfer learning to accurately identify and classify various defect types, such as vacancies and substitutional atoms, and to quantify their concentrations. This information is crucial for understanding and tailoring the properties of 2D materials for applications in electronics, optoelectronics, and catalysis. For example, I will show how our models can characterize the topological distribution of point defects in MXene transition metal carbides, providing valuable insights for optimizing their performance in energy storage and separation science. Finally, I will discuss our progress toward autonomous device-scale microscopy [3,4]. We are fundamentally redesigning electron microscopes around the principles of machine reasoning, enabling automation beyond basic tasks like sample navigation and data acquisition to include sophisticated experimental design. This approach paves the way for truly reproducible and massively scaled analysis campaigns. I will emphasize the importance of autonomous microscopy platforms for high-throughput materials discovery and characterization, facilitating the rapid screening of materials for a broad range of applications and accelerating the development of next-generation technologies.

36 MATERIALS SCIENCE↗

New Perspectives on the Exoplanet Radius Gap from a Mathematica Tool and Visualized Water Equation of State

Recent astronomical observations obtained with the Kepler and TESS missions and their related ground-based follow-ups revealed an abundance of exoplanets with a size intermediate between Earth and Neptune (1 R ⊕ ≤ R ≤ 4 R ⊕ ). A low occurrence rate of planets has been identified at around twice the size of Earth (2 × R ⊕ ), known as the exoplanet radius gap or radius valley. We explore the geometry of this gap in the mass–radius diagram, with the help of a Mathematica plotting tool developed with the capability of manipulating exoplanet data in multidimensional parameter space, and with the help of visualized water equations of state in the temperature–density (T–ρ) graph and the entropy–pressure (s–P) graph. We show that the radius valley can be explained by a compositional difference between smaller, predominantly rocky planets (<2 × R ⊕ ) and larger planets (>2 × R ⊕ ) that exhibit greater compositional diversity including cosmic ices (water, ammonia, methane, etc.) and gaseous envelopes. In particular, among the larger planets (>2 × R ⊕ ), when viewed from the perspective of planet equilibrium temperature (T eq ), the hot ones (T eq ≳ 900 K) are consistent with ice-dominated composition without significant gaseous envelopes, while the cold ones (T eq ≲ 900 K) have more diverse compositions, including various amounts of gaseous envelopes.

79 ASTRONOMY AND ASTROPHYSICS↗

Evidence-based Graph Adversary Mapping (EGRAM) [Poster]

Cybersecurity companies such as CrowdStrike, Dragos, Microsoft and Unit 42 categorize Advanced Persistent Threats (APTs) using their own naming schemes. As a result, these APTs are mapped to different malware sources and campaigns, all from differing sources, leading to inconsistent mapping. Inconsistent mapping causes confusion and adds further obscurity around these groups, making it difficult to track and mitigate APT cyberattacks. The Evidence-based Graph Adversary Mapping (EGRAM) tool remediates the mapping challenge by collecting, updating and converting adversary data and their sources into a valid, codified STIX v2.1 bundle which is then stored in a Neo4j graph database. It utilizes graph traversal methods and centrality analysis to generate actionable information as a Structured Threat Intelligence Graph (STIG), based on user queries. EGRAM exists as Python code and a Jupyter Notebook that acts as a searchable, evidence-based, source of intelligence for APT groups’ artifacts and cyber campaigns.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

3D-equivariant graph neural networks for protein model quality assessment

Quality assessment (QA) of predicted protein tertiary structure models plays an important role in ranking and using them. With the recent development of deep learning end-to-end protein structure prediction techniques for generating highly confident tertiary structures for most proteins, it is important to explore corresponding QA strategies to evaluate and select the structural models predicted by them since these models have better quality and different properties than the models predicted by traditional tertiary structure prediction methods. We develop EnQA, a novel graph-based 3D-equivariant neural network method that is equivariant to rotation and translation of 3D objects to estimate the accuracy of protein structural models by leveraging the structural features acquired from the state-of-the-art tertiary structure prediction method—AlphaFold2. We train and test the method on both traditional model datasets (e.g. the datasets of the Critical Assessment of Techniques for Protein Structure Prediction) and a new dataset of high-quality structural models predicted only by AlphaFold2 for the proteins whose experimental structures were released recently. Our approach achieves state-of-the-art performance on protein structural models predicted by both traditional protein structure prediction methods and the latest end-to-end deep learning method—AlphaFold2. It performs even better than the model QA scores provided by AlphaFold2 itself. The results illustrate that the 3D-equivariant graph neural network is a promising approach to the evaluation of protein structural models. Integrating AlphaFold2 features with other complementary sequence and structural features is important for improving protein model QA.

59 BASIC BIOLOGICAL SCIENCES↗

Integer Sequences from Configurations in the Hausdorff Metric Geometry via Edge Covers of Bipartite Graphs

The Hausdorff metric provides a way to measure the distance between nonempty compact sets in $\mathbb{R}^N$, from which we can build a geometry of sets. This geometry is very different than the standard Euclidean geometry and provides many interesting results. In this paper we focus on line segments in this geometry, where pairs of disjoint sets $A$ and $B$ satisfying certain distance conditions have the property that there are exactly $m$ different sets on the line segment $\overline{AB}$ at every distance from $A$, where $m$ can assume many values different than one. We provide new families of sets that generate previously unrecorded integer sequences via these values of $m$ by connecting the values of $m$ to the number of edge coverings of a graph corresponding to the sets $A$ and $B$.

97 MATHEMATICS AND COMPUTING↗

Towards Enhancing Coding Productivity for GPU Programming Using Static Graphs

The main contribution of this work is to increase the coding productivity of GPU programming by using the concept of Static Graphs. GPU capabilities have been increasing significantly in terms of performance and memory capacity. However, there are still some problems in terms of scalability and limitations to the amount of work that a GPU can perform at a time. To minimize the overhead associated with the launch of GPU kernels, as well as to maximize the use of GPU capacity, we have combined the new CUDA Graph API with the CUDA programming model (including CUDA math libraries) and the OpenACC programming model. We use as test cases two different, well-known and widely used problems in HPC and AI: the Conjugate Gradient method and the Particle Swarm Optimization. In the first test case (Conjugate Gradient) we focus on the integration of Static Graphs with CUDA. In this case, we are able to significantly outperform the NVIDIA reference code, reaching an acceleration of up to 11x thanks to a better implementation, which can benefit from the new CUDA Graph capabilities. In the second test case (Particle Swarm Optimization), we complement the OpenACC functionality with the use of CUDA Graph, achieving again accelerations of up to one order of magnitude, with average speedups ranging from 2x to 4x, and performance very close to a reference and optimized CUDA code. Our main target is to achieve a higher coding productivity model for GPU programming by using Static Graphs, which provides, in a very transparent way, a better exploitation of the GPU capacity. The combination of using Static Graphs with two of the current most important GPU programming models (CUDA and OpenACC) is able to reduce considerably the execution time w.r.t. the use of CUDA and OpenACC only, achieving accelerations of up to more than one order of magnitude. Finally, we propose an interface to incorporate the concept of Static Graphs into the OpenACC Specifications.

58 GEOSCIENCES↗

Disruption-Robust Community Detection Using Consensus Clustering in Complex Networks

Topological (graph-theoretic) analysis of critical infrastructure networks provides insight on several aspects of resilience. Graph clustering or community detection, which identifies densely connected components in a graph, has been employed for analysis. In this paper, we propose employing consensus clustering, which is a technique to determine consensus from a collection of different clusters on an input, such that the resulting clustering is robust to disruptions, where a disruption is represented as loss of one or more vertices or edges in the graph. Using two critical infrastructure networks as case studies, we empirically demonstrate the need to compute consensus clustering in order to address the drastic changes in the topology due to disruptions in the network.

Hussain, Md Taufique↗

Applying machine learning and quantum chemistry to predict the glass transition temperatures of polymers

Glass transition temperature (T g ) is important for understanding the physical and mechanical properties of a polymer material because it relates to the thermal energy required to transition between a hard glassy state and a soft rubbery one. Over the years, various models have been developed for predicting this thermal property from molecular structure to aid in designing novel polymers in selected classes. This work builds on those efforts by utilizing both machine learning (ML) and quantum chemistry (QC) techniques to develop models that can predict T g values from the molecular structure under different data availability scenarios and for a wide variety of polymer types. For the ML model, a graph convolutional network (GCN) was used to map topological polymer features; this model was trained against a dataset of more than 7500 T g values and resulted in a root mean square error (RMSE) of 38.1 °C. The QC-based regression model was trained on 83 T g values and produced an RMSE of 34.5 °C. In conclusion, this work demonstrated that while both model techniques produce accurate predictions and are suitable for different data availability scenarios, the QC-based regression model offered a more interpretable model framework with significantly less training data.

36 MATERIALS SCIENCE↗