Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Detecting Masquerade Attacks in Controller Area Networks Using Graph Machine Learning

Modern vehicles rely on a myriad of electronic control units (ECUs) interconnected via controller area networks (CANs) for critical operations. Despite their ubiquitous use and reliability, CANs are susceptible to sophisticated cyberattacks, particularly masquerade attacks, which inject false data that mimic legitimate messages at the expected frequency. These attacks pose severe risks such as unintended acceleration, brake deactivation, and rogue steering. Traditional intrusion detection systems (IDS) often struggle to detect these subtle intrusions due to their seamless integration into normal traffic. This paper introduces a novel framework for detecting masquerade attacks in the CAN bus using graph machine learning (ML). We hypothesize that the integration of shallow graph embeddings with time series features derived from CAN frames enhances the detection of masquerade attacks. We show that by representing CAN bus frames as message sequence graphs (MSGs) and enriching each node with contextual statistical attributes from time series, we can enhance detection capabilities across various attack patterns compared to using graph-based features only. Our method ensures a comprehensive and dynamic analysis of CAN frame interactions, improving robustness and efficiency. Extensive experiments on the ROAD dataset validate the effectiveness of our approach, demonstrating statistically significant improvements in the detection rates of masquerade attacks compared to a baseline that uses graph-based features only as confirmed by Mann-Whitney U and Kolmogorov-Smirnov tests (p < 0.05) .

Marfo, William [Univ. of Texas, El Paso, TX (Unite↗

Structure-Informed Graph Learning of Networked Dependencies for Online Prediction of Power System Transient Dynamics

Online transient analysis plays an increasingly important role in dynamic power grids as the renewable generation continues growing. Traditional numerical methods for transient analysis not only are computationally intensive but also require precise contingency information as input, and therefore, are not suitable for online applications. Existing online transient assessment studies focus on the determination of post-contingency system stability or stability margin. Here, this paper develops a novel graph-learning framework, Deep-learning Neural Representation or DNR, for online prediction, of the time-series trajectories of the system states using initial system responses that can be measured by phasor measurement units (PMUs). The proposed DNR framework consists of two sequential modules: a Network Constructor that captures network dependencies among generators, and a Dynamics Predictor that predicts the system trajectories. The key to improved prediction performance is the introduction of the spatio-temporal message-passing operations into graph neural networks with structural knowledge. Its effectiveness and scalability are validated through comparative studies, demonstrating the prediction performance under different contingency scenarios for systems of different sizes. This framework provides a solution to online predicting post-fault system dynamics based on real-time PMU measurements. Additionally, it can also be applied to facilitate the offline transient simulation without simulating the entire trajectories.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Differentially Private Synthesis and Sharing of Network Data Via Bayesian Exponential Random Graph Models

Abstract Network data often contain sensitive relational information. One approach to protecting sensitive information while offering flexibility for network analysis is to share synthesized networks based on the information in originally observed networks. We employ differential privacy (DP) and exponential random graph models (ERGMs) and propose the DP-ERGM method to synthesize network data. We apply DP-ERGM to two real-world networks. We then compare the utility of synthesized networks generated by DP-ERGM, the DyadWise Randomized Response (DWRR) approach, and the Synthesis through Conditional distribution of Edge given nodal Attribute (SCEA) approach. In general, the results suggest that DP-ERGM preserves the original information significantly better than two other approaches in network structural statistics and inference for ERGMs and latent space models. Furthermore, DP-ERGM satisfies node DP through modeling the global network structure with ERGM, a stronger notion of privacy than the edge DP under which DWRR and SCEA operate.

graph synthesis↗

De Sitter diagrammar and the resummation of time

Light scalars in inflationary spacetimes suffer from logarithmic infrared divergences at every order in perturbation theory. This corresponds to the scalar field values in different Hubble patches undergoing a random walk of quantum fluctuations, leading to a simple toy “landscape” on superhorizon scales, in which we can explore questions relevant to eternal inflation. However, for a sufficiently long period of inflation, the infrared divergences appear to spoil computability. Some form of renormalization group approach is thus motivated to resum the log divergences of conformal time. Such a resummation may provide insight into De Sitter holography. We present here a novel diagrammatic analysis of these infrared divergences and their resummation. Basic graph theory observations and momen- tum power counting for the in-in propagators allow a simple and insightful determination of the leading-log contributions. One thus sees diagrammatically how the superhorizon sector consists of a semiclassical theory with quantum noise evolved by a first-order, interacting classical equation of motion. This rigorously leads to the “Stochastic Inflation” ansatz developed by Starobinsky to cure the scalar infrared pathology nonperturbatively. Our approach is a controlled approximation of the underlying quantum field theory and is systematically improvable.

79 ASTRONOMY AND ASTROPHYSICS↗

CP‐SyNet: A tool for generating customised cyber‐power synthetic network for distribution systems with distributed energy resources

Abstract The integration of distributed energy resources and advancement in information technology has enabled the transition of traditional power distribution systems to active cyber‐physical distribution systems. A growing amount of research has been done on the modelling, analysis, and optimisation of power distribution system behaviour. However, existing publicly available distribution test feeders are limited in numbers and have minimal features. Furthermore, these test feeders do not include cyber models and are not customisable. To bridge this gap, we propose and develop Cyber‐physical synthetic distribution system network (CP‐SyNet), a tool for generating customisable cyber‐physical synthetic distribution test feeders. CP‐SyNet generates three‐phase unbalanced test feeders according to users' requirements, while simultaneously considering both the cyber side and the physical side of the network for cyber‐physical analysis. The physical test network is developed using a graph‐theoretical approach that employs information from existing test feeders. The cyber side considers an equivalent communication network by transforming the physical topology into possible and feasible simulated network. Two examples are presented to demonstrate the feasibility of the proposed framework to generate cyber‐physical test feeders.

Wang, Lusha↗

A variant selection framework for genome graphs

Abstract Motivation Variation graph representations are projected to either replace or supplement conventional single genome references due to their ability to capture population genetic diversity and reduce reference bias. Vast catalogues of genetic variants for many species now exist, and it is natural to ask which among these are crucial to circumvent reference bias during read mapping. Results In this work, we propose a novel mathematical framework for variant selection, by casting it in terms of minimizing variation graph size subject to preserving paths of length α with at most δ differences. This framework leads to a rich set of problems based on the types of variants [e.g. single nucleotide polymorphisms (SNPs), indels or structural variants (SVs)], and whether the goal is to minimize the number of positions at which variants are listed or to minimize the total number of variants listed. We classify the computational complexity of these problems and provide efficient algorithms along with their software implementation when feasible. We empirically evaluate the magnitude of graph reduction achieved in human chromosome variation graphs using multiple α and δ parameter values corresponding to short and long-read resequencing characteristics. When our algorithm is run with parameter settings amenable to long-read mapping (α = 10 kbp, δ = 1000), 99.99% SNPs and 73% SVs can be safely excluded from human chromosome 1 variation graph. The graph size reduction can benefit downstream pan-genome analysis. Availability and implementation https://github.com/AT-CG/VF. Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

Rare Higgs Processes at CMS and Precision Timing Detector Studies for HL-LHC CMS Upgrade

This thesis describes the search for two rare Higgs processes. The first analysis describes the CMS Run 2 search for $H$ $\rightarrow$ $\mu$$\mu$ decays, with 137.3 fb$^{-1}$ of data at $\sqrt{s}$ = 13 TeV. The analysis targeted four different Higgs production modes: the gluon fusion (ggH), the vector boson fusion (VBF), the Higgs-strahlung process (VH), and the production in association with a pair of top quarks (ttH). Each category used a dedicated machine learning based classifier to separate the signal from the background processes. A combined fit from all these categories saw a slight excess in the data corresponding to 3.0 standard deviations at $M$$_{H}$ = 125.38 GeV, and gave the first evidence for the Higgs boson decay to second-generation fermions. The best-fit signal strength and the corresponding 68% CL interval was found to be +0.17?????? = 1.19 $_{-0.39}^{+0.41}$ (stat)$_{-0.16}^{+0.17}$(syst) at $M$$_{H}$ = 125.38 GeV. The second analysis describes the CMS Run 2 search for 𝐻𝐻 → 𝑏𝑏𝑏𝑏 with highly boosted Higgs bosons. This analysis used a dedicated jet identification algorithm based on graph neural networks (ParticleNet) to identify boosted H→ bb jets. This search targeted the gluon fusion and the vector boson fusion HH production modes, and put constraints on the allowed values of the various Higgs couplings as: 𝜅𝜆 ∈ [−9.9, 16.9] when 𝜅𝑉 = 1, 𝜅2𝑉 = 1; 𝜅𝑉 ∈ [−1.17, −0.79] ∪ [0.81, 1.18] when 𝜅𝜆 = 1, 𝜅2𝑉 = 1; 𝜅2𝑉 ∈ [0.62, 1.41] when 𝜅𝜆 = 1, 𝜅𝑉 = 1. A scenario with 𝜅2𝑉 = 0 was excluded with a significance of 6.3 standard deviations for the first time, when other H couplings are fixed to their SM values. The combined observed (expected) 95% upper limit on the HH production cross section was found to be 9.9 (5.1) × SM. Finally, this thesis also discusses the planned MIP Timing Detector (MTD) upgrade for CMS at the HL-LHC. The MTD will be a time-of-flight (TOF) detector, designed to provide a precision timing information for charged particles using SiPMs + LYSO scintillating crystals, with a time resolution of ∼30 ps. This thesis describes several R&D tests that have been performed for characterizing the sensor properties (time resolution, light yield, etc.) and optimizing the sensor design geometry. This thesis also contains a description of mock test setups for cooling the sensors, since it is known to be an effective way of mitigating the increased dark current rates in the sensors due to radiation damage.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Impact-Driven Sampling Strategies for Hybrid Attack Graphs

Cyber-Physical Systems (CPSs) have a large input space, with discrete and continuous elements across multiple layers. Hybrid Attack Graph (HAG) provide a flexible and efficient approach to generate attack sequences for a CPS. Analysis and testing of large-scale HAGs are prohibitively costly. We propose a dimension reduction via property-preserving multi-layer graph sampling algorithms. Existing property-preserving graph sampling approaches generate a representative subgraph of an original large-sized graph while preserving the key properties, such as node and edge distribution, clustering coefficients, and betweenness. On the other hand, we propose impact-driven sampling strategies to transform the input data to a lower-dimensional representation while retaining key properties of the data.

Subasi, Omer↗

Improving ProtoDUNE pion cross-section measurements with NuGraph Michel-electron tagging

Understanding hadron-argon interactions is essential for precise neutrino energy reconstruction and final-state interaction modeling in liquid-argon time projection chamber (LArTPC) experiments such as DUNE. In particular, pion absorption and charge-exchange processes constitute significant sources of systematic uncertainty in neutrino oscillation measurements. ProtoDUNE-SP, a large-scale LArTPC prototype operated at the CERN Neutrino Platform and exposed to charged-particle test beams in the few-GeV range, enables direct measurements of these processes. This work focuses on the measurement of differential cross sections for pion absorption and charge exchange using the 2 GeV/c pion beam data from the ProtoDUNE-SP run. A key component of this analysis is the identification of Michel electrons from $\pi \rightarrow \mu \rightarrow e$ decay chains, which helps separate different interaction topologies and improves background rejection. Michel electron identification will also assist in reliably calibrating the electromagnetic response in ProtoDUNE-SP data and for the future DUNE detectors. In this analysis, we apply NuGraph to identify Michel electrons. NuGraph is a graph neural network that models detector hits as nodes connected by spatial and temporal edges for particle and topology classification in LArTPC detectors. We first benchmark NuGraph’s Michel electron classification performance using ICEBERG data, a small-scale LArTPC prototype used for DUNE electronics and reconstruction development, and then transfer the approach to ProtoDUNE-SP. This poster presents the analysis strategy, NuGraph-based classification studies, and discusses how these developments are expected to improve the pion cross-section measurement.

Razafinime, Soamasina Herilala [Cincinnati U.] (OR↗

Retrieval Augmented Generation for Robust Cyber Defense

In cybersecurity, the ability to efficiently analyze and respond to vulnerabilities, weaknesses, attack patterns, and threat tactics is critical for effective defense strategies. With the increasing complexity and volume of cybersecurity data, traditional methods of querying and retrieving information are often inadequate. To address this challenge, we implemented Retrieval-Augmented Generation (RAG) systems—CyRAG and GraphCyRAG—that integrate large language models (LLMs) with both structured data from relational databases and knowledge graphs such as Neo4j. CyRAG is designed to handle structured data, focusing on CVE (Common Vulnerabilities and Exposures) and CWE (Common Weakness Enumeration) entities to generate accurate and context-rich responses. In contrast, GraphCyRAG leverages Neo4j knowledge graphs to retrieve interconnected information from CVE, CWE, CAPEC (Common Attack Pattern Enumeration and Classification), and ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) datasets. By utilizing Neo4j’s graph-based framework, GraphCyRAG enables deeper traversal of relationships between vulnerabilities and attack patterns, providing cybersecurity analysts with more comprehensive insights into potential attack vectors and mitigation strategies. Our preliminary results demonstrate that integrating knowledge graphs with RAG significantly enhances both the accuracy and depth of threat analysis, allowing for the retrieval of dynamic, real-time data and the generation of contextually aware responses. This approach helps analysts uncover hidden relationships between cyber entities, predict exploit paths, and prioritize mitigation efforts effectively. The integration of RAG with cybersecurity knowledge graphs represents a significant advancement in cybersecurity threat intelligence, enabling more informed decision-making and stronger defense strategies.

97 MATHEMATICS AND COMPUTING↗

Local structure graph models with higher-order dependence

Local structure graph models (LSGMs) describe random graphs and networks as a Markov random field (MRF)—each graph edge has a specified conditional distribution dependent on explicit neighbourhoods of other graph edges. Centered parameterizations of LSGMs allow for direct control and interpretation of parameters for large- and small-scale structures (e.g., marginal means vs. dependence). Here, we extend this parameterization to account for triples of dependent edges and illustrate the importance of centered parameterizations for incorporating covariates and interpreting parameters. Using a MRF framework, common exponential random graph models are also shown to induce conditional distributions without centered parameterizations and thereby have undesirable features. This work attempts to advance graph models through conditional model specifications with modern parameterizations, covariates and higher-order dependencies.

97 MATHEMATICS AND COMPUTING↗

Graph Metric Learning Quantifies Morphological Differences between Two Genotypes of Shoot Apical Meristem Cells in Arabidopsis

We present a method for learning “spectrally descriptive” edge weights for graphs. We generalize a previously known distance measure on graphs (Graph Diffusion Distance), thereby allowing it to be tuned to minimize an arbitrary loss function. Because all steps involved in calculating this modified GDD are differentiable, we demonstrate that it is possible for a small neural network model to learn edge weights which minimize loss. We apply this method to discriminate between graphs constructed from shoot apical meristem images of two genotypes of Arabidopsis thaliana specimens: wild-type and trm678 triple mutants with cell division phenotype. Training edge weights and kernel parameters with contrastive loss produces a learned distance metric with large margins between these graph categories. We demonstrate this by showing improved performance of a simple k-nearest-neighbors classifier on the learned distance matrix. We also demonstrate a further application of this method to biological image analysis. Once trained, we use our model to compute the distance between the biological graphs and a set of graphs output by a cell division simulator. Comparing simulated cell division graphs to biological ones allows us to identify simulation parameter regimes which characterize mutant vs. wild-type Arabidopsis cells. We find that trm678 mutant cells are characterized by increased randomness of division planes and decreased ability to avoid previous vertices between cell walls.

59 BASIC BIOLOGICAL SCIENCES↗

Efficient graph representation framework for chemical molecule similarity tasks

Graph data has emerged in numerous scientific domains and machine learning techniques have been widely used for analysis and learning of diverse data for prediction and decision. Machine learning techniques can readily address complex problems by leveraging their structural information. But graphs cannot be directly used for existing machine learning algorithms unless encoded as vectors. The problem of efficient representation of graphs is a substantial challenge in graph machine learning. In this paper, we propose a novel two-stage framework for the representation of chemical molecule graphs based on the strengths of Graph Isomorphism Networks (GINs) and Siamese autoencoders. In the first stage, the GIN model is constructed and trained using the structural information of chemical molecule graphs. Node attributes, edge attributes, and edge indices are used as input data, while graph attributes are used as labels. The GIN model effectively captures the structural characteristics of graphs and can accurately predict graph attributes, i.e., molecular properties. It also generates Graph Embeddings, represented as vectors that encode the structural information of graphs. In the second stage, Graph Embedding vectors are further optimized for downstream similarity tasks while preserving the graph structural information. The Siamese autoencoder is constructed and trained, which reduces the dimensionality of the Graph Embedding vectors, while maximizing the preservation of structural information in the original high-dimensional vectors. The resulting low-dimensional Graph Embeddings can be effectively utilized for tasks such as approximate nearest neighbor search. The experimental results demonstrate the effectiveness of our proposed framework in accurately predicting graph similarity.

Ma, Jiaji↗

Search for Nonresonant Pair Production of Highly Energetic Higgs Bosons Decaying to Bottom Quarks

A search for nonresonant Higgs boson ($H$) pair production via gluon and vector boson ($V$) fusion is performed in the four-bottom-quark final state, using proton-proton collision data at 13 TeV corresponding to 138 fb$^{−1}$ collected by the CMS experiment at the LHC. The analysis targets Lorentz-boosted $H$ pairs identified using a graph neural network. It constrains the strengths relative to the standard model of the $H$ self-coupling and the quartic VVHH couplings, $κ_{2V}$, excluding $κ_{2V} = 0$ for the first time, with a significance of 6.3 standard deviations when other H couplings are fixed to their standard model values.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

DeepCare: Improving Patient Care using Deep Learning on Electronic Health Records

Coordinating patient care using electronic health records (EHR) data presents an exciting but formidable opportunity in data extraction, analysis and modeling. Traditional methods use a manual feature driven approach to model patients with age, family history and symptoms to predict disease outcomes. We propose a novel approach to model patients based on their streaming electronic health records data combined with information from medical knowledge bases, which has been gained over years of medical research. Using a combination of representation learning and long short term memory (LSTM) networks we plan to model patient evolution over time, leading to more accurate and individualized predictive models for patient’s diseases. Our approach will be transformative in providing critical decision support for patient care, enabling accurate understanding and evolution of diseases in patients.

60 APPLIED LIFE SCIENCES↗

Knowledge-guided graph machine learning for spatially distributed prediction of daily discharge and nitrogen export dynamics

Spatially distributed prediction of streamflow and nitrogen export dynamics is essential for precision management of agricultural watersheds. While temporal deep learning models such as Long Short-Term Memory (LSTM) have shown strong performance at basin scales, their ability to generalize spatially is limited by insufficient representation of spatial dependencies and flow paths, particularly under data-scarce conditions. To address this gap, we propose HydroGraphNet, a knowledge-guided graph machine learning framework that integrates process-based knowledge and explicit spatial learning into temporal modeling. This framework incorporates directed graph topology to encode watershed connectivity and upstream inflows, with mass balance constraints to improve physical consistency. To enhance generalization in sparsely monitored regions, HydroGraphNet is pretrained on synthetic data generated by the SWAT+ (Soil and Water Assessment Tool Plus) model. We evaluated HydroGraphNet in the Upper Sangamon River Basin (44 HUC-12 subwatersheds, 2001–2020) against two LSTM baselines: a lumped basin-level model and a distributed variant. When benchmarked on SWAT+ simulations in pretraining, HydroGraphNet improved test NSEs by 8.9% (discharge) and 13.7% (NO₃–N load) in temporal extrapolation, and by 27.1% and 34.7% in spatial extrapolation, relative to the Lumped LSTM baseline. After fine-tuning with USGS monitoring data, the model achieved mean test NSE (KGE) scores of 0.768 (0.861) for discharge and 0.626 (0.664) for NO₃–N load, substantially outperforming baselines. Attribution analysis further highlighted the importance of upstream inflow representation and graph-based spatial learning in capturing cross-subwatershed dependencies. The model also reproduced seasonal hydrological and biogeochemical patterns consistent with known processes, demonstrating its robustness and process fidelity for spatially distributed prediction. Altogether, HydroGraphNet advances the integration of physical knowledge and spatially explicit learning in hydrological modeling, offering a generalizable framework for distributed modeling to support spatially targeted water quality management in data-scarce watersheds.

54 ENVIRONMENTAL SCIENCES↗

Applications of the Dulmage–Mendelsohn decomposition for debugging nonlinear optimization problems

Nonlinear modeling and optimization is a valuable tool for aiding decisions by engineering practitioners, but programming an optimization problem based on a complex electrical, mechanical, or chemical process is a time-consuming and error-prone activity. Therefore, there is a need for model analysis and debugging tools that can detect and diagnose modeling errors. One such tool is the Dulmage–Mendelsohn decomposition, which identifies structurally under- and over-determined subsets in systems of equations and variables by partitioning the bipartite graph of the system. This work provides the necessary background to understand the Dulmage–Mendelsohn decomposition and its application to the analysis of nonlinear optimization problems, demonstrates its use in diagnosing a variety of modeling errors, and introduces software implementations for analyzing nonlinear optimization problems in the Pyomo and JuMP algebraic modeling languages.

42 ENGINEERING↗

Efficient Hierarchical State Vector Simulation of Quantum Circuits via Acyclic Graph Partitioning

Early but promising results in quantum computing have been enabled by the concurrent development of quantum algorithms, devices, and materials. Classical simulation of quantum programs has enabled the design and analysis of algorithms and implementation strategies targeting current and anticipated quantum device architectures. In this paper, we present a graph-based approach to achieve efficient quantum circuit simulation. Our approach involves partitioning the graph representation of a given quantum circuit into sub-graphs/circuits that exhibit better data locality. Simulation of each sub-circuit is organized hierarchically, with the iterative construction and simulation of smaller state vectors, improving overall performance. Also, this partitioning reduces the number of passes through data, improving the total computation time. We present three partitioning strategies and observe that acyclic graph partitioning typically results in the best time-to-solution. In contrast, other strategies reduce the partitioning time at the expense of potentially increased simulation times. Experimental evaluation demonstrates the effectiveness of our approach.

Fang, Bo↗