Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Equivariant, safe and sensitive — graph networks for new physics

This study introduces a novel Graph Neural Network (GNN) architecture that leverages infrared and collinear (IRC) safety and equivariance to enhance the analysis of collider data for Beyond the Standard Model (BSM) discoveries. By integrating equivariance in the rapidity-azimuth plane with IRC-safe principles, our model significantly reduces computational overhead while ensuring theoretical consistency in identifying BSM scenarios amidst Quantum Chromodynamics backgrounds. The proposed GNN architecture demonstrates superior performance in tagging semi-visible jets, highlighting its potential as a robust tool for advancing BSM search strategies at high-energy colliders.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

HydraGNN_OPF_GFM_2026 - Ensemble of predictive graph foundation models for power grid applications

This dataset supports research on graph foundation models for optimal power flow (OPF) on electric grids using HydraGNN. It contains heterogeneous graph representations of PGLib-OPF cases spanning systems from 14 to 13,659 buses, together with packed HDF5 datasets for pretraining, feasibility classification, and N-1 contingency analysis. The release includes OPF solution data, downstream fine-tuning datasets, pretrained HeteroSAGE and HeteroHEAT model checkpoints, hyperparameter-optimization summaries across multiple heterogeneous GNN architectures, and aggregated fine-tuning results for sample-efficiency studies. The dataset is designed to enable scalable training, evaluation, and transfer-learning studies for OPF surrogate modeling, including node-level AC-OPF solution prediction, graph-level prediction, feasibility classification, operating-condition generalization, and contingency-response tasks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep structural clustering for single-cell RNA-seq data jointly through autoencoder and graph neural network

Abstract Single-cell RNA sequencing (scRNA-seq) permits researchers to study the complex mechanisms of cell heterogeneity and diversity. Unsupervised clustering is of central importance for the analysis of the scRNA-seq data, as it can be used to identify putative cell types. However, due to noise impacts, high dimensionality and pervasive dropout events, clustering analysis of scRNA-seq data remains a computational challenge. Here, we propose a new deep structural clustering method for scRNA-seq data, named scDSC, which integrate the structural information into deep clustering of single cells. The proposed scDSC consists of a Zero-Inflated Negative Binomial (ZINB) model-based autoencoder, a graph neural network (GNN) module and a mutual-supervised module. To learn the data representation from the sparse and zero-inflated scRNA-seq data, we add a ZINB model to the basic autoencoder. The GNN module is introduced to capture the structural information among cells. By joining the ZINB-based autoencoder with the GNN module, the model transfers the data representation learned by autoencoder to the corresponding GNN layer. Furthermore, we adopt a mutual supervised strategy to unify these two different deep neural architectures and to guide the clustering task. Extensive experimental results on six real scRNA-seq datasets demonstrate that scDSC outperforms state-of-the-art methods in terms of clustering accuracy and scalability. Our method scDSC is implemented in Python using the Pytorch machine-learning library, and it is freely available at https://github.com/DHUDBlab/scDSC.

Gan, Yanglan↗

Genomic Metrics Applied to Rhizobiales ( Hyphomicrobiales ): Species Reclassification, Identification of Unauthentic Genomes and False Type Strains

Taxonomic decisions within the order Rhizobiales have relied heavily on the interpretations of highly conserved 16S rRNA sequences and DNA–DNA hybridizations (DDH). Currently, bacterial species are defined as including strains that present 95–96% of average nucleotide identity (ANI) and 70% of digital DDH (dDDH). Thus, ANI values from 520 genome sequences of type strains from species of Rhizobiales order were computed. From the resulting 270,400 comparisons, a ≥95% cut-off was used to extract high identity genome clusters through enumerating maximal cliques. Coupling this graph-based approach with dDDH from clusters of interest, it was found that: (i) there are synonymy between Aminobacter lissarensis and Aminobacter carboxidus, Aurantimonas manganoxydans and Aurantimonas coralicida, “Bartonella mastomydis,” and Bartonella elizabethae, Chelativorans oligotrophicus, and Chelativorans multitrophicus, Rhizobium azibense, and Rhizobium gallicum, Rhizobium fabae, and Rhizobium pisi, and Rhodoplanes piscinae and Rhodoplanes serenus; (ii) Chelatobacter heintzii is not a synonym of Aminobacter aminovorans; (iii) “Bartonella vinsonii” subsp. arupensis and “B. vinsonii” subsp. berkhoffii represent members of different species; (iv) the genome accessions GCF_003024615.1 (“Mesorhizobium loti LMG 6125T”), GCF_003024595.1 (“Mesorhizobium plurifarium LMG 11892T”), GCF_003096615.1 (“Methylobacterium organophilum DSM 760T”), and GCF_000373025.1 (“R. gallicum R-602 spT”) are not from the genuine type strains used for the respective species descriptions; and v) “Xanthobacter autotrophicus” Py2 and “Aminobacter aminovorans” KCTC 2477T represent cases of misuse of the term “type strain”. Aminobacter heintzii comb. nov. and the reclassification of Aminobacter ciceronei as A. heintzii is also proposed. To facilitate the downstream analysis of large ANI matrices, we introduce here ProKlust (“Prokaryotic Clusters”), an R package that uses a graph-based approach to obtain, filter, and visualize clusters on identity/similarity matrices, with settable cut-off points and the possibility of multiple matrices entries.

59 BASIC BIOLOGICAL SCIENCES↗

Oak Ridge National Laboratory Technical Input for the Nuclear Regulatory Commission Review of the 2017 Edition of ASME Section III, Division 5, ‘High Temperature Reactors’

To assist the Nuclear Regulatory Commission in its decision making on endorsement of the American Society for Mechanical Engineers Boiler and Pressure Vessel Code Section III, Division 5 (2017 Edition) for development of advanced non-light water reactors, the following Division 5 portions were reviewed: Article HBB-2000 Material; Article HCB-2000 Material; Article HGB-2000 Material; Mandatory Appendix HBB-I-14 Tables and Figures; and, Nonmandatory Appendix HBB-U Guidelines for Restricted Material Specifications to Improve Performance in Certain Service Applications. In addition to the 2017 Edition, the same parts of the 2019 Edition have also been reviewed as indicated in various sections of the report. This review was conducted by a collaboration of national laboratory and private sector participants with significant industrial experience, including some heavy lifting and deep diving from Clarus Consulting, LLC., all intended to achieve an objective, independent, and practical perspective. The report provides recommendations, descriptions of the evaluation methods, and the source references for the data used. To build confidence required for endorsement of the Code, this review was conducted as a verification and validation of the above Code contents. The objective of verification is to ensure that the Code is free of error – direct or implied; contains the information needed for its use, including proper coverage of the Code-specified materials for the intended application, and completeness and adequacy of references to other portions of the Code. The objective of validation is to authenticate that the Code tabulations and graphs represent design inputs consistent with what are determined using rules and methods specified by the Code. The authentication process used data that were assembled and/or generated independent of Code development, while the methods of analysis followed Code-specified methods where appropriate. The designated portions for this review cover the five alloys codified for high temperature reactor applications in Division 5, i.e. 316 SS, 304 SS, 800H, 2¼Cr-1Mo, and 9Cr-1Mo-V, regarding their general requirements, permitted specifications and design stress intensity values for pressure-retaining applications, deterioration in service, fatigue acceptance test, permissible weld materials, tensile and yield strength, expected minimum stress-to-rupture values (including for Alloy 718), weld stress rupture factors, permissible materials for bolting use, and restricted specifications in certain service applications. Additionally, stress intensity values for bolting materials including 316 SS, 304 SS and alloy 718 were reviewed. Analysis and discussion are also provided on contents outside of these designated Code portions where it was deemed relevant and necessary to develop a technically sound understanding of issues relating to the designated portions. Due to unavailability of sufficient test data on welds during the review period, the weld stress rupture factors in Tables HBB-I-10.14A to E, which cover a total of ten tables for the five alloys welded with twenty-eight different weld metals (some with similar properties), have been deferred to a future review effort. The review identified mainly two types of issues. The first type includes instances where the Code is found factually incomplete or incorrect, such as obsolete materials specifications listings, missing tabulation of stresses for bolting. Changes to the Code are recommended in these cases. The second type of issue includes instances where the Code tabulations and graphs are found to be less conservative than the review analysis results. In these cases, recommendations are made for further review and consideration where the difference in conservatism exceeds 10%, which is our threshold for questioning technical adequacy, meriting a risk assessment by the Nuclear Regulatory Commission and/or reactor designers. It is noted that this effort has been executed using all available data and established methods of analysis, including methods and criteria specified and used by the Code. As such, the findings that are presented in quantitative detail, in a format for convenient comparison with the Code, and with identification of where further review is recommended, should provide a sound technical basis for decisions about quantifying the implications of the reduced design margins and technical adequacy/inadequacy to form a basis for conditioning specific Code tabulation values on endorsement. Recommendations for specific changes to the Code, however, entail design conservatism considerations beyond the scope of this review effort, and are not made in this report.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Block Island Noise Modeling Data

Noise propagation near Block Island was simulated to assess environmental impacts of impact pile driving during wind turbine construction. Computational models complement field-recorded acoustic data, providing insights into sound attenuation, spectral variability, and propagation dynamics. The dataset includes: 1. Propagation Models: Simulated underwater sound fields documenting sound pressure and directional variability across frequency bands and distances. 2. Spectral Analysis (LTSA): Long-term averages and processed outputs calculating acoustic intensity over time. 3. Visualization Files: Graphs, 2D/3D simulation results, and reference calculations used in sound modeling.

17 WIND ENERGY↗

A Scalable Parallel Hypergraph Generator (HyGen)

Graphs are extensively used to model real-world complex systems. An edge in a graph can model pairwise relationships. However, multiway relationships (connections between three or more vertices) are common in many complex systems such as cellular process, image segmentation, and circuit design. A graph edge cannot model multiway relationships. A hypergraph, which can connect more than two vertices, is thus a better option to model multiway relationships. A large-scale hypergraph analysis has the potential to find useful insights from a complex system and assist in knowledge discovery. Currently a limited number of hypergraphs exists that are representative of real-world datasets. Moreover, real-world hypergraph datasets are small in size and inadequate to incorporate future needs. A graph generator that can produce large-scale synthetic hypergraphs can solve the above mentioned problems. In this paper, we present a scalable parallel hypergraph generator (HyGen) based on the Message Passing Interface (MPI) standard. To generate hypergraphs, HyGen takes the following parameter values as inputs: i) number of vertices, ii) number of hyperedges, iii) number of clusters, iv) vertex distribution, v) hyperedge distribution, vi) local cluster cardinality, and vii) global cluster cardinality. We have demonstrated that HyGen can generate hypergraphs of various sizes in a scalable fashion. HyGen takes approximately four minutes to generate a hypergraph with 4.8 million vertices, 1.6 million hyperedges, and 800 clusters using 1,024 processes on a leadership class computing platform. Our strong and weak scaling experiments on supercomputers demonstrate that HyGen can quickly create large-scale hypergraphs in a parallel manner, thus providing a useful capability for hypergraph analysis.

Hasan, S M Shamimul↗

Exact-WKB, complete resurgent structure, and mixed anomaly in quantum mechanics on S 1

We investigate the exact-WKB analysis for quantum mechanics in a periodic potential, with N minima on S 1 . We describe the Stokes graphs of a general potential problem as a network of Airy-type or degenerate Weber-type building blocks, and provide a dictionary between the two. The two formulations are equivalent, but with their own pros and cons. Exact-WKB produces the quantization condition consistent with the known conjectures and mixed anomaly. The quantization condition for the case of N-minima on the circle factorizes over the Hilbert sub-spaces labeled by discrete theta angle (or Bloch momenta), and is consistent with ’t Hooft anomaly for even N and global inconsistency for odd N. By using Delabaere-Dillinger-Pham formula, we prove that the resurgent structure is closed in these Hilbert subspaces, built on discrete theta vacua, and by a transformation, this implies that fixed topological sectors (columns of resurgence triangle) are also closed under resurgence.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Integrated data-driven and experimental approaches to accelerate lead optimization targeting SARS-CoV- 2 main protease

Identification of potential therapeutic candidates can be expedited by integrating computational modeling with domain aware machine learning (ML) approaches followed by experimental validation. Generative deep learning models have been recently developed that can generate thousands of new candidates, but their physiochemical properties are typically not optimized. Using our deep learning models and a scaffold as a starting point, we generated tens of thousands of compounds for SARS-CoV-2 M pro that preserve the core scaffold. Here we utilized and implemented several computational tools such as structural alert and toxicity analysis, high throughput virtual screening, ML-based 3D quantitative structure–activity relationships, multi-parameter optimization, and graph neural networks on libraries of generated candidates to predict biological activity and binding affinity a priori. From these collective computational results, eight promising candidates were identified and tested experimentally using Native Mass Spectrometry (MS) and FRET-based functional assays. Two compounds, with quinazoline-2-thiol and acetylpiperidine core moiety showed IC 50 values in the low micromolar range: 2.95±0.0017 µM and 3.41±0.0015 µM, respectively. The molecular dynamics simulations further highlight that binding of these compounds results in allosteric modulations in the chain B and the interface domains of the M pro . The key fragments from these top hits can be used as input for closed loop lead optimization in the integrated pipeline.

60 APPLIED LIFE SCIENCES↗

Community detection robustness of graph neural networks

Graph neural networks (GNNs) are increasingly widely used for community detection in attributed networks. They combine structural topology with node attributes through message passing and pooling. However, their robustness or lack thereof with respect to different perturbations and targeted attacks in conjunction with community detection tasks is not well understood. To shed light on latent mechanisms behind GNN sensitivity on community detection tasks, we conduct a systematic computational evaluation of six widely adopted GNN architectures graph convolutional network, graph attention network, graph sample and aggregate (GraphSAGE), differentiable pooling (DiffPool), minimum cut pooling (MinCUT), and deep modularity networks (DMoN). The analysis covers three perturbation categories: node attribute manipulations, edge topology distortions, and adversarial attacks. We use element-centric similarity as the evaluation metric on synthetic benchmarks and real-world citation networks. Our findings indicate that supervised GNNs tend to achieve higher baseline accuracy, while unsupervised methods, particularly DMoN, maintain stronger resilience under targeted and adversarial perturbations. Furthermore, robustness appears to be strongly influenced by community strength, with well-defined communities reducing performance loss. Across all models, node attribute perturbations associated with targeted edge deletions and shifts in attribute distributions tend to cause the largest degradation in community recovery. These findings highlight important trade-offs between accuracy and robustness in GNN-based community detection and offer insights into selecting architectures resilient to noise and adversarial attacks.

Goel, Jaidev [Virginia Polytechnic Inst. and State↗

Mortar: An Open Testbed for Portable Building Analytics

Access to large amounts of real-world data has long been a barrier to the development and evaluation of analytics applications for the built environment. Open datasets exist, but they are limited in their span (how much data is available) and context (what kind of data is available and how it is described). Evaluation of such analytics is also limited by how the analytics themselves are implemented, often using hard-coded names of building components, points and locations, or unique input data formats. To advance the methodology for how such analytics are implemented and evaluated, we present Mortar: an open testbed for portable building analytics, currently spanning 90 buildings and containing over 9.1 billion data points. All buildings in the testbed are described using Brick, a recently developed metadata schema, providing rich functional descriptions of building assets and subsystems. We also propose a simple architecture for writing portable analytics applications that are robust to the diversity of buildings and can configure themselves based on context. We demonstrate the utility of Mortar by implementing 11 applications from the literature.

47 OTHER INSTRUMENTATION↗

Radiological Monitoring Plan for the Oak Ridge Y-12 National Security Complex: Surface Water

DOE Order 458.1 requires that dose estimates consider contributions from all facilities. In the Y-12 Radiological Monitoring Plan (RMP), surface water is monitored at points that reflect individual facilities, as well as at points that reflect the combined contributions of all facilities. This monitoring plan does not consider other potential routes (i.e., airborne releases and food chains). Thus, a complete determination of total effective dose (TED) cannot be made based on this plan alone. The other routes from Y-12, and all routes from other DOE facilities on the Oak Ridge Reservation (e.g., Oak Ridge National Laboratory (ORNL) and The Heritage Center), must be considered in order to satisfy DOE Order 458.1 requirements. Determination of TED from all sites and pathways is done through the use of dose-assessment models and is documented in the Annual Site Environmental Report. This monitoring plan provides adequate monitoring goals for Y-12 surface water releases to provide input of sufficient sensitivity and accuracy to reliably determine the Y-12 surface water component of the TED. The routine radiological monitoring program is designed to monitor effluents at four types of locations: (1) treatment facilities, (2) other point and area source discharges, (3) instream locations, and (4) production building roof run-off. With this sampling and analysis program, data will be obtained on primary point sources as well as on locations that represent the composite of other potential sources. This plan will be reviewed periodically to determine necessary modifications to the sampling frequencies, parameters, and locations. Modifications, if any, will be based on the analysis of the previous data and its effectiveness in satisfying the objectives of this plan. Appendix A contains graphs of the sum of the DCS fractions for locations and frequencies contained in a previous version of this plan. The data was collected from January 2009 through December 2019. Each sample was analyzed, and each result was divided by the appropriate DCS to compute a DCS fraction. These fractions were summed for all isotopes. According to DOE –STD-1196-2011, the annual average of these sums should be below 1.

54 ENVIRONMENTAL SCIENCES↗

Exploring Multilayer Network Models to Build a Scientific Basis for Integrated Deterrence: Final Report

The emerging multipolar international security environment represents a fundamental restructuring of global nuclear balance of power to include two nuclear peer competitors, growing non-peer nuclear threats, and concerns of nuclear latency from both allies and adversaries. Conflicts in the grey zone, cyber operations, mis- and disinformation campaigns, and emerging disruptive technologies like drones, and hypersonic missiles are becoming more prevalent. These present a risk of cross-domain and multi-domain conflicts that may not follow known escalatory patterns. In order to prepare for the new deterrence environment, it is critical to have quantitative and qualitative understandings of these cross-domain conflicts, their potential for escalation, and which systems they may impact. To that end, our team created a Multi-Layer Network (MLN) model of ‘integrated deterrence’ where instruments of national power are modeled as individual network graph layers that include efforts from all domains. We then evaluate the potential for escalation against escalation scenarios. Analysis of the escalation scenarios is then used to identify insights of potential risk and escalation within integrated deterrence.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Digital Twin for Optimizing Real-time Economy of the Integrated Energy Systems

Economic and safe operation of integrated energy systems (IES) requires real-time optimization (RTO) of the control and actions conducted on each system component. In this regard, digital twins (DTs), which consist of a physical system, a virtual system, and the data communication that occurs between the two, are essential for effective RTO. Through the data warehouse, the virtual system is constantly updated with real-time data from the physical system, and functions as the model in the optimization framework. The reduced-order model of the dynamic process model in the virtual system is used in the optimization framework. The optimization results are then returned, via the data warehouse, as control actions to the physical system. This work demonstrates the software capabilities of DT assets for an IES in the context of preparing a DT for an experimental system comprised of Idaho National Laboratory (INL)’s Thermal Energy Delivery System and battery system. For the virtual demonstration, the DTs encompass (1) a physical system, including the Modelica models of the Thermal Energy Delivery System and the battery system; (2) virtual optimization via the Optimization of Real-Time Capacity Allocation (ORCA) platform; and (3) the open-source data warehouse software DeepLynx. This work assesses the performance of ORCA, which utilizes a reduced-order model built using the Risk Analysis Virtual Environment (RAVEN) and trained on the Modelica models and real-time data pipeline through the graph database hosted in DeepLynx. The proposed optimization workflow will be an RTO model based on DTs and the data they generate.

25 ENERGY STORAGE↗

The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species

Abstract Bridging the gap between genetic variations, environmental determinants, and phenotypic outcomes is critical for supporting clinical diagnosis and understanding mechanisms of diseases. It requires integrating open data at a global scale. The Monarch Initiative advances these goals by developing open ontologies, semantic data models, and knowledge graphs for translational research. The Monarch App is an integrated platform combining data about genes, phenotypes, and diseases across species. Monarch's APIs enable access to carefully curated datasets and advanced analysis tools that support the understanding and diagnosis of disease for diverse applications such as variant prioritization, deep phenotyping, and patient profile-matching. We have migrated our system into a scalable, cloud-based infrastructure; simplified Monarch's data ingestion and knowledge graph integration systems; enhanced data mapping and integration standards; and developed a new user interface with novel search and graph navigation features. Furthermore, we advanced Monarch's analytic tools by developing a customized plugin for OpenAI’s ChatGPT to increase the reliability of its responses about phenotypic data, allowing us to interrogate the knowledge in the Monarch graph using state-of-the-art Large Language Models. The resources of the Monarch Initiative can be found at monarchinitiative.org and its corresponding code repository at github.com/monarch-initiative/monarch-app.

60 APPLIED LIFE SCIENCES↗

A Data Processing Pipeline for Socio-Technical Network Analysis [Slides]

With the rapid adoption of emerging technologies, there is a need to catalog and model sociotechnical interdependencies that have been historically used to influence the operation of Critical Infrastructure networks including the impacts of mergers and acquisitions, hostile takeovers, and foreign investment. Our research intends to address this need with two primary contributions. First, we have developed a data curation and processing pipeline to generate sociotechnical networks extracted from a variety of data sources including SEC filings and infrastructure asset databases. The pipeline, implemented in Apache Airflow, extracts and normalizes the representation of entities and relations, specified within ontologies. Second, networks produced by our pipeline enable the development of graph-theoretic metrics that consider the properties of network components in addition to its topology. Measures of network complexity, such as degree distribution, reachability analyses, temporal analysis, and community detection may be adapted to indicate adversarial organizational influence. Our intent is to provide an extensible, machine-actionable approach to quickly communicate such models, reproduce previous results, and adapt them to new, unanticipated situations.

97 MATHEMATICS AND COMPUTING↗

Harnessing graph convolutional neural networks for identification of glassy states in metallic glasses

Graph Convolutional Neural Networks (GCNNs) have emerged as powerful tools for analyzing materials. In this study, we employ GCNNs to examine structural characteristics of CuZr metallic glasses (MGs) and identify their states. We use molecular dynamics to simulate the quenching process of CuZr, using cooling rates ranging from 10 9 to 10 15 K/s, to produce six unique glassy states. For each state, we create a dataset comprising 1,800 distinct samples. We evaluate the effectiveness of various GCNNs, including Graph Attention Neural Network (GANN), Graph Sample and AggreGatE (GraphSAGE), Graph Isomorphism Network (GIN), and Relational Graph Convolutional Neural Network (RGCN). GANN and GraphSAGE demonstrate comparable performance, achieving an overall accuracy of 81% in classifying the MG states. Furthermore, these results underscore the potential of GCNNs to detect subtle structural variances in disordered materials and point to broader application of deep learning in the analysis of MGs and other amorphous substances.

36 MATERIALS SCIENCE↗