Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “graph processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Graph neural networks for CO 2 solubility predictions in Deep Eutectic Solvents

Deep Eutectic Solvents (DESs) are a promising class of solvents for CO 2 capture. DESs are complex mixtures that can be designed to optimize CO solubility and overall capture process efficiency. However, the vast design landscape of DES mixtures makes experimental investigation prohibitive; as such, there is a need for computational models that can quickly and efficiently navigate the design space and inform data collection efforts. In this work, we propose Graph Neural Network (GNN) models for predicting CO 2 solubility for DESs; the GNN leverages a mixture graph representation that captures the molecular structure of the DES components as well as their intermolecular interactions. Here, we compare the GNN framework against alternative architectures (neural networks, graph convolution networks, and random forests) and data representations (molecular fingerprints, sigma profiles, and graphs). We show that the proposed approach offers superior predictive performance; specifically, we show that solubility can be predicted reliably directly from molecular structure (without the need of using sigma profiles as proposed in previous studies). This result is important, as obtaining sigma profiles requires expensive density functional theory computations. We also explored the ability of GNNs to predict solubility for new DES mixtures and operating conditions. We found that the model extrapolates across temperature reliably. However, we also found deficiencies in the ability of the model to predict solubility for DES mixtures, pressures, and molar ratio not included in the training sets; we show that this is due to an inherent lack of chemical diversity in datasets available in the literature. The proposed computational capabilities can thus help navigate the design space of DES and inform data collection efforts. Our models, data, and benchmarks are shared as Python code implemented in Jupyter notebooks.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Data-Driven Template Discovery Using Graph Convolutional Neural Networks

Modeling adversarial activities is a critical component of developing high-con?dence indicators of efforts to acquire, fabricate, proliferate, and/or deploy weapons of mass terror (WMTs). Current approaches to generating representative patterns of interest (a.k.a templates) from the real-world domains involve a Subject Matter Expert (SME)-guided manual process. The goal of Data-Driven Template Discovery (DDTD) is to use a (potentially small) set of SME generated templates to discover other previously unknown and interesting templates in an attributed graph. A template is an activity pattern describing a set of interactions among a group of nodes in the graph. The motivation behind DDTD is to expand the original set of templates, without having SMEs craft all the templates by hand. DDTD also provides seed templates to SMEs, to help them construct larger, high-?delity, and scenario-oriented templates. In these cases, obtaining a larger set of templates that are related (contain similar signals) to the original set is of great value. In this work, we propose to use Graph Convolutional Neural Networks (GCNs) to discover new templates that are heavily related to the original set. GCNs are a family of Neural Network (NN) architectures especially designed to work directly on graphs. In contrast to the traditional NNs, that require considerable amounts of labeled data, GCNs do not require a big labeled training set because they can directly leverage the graph structure instead. This property makes GCNs the perfect tool for creating activity templates.

Joaristi, Mikel↗

GCoD: Graph Convolutional Network Acceleration via Dedicated Algorithm and Accelerator Co-Design

Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art graph learning model. However, it remains notoriously challenging to inference GCNs over large graph datasets, limiting their application to large real-world graphs and hindering the exploration of deeper and more sophisticated GCN graphs. This is because real-world graphs can be extremely large and sparse. Furthermore, the node degree of GCNs tends to follow the power-law distribution and therefore have highly irregular adjacency matrices, resulting in prohibitive inefficiencies in both data processing and movement and thus substantially limiting the achievable GCN acceleration efficiency. To this end, this paper proposes the first GCN algorithm and accelerator Co-Design framework dubbed GCoD which can largely alleviate the aforementioned GCN irregularity and boost GCNs' inference efficiency. Specifically, on the algorithm level, GCoD integrates a divide and conquer GCN training strategy that polarizes the graphs to be either denser or sparser in local neighborhoods without compromising the model accuracy, resulting in graph adjacency matrices that (mostly) have merely two levels of workload and enjoys largely enhanced regularity and thus ease of acceleration. On the hardware level, we further develop a dedicated two-pronged accelerator with a separated engine to process each of the aforementioned workloads, further boosting the overall utilization and acceleration efficiency. Extensive experiments and ablation studies validate that our GCoD consistently outperforms state-of-the-art designs in terms of accelerator efficiency while maintaining or even improving the task accuracy. Additionally, we visualize GCoD trained graph adjacency matrices to better understand its advantages. All codes and pre-trained models will be released upon acceptance.

You, Haoran↗

Oak Ridge National Laboratory Technical Input for the Nuclear Regulatory Commission Review of the 2017 Edition of ASME Section III, Division 5, ‘High Temperature Reactors’

To assist the Nuclear Regulatory Commission in its decision making on endorsement of the American Society for Mechanical Engineers Boiler and Pressure Vessel Code Section III, Division 5 (2017 Edition) for development of advanced non-light water reactors, the following Division 5 portions were reviewed: Article HBB-2000 Material; Article HCB-2000 Material; Article HGB-2000 Material; Mandatory Appendix HBB-I-14 Tables and Figures; and, Nonmandatory Appendix HBB-U Guidelines for Restricted Material Specifications to Improve Performance in Certain Service Applications. In addition to the 2017 Edition, the same parts of the 2019 Edition have also been reviewed as indicated in various sections of the report. This review was conducted by a collaboration of national laboratory and private sector participants with significant industrial experience, including some heavy lifting and deep diving from Clarus Consulting, LLC., all intended to achieve an objective, independent, and practical perspective. The report provides recommendations, descriptions of the evaluation methods, and the source references for the data used. To build confidence required for endorsement of the Code, this review was conducted as a verification and validation of the above Code contents. The objective of verification is to ensure that the Code is free of error – direct or implied; contains the information needed for its use, including proper coverage of the Code-specified materials for the intended application, and completeness and adequacy of references to other portions of the Code. The objective of validation is to authenticate that the Code tabulations and graphs represent design inputs consistent with what are determined using rules and methods specified by the Code. The authentication process used data that were assembled and/or generated independent of Code development, while the methods of analysis followed Code-specified methods where appropriate. The designated portions for this review cover the five alloys codified for high temperature reactor applications in Division 5, i.e. 316 SS, 304 SS, 800H, 2¼Cr-1Mo, and 9Cr-1Mo-V, regarding their general requirements, permitted specifications and design stress intensity values for pressure-retaining applications, deterioration in service, fatigue acceptance test, permissible weld materials, tensile and yield strength, expected minimum stress-to-rupture values (including for Alloy 718), weld stress rupture factors, permissible materials for bolting use, and restricted specifications in certain service applications. Additionally, stress intensity values for bolting materials including 316 SS, 304 SS and alloy 718 were reviewed. Analysis and discussion are also provided on contents outside of these designated Code portions where it was deemed relevant and necessary to develop a technically sound understanding of issues relating to the designated portions. Due to unavailability of sufficient test data on welds during the review period, the weld stress rupture factors in Tables HBB-I-10.14A to E, which cover a total of ten tables for the five alloys welded with twenty-eight different weld metals (some with similar properties), have been deferred to a future review effort. The review identified mainly two types of issues. The first type includes instances where the Code is found factually incomplete or incorrect, such as obsolete materials specifications listings, missing tabulation of stresses for bolting. Changes to the Code are recommended in these cases. The second type of issue includes instances where the Code tabulations and graphs are found to be less conservative than the review analysis results. In these cases, recommendations are made for further review and consideration where the difference in conservatism exceeds 10%, which is our threshold for questioning technical adequacy, meriting a risk assessment by the Nuclear Regulatory Commission and/or reactor designers. It is noted that this effort has been executed using all available data and established methods of analysis, including methods and criteria specified and used by the Code. As such, the findings that are presented in quantitative detail, in a format for convenient comparison with the Code, and with identification of where further review is recommended, should provide a sound technical basis for decisions about quantifying the implications of the reduced design margins and technical adequacy/inadequacy to form a basis for conditioning specific Code tabulation values on endorsement. Recommendations for specific changes to the Code, however, entail design conservatism considerations beyond the scope of this review effort, and are not made in this report.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Finding diverse ways to improve algebraic connectivity through multi-start optimization

The algebraic connectivity, also known as the Fiedler value, is a spectral measure of network connectivity that can be increased through edge addition. We present an algorithm for producing many diverse ways to add a fixed number of edges to a network to achieve a near optimal Fiedler value. Previous Fielder value optimization algorithms (i.e. the greedy algorithm) output only one solution. Obtaining a single solution is rarely good enough for real-world network redesign problems, as practical constraints (political, physical or financial) may prevent implementation. Our algorithm takes a multi-start optimization approach, adding a random initial edge and then applies a greedy heuristic to improve the Fiedler value. The random choice moves us to a new region of the search space, enabling discovery of diverse solutions. Additionally, we present a Determinantal Point Process framework for quantifying diversity. We then apply a Markov chain Monte Carlo technique to sift through the large number of output solutions and locate a smaller, more manageable collection of highly diverse solutions that can be presented to network redesign engineers. We demonstrate the effectiveness of our algorithm on real-world graphs with varied structures.

97 MATHEMATICS AND COMPUTING↗

A Mass‐Conserving‐Perceptron for Machine‐Learning‐Based Modeling of Geoscientific Systems

Although decades of effort have been devoted to building Physical-Conceptual (PC) models for predicting the time-series evolution of geoscientific systems, recent work shows that Machine Learning (ML) based Gated Recurrent Neural Network technology can be used to develop models that are much more accurate. However, the difficulty of extracting physical understanding from ML-based models complicates their utility for enhancing scientific knowledge regarding system structure and function. Here, we propose a physically interpretable Mass-Conserving-Perceptron (MCP) as a way to bridge the gap between PC-based and ML-based modeling approaches. The MCP exploits the inherent isomorphism between the directed graph structures underlying both PC models and GRNNs to explicitly represent the mass-conserving nature of physical processes while enabling the functional nature of such processes to be directly learned (in an interpretable manner) from available data using off-the-shelf ML technology. As a proof of concept, we investigate the functional expressivity (capacity) of the MCP, explore its ability to parsimoniously represent the rainfall-runoff (RR) dynamics of the Leaf River Basin, and demonstrate its utility for scientific hypothesis testing. To conclude, we discuss extensions of the concept to enable ML-based physical-conceptual representation of the coupled nature of mass-energy-information flows through geoscientific systems.

58 GEOSCIENCES↗

Graph Convolutional Network-Based Topology Embedded Deep Reinforcement Learning for Voltage Stability Control

Topological variations in power system is a common phenomenon and can impose significant challenges to traditional controllers of power system. Recent study revealed the strength of deep reinforcement learning (DRL) based approaches in power system preventive and corrective control. But topological variations are difficult to capture using classical fully connected neural network (FCN) model and has not been explicitly modeled in previous work. Hence, we develop a Graph Convolutional Network (GCN) based DRL framework to tackle topology changes in control design of power system. The GCN model exploits the graph structure of the power network and helps the DRL agent to embed the topology information during learning process. Our GCN based approach is evaluated using the IEEE-39 bus system and it outperforms the FCN-based DRL scheme in terms of training convergence and control performance considering grid topology changes.

Hossain, Ramij Raja↗

Automating Log Synthesis and Visualization with Python and Splunk

The goal of this project is to automate log analysis by utilizing Splunk, Bash, and Python together. Simplifying the monitoring and analysis of network traffic was the main goal. In order to accomplish this, a Bash script was created to use 'tcpdump' to automate network sniffing. It also included a 24-hour file rotation mechanism to effectively manage the pcap files that were generated. After that, a Python script was written to read these pcap files and retrieve pertinent data about network traffic. After processing the collected data, Splunk is used to summarize the important metrics and visualize said information with relevant graphs.

99 GENERAL AND MISCELLANEOUS↗

HEPOM: Using Graph Neural Networks for the Accelerated Predictions of Hydrolysis Free Energies in Different pH Conditions

Hydrolysis is a fundamental family of chemical reactions where water facilitates the cleavage of bonds. The process is ubiquitous in biological and chemical systems, owing to water’s remarkable versatility as a solvent. However, accurately predicting the feasibility of hydrolysis through computational techniques is a difficult task, as subtle changes in reactant structure like heteroatom substitutions or neighboring functional groups can influence the reaction outcome. Furthermore, hydrolysis is sensitive to the pH of the aqueous medium, and the same reaction can have different reaction properties at different pH conditions. In this work, we have combined reaction templates and high-throughput ab initio calculations to construct a diverse data set of hydrolysis free energies. The developed framework automatically identifies reaction centers, generates hydrolysis products, and utilizes a trained graph neural network (GNN) model to predict ΔG values for all potential hydrolysis reactions in a given molecule. The long-term goal of the work is to develop a data-driven, computational tool for high-throughput screening of pH-specific hydrolytic stability and the rapid prediction of reaction products, which can then be applied in a wide array of applications including chemical recycling of polymers and ion-conducting membranes for clean energy generation and storage.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Structure-aware graph neural network based deep transfer learning framework for enhanced predictive analytics on diverse materials datasets

Abstract Modern data mining methods have demonstrated effectiveness in comprehending and predicting materials properties. An essential component in the process of materials discovery is to know which material(s) will possess desirable properties. For many materials properties, performing experiments and density functional theory computations are costly and time-consuming. Hence, it is challenging to build accurate predictive models for such properties using conventional data mining methods due to the small amount of available data. Here we present a framework for materials property prediction tasks using structure information that leverages graph neural network-based architecture along with deep-transfer-learning techniques to drastically improve the model’s predictive ability on diverse materials (3D/2D, inorganic/organic, computational/experimental) data. We evaluated the proposed framework in cross-property and cross-materials class scenarios using 115 datasets to find that transfer learning models outperform the models trained from scratch in 104 cases, i.e., ≈90%, with additional benefits in performance for extrapolation problems. We believe the proposed framework can be widely useful in accelerating materials discovery in materials science.

Chemistry↗

Expediting field-effect transistor chemical sensor design with neuromorphic spiking graph neural networks

Improving the sensitive and selective detection of analytes in a variety of applications requires accelerating the rational design of field-effect transistor (FET) chemical sensors. Achieving high-performance detection relies on identifying optimal probe materials that can effectively interact with target analytes, a process traditionally driven by chemical intuition and time-consuming trial-and-error methods. To address the difficulties in probe screening for FET sensor development, this work presents a methodology that combines neuromorphic machine learning (ML) architectures, specifically a hybrid spiking graph neural network (SGNN), with an enriched dataset of physicochemical properties through semi-automated data extraction using large language models. Achieving a classification accuracy of 0.89 in predicting sensor sensitivity categories, the SGNN model outperformed traditional ML techniques by leveraging its ability to capture both global physicochemical properties and sparse topological features through a hybrid modeling framework. Next-generation sensor design was informed by the actionable insights into the connections between material properties and sensing performance offered by the SGNN framework. Through virtual screening for the detection of per- and polyfluoroalkyl substances (PFAS) as a use case, the effectiveness of the SGNN model was further validated. Density functional theory simulations confirmed graphene as a promising active material for PFAS detection as suggested by the SGNN framework. By bridging gaps in predictive modeling and data availability, this integrated approach provides a strong foundation for accelerating advancements in FET sensor design and innovation.

Ferreira, Rodrigo Pires [Univ. of Chicago, IL (Uni↗

Understanding Metal–Organic Framework Nucleation from a Solution with Evolving Graphs

A mechanistic understanding of metal–organic framework (MOF) synthesis and scale-up remains underexplored due to the complex nature of the interactions of their building blocks. In this work, we investigate the collective assembly of building units at the early stages of MOF nucleation, using MIL-101(Cr) as a prototypical example. Using large-scale molecular dynamics simulations, we observe that the choice of solvent (water and N,N-dimethylformamide), the introduction of ions (Na+ and F–) and the relative populations of MIL-101(Cr) half-secondary building unit (half-SBU) isomers have a strong influence on the cluster formation process. Additionally, the shape, size, nucleation and growth rates, crystallinity, and short and long-range order largely vary depending on the synthesis conditions. We evaluate these properties as they naturally emerge when interpreting the self-assembly of MOF nuclei as the time evolution of an undirected graph. Solution-induced conformational complexity and ionic concentration have a dramatic effect on the morphology of clusters emerging during assembly. While pure solvents lead to the rapid formation of a small number of large clusters, the presence of ions in aqueous solutions results in smaller clusters and slower nucleation. This diversity is captured by the key features of the graph representation. Principle component analysis on graph properties reveals that only a small number of molecular descriptors is needed to deconvolute MOF self-assembly. Furthermore, descriptors such as the average coordination number between half-SBUs and fractal dimension are of particular interest as they can be can be followed experimentally by techniques like by time-resolved spectroscopy. Ultimately, graph theory emerges as an approach that can be used to understand complex processes revealing molecular descriptors accessible by both simulation and experiment.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Graph link prediction in computer networks using Poisson matrix factorisation

Graph link prediction is an important task in cybersecurity: relationships between entities within a computer network, such as users interacting with computers or system libraries and the corresponding processes that use them, can provide key insights into adversary behaviour. Poisson matrix factorisation (PMF) is a popular model for link prediction in large networks, particularly useful for its scalability. In this article PMF is extended to include scenarios that are commonly encountered in cybersecurity applications. Specifically, an extension is proposed to explicitly handle binary adjacency matrices and include known categorical covariates associated with the graph nodes. A seasonal PMF model is also presented to handle seasonal networks. To allow the methods to scale to large graphs, variational methods are discussed for performing fast inference. The results show an improved performance over the standard PMF model and other statistical network models.

97 MATHEMATICS AND COMPUTING↗

Development of Multimodal Few-Shot Analytics for Electron Micrographs

Recent advances in materials data analytics have provided new avenues for determining process-structure-property (PSP) linkages in a variety of materials. Machine learning techniques including few-shot learning have increased the efficiency of classifying microscopy images for the purposes of material characterization. Attempts at creating a multimodal approach can provide further improvements to current models and help extract more salient features from data. In this vein, raw spectrum data was taken to provide an additional modality to our current pyCHIP classifier. Modifications in segmentation also show potential in improving the accuracy of the pyCHIP classifier. Classifier output was analyzed using network graphs and unsupervised clustering algorithms such as spectral clustering to detect better segmentation methods than the current “chipping” approach. We suggest that the chip selection process can be automated in the future using a combination of these techniques to enable high-throughput analyses.

36 MATERIALS SCIENCE↗

GMFOLD: Subgraph matching for high-throughput DNA-aptamer secondary structure classification and machine learning interpretability

Aptamers are oligonucleotide receptors that bind to their targets with high affinity. Here, we consider aptamers comprised of single-stranded DNA that undergo target-binding-induced conformational changes, giving rise to unique secondary and tertiary structures. Given a specific aptamer primary sequence, there are well-established computational tools (notably mfold) to predict the secondary structure via free energy minimization algorithms. While mfold generates secondary structures for individual sequences, there is a need for a high-throughput process whereby thousands of DNA structures can be predicted in real-time for use in an interactive setting, when combined with aptamer selections that generate candidate pools that are too large to be experimentally interrogated. We developed a new Python code for high-throughput aptamer secondary structure determination (GMfold). GMfold uses subgraph matching methods to group aptamer candidates by secondary structure similarities. We also improve an open-source code, SeqFold, to incorporate subgraph matching concepts. We represent each secondary structure as a lowest-energy bipartite subgraph matching of the DNA graph to itself. These new tools enable thousands of DNA sequences to be compared based on their secondary structures, using machine-learning algorithms. This process is advantageous when analyzing sequences that arise from aptamer selections via systematic evolution of ligands by exponential enrichment (SELEX). This work is a building block for future machine-learning-informed DNA-aptamer selection processes to identify aptamers with improved target affinity and selectivity and advance aptamer biosensors and therapeutics.

Aptamer↗

Conceptual design of inverted core lead bismuth eutectic fast reactor for marine applications

The development of an inverted core fast reactor aims to generate 60 MWth for about 30 Effective Full Power Years without refueling. The reactor design is a transportable reactor using UO{sub 2} fuel and lead-bismuth-eutectic cooled designed for marine applications and is intended to improve the reactor performances compared to the normal core design: better condition for passive cooling system capability by lower core pressure drop, taking advantage of potential power uprate from the lower maximum fuel temperature. Systematic design processes are presented in this work: fuel pin geometry selection, fuel assembly (FA) design, and core design. A relationship between pressure drops, coolant velocity, maximum fuel temperature, coolant channel diameter, and fuel volume fraction was introduced in a single graph used as a tool to select fuel pin geometry. Fuel fabrication capability also took place in consideration of FA design which led to 7 holes per FA, and two-dimensional temperature distribution studies were also carried out. Core design processes including radial zoning, axial zoning, and core optimization were conducted using Monte Carlo code MCS, which is UNIST CORE laboratory in-house code. The current core design uses 3 fuel enrichment levels and 3 FA types to control the local power distribution and power shift during its lifetime. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Harnessing graph convolutional neural networks for identification of glassy states in metallic glasses

Graph Convolutional Neural Networks (GCNNs) have emerged as powerful tools for analyzing materials. In this study, we employ GCNNs to examine structural characteristics of CuZr metallic glasses (MGs) and identify their states. We use molecular dynamics to simulate the quenching process of CuZr, using cooling rates ranging from 10 9 to 10 15 K/s, to produce six unique glassy states. For each state, we create a dataset comprising 1,800 distinct samples. We evaluate the effectiveness of various GCNNs, including Graph Attention Neural Network (GANN), Graph Sample and AggreGatE (GraphSAGE), Graph Isomorphism Network (GIN), and Relational Graph Convolutional Neural Network (RGCN). GANN and GraphSAGE demonstrate comparable performance, achieving an overall accuracy of 81% in classifying the MG states. Furthermore, these results underscore the potential of GCNNs to detect subtle structural variances in disordered materials and point to broader application of deep learning in the analysis of MGs and other amorphous substances.

36 MATERIALS SCIENCE↗

Development of message passing-based graph convolutional networks for classifying cancer pathology reports

Abstract Background Applying graph convolutional networks (GCN) to the classification of free-form natural language texts leveraged by graph-of-words features (TextGCN) was studied and confirmed to be an effective means of describing complex natural language texts. However, the text classification models based on the TextGCN possess weaknesses in terms of memory consumption and model dissemination and distribution. In this paper, we present a fast message passing network (FastMPN), implementing a GCN with message passing architecture that provides versatility and flexibility by allowing trainable node embedding and edge weights, helping the GCN model find the better solution. We applied the FastMPN model to the task of clinical information extraction from cancer pathology reports, extracting the following six properties: main site, subsite, laterality, histology, behavior, and grade. Results We evaluated the clinical task performance of the FastMPN models in terms of micro- and macro-averaged F1 scores. A comparison was performed with the multi-task convolutional neural network (MT-CNN) model. Results show that the FastMPN model is equivalent to or better than the MT-CNN. Conclusions Our implementation revealed that our FastMPN model, which is based on the PyTorch platform, can train a large corpus (667,290 training samples) with 202,373 unique words in less than 3 minutes per epoch using one NVIDIA V100 hardware accelerator. Our experiments demonstrated that using this implementation, the clinical task performance scores of information extraction related to tumors from cancer pathology reports were highly competitive.

59 BASIC BIOLOGICAL SCIENCES↗