Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Graph processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Anatomical Fraction segmentation in the biomass bales

The U.S. Department of Energy Bioenergy Technologies Office (BETO) is committed to the development of sustainable, nationwide, commercial biofuel production to displace petroleum-derived fuels, increase domestic energy production, and encourage the creation of a domestic bioenergy and bioproducts industry. Operating a commercial-scale bioenergy operation requires significant technological advancements for determining biomass feedstock quality during the preprocessing stage. Usage of non-food feedstocks- e.g., corn-stover, pine residue, and forest residue - at the biorefineries reduces strain on the food supply chain. But the non-food feedstock heterogeneity- physical size, shape, and chemical composition-poses a significant challenge during milling, conveyance, feeding, and biofuel conversion processes. 3D X-Ray imaging (CT) is capable of distinguishing material features by detecting density and compositional differences. But the inherent complexity introduced during harvesting and bailing makes the reconstruction and interpretation of baled biomass materials from x-ray data time consuming, laborious, and expensive. Feedstocks are low-density materials that produce small contrast differences in the x-ray images. This paper focuses on using the shape and texture properties, using 3D image processing techniques like 3D skeletonization, directional statistics to characterize and extract volumetric content of the different tissue samples in the biomass bales.

09 BIOMASS FUELS↗

Attention-Augmented Parametric Kernel Graph Neural Network (APKGNN) for Node Classification

We present a new graph neural network, the Attention-based Parametric-Kernel augmented Graph Neural Network (APKGNN), developed for node classification tasks. Despite extensive work on modeling multi-faceted relationships between connected nodes of a graph, the effect of attention on edge features mapped to relationships has not yet been analyzed through learning representation. This study derives such an attention vector by first calculating node features corresponding to endpoints of an edge and then aggregating these with extracted local intrinsic patches of a given graph to generate augmented local patch vectors. This process uses a parametric kernel based on Gaussian mixture models (GMMs) to embed local neighborhoods of the graph in local patches. The patch vectors then convolve with the above node features to produce an updated node representation. We show that this new learning representation (APKGNN) achieves higher node classification accuracy on tasks - both standard benchmarks (Cora, PubMed, Citeseer) and new experimental short text corpora where nodes correspond to text documents and words. This implementation of the GNN convolution layer outperforms state-of-the-art (SOTA) algorithms, achieving higher training, validation, and test accuracy by a significant margin on three standard benchmark data sets under both SOTA experimental settings and those for new testbeds.

Bose, Avishek↗

Artificial Intelligence Designer of Materials and Processes for Advanced Power Generation

In this presentation, ‘deep-freeze’ graphs, ‘convoluted filtering’ networks, ‘mirror-image’ graphs, and adversarial ensemble methods are utilized to support inversion modeling for optimization of the complex compositions and complex processes in design of high-performing alloys, with their properties tailored to the energy application specifications.

Romanov, Vyacheslav↗

Exploring the Value of Nodes with Multicommunity Membership for Classification with Graph Convolutional Neural Networks

Sampling is an important step in the machine learning process because it prioritizes samples that help the model best summarize the important concepts required for the task at hand. The process of determining the best sampling method has been rarely studied in the context of graph neural networks. In this paper, we evaluate multiple sampling methods (i.e., ascending and descending) that sample based off different definitions of centrality (i.e., Voterank, Pagerank, degree) to observe its relation with network topology. We find that no sampling method is superior across all network topologies. Additionally, we find situations where ascending sampling provides better classification scores, showing the strength of weak ties. Two strategies are then created to predict the best sampling method, one that observes the homogeneous connectivity of the nodes, and one that observes the network topology. In both methods, we are able to evaluate the best sampling direction consistently.

97 MATHEMATICS AND COMPUTING↗

Improved Distributed-memory Triangle Counting by Exploiting the Graph Structure

Graphs are ubiquitous in modeling complex systems and representing interactions between entities to uncover structural information of the domain. Traditionally, graph analytics workloads are challenging to efficiently scale (both strong and weak cases) on distributed memory due to the irregular memory-access driven nature (with little or no computations) of the methods. The structure of graphs and their relative distribution over the processing elements poses another level of complexity, making it difficult to attain sustainable scalability across platforms. In this paper, we discuss enhancements to TriC, a distributed-memory implementation of graph triangle counting using Message Passing Interface (MPI), which was featured in the 2020 Graph Challenge competition. We have made some incremental enhancements to TriC, primarily adopting a user-defined buffering strategy to overcome the startup problem for large graphs (by fixing the memory for intermediate data), and experimenting with probabilistic data structures such as bloom filter to improve the query response time for assessing edge existence, at the expense of increasing the overall false positive rate. These adjustments have led to a modest improvements in most cases, as compared to the previous version.

Graph Analytics, HPC↗

Stream Temperature Prediction in a Shifting Environment: Explaining the Influence of Deep Learning Architecture

Stream temperature is a fundamental control on ecosystem health. Recent efforts incorporating process guidance into deep learning models for predicting stream temperature have been shown to outperform existing statistical and physical models. This performance is in part because deep learning architectures can actively learn spatiotemporal relationships that govern how water and energy propagate through a river network. However, exploration of how spatiotemporal awareness and process guidance influence a model's generalizability under shifting environmental conditions such as climate change is limited. Here, we use Explainable Artificial Intelligence (XAI) to interrogate how differing deep learning architectures affect a model's learned spatial and temporal dependencies, and how those learned dependencies affect a model's ability to maintain high accuracy when applied to unseen environmental conditions. Using the Delaware River Basin in the northeastern United States as a test case, we compare two spatiotemporally aware process–guided deep learning models for predicting stream temperature (a recurrent graph convolution network—RGCN, and a temporal convolution graph model—Graph WaveNet). Both models achieve equally high predictive performance when testing data are well represented in the training data (test root mean squared errors of 1.64°C and 1.65°C); however, Graph WaveNet significantly outperforms RGCN in 4 out of 5 experiments where test partitions represent different types of unseen environmental conditions. XAI results show that the architecture of Graph WaveNet leads to learned spatial relationships with greater fidelity to physical processes, and that this fidelity improves the generalizability of the model when applied to shifting and/or unseen environmental conditions.

54 ENVIRONMENTAL SCIENCES↗

A review of non-cognitive applications for neuromorphic computing

Abstract Though neuromorphic computers have typically targeted applications in machine learning and neuroscience (‘cognitive’ applications), they have many computational characteristics that are attractive for a wide variety of computational problems. In this work, we review the current state-of-the-art for non-cognitive applications on neuromorphic computers, including simple computational kernels for composition, graph algorithms, constrained optimization, and signal processing. We discuss the advantages of using neuromorphic computers for these different applications, as well as the challenges that still remain. The ultimate goal of this work is to bring awareness to this class of problems for neuromorphic systems to the broader community, particularly to encourage further work in this area and to make sure that these applications are considered in the design of future neuromorphic systems.

97 MATHEMATICS AND COMPUTING↗

Auto Procedure Parsing: A Natural Language Processing Approach

Nuclear Power Plant (NPP) operating procedure is “a set of rules that describes how actions on the plant should be made if a certain system goal should be accomplished. U.S. NPPs use paper-based procedures (PBPs). PBPs are difficult to use. Common errors with PBPs are: following the wrong procedure, omit a step etc. Computer based procedures (CPBs) offer great improvement in ensuring plant safety. Some studies rely on experienced operators to understand the procedure content and then reorganize the procedure with digitally executable capabilities. Other studies utilize the procedure format to design rules to extract information from the operating procedures. Existing studies in procedure parsing are manual, laborious and extract limited information. This study aims to automatically extract critical information from operating procedures for generating computer interpretable representation of procedures. Such representation can further be used for automatic dynamic human reliability analysis and CPB design etc.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Graph-based featurization methods for classifying small molecule compounds

For over a decade, drug-induced liver injury (DILI) has posed significant drawbacks in the synthesis and development of drugs and remains a consequential concern. With finite success within the existing preclinical models, DILI is one of the main causes of drug withdrawal or termination from the market. Particularly, this withdrawal occurs during the late stages of drug development (Kullak-Ublick, 2017). Since DILI is difficult to diagnose and treat, it has become an obstacle in the drug production market that in turn affects clinicians, pharmaceutical companies, and consumers. We propose a method for learning features of DILI-positive drugs based on the graphical relationships and patterns they possess within a network of biological databases. We also train various statistical and machine learning models on these learned features in order to classify the drugs as DILI-positive or negative. Our methods include Random Forest, Neural networks, and logistic regression classification. We utilize labeled DILI-positive and DILI-negative datasets, which were developed by the FDA and the National center for toxicological research, as well as additional literature datasets (Thakkar, 2020) in order to validate our results and assess our featurization and model accuracy.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

ARCH: Large-scale knowledge graph via aggregated narrative codified health records analysis

Objective: Electronic health record (EHR) systems contain a wealth of clinical data stored as both codified data and free-text narrative notes (NLP). The complexity of EHR presents challenges in feature representation, information extraction, and uncertainty quantification. Here, to address these challenges, we proposed an efficient Aggregated naRrative Codified Health (ARCH) records analysis to generate a large-scale knowledge graph (KG) for a comprehensive set of EHR codified and narrative features. Methods: Using data from 12.5 million Veterans Affairs patients, ARCH first derives embedding vectors and generates similarities along with associated p-values to measure the strength of relatedness between clinical features with statistical certainty quantification. Next, ARCH performs a sparse embedding regression to remove indirect linkage between features to build a sparse KG. Finally, ARCH was validated on various clinical tasks, including detecting known relationships between entity pairs, predicting drug side effects, disease phenotyping, as well as sub-typing Alzheimer’s disease patients. Results: ARCH produces high-quality clinical embeddings and KG for over 60,000 codified and narrative EHR concepts. The KG and embeddings are visualized in the R-shiny powered web-API.3 ARCH achieved high accuracy in detecting EHR concept relationships, with AUCs of 0.926 (codified) and 0.861 (NLP) for similar EHR concepts, and 0.810 (codified) and 0.843 (NLP) for related pairs. It detected drug side effects with a 0.723 AUC, which improved to 0.826 after fine-tuning. Using both codified and NLP features, the detection power increased significantly. Compared to other methods, ARCH has superior accuracy and enhances weakly supervised phenotyping algorithms’ performance. Notably, it successfully categorized Alzheimer’s patients into two subgroups with varying mortality rates. Conclusion: The proposed ARCH algorithm generates large-scale high-quality semantic representations and knowledge graph for both codified and NLP EHR features, useful for a wide range of predictive modeling tasks.

Electronic health records↗

Direction-optimizing Label Propagation and its Application for Community Detection

Label Propagation is a machine learning algorithm typically used for classification. It has also been found to be an effective method for detecting communities in networks. It has two attractive features as a community detection method: it has nearly linear runtime and it requires no \textit{a priori} community information. We propose a new Direction Optimizing Label Propagation Algorithm (DOLPA) that relies on the use of {\em frontiers} and alternates between label {\em push} and label {\em pull} operations to enhance the performance of LPA. Specifically, DOLPA has parameters for tuning the processing order of vertices in a graph. This reduces the number of edges visited and improves the quality of solution. We apply DOLPA to community detection and present the design and implementation of the algorithm as well as its shared-memory parallelization using OpenMP. Empirically, we evaluate our algorithm using synthetic graphs as well as real-world networks. Compared with the state-of-art \textit{Parallel Label Propagation} algorithm, we achieve at least two times the F-Score while reducing the runtime by 50\% for synthetic graphs with overlapping communities. We also compare DOLPA against the state-of-art parallel implementation of the Louvain method using the same graphs and show that DOLPA achieves about three times the F-Score at 10\% the runtime. On real-world graphs, we get a speedup of up to $10\times$ using 64 threads.

Liu, X↗

Topological Analysis of The SPOKE Graph

The SPOKE graph [2, 6] is a sparse decorated semantic graph representing a collection of knowledge collected in many scientific databases from the fields of healthcare, biochemistry, chemistry, biology, et cetera. This knowledge graph is stored as a relational dataset decorated with metadata on each constituent vertex and edge. Formally, the graph is G(V, E, D), where V is a set of n vertices V := {1, ..., n} and edges of the form (i, j) ϵ E for i, j ϵ V, and table D that for any item in V υ E stores unstructured data such as vertex/edge type, nature of a relationship, et cetera. D(i) = {data involving vertex i ϵ V}, and D(i, j) = {data involving edge (i, j) ϵ E}. Here, we treat the graph as undirected in the sense that a direct relationship for (i, j) causes a (possibly opposite) reverse direct relationship for (j, i). The SPOKE graph G(V, E, D) is formed by processing a collection of relational datasets from medicine, chemistry, and biology, connecting many entities. Here, we analyze an instance from 2019, Spoke-20190707, where a graph file contains 6.16M edges and associated metadata and a vertex file contains 2.15M vertices and the associated metadata. There are 12 different types of vertex entities; all edge types used are implicit (see §2). There is other metadata in D on edges and vertices, but we just use the topology and the vertex labels in this report. SPOKE is growing as more knowledge is gained and more datasets are added. SPOKE is likely to grow 10x during the next phase of this project, and we therefore would like to consider topoligical analysis techniques that are scalable to several orders of magnitude larger than the current dataset (say >1B edges).

59 BASIC BIOLOGICAL SCIENCES↗

Meta-Learning Enhanced Physics-Informed Graph Attention Convolutional Network for Distribution Power System State Estimation

Promptly perceiving distribution system states is challenged by frequent topology changes and uncertain power injections. To address these issues, a Meta-learning enhanced physics-informed graph attention convolutional network (Meta-PIGACN) model is proposed to handle topological variability in distribution system state estimation (DSSE). Specifically, physics information is integrated into the graph convolutional network, enabling a physics-informed edge-weighting process that incorporates physical information to control the aggregation of neighboring nodes. Besides, the graph attention mechanism automatically adjusts the importance of different neighboring nodes, allowing the capture and preservation of inherent system features across varying topologies, thereby improving state estimation accuracy. Furthermore, meta-learning is proposed to acquire empirical knowledge across multiple topologies so that the model can rapidly adapt to new configurations through iterative gradient descent updates even in large-scale systems. In conclusion, the simulation results based on the 33/118/1746-node distribution systems show the high accuracy and efficiency of the proposed model.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Neuromorphic Graph Algorithms: Cycle Detection, Odd Cycle Detection, and Max Flow

Neuromorphic computing is poised to become a promising computing paradigm in the post Moore’s law era due to its extremely low power usage and inherent parallelism. Spiking neural networks are the traditional use case for neuromorphic systems, and have proven to be highly effective at machine learning tasks such as control problems. More recently, neuromorphic systems have been applied outside of the arena of machine learning, primarily in the field of graph algorithms. Neuromorphic systems have been shown to perform graph algorithms faster and with lower power consumption than their traditional (GPU/CPU) counterparts, and are hence an attractive option for a co-processing unit in future high performance computing systems, where graph algorithms play a critical role. In this paper, we present a neuromorphic implementation of cycle detection, odd cycle detection, and the Ford-Fulkerson max-flow algorithm. We further evaluate the performance of these implementations using the NEST neuromorphic simulator by using spike counts and simulation time as proxies for energy consumption and run time. In addition to gains inherent in neuromorphic systems, we show that within the neuromorphic implementations early stopping criteria can be implemented to further improve performance.

Kay, Bill↗

Navigating Large Chemical Spaces Using Graph Theory and Integer Programming

Navigating and analyzing large chemical spaces are necessary to accelerate the design and discovery of new molecules and chemical processes. In this work, we introduce a computational framework that integrates graph theory and integer programming to enable the efficient navigation of large chemical spaces. Our framework represents the chemical space as a graph, wherein nodes represent molecules and edges represent the degree of similarity or connectivity based on domain-specific information. Using the graph representation, we identify representative molecules by computing the so-called minimum dominating set (MDS), which in our context is the minimum set of molecules that is connected to all other molecules. We present a suite of solution strategies for the MDS problem including heuristic and rigorous integer programming (IP) approaches. We show that these approaches allow us to capture physicochemical properties and domain-specific logic and constraints, facilitating the identification of molecules with the target properties. We demonstrate the effectiveness of the proposed approach by navigating the chemical space of per- and polyfluoroalkyl substances (PFAS); this comprises approximately 15,000 molecular structures. We compare our framework against traditional dimensionality reduction and clustering methods such as t-SNE and K-means clustering.

Chemical structure↗

Solid State Quantum Refrigeration Superconducting, Absorption and Measurement Based (Final Technical Report)

During this DOE grant, DE-SC0017890, in place for the past six years, all proposed research was carried out and published in peer-reviewed papers, as well as other projects that emerged during the research. In that effort the research team accomplished all proposed research, as well as many closely related research projects discovered and conceived of during the grant. These works included “Efficient Quantum Measurement Engines”, a work published in Physical Review Letters, giving a theory of quantum measurement-based engines, which uses quantum measurement as a resource. These engines are designed to efficiently convert energy from the stochastic quantum measurement process into useful work. Further publications include “Experimental Realization of a Quantum Dot Energy Harvester”, a joint theory and experimental work in collaboration with the group of Charles Smith in Cambridge, UK, as well as long time theoretical collaborators, Rafael Sánchez and Björn Sothmann. This work, featured as an Editor’s Suggestion in Physical Review Letters, realized an earlier theoretical proposal of ours, whereby two resonant tunneling quantum dots are connected to a central electronic cavity that is heated by a hot energy source. We also published “Superconducting Quantum Refrigerator: Breaking and Rejoining Cooper Pairs with Magnetic Field Cycles” a work done in collaboration with experimentalist Francesco Giazotto from ENS Pisa, Italy, which also resulted in a patent. This paper, published in Phys. Rev. Applied, advanced the concept of a cyclic fridge based on the normal/superconducting phase transition together with layered materials separated by tunnel junctions. We also completed the proposed research on a heat transistor, publishing “Thermal transistor and thermometer based on Coulomb-coupled conductors”, carried out as a collaboration between my group and theorists Splettstoesser (Lund U., Sweden), Sothmann (U. Duisburg-Essen, Germany), and Sánchez (U. Autónoma de Madrid, Spain). We carried out an analysis of a quantum coupled to a quantum point contact as a sensitive thermometer and heat transistor. We found the optimal statistical estimator for the temperature and compared it with experiments on the same type of devices. We also investigated autonomous quantum absorption refrigerators using quantum dots to cool by using a very hot thermal reservoir to drive heat between two other reservoirs. In the article “Quantifying the quantum heat contribution from a driven superconducting circuit”, we demonstrated that for a driven superconducting circuit, we showed heat flow provided by a hot source to the qubit can be switched on and off by varying external parameters, the frequency and the intensity of the driving. In the work “Stochastic thermodynamic cycles of a mesoscopic thermoelectric engine”, we reconsidered the autonomous thermoelectric heat engine in terms of underlying cycles. Rather than periodic behavior, the cycles were stochastic in nature. Nevertheless, by undertaking a graph theoretical analysis of the elementary transport processed, great quantitative and qualitative insight could be found. We also considered the quantum measurement process and showed that a quantum version of Maxwell’s demon could be related to the work extraction of a quantum system, closely related to arrow-of-time measures for quantum measurement, as described in our article “Thermodynamics of quantum measurement and Maxwell's demon's arrow of time”. This work was selected in Phys. Rev. A as an Editor’s Suggestion. A recent preprint titled “Cyclic Superconducting Quantum Refrigerators Using Guided Fluxon Propagation” accomplished an important piece of this grant: to propose a new kind of quantum refrigerator using the dynamics of fluxons in a type II superconductor. This invention envisioned a race-track type geometry where fluxons are confined. By applying a gradient of magnetic field together with electric current in a Corbino geometry, the circulating fluxons can actively cool a cold reservoir, realizing a new type of cyclic superconducting refrigerator. We also investigated the possibility of thermal control from different points of view. The application of quantum measurement to the system gives a new kind of control on the system of interest – we have pioneered this approach and shown that measurement can boost the thermal power of quantum engines as described in “Continuous measurement boosted adiabatic quantum thermal machines”. The ability to have heat flows on demand is an outstanding challenge, and we have provided new solutions to this problem in Thermal control across a chain of electronic nanocavities” for a chain of electron cavities using gating voltage control. The control methods using qubit/qubit coupling to create absorption fridges at their most fundamental level have also been developed.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Phlex: Parallel, Hierarchical, and Layered EXecution of data-processing algorithms

Phlex is a computing framework supporting the parallel, hierarchical, and layered execution of data-processing algorithms. It is based on the functional-programming paradigm, thus guaranteeing thread-safety when invoking user-defined pure functions. Phlex allows users to specify arbitrary graph-based hierarchies of data organization, enabling more flexible processing of data as required by the constraints of the program.

Knoepfel, KyleJ. [Fermi National Accelerator Labor↗