Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Knowledge graph”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

GraphAide: Advanced Graph-Assisted Query and Reasoning System

Curating knowledge from multiple siloed sources that contain both structured and unstructured data is a major challenge in many real-world applications. Pattern matching and querying represent fundamental tasks in modern data analytics that leverage this curated knowledge. The development of such applications necessitates overcoming several research challenges, including data extraction, named entity recognition, data modeling, and designing query interfaces. Moreover, the explainability of these functionalities is critical for their broader adoption. The emergence of Large Language Models (LLMs) has accelerated the development lifecycle of new capabilities. Nonetheless, there is an ongoing need for domain-specific tools tailored to user activities. The creation of digital assistants has gained considerable traction in recent years, with LLMs offering a promising avenue to develop such assistants utilizing domain-specific knowledge and assumptions. In this context, we introduce an advanced query and reasoning system, GraphAide, which constructs a knowledge graph (KG) from diverse sources and allows to query and reason over the resulting KG. GraphAide harnesses both the KG and LLMs to rapidly develop domain-specific digital assistants. It integrates design patterns from retrieval augmented generation (RAG) and the semantic web to create an agentic LLM application. GraphAide underscores the potential for streamlined and efficient development of specialized digital assistants, thereby enhancing their applicability across various domains.

Purohit, Sumit [BATTELLE (PACIFIC NW LAB)] (ORCID:↗

Topological Analysis of The SPOKE Graph

The SPOKE graph [2, 6] is a sparse decorated semantic graph representing a collection of knowledge collected in many scientific databases from the fields of healthcare, biochemistry, chemistry, biology, et cetera. This knowledge graph is stored as a relational dataset decorated with metadata on each constituent vertex and edge. Formally, the graph is G(V, E, D), where V is a set of n vertices V := {1, ..., n} and edges of the form (i, j) ϵ E for i, j ϵ V, and table D that for any item in V υ E stores unstructured data such as vertex/edge type, nature of a relationship, et cetera. D(i) = {data involving vertex i ϵ V}, and D(i, j) = {data involving edge (i, j) ϵ E}. Here, we treat the graph as undirected in the sense that a direct relationship for (i, j) causes a (possibly opposite) reverse direct relationship for (j, i). The SPOKE graph G(V, E, D) is formed by processing a collection of relational datasets from medicine, chemistry, and biology, connecting many entities. Here, we analyze an instance from 2019, Spoke-20190707, where a graph file contains 6.16M edges and associated metadata and a vertex file contains 2.15M vertices and the associated metadata. There are 12 different types of vertex entities; all edge types used are implicit (see §2). There is other metadata in D on edges and vertices, but we just use the topology and the vertex labels in this report. SPOKE is growing as more knowledge is gained and more datasets are added. SPOKE is likely to grow 10x during the next phase of this project, and we therefore would like to consider topoligical analysis techniques that are scalable to several orders of magnitude larger than the current dataset (say >1B edges).

59 BASIC BIOLOGICAL SCIENCES↗

Machine Learning for the Validation of Expert-Elicited Causal Risk Diagrams

Exposure to spaceflight poses risk to human health in complex ways. To help manage this risk, the Human Systems Risk Board (HSRB) at the National Aeronautics and Space Administration (NASA) maintains a set of causal diagrams that attempt to explain how spaceflight hazards generate health risks and lead to adverse outcomes both in-mission, immediately post-mission, and over the long term. These causal risk diagrams are formulated as directed acyclic graphs (DAGs) and can function as knowledge graphs of connected risks and outcomes. These DAGs have proven useful for communication, and, through network analysis, have allowed for the identification of structurally important factors in the risk network. However, the utility these DAGs provide is directly proportional to their verisimilitude, making assessment of this trait using empirical data – whether from actual human spaceflight or various spaceflight analogue exposures and model organisms – a high priority. In this research we explore the use of machine learning algorithms to learn DAG structure from empirical data as a means of evaluating human-elicited DAG structures. To do so, we test several different graph structure-learning algorithms on data concerning changes in the bones of rats and mice after exposure to either spaceflight or a spaceflight analogue. We explore potential methods for indexing the similarity between each algorithm’s output DAG with all the others and with that of the expert-elicited DAG. We discuss next steps in this ongoing line of research and open science initiatives underway to complete them.

directed acyclic graphs↗

What Is the Agent Doing? Visualizing Agentic AI Querying Workflows

We explore how visualizations can help users understand what an AI agent is doing as it builds and runs queries over data. As part of the LinkQ system, a natural language interface for querying knowledge graphs with a large language model (LLM), we designed two complementary views: A State Diagram that shows where the agent is within a larger workflow, and a Live Action Display that gives real-time updates about the agent's current task. In a study with 14 practitioners, we found that these visuals helped participants build stronger mental models of the agent's behavior while also increasing their confidence in the system. However, we also observed that users sometimes trusted incorrect outputs simply because the agent appeared to be doing the "right" thing. Our findings point to both the value and risk of visualizing agent behavior in interactive AI systems.

97 MATHEMATICS AND COMPUTING↗

Directed Acyclic Graph Guidance Documentation

For over a decade, the National Aeronautics and Space Administration (NASA) has tracked and configuration-managed approximately 30 risks to astronaut health and performance that occur before, during and after spaceflight. The Human System Risk Board (HSRB), a Health and Medical Technical Authority (HMTA) Board at NASA Johnson Space Center, is the entity responsible for identifying, assessing, analyzing, and monitoring the official understanding of the risk or risk posture for each of the Human System Risks and determining – based on evaluation of the available evidence – when that risk posture changes. The ultimate purpose of tracking and researching these risks is to find ways to reduce the risk that astronaut crews face during spaceflight. Historically, research, development and operations relevant to one risk have been conducted in isolation from other risks; these individual risk ‘silos’ enabled initial characterization of each specific risk. In spaceflight however, the impact of exposure to risk for astronaut crews is cumulative, and not independent of exposures or other risks, as all the adverse effects of the spaceflight environment begin at launch, continue throughout the duration of the mission and in some cases across the lifetime of the crews. In January of 2020, the HSRB at NASA embarked on a pilot project designed to assess the potential value of causal diagramming as a tool to facilitate understanding these cumulative and interdependent effects as applied within Human System Risk management. This process uses directed acyclic graphs as a means of formalizing a shared mental model of the causal flow of risk among Risk Board stakeholders. Initially this model was to improve communication among those stakeholders, but the potential value exceeds communication alone. Formalization of the process for creating these causal diagrams will enable the creation of a composite risk network that is vetted by members of the NASA community and configuration managed. The causal diagrams are formulated as directed acyclic graphs (DAGs) to function as a type of knowledge graph for reference for the board and its stakeholders. This document outlines the pilot process, the standardized approaches, and guidance for risk custodian teams when creating and updating DAGs as a part of the NASA Human System Risk Management process.

Risk↗

Directed Acyclic Graphs: A Tool for Understanding the NASA Human Spaceflight System Risks - Human System Risk Board

For over a decade, the National Aeronautics and Space Administration (NASA) has tracked and configuration-managed approximately 30 risks to astronaut health and performance that occur before, during and after spaceflight. The Human System Risk Board (HSRB), a Health and Medical Technical Authority (HMTA) Board at NASA Johnson Space Center, is the entity responsible for identifying, assessing, analyzing, and monitoring the official understanding of the risk or risk posture for each of the Human System Risks and determining – based on evaluation of the available evidence – when that risk posture changes. The ultimate purpose of tracking and researching these risks is to find ways to reduce the risk that astronaut crews face during spaceflight. Historically, research, development and operations relevant to one risk have been conducted in isolation from other risks; these individual risk ‘silos’ enabled initial characterization of each specific risk. In spaceflight however, the impact of exposure to risk for astronaut crews is cumulative, and not independent of exposures or other risks, as all the adverse effects of the spaceflight environment begin at launch, continue throughout the duration of the mission and in some cases across the lifetime of the crews. In January of 2020, the HSRB at NASA embarked on a pilot project designed to assess the potential value of causal diagramming as a tool to facilitate understanding of these cumulative and interdependent effects as applied within Human System Risk management. This process uses directed acyclic graphs as a means of formalizing a shared mental model of the causal flow of risk among Risk Board stakeholders. Initially this model was to improve communication among those stakeholders, but the potential value exceeds communication alone. The causal diagrams are formulated as directed acyclic graphs (DAGs) to function as a type of knowledge graph for reference for the board and its stakeholders. This document is a sister document to NASA/TM 20220006812 Directed Acyclic Graph Guidance Documentation (1). In that document, the basic guidance for creating and standardizing directed acyclic graphs as tools for cross-risk analysis is provided. This document contains the initial configuration managed DAGs that were created as a result of applying those principles. These initial versions were accepted by the HSRB in January of 2022. Each of the Human System Risks are represented by a DAG that has been reviewed by the larger Human Health and Performance community at NASA including life scientists, physical scientists, physicians, nurses, pharmacists, exercise specialists and more. These results show the starting point for Human System Risk DAGs as shared mental models and communication aids across the boundaries of the various expertise needed to understand and mitigate the human risks in spaceflight. Because they are a starting point, each of these DAGs can be expected to change over time as new or refined evidence becomes available. The process for updating these DAGs can be found in the JSC-66705 Human System Risk Management Plan (2) that is publicly available on the NASA Technical Reports Server.

Erik L. Antonsen↗

Human System Risk Communication: Directed Acyclic Graphs

- The Human System Risk Board (HSRB) is responsible for the management of a portfolio of 30 human system risks that NASA tracks and configuration manages to mitigate for future crewed exploration missions. - The HSRB has been exploring the concept of causal diagrams (in the form of Directed Acyclic Graphs or DAGs) as an approach to creating knowledge graphs for each risk to enable shared mental models of causal flow from spaceflight hazards to mission outcomes among HSRB Stakeholders. - These diagrams are intended to improve insight and communication of risk across the myriad subject matter experts and management interested in human system risk reduction. This includes program managers, systems engineers, and operators in addition to the Human Health and Performance Directorate. - The DAG project was intended to create the foundation for composition of the 30 baselined DAGs into a single risk network and software is being developed in parallel to enable this forward work.

directed acrylic graph↗

Synthetic Data and Graph Generation for Modeling Adversarial Activity (Final Project Report)

The Data and Graph Generation for Modeling Adversary Activity (MAA) project developed a methodology along with scalable graph modeling and generation tools to produce realistic large-scale background activity graphs with embedded adversarial activity pathways. The technical report presents PNNL methodology, released datasets, lessons learned, and recommendations to develop graph analytic algorithms for structure-only and attributed knowledge graphs.

97 MATHEMATICS AND COMPUTING↗

Gene-Metabolite Association Prediction with Interactive Knowledge Transfer Enhanced Graph for Metabolite Production

Identifying gene targets for enhancing metabolite production in metabolic engineering is challenging due to the vast research literature and the approximation in genome-scale metabolic model (GEM) simulations. Here, to address this, we propose the Gene-Metabolite Association Prediction task, which automates gene discovery for given metabolite-gene pairs, accompanied by a benchmark dataset of 2474 metabolites and 1947 genes for Saccharomyces cerevisiae (SC) and Issatchenkia orientalis (IO). This task is complicated by incomplete metabolic graphs and metabolic heterogeneity. We introduce an Interactive Knowledge Transfer mechanism based on Metabolism Graphs (IKT4Meta) to enhance prediction accuracy by integrating cross-metabolism knowledge. Using Pretrained Language Models (PLMs) to generate inter-graph links mitigates heterogeneity issues, while intra-graph links are propagated via these anchors. Gene-metabolite predictions are then performed on the enriched graphs integrating multiple microorganisms’ knowledge. Experiments show that IKT4Meta outperforms baselines by up to 12.3% in link prediction.

59 BASIC BIOLOGICAL SCIENCES↗

Session Introduction: Graph Representations and Algorithms in Biomedicine

Connectivity is a fundamental property of biological systems: on the cellular level, proteins interact with each other to form protein-protein interaction networks (PPIs); on the organism level, neurons are arranged in a network; and on a community-level, species can have complex relationships with one another that drive the development and balance of an ecosystem. Graphs, representations of systems consisting of entities as vertices and their connections as edges, are a useful structure to characterize many such systems. Such models can be used to understand biological systems that naturally have a network structure, including PPIs, biological neurons, and ecosystems. In today’s information age, graph representations and algorithms (often in combination with machine learning techniques) are used to organize massive amounts of related data, much of which may be heterogeneous or unstructured, and identify patterns that represent novel biological insights. PSB’s 2023 session “Graph Representations and algorithms in Biomedicine,” encompasses modern developments in graph theory and its applications to various fields of biomedicine. This session includes a wide range of research - knowledge graphs built from text-mined health data, heterogeneous networks using multi-omic databases, and graphs refined to represent uncertainty or improve memory usage.

Chrisman, Brianna S.↗

Generating and Analyzing Program Call Graphs using Ontology

Call graph or caller-callee relationships have been used for various kinds of static program analysis, performance analysis and profiling, and for program safety or security analysis such as detecting anomalies of program execution or code injection attacks. However, different tools generate call graphs in different formats, which prevents efficient reuse of call graph results. In this paper, we present an approach of using ontology and resource description framework (RDF) to create knowledge graphs for specifying call graphs to facilitate the construction of full-fledged and complex call graphs of computer programs, realizing more interoperable and scalable program analyses than conventional approaches. We create a formal ontology-based specification of call graph information to capture concepts and properties of both static and dynamic call graphs so different tools can collaboratively contribute to more comprehensive analysis results. Our experiments show that ontology enables merging of call graphs generated from different tools and flexible queries using a standard query interface. Index Terms—Callgraph, ontology, knowl

Dorta, E.↗

QLiG: Query Like a Graph For Subgraph Matching

A graph is a natural and flexible modeling approach to represent entities and relationships between them in real-world. A Knowledge Graphs (KG) is a specialized graph with formal and structured representation of facts, relationships, annotated with semantic descriptions. Subgraph matching is one of the fundamental graph problems to identify relationships, interactions and activities of interest within a large graph. A query specification is a collection of abstract components, operations, and constraints to express a pattern. The specification can be implemented in different ways based on underlying data model. Various graph query specifications have been developed over the years and have led to the development of different open-sourced and vendor-specific query languages. Such specification are modeled as an extension of relational algebra used to develop relational query languages such as SQL. Such relational concepts do not inherently support graph queries. There is a need to represent graph queries in terms on graph-based components to expedite query construction by non-database experts. We present a graph-based query approach QLiG (pronounced cleeg), to perform subgraph matching in Labeled Property Graph. We present the query specifications, salient features, and a use case to show functional examples.

Purohit, Sumit↗

Node-degree aware edge sampling mitigates inflated classification performance in biomedical random walk-based graph representation learning

Motivation: Graph representation learning is a family of related approaches that learn low-dimensional vector representations of nodes and other graph elements called embeddings. Embeddings approximate characteristics of the graph and can be used for a variety of machine-learning tasks such as novel edge prediction. For many biomedical applications, partial knowledge exists about positive edges that represent relationships between pairs of entities, but little to no knowledge is available about negative edges that represent the explicit lack of a relationship between two nodes. For this reason, classification procedures are forced to assume that the vast majority of unlabeled edges are negative. Existing approaches to sampling negative edges for training and evaluating classifiers do so by uniformly sampling pairs of nodes. Results: We show here that this sampling strategy typically leads to sets of positive and negative examples with imbalanced node degree distributions. Using representative heterogeneous biomedical knowledge graph and random walk-based graph machine learning, we show that this strategy substantially impacts classification performance. If users of graph machine-learning models apply the models to prioritize examples that are drawn from approximately the same distribution as the positive examples are, then performance of models as estimated in the validation phase may be artificially inflated. We present a degree-aware node sampling approach that mitigates this effect and is simple to implement. Availability and implementation: Our code and data are publicly available at https://github.com/monarch-initiative/negativeExampleSelection.

59 BASIC BIOLOGICAL SCIENCES↗

Graph Convolutional Network-Strengthened Topic Modeling for Scientific Papers

Machine learning has been woven into statistics to modernize topic modeling over textual documents written in natural language, and scientific paper search and recommendation can consequently offer higher accuracy instead of counting on traditional keyword-based search. However, topic distribution of a paper resulted from existing topic modeling techniques only relies on the statistics of words contained in the paper itself. We argue that community users’ views of a paper may also provide insights at the time of recommendation. For example, if a paper on fake image detection has been cited heavily by machine learning papers, such a feature should be absorbed in the embedding of this paper, so that it can be recommended for future query on machine learning. In this paper, we present a Graph Convolutional Network-strengthened Topic Modeling (GCN-TM) method, which employs GCN technique to refine topic modeling of scientific papers. A citation-oriented knowledge graph is constructed, and topic modeling is mapped to feature embedding of the comprising papers. On top of its own topics carried in its content, each paper learns topics from its neighbors and revise its embedding accordingly. Our empirical studies over real-life scientific literature has proved the necessity and effectiveness of our proposed approach.

Jia Zhang↗

BrickQA: Bridging the Semantic Gap in Building Operations with Dynamic Graph Exploration

While standardized ontologies like the Brick schema address data heterogeneity in Building Automation Systems (BAS), accessing this semantic data remains a challenge as domain experts often lack the expertise to formulate complex SPARQL queries. To bridge this gap, we present BrickQA, a Large Language Model (LLM)-based framework that translates natural language into executable SPARQL queries through structured query decomposition, dynamic schema exploration, and inline validation. BrickQA utilizes an iterative reasoning agent to actively navigate graph topology through dynamic exploration actions without requiring exhaustive context injection or model fine-tuning. This approach effectively mitigates hallucinations, particularly in large-scale building knowledge graphs. Empirical evaluation on BuildingQA, a standardized benchmark, demonstrates that BrickQA significantly outperforms ReAct baselines, delivering a 0.291–0.355 absolute F1 improvement while achieving 3 × –12.7 × higher token cost-efficiency. Beyond these metrics, the framework maintains structural fidelity across heterogeneous buildings and remains resilient to ambiguous queries without requiring site-specific fine-tuning. Furthermore, a case study on operational analytics validates the framework’s capability to handle temporal and aggregation constraints, effectively transforming abstract semantic models into actionable facility management insights.1

Ko, Yun-Dam↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, whole organism, behavior; tabular, imagery). Open Science is the concept that the more people have access to scientifically curated data, the more knowledge will be gained. This led NASA to start the development of GeneLab in 2015. GeneLab houses spaceflight and space-analog multi-omics datasets from plant, rodent, small animal, and microbial experiments. The success and knowledge gained from GeneLab led to a new alliance of NASA “Open Science Data Repositories” (OSDR), which include the Ames Life Sciences Data Archive (ALSDA) and the NASA Biological Institutional Scientific Collection (NBISC). Both are adopting the GeneLab data system, so data are more findable, accessible, interoperable, and reusable (FAIR). OSDR systems provide users the ability to upload, download, search, share, analyze, and visualize. Open Science also needs strong confidence in the data, which is gained through building science communities. With ~400 current members, GeneLab and ALSDA formed Analysis Working Groups (AWGs) to provide feedback on processing pipelines, metadata curation standards (for ‘omics and phenotypic-physiological-behavioral assays), and to collaborate in effectively reusing data. The AWG also led to the development of the Radiation Biology Ontology (RBO), ensuring radiation metadata are efficiently captured, connected, and interoperable. Feedback from the AWG provided design input toward the new single point-of-entry data submission portal for all investigators to submit, curate, and share their research data. Space biological data is now maximally open access, collected-curated with rich metadata, and formatted for interoperability to enable systems biology, meta-analysis, knowledge graphs, machine learning, modeling, and other reuse approaches. With potential for further federation of OSDR for data mining with traditional biological and medical databases (NIH, NCI, EBI, etc.), a new era for space biology has begun to support the knowledge discovery necessary for Lunar and Martian missions.

Ryan T Scott↗

FAIR to WISE (F2W) v1.0.0

FAIR to WISE (F2W) is an iterative, large-language model (LLM) driven pipeline that turns unstructured research PDFs into structured, queryable knowledge graphs (KGs). Core features include schema-driven extraction to a LinkML model; full provenance capture; ontology-grounded enrichment (e.g., chemical validation and ChEBI lookup); graph construction to JSON-LD with stable IDs; and KG-RAG question answering with evidence-aware retrieval. The system is engineered for reproducibility and accessibility (open-source Ollama models, temperature=0, NVTX/Nsight profiling) with robust QA (relation verification, deduplication, and deterministic outputs). Primary uses are literature-to-KG automation, knowledge-grounded Q&A, and experimental steering support. We demonstrate the approach in organic photovoltaics, where the pipeline ingests papers, builds a domain KG, and evaluates answers against expert competency questions to guide experimental planning and interpretation. Compared with off-the-shelf LLMs and ad-hoc NLP tools, F2W addresses ontology gaps and reduces hallucination risk by grounding responses in extracted evidence and enforcing schema constraints; it also offers deterministic, provenance-linked outputs and open, cost-aware deployment. Evidence-aware ranking further improves answer quality over pure vector search.

Abramov, David [Lawrence Berkeley National Laborat↗

Vulcan-Forge: Architecture and Design of a Multi-Modal Forensic Analysis Plugin for CALDERA

Forge and VULCAN together describe an open-architecture cybersecurity analysis ecosystem that unifies forensic artifact processing, detection engineering, and vulnerability intelligence within integrated platforms. Forge operates as a plugin for MITRE CALDERA, ingesting diverse evidence formats—including EVTX, PCAP/PCAPNG, CSV, JSON, YAML, XML, binaries, and archives—to construct a unified artifact graph enriched with severity scoring, TLP classification, and audit trails. It provides subsystems for artifact parsing, streaming structured-data visualization, NetworkMiner-based packet inspection, PE/.NET binary analysis, and LLM-assisted triage and rule generation, with outputs validated against CCCS-YARA and pySigma schemas. VULCAN complements this by serving as a cybersecurity analyst platform that integrates a Neo4j knowledge graph, Qdrant vector retrieval, SSVC-based triage, and a local LLM to deliver CVE intelligence and forensic analysis through a multi-source ingest pipeline drawing from NVD, CISA KEV, EPSS, MITRE ATT&CK, and CAPEC. Together, they bridge structured threat intelligence with automated forensic analysis and detection workflows.

97 MATHEMATICS AND COMPUTING↗