Engineering Papers⌕ Search

Engineering topics

Heller, Natalie C.

Publications and source records attributed to Heller, Natalie C..

The Utility in Conjoint Analysis as a Fast Expert Elicitation Technique

This paper presents an interface and analysis technique for quickly conducting expert elicitation with the goal of determining entity importance. Our interface deploys a two-alternative choice experiment that is capable of representing knowledge graphs in an easy to interpret fashion for users with limited experience with knowledge graphs. Our analysis methodology takes advantage of conjoint analysis techniques and provides entity weights for many SMEs simultaneously. The results largely align with individual participant fits.

Conjoint Analysis, interface, Expert Elicitation, ↗

Hypergraph Models of Biological Networks to Identify Genes Critical to Pathogenic Viral Response

Motivation: Representing biological networks as graphs is a powerful approach to reveal underlying patterns, signatures, and critical components from high-throughput biomolecular data. However, graphs do not natively capture the multi-way relationships present among genes and proteins in biological systems such as protein complexes, metabolic reactions, and signal transduction pathways. Hypergraphs are generalizations of graphs that naturally model multi-way interactions in data, and we therefore seek to understand how they can more faithfully identify, and potentially predict, complex relationships in genomic expression data sets. Results: We compiled a novel data set of transcriptional host response to pathogenic viral infections and formulated relationships between genes as a hypergraph where hyperedges are differentially expressed genes and vertices represent conditions. We find that hypergraph betweenness centrality is a superior method for identification of genes important to viral response when compared with graph centrality. Our results demonstrate the utility of using hypergraphs to represent complex biological systems, and highlight potentially interesting biological results about host response to highly pathogenic viruses.

systems biology, hypergraph, viral infection, biol↗

Using Graph Edit Distance for Noisy Subgraph Matching of Semantic Property Graphs

The subgraph matching problem is a fundamental problem in graph theory that is known to be NP-complete. In this study, performers were asked to develop algorithms to search for semantic property graphs that were subgraphs of a large knowledge graph. The templates provided contained structural information about the subgraphs and some attributes for each node and edge. There also exists a similarity measure between a set of attribute values that occurs on every node and edge. Algorithms performed well in the case where an exact match existed, but performers were also provided templates that had noise added such that there existed no match in the knowledge graph. Performers were asked to find the closest matches to those noisy subgraphs. To evaluate performance on this task, we developed a version of the graph edit distance algorithm to measure the cost of editing the template graph so that it is isomorphic in structure and attributes to the performer submission.

Ebsch, Christopher L.↗

Evaluation of Alignment: Precision, Recall, Weighting and Limitations

In the real world, data does not come neatly packaged. Instead, it typically comes as small updates from many sources with different conventions. Building a single, cohesive knowledge-base to work from requires merging small updates from many different sources. This paper outlines methods we have investigated for scoring merging routines. Given a challenge problem consisting of a large knowledge-base and a set of smaller documents, algorithms are asked to identify alignment points between the smaller document and the knowledge base. This paper surveys options for evaluating such algorithms, providing notes on strengths, weaknesses and considerations for interpretation.

Cottam, Joseph A.↗

Multi-Channel Entity Alignment via Name Uniqueness Estimation

When searching for adversarial activity within multiple networks, one of the greatest challenges is how to accurately align entities across different channels of information. This task becomes increasingly difficult when minimal additional information is known about each individual besides a name. Within this study, we analyze name rarity and how it can be used to align people on three distinct data channels: Venmo financial transactions, Reddit online discussions, and a bibliographic data source of academic writings. We explore how the uniqueness of a name can be used to decide if a person is likely the same as another across networks, in the absence of any additional ground truth. While 100 percent confidence cannot be gained, we can use this information to clarify when a possible alignment is more or less likely to be the same individual, increasing our confidence of accurately detecting adversarial behavioral patterns. From the data collected, we found that 0.1% of people had the same name across data sets, and 22.5% of those names are considered rare by our threshold. In our study, we also examine the accuracy of our method and show how real names can be extracted from account usernames, and compared in a similar manner.

Orren, Miquette J.↗