Engineering Papers⌕ Search

Engineering topics

Herrmannova, Dasha

Publications and source records attributed to Herrmannova, Dasha.

Methodology to Compare Twitter Reaction Trends between Disinformation Communities, to COVID related Campaign Events at Different Geospatial Granularities

With still ongoing COVID pandemic, there is an immediate need for a deeper understanding of how Twitter discussions (or chatters) in disinformation spreading communities get triggered. More specifically, the value is in monitoring how such trigger events in Twitter discussion do align with the timelines of relevant influencing events in the society (indicated in this work as campaign events). For campaign events in regards to COVID pandemic, we consider both NPI (Nonpharmaceutical Interventions) campaigns and disinformation spreading campaigns together. In this short paper we have presented a novel methodology to quantify, compare and relate two Twitter disinformation communities, in terms of their reaction patterns to the timelines of major campaign events. We have also analyzed these campaigns at their three geospatial granularity contexts: local county, state, and country/ federal. We have conducted a novel dataset collection on campaigns (NPI + Disinformation) at these different geospatial granularities. Then, with collected dataset on Twitter disinformation communities, we have performed a case study to validate our proposed methodology.

De, Debraj↗

Smoky Mountain Data Challenge 2021: An Open Call to Solve Scientific Data Challenges Using Advanced Data Analytics and Edge Computing

The 2021 Smoky Mountains Computational Sciences and Engineering Conference enlists scientists from across Oak Ridge National Laboratory (ORNL) and industry to be data sponsors and help create data analytics and edge computing challenges for eminent datasets in a variety of scientific domains. This work describes the significance of each of the eight datasets and their associated challenge questions. The challenge questions for each dataset were required to cover multiple difficulty levels. An international call for participation was sent to students, asking them to form teams of up to six people and apply novel data analytics and edge computing methods to solve these challenges.

Devineni, Pravallika↗

Visual Understanding of COVID-19 Knowledge Graph for Predictive Analysis

This study aims to effectively analyze and visualize the concept to concept network derived from the COVID-19 Open Research Dataset (CORD-19) dataset, where we have more than 48,000 concepts with more than 300,000 relationships between concepts. In analyzing networks, we focus on finding relationship patterns between the coronavirus disease 2019 (COVID-19) concepts and other concepts. Given the node and edge datasets, we construct directional graphs and calculate all pair shortest paths based on multiple edge weight schemes. However, statistical metrics are not sufficient to identify specific relationships represented in the network. Therefore, we also propose a visual analytics approach to effectively understand the knowledge graph. Our highly interactive visual analytics allows users to effectively analyze the evolving graphs and (COVID-19) concept nodes and other nodes related to the COVID-19 nodes. We envision that this study will pave the path to develop strategies to provide more accurate and scalable predictive analysis on knowledge graphs related to CORD19 and other biomedical knowledge graphs.

Lim, Seung-Hwan↗

A Mixed-Method Design Approach for Empirically Based Selection of Unbiased Data Annotators

Implicit bias embedded in the annotated data is by far the greatest impediment in the effectual use of supervised machine learning models in tasks involving race, ethics, and geopolitical polarization. For societal good and demonstrable positive impact on wider society, it is paramount to carefully select data annotators and rigorously validate the annotation process. Current approaches to selecting annotators are not sufficiently grounded in scientific principles and are limited at the policy-guidance level, thereby rendering them unusable for machine learning practitioners. This work proposes a new approach based on the mixed-methods design that is functional, adaptable, and simpler to implement in selecting unbiased annotators for any machine learning problem. By demonstrating it on a real-world geopolitical problem, we also identified and ranked key inane profile characteristics towards an empirically-based selection of unbiased data annotators.

Thakur, Gautam Malviya↗

Overview of the Second Workshop on Scholarly Document Processing

With the ever-increasing pace of research and high volume of scholarly communication, scholars face a daunting task. Not only must they keep up with the growing literature in their own and related fields, scholars increasingly also need to rebut pseudo-science and disinformation. These needs have motivated an increasing focus on computational methods for enhancing search, summarization, and analysis of scholarly documents. However, the various strands of research on scholarly document processing remain fragmented. To reach out to the broader NLP and AI/ML community, pool distributed efforts in this area, and enable shared access to published research, we held the 2nd Workshop on Scholarly Document Processing (SDP) at NAACL 2021 as a virtual event (https://sdproc.org/2021/). The SDP workshop consisted of a research track, three invited talks, and three Shared Tasks (LongSumm 2021, SCIVER, and 3C). The program was geared towards the application of NLP, information retrieval, and data mining for scholarly documents, with an emphasis on identifying and providing solutions to open challenges.

Beltagy, Iz↗

Overview of the 2021 SDP 3C Citation Context Classification Shared Task

This paper provides an overview of the 2021 3C Citation Context Classification shared task. The second edition of the shared task was organised as part of the 2nd Workshop on Scholarly Document Processing (SDP 2021). The task is composed of two subtasks: classifying citations based on their (Subtask A) purpose and (Subtask B) influence. As in the previous year, both tasks were hosted on Kaggle and used a portion of the new ACT dataset. A total of 22 teams participated in Subtask A, and 19 teams competed in Subtask B. All the participated systems were ranked based on their achieved macro f-score. The highest scores of 0.26973 and 0.60025 were reported for sub-task A and B, respectively.

Kunnath, Suchetha N.↗

Challenges in Automated Detection of COVID-19 Misinformation

The COVID-19 pandemic has made the dangers of the spread of misinformation obvious but despite much global effort to curbing its spread, fake information about the pandemic keeps proliferating. In this paper, we address the development of automated methods for verification of claims about COVID-19 and discuss the challenges associated with this task. We focus on labeled data collection, limitations of existing models, and difficulties of applying misinformation detection models in practical applications. Our initial analysis indicates label imbalance may be a particular challenge for developing claim verification models and we discuss options for alleviating this issue.

Herrmannova, Dasha↗

Smoky Mountain Data Challenge 2020: An Open Call to Solve Data Problems in the Areas of Neutron Science, Material Science, Urban Modeling and Dynamics, Geophysics, and Biomedical Informatics

The 2020 Smoky Mountains Computational Sciences and Engineering Conference enlists research scientists from across Oak Ridge National Laboratory (ORNL) to be data sponsors and help create data analytics challenges for eminent data sets at the laboratory. This work describes the significance of each of the seven data sets and their as- sociated challenge questions. The challenge questions for each data set were required to cover multiple difficulty levels. An international call for participation was sent to students, and researchers asking them to form teams of up to four people to apply novel data analytics techniques to these data sets.

Parete-Koon, Suzanne↗

Scalable Knowledge Graph Analytics at 136 Petaflop/s

We are motivated by newly proposed methods for data mining large-scale corpora of scholarly publications, such as the full biomedical literature, which may consist of tens of millions of papers spanning decades of research. In this setting, analysts seek to discover how concepts relate to one another. They construct graph representations from annotated text databases and then formulate the relationship-mining problem as one of computing all-pairs shortest paths (APSP), which becomes a significant bottleneck. In this context, we present a new high-performance algorithm and implementation of the Floyd-Warshall algorithm for distributed-memory parallel computers accelerated by GPUs, which we call DSNAPSHOT (Distributed Accelerated Semiring All-Pairs Shortest Path). For our largest experiments, we ran DSNAPSHOT on a connected input graph with millions of vertices using 4, 096nodes (24,576GPUs) of the Oak Ridge National Laboratory's Summit supercomputer system. We find DSNAPSHOT achieves a sustained performance of 136×1015 floating-point operations per second (136petaflop/s) at a parallel efficiency of 90% under weak scaling and, in absolute speed, 70% of the best possible performance given our computation (in the single-precision tropical semiring or “min-plus” algebra). Looking forward, we believe this novel capability will enable the mining of scholarly knowledge corpora when embedded and integrated into artificial intelligence-driven natural language processing workflows at scale.

Kannan, Ramakrishnan {ramki}↗