Engineering Papers⌕ Search

Engineering topics

Eaton, Bryan

Publications and source records attributed to Eaton, Bryan.

A Taxonomic Classification Approach for Global Spatio-temporal Data

The World Bank, World Health Organization, and other major vendors collectively provide thousands of global time series datasets that focus on issues of the environment, public health, economics, violence, education, and national security. Sorting these data into meaningful information requires the use of data mining techniques to cluster trends into an orderly and manageable number of cases. The World SpatioTemporal Analytics and Mapping (WSTAMP) project database (wstamp.ornl.gov) was developed to spatiotemporally harmonize global vendor data (23,300+ attributes, 200+ locations, 50+ years). Within the WSTAMP analytical environment, Dynamic Time Warping (DTW) has been a highly effective data-driven approach for clustering and mapping these time series into national spatiotemporal behavior maps. Two significant properties have surfaced from this work. First, several recognizable cluster patterns have emerged and persist across a range of locations, attributes, and time frames (e.g., increasing, decreasing, rebounding, peak, oscillating). Secondly, practitioners engaging WSTAMP have noted the explanatory and anticipatory value of these patterns and articulated particular interest in detecting them within the spatiotemporal cube. This need was addressed by shifting DTW-based clustering from an open ended, data-driven implementation to a taxonomic pattern matching approach. This paper presents the method including implementation strategies for visualization and human computer interaction and applies the approach to a sample data set and concludes with next steps.

Stewart, Robert↗

Accelerated Assessment of Critical Infrastructure in Aiding Recovery Efforts During Natural and Human-made Disaster

Relief and recovery from disasters (both natural and human-made) require a coordinated approach across several federal and state government agencies. In order to achieve optimal resource allocation and deployment of first responders, accurate and timely assessment of the impact and extent of destruction are the cornerstones to any recovery effort. Ideally, this knowledge should be gathered and shared within the first 0-24 hours (termed as "Acute Phase" by the U.S. CDC guideline) for informed decision-making. But achieving this poses significant challenges for the data collection and data harmonization processes, particularly when voluminous data are being generated from diverse and distributed sources during the disaster responses. To this end, this work developed a scalable and efficient workflow to dynamically collect and harmonize crowd-sourced geographic multi-modal data, and then assess critical infrastructure (CI) damaged during disaster events. We demonstrate the application of our framework with two real-world experiences in addressing post-disaster recovery efforts - for the Bahamas (Natural - due to Hurricane Dorian, 2019) and Beirut (Human-made - due to explosion caused by the ammonium nitrate stored in a warehouse, 2020). We have illustrated that a coordinated effort is needed for planning as well as for execution to achieve informed decision making.

Thakur, Gautam Malviya↗

A Mixed-Method Design Approach for Empirically Based Selection of Unbiased Data Annotators

Implicit bias embedded in the annotated data is by far the greatest impediment in the effectual use of supervised machine learning models in tasks involving race, ethics, and geopolitical polarization. For societal good and demonstrable positive impact on wider society, it is paramount to carefully select data annotators and rigorously validate the annotation process. Current approaches to selecting annotators are not sufficiently grounded in scientific principles and are limited at the policy-guidance level, thereby rendering them unusable for machine learning practitioners. This work proposes a new approach based on the mixed-methods design that is functional, adaptable, and simpler to implement in selecting unbiased annotators for any machine learning problem. By demonstrating it on a real-world geopolitical problem, we also identified and ranked key inane profile characteristics towards an empirically-based selection of unbiased data annotators.

Thakur, Gautam Malviya↗

Challenges in Automated Detection of COVID-19 Misinformation

The COVID-19 pandemic has made the dangers of the spread of misinformation obvious but despite much global effort to curbing its spread, fake information about the pandemic keeps proliferating. In this paper, we address the development of automated methods for verification of claims about COVID-19 and discuss the challenges associated with this task. We focus on labeled data collection, limitations of existing models, and difficulties of applying misinformation detection models in practical applications. Our initial analysis indicates label imbalance may be a particular challenge for developing claim verification models and we discuss options for alleviating this issue.

Herrmannova, Dasha↗