Engineering Papers⌕ Search

Engineering topics

Fan, Junchuan

Publications and source records attributed to Fan, Junchuan.

Conflation of Geospatial POI Data and Ground-level Imagery

The code solves the problem of conflating POI (Points of Interest) geospatial data and geo-tagged ground-level imagery data. The main challenges with fusion or conflation of these two types of data has been that these two data entities are represented not only in several ways in current state-of-the-art, but also there is mismatch of representation format of these two data types. This source code/software brings POI datapoints and ground-level imagery datapoints into same representation format, and then generates valuable Knowledge Graph combining POI and images data (so that it can be used for multitude of applications).

De, Debraj↗

Towards POI-based large-scale land use modeling: spatial scale, semantic granularity, and geographic context

The combination of spatial distribution, semantic characteristics, and sometimes temporal dynamics of POIs inside a geographic region can capture its unique land use characteristics. Most previous studies on POI-based land use modeling research focused on one geographic region and select one spatial scale and semantic granularity for land use characterization. There is a lack of understanding on the impact of spatial scale, semantic granularity, and geographic context on POI-based land use modeling, particularly large-scale land use modeling. In this study, we developed a scalable POI-based land use modeling framework and examined the impact of these three factors on POI-based land use characterization using data from three geographic regions. We developed a unified semantic representation framework for POI semantics that can help fuse heterogeneous POI data sources. Then, by combining POIs with a neural network language model, we developed a spatially explicit approach to learn the embedding representation of POIs and AOIs. We trained multiple supervised classifiers using AOI embeddings as input features to predict AOI land use at different semantic granularities. The classification performance of different land use classes was analyzed and compared across three geographic regions to identify the semantic representativeness of POI-based AOI embedding and the impact of geographic context.

58 GEOSCIENCES↗

Accelerated Assessment of Critical Infrastructure in Aiding Recovery Efforts During Natural and Human-made Disaster

Relief and recovery from disasters (both natural and human-made) require a coordinated approach across several federal and state government agencies. In order to achieve optimal resource allocation and deployment of first responders, accurate and timely assessment of the impact and extent of destruction are the cornerstones to any recovery effort. Ideally, this knowledge should be gathered and shared within the first 0-24 hours (termed as "Acute Phase" by the U.S. CDC guideline) for informed decision-making. But achieving this poses significant challenges for the data collection and data harmonization processes, particularly when voluminous data are being generated from diverse and distributed sources during the disaster responses. To this end, this work developed a scalable and efficient workflow to dynamically collect and harmonize crowd-sourced geographic multi-modal data, and then assess critical infrastructure (CI) damaged during disaster events. We demonstrate the application of our framework with two real-world experiences in addressing post-disaster recovery efforts - for the Bahamas (Natural - due to Hurricane Dorian, 2019) and Beirut (Human-made - due to explosion caused by the ammonium nitrate stored in a warehouse, 2020). We have illustrated that a coordinated effort is needed for planning as well as for execution to achieve informed decision making.

Thakur, Gautam Malviya↗

Conflation of Geospatial POI Data and Ground-level Imagery via Link Prediction on Joint Semantic Graph

With the proliferation of smartphone cameras and social networks, we have rich, multi-modal data about points of interest (POIs) - like cultural landmarks, institutions, businesses, etc. - within a given areas of interest (AOI) (e.g., a county, city or a neighborhood) available to us. Data conflation across multiple modalities of data sources is one of the key challenges in maintaining a geographical information system (GIS) which accumulate data about POIs. Given POI data from nine different sources, and ground-level geo-tagged and scene-captioned images from two different image hosting platforms, in this work we explore the application of graph neural networks (GNNs) to perform data conflation, while leveraging a natural graph structure evident in geospatial data. The preliminary results demonstrate the capacity of a GNN operation to learn distributions of entity (POIs and images) features, coupled with topological structure of entity's local neighborhood in a semantic nearest neighbor graph, in order to predict links between a pair of entities.

Gurav, Rutuja↗

MapSpace: POI-based Multi-Scale Global Land Use Modeling

Accurate and up-to-date land use maps are important to the study of human-environment interactions, urban morphology, environmental justice, etc. Traditional land use mapping approaches involve several surveys and expert knowledge of the region to be mapped. While traditional approaches generate accurate and authoritative maps, it is expensive and takes a long time to develop a new version of map. Besides, such maps have region-specific spatial embedding, making them difficult to benchmark and compare against other land use maps. This work introduces a scalable POI-based land use modeling approach to generate global land use maps at multiple spatial scales and different semantic granularities. In addition, our land use maps adhere to a unified land use categories and can be compared for accuracy and precision.

Thakur, Gautam Malviya↗

Understanding the Drivers of Mobility during the COVID-19 Pandemic in Florida, USA Using a Machine Learning Approach

As of March 2021, the State of Florida, U.S.A. had accounted for approximately 6.67% of total COVID-19 (SARS-CoV-2 coronavirus disease) cases in the U.S. The main objective of this research is to analyze mobility patterns during a three month period in summer 2020, when COVID-19 case numbers were very high for three Florida counties, Miami-Dade, Broward, and Palm Beach counties. To investigate patterns, as well as drivers, related to changes in mobility across the tri-county region, a random forest regression model was built using sociodemographic, travel, and built environment factors, as well as COVID-19 positive case data. Mobility patterns declined in each county when new COVID-19 infections began to rise, beginning in mid-June 2020. While the mean number of bar and restaurant visits was lower overall due to closures, analysis showed that these visits remained a top factor that impacted mobility for all three counties, even with a rise in cases. Our modeling results suggest that there were mobility pattern differences between counties with respect to factors relating, for example, to race and ethnicity (different population groups factored differently in each county), as well as social distancing or travel-related factors (e.g., staying at home behaviors) over the two time periods prior to and after the spike of COVID-19 cases.

60 APPLIED LIFE SCIENCES↗

Understanding collective human movement dynamics during large-scale events using big geosocial data analytics

Conventional approaches for modeling human mobility pattern often focus on human activity and movement dynamics in their regular daily lives and cannot capture changes in human movement dynamics in response to large-scale events. With the rapid advancement of information and communication technologies, many researchers have adopted alternative data sources (e.g., cell phone records, GPS trajectory data) from private data vendors to study human movement dynamics in response to large-scale natural or societal events. Big geosocial data such as georeferenced tweets are publicly available and dynamically evolving as real-world events are happening, making it more likely to capture the real-time sentiments and responses of populations. However, precisely-geolocated geosocial data is scarce and biased toward urban population centers. In this research, we developed a big geosocial data analytical framework for extracting human movement dynamics in response to large-scale events from publicly available georeferenced tweets. The framework includes a two-stage data collection module that collects data in a more targeted fashion in order to mitigate the data scarcity issue of georeferenced tweets; in addition, a variable bandwidth kernel density estimation(VB-KDE) approach was adopted to fuse georeference information at different spatial scales, further augmenting the signals of human movement dynamics contained in georeferenced tweets. To correct for the sampling bias of georeferenced tweets, we adjusted the number of tweets for different spatial units (e.g., county, state) by population. To demonstrate the performance of the proposed analytic framework, we chose an astronomical event that occurred nationwide across the United States, i.e., the 2017 Great American Eclipse, as an example event and studied the human movement dynamics in response to this event. Finally, this analytic framework can easily be applied to other types of large-scale events such as hurricanes or earthquakes.

54 ENVIRONMENTAL SCIENCES↗

Incorporating space and time into random forest models for analyzing geospatial patterns of drug-related crime incidents in a major U.S. metropolitan area

The opioid crisis has hit American cities hard, and research on spatial and temporal patterns of drug-related activities including detecting and predicting clusters of crime incidents involving particular types of drugs is useful for distinguishing hot zones where drugs are present that in turn can further provide a basis for assessing and providing related treatment services. In this study, we investigated spatiotemporal patterns of more than 52,000 reported incidents of drug-related crime at block group granularity in Chicago, IL between 2016 and 2019. We applied a space-time analysis framework and machine learning approaches to build a model using training data that identified whether certain locations and built environment and sociodemographic factors were correlated with drug-related crime incident patterns, and establish the top contributing factors that underlaid the trends. Space and time, together with multiple driving factors, were incorporated into a random forest model to analyze these changing patterns. We accommodated both spatial and temporal autocorrelation in the model learning process to assist with capturing the changes over time and tested the capabilities of the space-time random forest model by predicting drug-related activity hot zones. Overall, we focused particularly on crime incidents that involved heroin and synthetic drugs as these have been key drug types that have highly impacted cities during the opioid crisis in the U.S.

97 MATHEMATICS AND COMPUTING↗