Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Crowdsourced Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A general spatial-temporal framework for short-term building temperature forecasting at arbitrary locations with crowdsourcing weather data

Weather forecasting has been a critical component to predict and control building energy consumption for better building energy management. Without accessibility to other data sources, the onsite observed temperatures or the airport temperatures are used in forecast models. In this paper, we present a novel approach by utilizing the crowdsourcing weather data from neighboring personal weather stations (PWS) to improve the weather forecast accuracy around buildings using a general spatial-temporal modeling framework. The final forecast is based on the ensemble of local forecasts for the target location using neighboring PWSs. Our approach is distinguished from existing literature in various aspects. First, we leverage the crowdsourcing weather data from PWS in addition to public data sources. In this way, the data is at much finer time resolution (e.g., at 5-minute frequency) and spatial resolution (e.g., arbitrary location vs grid). Second, our proposed model incorporates spatial-temporal correlation information of weather variables between the target building and a set of neighboring PWSs so that underlying correlations can be effectively captured to improve forecasting performance. Here, we demonstrate the performance of the proposed framework by comparing to the benchmark models on temperature forecasting for a building located at an arbitrary location at San Antonio, Texas, USA. In general, the proposed model framework equipped with machine learning technique such as Random Forest can improve forecasting by 50% compares with persistent model and has 90% chance to outperform airport forecast in short-term forecasting. In a real-time setting, the proposed model framework can provide more accurate temperature forecasting results compared with using airport temperature forecast for most forecast horizon. Moreover, we analyze the sensitivity of model parameters to gain insights on how crowdsourcing data from the neighboring personal weather stations impacts forecasting performance. Finally, we implement our model in other cities such as Syracuse and Chicago to test the model's performance in different landforms and climate types.

54 ENVIRONMENTAL SCIENCES↗

Improve Learning from Crowds via Generative Augmentation

Crowdsourcing provides an efficient label collection schema for supervised machine learning. However, to control annotation cost, each instance in the crowdsourced data is typically annotated by a small number of annotators. This creates a sparsity issue and limits the quality of machine learning models trained on such data. In this paper, we study how to handle sparsity in crowdsourced data using data augmentation. Specifically, we propose to directly learn a classifier by augmenting the raw sparse annotations. We implement two principles of high-quality augmentation using Generative Adversarial Networks: 1) the generated annotations should follow the distribution of authentic ones, which is measured by a discriminator; 2) the generated annotations should have high mutual information with the ground-truth labels, which is measured by an auxiliary network. Extensive experiments and comparisons against an array of state-of-the-art learning from crowds methods on three real-world datasets proved the effectiveness of our data augmentation framework. It shows the potential of our algorithm for low-budget crowdsourcing in general.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Automatic Traffic Queue-End Identification using Location-Based Waze User Reports

Traffic queues, especially queues caused by non-recurrent events such as incidents, are unexpected to high-speed drivers approaching the end of queue (EOQ) and become safety concerns. Though the topic has been extensively studied, the identification of EOQ has been limited by the spatial-temporal resolution of traditional data sources. This study explores the potential of location-based crowdsourced data, specifically Waze user reports. It presents a dynamic clustering algorithm that can group the location-based reports in real time and identify the spatial-temporal extent of congestion as well as the EOQ. The algorithm is a spatial-temporal extension of the density-based spatial clustering of applications with noise (DBSCAN) algorithm for real-time streaming data with an adaptive threshold selection procedure. Here, the proposed method was tested with 34 traffic congestion cases in the Knoxville, Tennessee area of the United States. It is demonstrated that the algorithm can effectively detect spatial-temporal extent of congestion based on Waze report clusters and identify EOQ in real-time. The Waze report-based detection are compared to the detection based on roadside sensor data. The results are promising: The EOQ identification time of Waze is similar to the EOQ detection time of traffic sensor data, with only 1.1 min difference on average. In addition, Waze generates 1.9 EOQ detection points every mile, compared to 1.8 detection points generated by traffic sensor data, suggesting the two data sources are comparable in respect of reporting frequency. The results indicate that Waze is a valuable complementary source for EOQ detection where no traffic sensors are installed.

99 GENERAL AND MISCELLANEOUS↗

The silicon citizen naturalist

Smartphone-wielding citizen scientists and an AI called FLORIST are transforming ecology at the continental scale. Here, in this issue of Cell, when Tibbs-Cortes et al. pair the crowdsourced data with controlled genetics, they discover how switchgrass times its flowering to outwit both frost and heat, depending on latitude.

Hudson, Matthew E. [University of Illinois at Urba↗

Dust Storms, Valley Fever, and Public Awareness

We discuss several issues raised by Comrie (2021, https://doi.org/10.1029/2021GH000504), which uses a crowdsourced data set to study dust storms and coccidioidomycosis (Valley fever). There is inconsistency in the term “dust storm” used by science communities. The dust data from National Oceanic and Atmospheric Administration Storm Events Database are from diverse sources, unsuitable for assessing dust-coccidioidomycosis relationships. Population exposure to dust or Coccidioides needs to consider the frequency, magnitude, and duration of dust events. Given abundant evidence that dust storms are a viable driver to transport pathogens, it is in best public interest to advocate dust storms may put people at risk for contracting Valley fever.

54 ENVIRONMENTAL SCIENCES↗

Automatic Lane-Level Road Network Extraction from Aerial Imagery for Transportation Digital Twins

Accurate road networks are essential for credible traffic microsimulation and transportation digital twins, yet high-definition maps are often difficult to obtain due to limited availability, high cost, or proprietary restrictions. Some build networks from crowdsourced data, such as OpenStreetMap, but these sources often contain geometric and semantic inconsistencies. Others create networks manually, a process that is labor-intensive and difficult to scale. To address these limitations, this work presents an end-to-end pipeline that automatically extracts georeferenced, lane-level road networks from publicly available high-resolution satellite imagery and converts them into simulation-ready assets. The developed end-to-end pipeline has three primary modules: (1) A computer-vision-based module first detects directed lane geometries and intersection layouts. (2) A heuristic-based topology construction module then identifies approach and exit legs and establishes conflict-free lane-to-lane connections. (3) Finally, an automatic simulation-building module converts the extracted network into standard formats, e.g., OpenDRIVE, and generates routable SUMO networks. The framework supports both complete network construction from scratch and local-scale refinement of existing networks through lane-count correction, transition recovery, and geometric regularization. The proposed pipeline provides a practical pathway to generate traffic simulation networks from satellite imagery, significantly reducing manual reconstruction effort and enabling scalable, continuously updated transportation digital twins.

Guo, Hetian [University of Georgia, Athens] (ORCID↗

Examining Rail Transportation Route of Crude Oil in the United States Using Crowdsourced Social Media Data

Safety issues associated with transporting crude oil by rail have been a concern since the boom of the U.S. domestic shale oil production in 2012. During the last decade, over 300 crude-oil-by-rail incidents have occurred in the United States. Some of them have caused adverse consequences including fire and hazardous materials leakage. However, only limited information on crude-on-rail routes and their associated risks is available to the public. To this end, this study proposed an unconventional way to reconstruct crude-on-rail routes using geotagged photos harvested from the Flickr website. The proposed method linked the geotagged photos of crude oil trains posted online with national railway networks to identify potential railway segments that those crude oil trains were traveling on. Here, a shortest path-based method was applied to infer the complete crude-on-rail routes, by utilizing the confirmed railway segments as well as their directional information. Validation of the inferred routes was performed using a public map and official crude oil incident data. The results suggested that the inferred routes based on geotagged photos had high coverage, with approximately 96% of the documented crude oil incidents aligned with the reconstructed crude-on-rail network. The inferred crude oil train routes were found to pass through several metropolitan areas of high population density, who were exposed to potential risk. These findings could improve situational awareness for policy makers and transportation planners. In addition, with the inferred routes, this study has established a good foundation for future crude oil train risk-analyses along the rail route.

42 ENGINEERING↗

Ultrahigh-resolution mass spectrometry data associated with the manuscript “A functional microbiome catalog crowdsourced from North American rivers"

This data package is associated with the publication “A functional microbiome catalog crowdsourced from North American rivers” submitted to Nature (Borton et al., 2024); (https://www.biorxiv.org/content/10.1101/2023.07.22.550117v1). Predicting elemental cycles and maintaining water quality under increasing anthropogenic influence requires understanding the spatial drivers of river microbiomes. However, the unifying microbial determinants governing river biogeochemistry are hindered by a lack of genome-resolved functional insights and sampling across multiple rivers. Here we employed a community science effort to accelerate the sampling of river microbiomes to create the Genome Resolved Open Watersheds database (GROWdb). GROWdb is a publicly available resource that paves the way for watershed predictive modeling and microbiome-based management practices. This resource profiled the identity, distribution, function, and expression of thousands of microbial genomes across rivers covering 90% of United States watersheds. We identified the most cosmopolitan microbiome members, while also revealing local drivers of strain endemism across ecological dimensions. We provide the first evidence that microbial functional trait expression followed the tenets of the River Continuum Concept, suggesting the structure and function of river microbiomes is predictable. The Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) data were one of many different data types used in establishing the ecological dimensions along which different microbes were detected .This data package only contains the processed FTICR-MS data associated with this manuscript; all other data is accessible via Zenodo (https://zenodo.org/records/8173287), GitHub (https://github.com/jmikayla1991/Genome-Resolved-Open-Watersheds-database-GROWdb), KBase (https://doi.org/10.25982/109073.30/1895615), and NCBI via Bioproject PRJNA946291.This dataset consists of (1) a file-level metadata (flmd) file; (2) a data dictionary (dd) file; (3) a readme; (4) three Fourier-transform ion cyclotron resonance mass spectrometry (FTICR-MS) processed data files (a ‘data’ file containing peak-by-sample observations, a ‘mol’ file containing peak metadata, and a transformation profile containing transformation-by-sample observations). All files are .csv or .pdf.

54 ENVIRONMENTAL SCIENCES↗

Crowd-based spatial risk assessment of urban flooding: Results from a municipal flood hotline in Detroit, MI

Climate change is increasing the frequency and intensity of extreme precipitation events, raising the risk of urban flood disasters. This study uses a crowd-sourced municipal call database to characterize the spatial distribution of flood risk in Detroit, MI. Call data including dates and addresses were obtained from the City of Detroit Department of Public Works for 2021. Calls were mapped and aggregated to census tract counts and merged with neighborhood-level data. Associations of predictors with flood calls were tested using spatial regression models. Flooding calls were located throughout the city but were concentrated in specific areas. Multivariate models of census tract level call counts indicated that increased poverty and Black, immigrant, and older residents were positively associated with flood calls, while increased elevation was associated with protective effects. Longer distances from waste water interceptors were associated with higher risk for calls. Crowd-sourced flood hotline call data can be used for effective spatial flood risk assessment. Though flooding occurs throughout the city of Detroit, infrastructural, neighborhood, and household factors influence flooding extent. Limitations included the self-reported nature of calls. Future modeling efforts might include input from local stakeholders to improve spatial risk assessment.

54 ENVIRONMENTAL SCIENCES↗

Urban Versus Lake Impacts on Heat Stress and Its Disparities in a Shoreline City

Abstract Shoreline cities are influenced by both urban‐scale processes and land‐water interactions, with consequences on heat exposure and its disparities. Heat exposure studies over these cities have focused on air and skin temperature, even though moisture advection from water bodies can also modulate heat stress. Here, using an ensemble of model simulations covering Chicago, we find that Lake Michigan strongly reduces heat exposure (2.75°C reduction in maximum average air temperature in Chicago) and heat stress (maximum average wet bulb globe temperature reduced by 0.86°C) during the day, while urbanization enhances them at night (2.75 and 1.57°C increases in minimum average air and wet bulb globe temperature, respectively). We also demonstrate that urban and lake impacts on temperature (particularly skin temperature), including their extremes, and lake‐to‐land gradients, are stronger than the corresponding impacts on heat stress, partly due to humidity‐related feedback. Likewise, environmental disparities across community areas in Chicago seen for skin temperature are much higher (1.29°C increase for maximum average values per $10,000 higher median income per capita) than disparities in air temperature (0.50°C increase) and wet bulb globe temperature (0.23°C increase). The results call for consistent use of physiologically relevant heat exposure metrics to accurately capture the public health implications of urbanization.

54 ENVIRONMENTAL SCIENCES↗

Linkages Between Mineral Element Composition of Soils and Sediments With Hyporheic Zone Dissolved Organic Matter Chemistry Across the Contiguous United States

The hyporheic zone is a hotspot for biogeochemical cycling where interactions with mineral metals preserve the release and biodegradation of organic matter (OM). A small fraction of OM can still be exchanged between localized sediments and the overlying water column, and recent evidence suggests there exists a longitudinal structuring in sediment dissolved OM (DOM) chemistry across the continental United States (CONUS). In this study, we tested a hypothesis that water extractable sediment DOM chemistry could be explained by sediment metal contents and integrative watershed scale features at the CONUS scale. Crowdsourced samples were characterized for high resolution mass spectrometry and coupled with sediment metals determined via x-ray fluorescence as well as with land cover and soil elemental information obtained from national databases. Our results highlight weak relationships between DOM chemistry and elemental composition at the CONUS scale indicating limited transferability of organo-metal linkages into multi-scale hydrobiogeochemical models.

58 GEOSCIENCES↗

Secondary Crash Identification using Crowdsourced Waze User Reports

Secondary crashes are crashes that occur as a result of the nonrecurrent congestion originating from primary crashes, and always have a greater impact on safety and traffic than a single crash. A better understanding of secondary crashes would benefit traffic incident management, and this requires accurate identification of secondary crashes. This study explores using crowdsourced Waze user reports to identify secondary crashes. Here, a network-based clustering algorithm is proposed to extract the primary crash cluster, including all user reports originating from the primary crash, and any crash that occurred within the cluster would be a secondary crash. This method works as a filter to select accurate primary–secondary relationships, thus precisely identifying secondary crashes. A case study is performed with crashes occurring from June to December 2019 on a 30-mi stretch of I-40 in Knoxville, TN. A static threshold method (crash duration and 10 mi) was used to preselect the potential primary–secondary crash pairs, and 75 out of 708 crashes were identified as potential secondary crashes. Based on the preselected primary–secondary crash pairs, 17 secondary crashes were obtained with the proposed method and the results were compared with one of the commonly used methods, the speed contour plot method. Though the proposed method captured fewer secondary crashes, it did identify several secondary crashes that could not be observed with the speed contour plot method. The results showed the applicability of the method and the potential of crowdsourced Waze user reports in secondary crash identification.

99 GENERAL AND MISCELLANEOUS↗

Labeling sequential data from noisy annotations

Crowdsourcing algorithms often work under the assumption that the data samples are independent. Recent work has shown that data dependence, such as temporal correlations in sequential data, can be leveraged to improve the label quality. Existing methods that exploit this special structure rely on third-order statistics of the annotator outputs to ensure the identifiability of key latent parameters, which are costly to acquire. This work proposes an approach for integrating crowdsourced annotations under the Dawid-Skene/Hidden Markov Model (DS-HMM) for sequential data based on second-order statistics, which naturally enjoys a lower sample complexity. An effective algorithm is proposed to tackle the challenging optimization problem associated with the proposed estimator. Numerical experiments showcase the effectiveness of the data labeling paradigm.

Marrinan, Timothy P.↗

Estimating building occupancy: a machine learning system for day, night, and episodic events

Building occupancy research increasingly emphasizes understanding the social and physical dynamics of how people occupy space. Opportunities in the open source domain including social media, Volunteered Geographic Information, crowdsourcing, and sensor data have proliferated, resulting in the exploration of building occupancy dynamics at varying spatiotemporal scales. At Oak Ridge National Laboratory, research into building occupancies through the development of a global learning framework that accommodates exploitation of open source authoritative sources, including governmental census and surveys, journal articles, real estate databases, and more, to report national and subnational building occupancies across the world continues through the Population Density Tables (PDT) project. This probabilistic learning system accommodates expert knowledge, experience, and open-source data to capture local, socioeconomic, and cultural information about human activity. It does so through a systematic process of data harmonization techniques in the development of observation models for over 50 building types to dynamically update baseline estimates and report probabilistic diurnal and episodic building occupancy estimates. This discussion will explore how PDT is implemented at scale and expanded based on the development of observation model classes and will explain how to interpret and spatially apply the reported probability occupancy estimates and uncertainty.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Determining the biogeochemical transformations of organic matter composition in rivers using molecular signatures

Inland waters are hotspots for biogeochemical activity, but the environmental and biological factors that govern the transformation of organic matter (OM) flowing through them are still poorly constrained. Here we evaluate data from a crowdsourced sampling campaign led by the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium to investigate broad continental-scale trends in OM composition compared to localized events that influence biogeochemical transformations. Samples from two different OM compartments, sediments and surface water, were collected from 97 streams throughout the Northern Hemisphere and analyzed to identify differences in biogeochemical processes involved in OM transformations. By using dimensional reduction techniques, we identified that putative biogeochemical transformations and microbial respiration rates vary across sediment and surface water along river continua independent of latitude (18°N–68°N). In contrast, we reveal small- and large-scale patterns in OM composition related to local (sediment vs. water column) and reach (stream order, latitude) characteristics. These patterns lay the foundation to modeling the linkage between ecological processes and biogeochemical signals. We further showed how spatial, physical, and biogeochemical factors influence the reactivity of the two OM pools in local reaches yet find emergent broad-scale patterns between OM concentrations and stream order. OM processing will likely change as hydrologic flow regimes shift and vertical mixing occurs on different spatial and temporal scales. As our planet continues to warm and the timing and magnitude of surface and subsurface flows shift, understanding changes in OM cycling across hydrologic systems is critical, given the unknown broad-scale responses and consequences for riverine OM.

rivers↗

It takes a village: using a crowdsourced approach to investigate organic matter composition in global rivers through the lens of ecological theory

Though community-based scientific approaches are becoming more common, many scientific efforts are conducted by small groups of researchers that together develop a concept, analyze data, and interpret results that ultimately translate into a publication. Here, we present a community effort that breaks these traditional boundaries of the publication process by engaging the scientific community from initial hypothesis generation to final publication. We leverage community-generated data from the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS) consortium to study organic matter composition through the lens of ecological theory. This community endeavor will use a suite of paired physical and chemical datasets collected from 97 river corridors across the globe. With our first step aimed at ideation, we engaged a community of scientists from 20 countries and 60 institutions, spanning disciplines and career stages by holding a virtual workshop (April 2021). In the workshop, participants generated content for questions, hypotheses, and proposed analyses based on the WHONDRS dataset. These ideation efforts resulted in several narratives investigating different questions led by different teams, which will be the basis for research articles in a Frontiers in Water collection. Currently, the community is collectively analyzing, interpreting, and synthesizing these data that will result in seven crowdsourced articles using a single, existing WHONDRS dataset. The use of a shared dataset across articles not only lowers barriers for broad participation by not requiring generation of new data, but also provides unique opportunities for emergent learning by connecting outcomes across studies. Here we will explain methods used to enable this community endeavor aimed to promote a greater diversity of thinking on river corridor biogeochemistry through community science.

Borton, Mikayla A.↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗