Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Crowdsourcing Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

23 records · Page 2

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Decision Making Under Uncertainty Human Subjects Data - Fire Evacuation Task

This dataset contains de-identified data from human subjects experiments, along with the images and code that were used to run the experiments (as a crowdsourced online study). In this study, participants were shown the probability of a house being in the burn zone of a wildfire. They were asked if they would stay in the house or evacuate in that scenario. The probability information was presented in different ways, including text and maps. The studies tested the impact of different visual cues on the participants' patterns of decisions.

Matzen, Laura E. [Sandia National Laboratories (SN↗

Assessing residential PM 2.5 concentrations and infiltration factors with high spatiotemporal resolution using crowdsourced sensors

Building conditions, outdoor climate, and human behavior influence residential concentrations of fine particulate matter (PM 2.5 ). To study PM 2.5 spatiotemporal variability in residences, we acquired paired indoor and outdoor PM 2.5 measurements at 3,977 residences across the United States totaling >10,000 monitor-years of time-resolved data (10-min resolution) from the PurpleAir network. Time-series analysis and statistical modeling apportioned residential PM 2.5 concentrations to outdoor sources (median residential contribution = 52% of total, coefficient of variation = 69%), episodic indoor emission events such as cooking (28%, CV = 210%) and persistent indoor sources (20%, CV = 112%). Residences in the temperate marine climate zone experienced higher infiltration factors, consistent with expectations for more time with open windows in milder climates. Likewise, for all climate zones, infiltration factors were highest in summer and lowest in winter, decreasing by approximately half in most climate zones. Large outdoor–indoor temperature differences were associated with lower infiltration factors, suggesting particle losses from active filtration occurred during heating and cooling. Absolute contributions from both outdoor and indoor sources increased during wildfire events. Infiltration factors decreased during periods of high outdoor PM 2.5 , such as during wildfires, reducing potential exposures from outdoor-origin particles but increasing potential exposures to indoor-origin particles. Time-of-day analysis reveals that episodic emission events are most frequent during mealtimes as well as on holidays (Thanksgiving and Christmas), indicating that cooking-related activities are a strong episodic emission source of indoor PM 2.5 in monitored residences.

54 ENVIRONMENTAL SCIENCES↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

Agile collaboration: Citizen science as a transdisciplinary approach to heliophysics

Citizen science connects scientists with the public to enable discovery, engaging broad audiences across the world. There are many attributes that make citizen science an asset to the field of heliophysics, including agile collaboration. Agility is the extent to which a person, group of people, technology, or project can work efficiently, pivot, and adapt to adversity. Citizen scientists are agile; they are adaptable and responsive. Citizen science projects and their underlying technology platforms are also agile in the software development sense, by utilizing beta testing and short timeframes to pivot in response to community needs. As they capture scientifically valuable data, citizen scientists can bring expertise from other fields to scientific teams. The impact of citizen science projects and communities means citizen scientists are a bridge between scientists and the public, facilitating the exchange of information. These attributes of citizen scientists form the framework of agile collaboration. In this paper, we contextualize agile collaboration primarily for aurora chasers, a group of citizen scientists actively engaged in projects and independent data gathering. Nevertheless, these insights scale across other domains and projects. Citizen science is an emerging yet proven way of enhancing the current research landscape. To tackle the next-generation’s biggest research problems, agile collaboration with citizen scientists will become necessary.

79 ASTRONOMY AND ASTROPHYSICS↗