Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Crowdsourcing Data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Integrating Machine Learning into a Crowdsourced Model for Earthquake-Induced Damage Assessment

On January 12th, 2010, a catastrophic 7.0M earthquake devastated the country of Haiti. In the aftermath of an earthquake, it is important to rapidly assess damaged areas in order to mobilize the appropriate resources. The Haiti damage assessment effort introduced a promising model that uses crowdsourcing to map damaged areas in freely available remotely-sensed data. This paper proposes the application of machine learning methods to improve this model. Specifically, we apply work on learning from multiple, imperfect experts to the assessment of volunteer reliability, and propose the use of image segmentation to automate the detection of damaged areas. We wrap both tasks in an active learning framework in order to shift volunteer effort from mapping a full catalog of images to the generation of high-quality training data. We hypothesize that the integration of machine learning into this model improves its reliability, maintains the speed of damage assessment, and allows the model to scale to higher data volumes.

crowdsourcing↗

Laboratory time series moisture manipulative experiment from sediment across the contiguous US: time series aerobic respiration and geochemistry (v2)

This dataset supports a broader study examining the effects of wetting and drying on hyporheic zone respiration across the contiguous United States (CONUS). The dataset provides data generated from a laboratory moisture manipulation experiment. The contents include time series aerobic respiration and moisture; dissolved oxygen; sediment geochemistry data; and field metadata (including qualitative information on instream and river corridor characteristics). Samples were collected as part of the WHONDRS CONUS-Scale Model-Sample Study (CM). This study was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. The data package associated with the CM study is available at https://data.ess-dive.lbl.gov/view/doi:10.15485/1923689. CM sampling began in April 2022 and ended in October 2023. This study uses subsamples from a subset of CM samples collected between June 2022 and June 2023. The original field samples were labeled as CM_###. Subsequent subsamples for this study were labeled as EC_###. The labels from the field samples and the EC subsamples can be mapped directly based on the digits following the prefix and underscore (i.e., EC_001 is a subsample from CM_001). See the critical details section below for more details on sample naming. This data package was originally published in August 2024. It was updated in February 2026 (v2; new and modified files). See the change history section in the readme for more details. For details on how to navigate this data package, see this infographic from the River Corridor SFA https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) field protocol; and a (6) a subfolder with sediment sample data from the incubation experiment. The sample data subfolder contains (1) dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC); (2) total nitrogen (TN); (3) adenosine triphosphate (ATP); (4) percent carbon and nitrogen; (5) effect size; (6) iron (II); (7) gravimetric moisture; (8) respiration rates and raw dissolved oxygen values; (9) specific conductance; (10) pH; (11) temperature; (12) a summary containing median values of each data type for each treatment (wet and dry); (13) methods codes; (14) FTICR-MS methods; and (15) a subfolder of 9.4 Tesla FTICR-MS data. This folder contains three subfolders, one containing the sediment .xml data files, one containing the sediment CoreMS output files, the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS). All files are .csv, .pdf, .R, .ref, or .xml.

54 ENVIRONMENTAL SCIENCES↗

Maximum respiration rates in hyporheic zone sediments are primarily constrained by organic carbon concentration and secondarily by organic matter chemistry

Abstract. River corridors are fundamental components of the Earth system, and their biogeochemistry can be heavily influenced by processes in subsurface zones immediately below the riverbed, referred to as the hyporheic zone. Within the hyporheic zone, organic matter (OM) fuels microbial respiration, and OM chemistry heavily influences aerobic and anaerobic biogeochemical processes. The link between OM chemistry and respiration has been hypothesized to be mediated by OM molecular diversity, whereby respiration is predicted to decrease with increasing diversity. Here we test the specific prediction that aerobic respiration rates will decrease with increases in the number of unique organic molecules (i.e., OM molecular richness, as a measure of diversity). We use publicly available data across the United States from crowdsourced samples taken by the Worldwide Hydrobiogeochemical Observation Network for Dynamic River Systems (WHONDRS) consortium. Our continental-scale analyses rejected the hypothesis of a direct limitation of respiration by OM molecular richness. In turn, we found that organic carbon (OC) concentration imposes a primary constraint over hyporheic zone respiration, with additional potential influences of OM richness. We specifically observed respiration rates to decrease nonlinearly with the ratio of OM richness to OC concentration. This relationship took the form of a constraint space with respiration rates in most systems falling below the constraint boundary. A similar, but slightly weaker, constraint boundary was observed when relating respiration rate to the inverse of OC concentration. These results indicate that maximum respiration rates may be governed primarily by OC concentration, with secondary influences from OM richness. Our results also show that other variables often suppress respiration rates below the maximum associated with the richness-to-concentration ratio. An important focus of future research will identify physical (e.g., sediment grain size), chemical (e.g., nutrient concentrations), and/or biological (e.g., microbial biomass) factors that suppress hyporheic zone respiration below the constraint boundaries observed here.

58 GEOSCIENCES↗

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

Perceived Costs and Benefits of ICON Science and Foundational Documents associated with “Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles"

This data package is associated with the publication "Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles" in Earth and Space Science (Goldman et al. 2022; https://doi.org/10.1029/2021EA002099). The manuscript is an introductory article for a special collection of commentary articles across 19 geoscience disciplines that explore the challenges and opportunities associated with the use of ICON science principles. These principles focus on research intentionally designed to be Integrated, Coordinated, Open, and Networked (ICON) with the goal of maximizing mutual benefit (among stakeholders) and cross-system transferability of science outcomes. This data package contains data, figures, and R scripts associated with the cost/benefit analysis presented in the manuscript. The writing teams involved in the special collection placed each letter of ICON on a plot with perceived cost on one axis and perceived benefit on the other to summarize their perceptions of pursuing each principle of ICON science. These data were subsequently quantified and analyzed. Files are saved as .csv, .R, and .pdf. This data package also contains (1) the public foundational and instructional documents that enabled the crowdsourced creation of the special collection; (2) file-level metadata (flmd) that lists each file in the data package with a description; (3) data dictionary (dd) that defines column headers that appear in csv files. Files are saved as .pdf and .csv.

54 ENVIRONMENTAL SCIENCES↗

Scripts and data associated with a manuscript linking soil and sediment elemental composition with dissolved organic matter chemistry across CONUS

This data package provides scripts and geochemical data for a manuscript titled “Linkages between mineral element composition of soils and sediments with hyporheic zone dissolved organic matter chemistry across the contiguous United States” (preprint: doi: 10.22541/essoar.169447343.31694990/v1). This data is associated with the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS, https://whondrs.pnnl.gov) and is an extension of the Summer 2019 Sampling campaign which crowdsourced samples from rivers and sediment across the continental United States. Data from this study can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719. The main objective of this manuscript was to couple sediment water extractable dissolved organic matter chemistry, defined by ultra-high resolution mass spectrometry, with localized sediment elemental composition and watershed scale soil elemental characteristics. This data package contains one main folder with four subfolders. The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) an R markdown to reproduce manuscript figures and analyses; (5) a pdf of instructions to reproduce NGS interpolations with ArcGIS software; and (6) a python script to reproduce NGS extrapolations with python. The four subfolders contain files required to reproduce NGS extrapolations include (1) ‘CONUS_boundaries’ containing boundary layers (.shp) for the Continental United States; (2) ‘ngs_project’ containing files (.shp) with point level NGS soil elemental data (Grossman et al., 2004); (3) ‘raster_outputs’ containing the interpolated raster output files for various soil elements; and (4) ‘NGS_Chemistry_Final’ contain final extracted soil elemental data.

54 ENVIRONMENTAL SCIENCES↗

Do Citizen Science Intense Observation Periods Increase Data Usability? A Deep Dive of the NASA GLOBE Clouds Data Set With Satellite Comparisons

The Global Learning and Observations to Benefit the Environment (GLOBE) citizen science program has recently conducted a series of month-long intensive observation periods (IOPs), asking the public to submit daily reports on cloud and sky conditions from all regions of Earth. This provides a wealth of crowdsourced observations from the ground, which complements other conventional scientific cloud data. In addition, the GLOBE reports are matched in space and time with geostationary and low Earth orbit satellites, which allows for a straightforward comparison of cloud properties, and minimizes the biases associated with mismatched sampling between participants and satellites. The matched GLOBE dataset is used to calculate the mean observed cloud cover by atmospheric level both worldwide and by region. The overall magnitudes of cloud cover between the GLOBE participants and the matched satellites agree within 10%, which is notable given the distinctly different natures of the data sources. The mean vertical cloud profiles show GLOBE reporting more low-level clouds and fewer high-level clouds than satellites. The low cloud disagreement is likely related to satellites missing low clouds when high clouds block their view. Conversely, the high cloud disagreement is related primarily to cloud opacity, as satellites may miss some optically thin clouds. Monte Carlo testing shows the results to be robust, and the tripled amount of IOP data reduces uncertainty by half. These findings also highlight ways in which citizen science IOP data may be used to support scientific research while accounting for their unique properties. Plain Language Summary: Citizen science is becoming an increasingly prominent aspect of scientific research, and so it important to study how citizen science data can be used effectively. For example, The GLOBE Program has recently conducted a series of special data-collecting events, or “challenges”, which gathered large numbers of reports on cloud and sky conditions. Because NASA GLOBE Clouds matches the participant reports with cloud observations from satellites, we can use these data to get a combined view of clouds from above and below. When looking at the average cloud cover for different atmospheric levels across Earth, we find that the GLOBE participants and the satellites agree quite closely. This is a surprising and fascinating find, given how different in nature volunteer ground reports are to satellite measurements. However, there are some small but notable disagreements between GLOBE participants and satellites about the distribution of cloud cover at different levels. In addition, by testing the data for uncertainty, we show that the results from the GLOBE data are reliable, and that more public participation improves the reliability. So, by carefully designing the analysis methodology, and by testing for the uncertainty of the data, citizen science can make a meaningful contribution to scientific research.

J. Brant Dodson↗

Machine-learning Solution for Automatic Spacesuit Motion Recognition and Measurement from Conventional Video

Extravehicular Activity (EVA) spacesuits exhibit unique movement patterns due to their design characteristics. Mobility assessments using traditional motion capture systems are cost prohibitive and not feasible for some training conditions (e.g., simulated lunar outdoor terrain). This paper aims to present the ongoing development of machine learning solutions to quantify suit motions from conventional videos without special sensors or hardware. Given the fast growth in deep/machine learning technologies, external expertise was sought from open-source communities. This was expected to accelerate development and provide more cost-effective, time-saving solutions. This work was selected for a NASA Crowdsourcing project through an agency-wide solicitation. Partnerships were formed with the NASA JSC Center of Excellence for Collaborative Innovation and an execution crowdsourcing platform partner to solicit framework developments from external contenders. NASA provided contenders with video clips of spacesuits and simultaneously measured motion capture data during EVA simulation tasks. The contenders used this data to train and develop generalized algorithms to predict motions. At the end of the crowdsourcing event, five solutions were selected from 250 submissions. Each submission was tested and scored using video clips not previously disclosed to the contenders. The scoring metrics measured how well the algorithm detected the suit shape, the 2D suit joint detection accuracy, and 3D joint detection accuracy. The winning solution was able to achieve roughly 85% prediction accuracy (weighted combination of scoring metrics). Overall, the algorithms could efficiently detect various types of spacesuits and motions across different EVA simulation environments such as the Neutral Buoyancy Lab (NBL). However, 3D joint identification is less reliable when parts of the suit were obstructed in the image. After continued improvements and validation, the fully developed system will enable EVA stakeholders to quantify suit kinematic patterns, which can help optimize suit, hardware, and task designs.

Linh Vu↗

Machine-learning Solution for Automatic Spacesuit Motion Recognition and Measurement from Conventional Video

Extravehicular Activity (EVA) spacesuits exhibit unique movement patterns due to their design characteristics. Mobility assessments using traditional motion capture systems are cost prohibitive and not feasible for some training conditions (e.g., simulated lunar outdoor terrain). This paper aims to present the ongoing development of machine learning solutions to quantify suit motions from conventional videos without special sensors or hardware. Preliminary work into this field was promising but given the fast growth in deep/machine learning technologies, external expertise was sought from open-source communities. Partnerships were formed with the NASA JSC Center of Excellence for Collaborative Innovation (CoCEI) and an execution crowdsourcing platform partner to solicit machine learning framework developments from external contenders. NASA provided contenders with images and video clips of spacesuits with simultaneously measured motion capture data during EVA simulation tasks. The contenders used this data to train and develop generalized algorithms to predict motions. At the end of the crowdsourcing event, the top five solutions were selected from 250 submissions. Each submission was tested and scored using video clips not previously disclosed to the contenders. The weighted scoring metrics measured how well the algorithm detected the suit shape, the 2D suit joint detection accuracy, and 3D joint detection accuracy. The winning solution was able to achieve roughly 85% prediction accuracy. Overall, the algorithms could efficiently detect various types of spacesuits and motions across different EVA environments such as the NASA Active Response Gravity Offload System (ARGOS). After continued improvements and validation, the fully developed system will enable EVA stakeholders to quantify suit kinematic patterns, which can help optimize suit, hardware, and task designs.

Linh Vu↗

Game Based Learning For Earth Science Applications Training

Current NASA Earth capacity development programs employ mechanisms ranging from online resource sharing, and virtual and in-person trainings to share knowledge. While these programs are highly successful at engaging individuals around the world – in 2018, over 8000 individuals and over 2000 institutions from all 50 US states and over 140 countries were engaged through over 150 projects and trainings – user feedback has highlighted the desire for expanded hands-on, practical experiences in incorporating NASA EO insights with localized data and actions. We aim to address this gap by leveraging the benefits of game-based learning to build user skills in integrating NASA and local EO data to guide decisions for climate resiliency and hazard planning. This project is being executed as a two-phase crowdsourced challenge: 1) Phase 1 will require a well-researched product concept that reflects an understanding of NASA’s Earth data and tools and user needs, and proposes an innovative and interactive game or extended reality experience to train users in identifying relevant NASA data and applying insights to their climate resiliency decisions; 2) Winners of Phase 1 will be provided seed funding to develop a working prototype of the product. We aim to award 1-3 final winners to support the development of more than one game, thereby ensuring that NASA's diverse audiences around the world can access training games that best suit their needs and capabilities. This EO training game project fits in the NASA Earth Science Applied Sciences Program’s Capacity Development Program, contributing to the program mission of “helping people around the world better understand [NASA’s Earth] data and find ways to use them” (https://appliedsciences.nasa.gov/what-we-do/capacity-building). The final training game will complement existing programmatic activities of workforce development, trainings, and collaborative projects, while providing the unique value of providing interactive experiences to users and collecting real-time data and feedback to improve NASA’s Earth applications’ products and services related to climate resilience.

Human centered design↗

Populating a Graph Database to Run a Usage-Based Discovery Tool

Most dataset discovery tools for Earth Observation data rely on descriptions and other metadata of the datasets, using keyword searches or attribute filtering to determine relevance. However, these descriptions often do not include the potential uses of the data. Thus, a user working on floods will rarely see few if any rainfall datasets show up in such a search. The Usage Based Discovery tool, on the other hand, offers usage instances to the user, either research articles or applications, along with the datasets that those usage instances used. This allows a user, particularly one new to the world of Earth Observation data, to investigate which datasets are used in similar cases. The information that powers Usage-Based Discovery is a graph database of relationships of usage to dataset and usage to topic, allowing the user to narrow their search for similar cases. In order to scale out to a graph database rich enough to provide a satisfactory user experience, we combine manual and automated processes to populate the graph. The initial content of the graph has been seeded primarily via human-aided data curation methods, using sites like Google Scholar. To scale up this effort, we’ve employed crowdsourcing. It is easy for anyone to contribute to our graph using their Open Researcher and Contributor Identifier for authorization. We’re now experimenting with Machine Learning and Natural Language Processing to help automate population of the graph, starting with the classification of research articles by topic. Finding adequate training data in the absence of a comprehensive and open research article API continues to be a significant challenge.

Vincent Inverso↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Decision Making Under Uncertainty Human Subjects Data - Fire Evacuation Task

This dataset contains de-identified data from human subjects experiments, along with the images and code that were used to run the experiments (as a crowdsourced online study). In this study, participants were shown the probability of a house being in the burn zone of a wildfire. They were asked if they would stay in the house or evacuate in that scenario. The probability information was presented in different ways, including text and maps. The studies tested the impact of different visual cues on the participants' patterns of decisions.

Matzen, Laura E. [Sandia National Laboratories (SN↗

Mining Twitter Data to Augment NASA GPM Validation

The Twitter data stream is an important new source of real-time and historical global information for potentially augmenting the validation program of NASA's Global Precipitation Measurement (GPM) mission. There have been other similar uses of Twitter, though mostly related to natural hazards monitoring and management. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. Twitter provides a large source of crowd for crowdsourcing. During a 24-hour period in the middle of the snow storm this past March in the U.S. Northeast, we collected more than 13,000 relevant precipitation tweets with exact geolocation. The overall objective of our project is to determine the extent to which processed tweets can provide additional information that improves the validation of GPM data. Though our current effort focuses on tweets and precipitation, our approach is general and applicable to other social media and other geophysical measurements. Specifically, we have developed an operational infrastructure for processing tweets, in a format suitable for analysis with GPM data; engaged with potential participants, both passive and active, to "enrich" the Twitter stream; and inter-compared "precipitation" tweet data, ground station data, and GPM retrievals. In this presentation, we detail the technical capabilities of our tweet processing infrastructure, including data abstraction, feature extraction, search engine, context-awareness, real-time processing, and high volume (big) data processing; various means for "enriching" the Twitter stream; and results of inter-comparisons. Our project should bring a new kind of visibility to Twitter and engender a new kind of appreciation of the value of Twitter by the science research communities.

validatio↗

Using Social Media and Mobile Devices to Discover and Share Disaster Data Products Derived From Satellites

Data products derived from Earth observing satellites are difficult to find and share without specialized software and often times a highly paid and specialized staff. For our research effort, we endeavored to prototype a distributed architecture that depends on a standardized communication protocol and applications program interface (API) that makes it easy for anyone to discover and access disaster related data. Providers can easily supply the public with their disaster related products by building an adapter for our API. Users can use the API to browse and find products that relate to the disaster at hand, without a centralized catalogue, for example floods, and then are able to share that data via social media. Furthermore, a longerterm goal for this architecture is to enable other users who see the shared disaster product to be able to generate the same product for other areas of interest via simple point and click actions on the API on their mobile device. Furthermore, the user will be able to edit the data with on the ground local observations and return the updated information to the original repository of this information if configured for this function. This architecture leverages SensorWeb functionality [1] presented at previous IGARSS conferences. The architecture is divided into two pieces, the frontend, which is the GeoSocial API, and the backend, which is a standardized disaster node that knows how to talk to other disaster nodes, and also can communicate with the GeoSocial API. The GeoSocial API, along with the disaster node basic functionality enables crowdsourcing and thus can leverage insitu observations by people external to a group to perform tasks such as improving water reference maps, which are maps of existing water before floods. This can lower the cost of generating precision water maps. Keywords-Data Discovery, Disaster Decision Support, Disaster Management, Interoperability, CEOS WGISS Disaster Architecture

Dust Mitigation Technology to Enable Survive the Night Capabilities

Introduction: As we return to the Moon, the lunar regolith (i.e. lunar dust) covering the surface will be an obstacle to nominal operations. Accounts from Apollo astronauts and analysis of hardware returned from the surface illustrate just how deleterious the dust can be [1]. During Apollo missions, the lunar dust adhered to hardware mechanically and electrostatically [2]. Surviving the Night: Mitigating the lunar dust will be critical to surviving the night. Going hand-in-hand with other extreme environment considerations, dust mitigation is critical to mission success. Dust Impacts on Other Systems: The lunar dust can have negative implications for power, thermal, mechanisms, and several other systems or sub-systems. For example, Apollo encountered marked degradation of performance in heat rejection systems for the lunar roving vehicle, science packages, and other components because of the lunar dust [1]. For power alone, dust can cause internal clogging for power connectors, heat rejection issues, excessive dust on reflective surfaces, reduced power output for solar arrays, and so on. Dust Mitigation Strategy: In addition to considering technology solutions, it is important for hardware, systems, and or components to have a dust mitigation strategy. At a high level, hardware that will encounter the lunar dust should consider these things when defining a dust mitigation strategy: • Understand Natural Environment • Understand Induced Environment • Understand Tolerance to Dust • Write Dust Requirements • Select Dust Mitigation Solutions • Test Hardware in Dusty Environment More information on each of these can be provided to hardware owners. Dust Mitigation Technology Development: NASA has a series of technologies that may be available for hardware that needs to survive the lunar night. Many of these solutions are leveraging dust mitigation technology development efforts from NASA’s Space Technology Mission Directorate (STMD), as well as efforts from ESDMD programs, industry, and academia. Through a series of STMD programs (both internal to NASA and through partnerships), there are several technologies in development as considerations as dust mitigation solutions for hardware. Within STMD, the Game Changing Development Program (GCD) has funded several internal dust mitigation projects including low to mid TRL development, demonstrations on CLPS landers of high TRL solutions, and creating standards and best practices for dust mitigation. STMD dust mitigation efforts also include a series of partnerships for developing technologies and advancing the state of dust mitigation at NASA. This includes the Lunar Surface Innovation Consortium (LSIC), Small Business Innovation Research, Early Stage Innovations (ESI), Space Technology Research Grants (STRG), Announcement of Collaboration Opportunities (ACOs) and Tipping Points (TPs), and Challenges and Crowdsourcing, among others. There are also a series of dust mitigation solutions that have been widely used terrestrially, or during Apollo. In recent years, several studies have produced more data on the efficacy of these potential solutions in the lunar environment. Dust Mitigation Solutions: Dust mitigation solutions generally fall into four categories: • Dust Tolerant Mechanisms • Passive Dust Mitigation Capabilities • Active Dust Mitigation Capabilities • Dust Measurement Capabilities There are a series of solutions that may prove beneficial for hardware that needs to survive the lunar night, including new technology development as well as proven, terrestrial solutions. This presentation will discuss in more detail what some of these solutions are for payloads going to the surface. References: [1] J. R. Gaier, NASA/TM—2005-213610, The Effects of Lunar Dust on EVA Systems During the Apollo Missions [2] T. J. Stubbs, et al. Impact of Dust on Lunar Exploration, 2005

dust mitigation↗