Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Crowdsourced data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Maximum respiration rates in hyporheic zone sediments are primarily constrained by organic carbon concentration and secondarily by organic matter chemistry

Abstract. River corridors are fundamental components of the Earth system, and their biogeochemistry can be heavily influenced by processes in subsurface zones immediately below the riverbed, referred to as the hyporheic zone. Within the hyporheic zone, organic matter (OM) fuels microbial respiration, and OM chemistry heavily influences aerobic and anaerobic biogeochemical processes. The link between OM chemistry and respiration has been hypothesized to be mediated by OM molecular diversity, whereby respiration is predicted to decrease with increasing diversity. Here we test the specific prediction that aerobic respiration rates will decrease with increases in the number of unique organic molecules (i.e., OM molecular richness, as a measure of diversity). We use publicly available data across the United States from crowdsourced samples taken by the Worldwide Hydrobiogeochemical Observation Network for Dynamic River Systems (WHONDRS) consortium. Our continental-scale analyses rejected the hypothesis of a direct limitation of respiration by OM molecular richness. In turn, we found that organic carbon (OC) concentration imposes a primary constraint over hyporheic zone respiration, with additional potential influences of OM richness. We specifically observed respiration rates to decrease nonlinearly with the ratio of OM richness to OC concentration. This relationship took the form of a constraint space with respiration rates in most systems falling below the constraint boundary. A similar, but slightly weaker, constraint boundary was observed when relating respiration rate to the inverse of OC concentration. These results indicate that maximum respiration rates may be governed primarily by OC concentration, with secondary influences from OM richness. Our results also show that other variables often suppress respiration rates below the maximum associated with the richness-to-concentration ratio. An important focus of future research will identify physical (e.g., sediment grain size), chemical (e.g., nutrient concentrations), and/or biological (e.g., microbial biomass) factors that suppress hyporheic zone respiration below the constraint boundaries observed here.

58 GEOSCIENCES↗

Crowdsourcing biocuration: The Community Assessment of Community Annotation with Ontologies (CACAO)

Experimental data about gene functions curated from the primary literature have enormous value for research scientists in understanding biology. Using the Gene Ontology (GO), manual curation by experts has provided an important resource for studying gene function, especially within model organisms. Unprecedented expansion of the scientific literature and validation of the predicted proteins have increased both data value and the challenges of keeping pace. Capturing literature-based functional annotations is limited by the ability of biocurators to handle the massive and rapidly growing scientific literature. Within the community-oriented wiki framework for GO annotation called the Gene Ontology Normal Usage Tracking System (GONUTS), we describe an approach to expand biocuration through crowdsourcing with undergraduates. This multiplies the number of high-quality annotations in international databases, enriches our coverage of the literature on normal gene function, and pushes the field in new directions. From an intercollegiate competition judged by experienced biocurators, Community Assessment of Community Annotation with Ontologies (CACAO), we have contributed nearly 5,000 literature-based annotations. Many of those annotations are to organisms not currently well-represented within GO. Over a 10-year history, our community contributors have spurred changes to the ontology not traditionally covered by professional biocurators. The CACAO principle of relying on community members to participate in and shape the future of biocuration in GO is a powerful and scalable model used to promote the scientific enterprise. It also provides undergraduate students with a unique and enriching introduction to critical reading of primary literature and acquisition of marketable skills.

59 BASIC BIOLOGICAL SCIENCES↗

Incorporating uncertainty for enhanced leaderboard scoring and ranking in data competitions

Data competitions have become a popular and cost-effective approach for crowdsourcing versatile solutions from diverse expertise. Current practice relies on the simple leaderboard scoring based on a given set of competition data for ranking competitors and distributing the prize. However, a disadvantage of this practice in many competitions is that a slight difference in the scores due to the natural variability of the observed data could result in a much larger difference in the prize amounts. In this article, we propose a new strategy to quantify the uncertainty in the rankings and scores from using different data sets that share common characteristics with the provided competition data. By using a bootstrap approach to generate many comparable data sets, the new method has four advantages over current practice. Furthermore, during the competition, it provides a mechanism for competitors to get feedback about the uncertainty in their relative ranking. After the competition, it allows the host to gain a deeper understanding of the algorithm performance and their robustness across representative data sets. It also offers a transparent mechanism for prize distribution to reward the competitors more fairly with superior and robust performance. Finally, it has the additional advantage of being able to explore what results might have looked like if competition goals evolved from their original choices. The implementation of the strategy is illustrated with a real data competition hosted by Topcoder on urban radiation search.

42 ENGINEERING↗

Perceived Costs and Benefits of ICON Science and Foundational Documents associated with “Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles"

This data package is associated with the publication "Integrated, Coordinated, Open, and Networked (ICON) Science to Advance the Geosciences: Introduction and Synthesis of a Special Collection of Commentary Articles" in Earth and Space Science (Goldman et al. 2022; https://doi.org/10.1029/2021EA002099). The manuscript is an introductory article for a special collection of commentary articles across 19 geoscience disciplines that explore the challenges and opportunities associated with the use of ICON science principles. These principles focus on research intentionally designed to be Integrated, Coordinated, Open, and Networked (ICON) with the goal of maximizing mutual benefit (among stakeholders) and cross-system transferability of science outcomes. This data package contains data, figures, and R scripts associated with the cost/benefit analysis presented in the manuscript. The writing teams involved in the special collection placed each letter of ICON on a plot with perceived cost on one axis and perceived benefit on the other to summarize their perceptions of pursuing each principle of ICON science. These data were subsequently quantified and analyzed. Files are saved as .csv, .R, and .pdf. This data package also contains (1) the public foundational and instructional documents that enabled the crowdsourced creation of the special collection; (2) file-level metadata (flmd) that lists each file in the data package with a description; (3) data dictionary (dd) that defines column headers that appear in csv files. Files are saved as .pdf and .csv.

54 ENVIRONMENTAL SCIENCES↗

Scripts and data associated with a manuscript linking soil and sediment elemental composition with dissolved organic matter chemistry across CONUS

This data package provides scripts and geochemical data for a manuscript titled “Linkages between mineral element composition of soils and sediments with hyporheic zone dissolved organic matter chemistry across the contiguous United States” (preprint: doi: 10.22541/essoar.169447343.31694990/v1). This data is associated with the Worldwide Hydrobiogeochemistry Observation Network for Dynamic River Systems (WHONDRS, https://whondrs.pnnl.gov) and is an extension of the Summer 2019 Sampling campaign which crowdsourced samples from rivers and sediment across the continental United States. Data from this study can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775 and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719. The main objective of this manuscript was to couple sediment water extractable dissolved organic matter chemistry, defined by ultra-high resolution mass spectrometry, with localized sediment elemental composition and watershed scale soil elemental characteristics. This data package contains one main folder with four subfolders. The main data folder contains (1) readme; (2) data dictionary (dd); (3) file-level metadata (flmd); (4) an R markdown to reproduce manuscript figures and analyses; (5) a pdf of instructions to reproduce NGS interpolations with ArcGIS software; and (6) a python script to reproduce NGS extrapolations with python. The four subfolders contain files required to reproduce NGS extrapolations include (1) ‘CONUS_boundaries’ containing boundary layers (.shp) for the Continental United States; (2) ‘ngs_project’ containing files (.shp) with point level NGS soil elemental data (Grossman et al., 2004); (3) ‘raster_outputs’ containing the interpolated raster output files for various soil elements; and (4) ‘NGS_Chemistry_Final’ contain final extracted soil elemental data.

54 ENVIRONMENTAL SCIENCES↗

Learning from Crowds by Modeling Common Confusions

Crowdsourcing provides a practical way to obtain large amounts of labeled data at a low cost. However, the annotation quality of annotators varies considerably, which imposes new challenges in learning a high-quality model from the crowdsourced annotations. In this work, we provide a new perspective to decompose annotation noise into common noise and individual noise and differentiate the source of confusion based on instance difficulty and annotator expertise on a per-instance-annotator basis. We realize this new crowdsourcing model by an end-to-end learning solution with two types of noise adaptation layers: one is shared across annotators to capture their commonly shared confusions, and the other one is pertaining to each annotator to realize individual confusion. To recognize the source of noise in each annotation, we use an auxiliary network to choose from the two noise adaptation layers with respect to both instances and annotators. Extensive experiments on both synthesized and real-world benchmarks demonstrate the effectiveness of our proposed common noise adaptation solution.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Radiation Detection Data Competition Report

In FY2018 through FY2020, NA-22, the Defense Nuclear Nonproliferation Research and Development Program, funded a Data Science project to develop and implement statistical methodology to effectively host data competitions with the goal of leveraging the opportunity provided by crowdsourcing. By accessing and engaging expertise from a broader research community, there is an opportunity to attract innovative solutions from a variety of different research disciplines to advance the ability to solve important non-proliferation problems. This report summarizes the key results of this project after hosting two data competitions focused on urban radiation detection. The first competition was focused on attracting participants from the U.S. national laboratories, while the second, hosted by TopCoder, was open to the broader international community and awarded prize money to the top 10 competitors. At the start of the project, there was strong interest from NA-22 to explore and develop the capability to host data competitions as a means of leveraging the broader community to solve important nuclear nonproliferation problems. Having a standard data set on which to compare different approaches based on clearly defined criteria was desirable to be able to evaluate the state of solutions for important problems. Initially, it was not clear that it would even be possible logistically and bureaucratically to host a competition with an international field of competitors and to award the prize money needed to attract solutions from top competitors. Happily, a path to host the competitions was ultimately found that allowed this powerful accelerator of improvements to be leveraged.

61 RADIATION PROTECTION AND DOSIMETRY↗

WHONDRS River Corridor Sediment and Water Geochemistry and In Situ Sensor Data from Machine-Learning-Informed Sites across the Contiguous United States (v6)

This dataset supports a broader study examining hyporheic zone respiration rates to improve predictive models at a contiguous United States (CONUS) scale. The CONUS-Scale Model-Sample Study (CM) was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the CONUS. New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Sampling began in April 2022 and ended in October 2023. In addition to the widely distributed CONUS sites, a more spatially focused sampling occurred in the Yakima River Basin, WA in summer 2022. Data from this more spatially intensive sampling occurred under the label “Second Spatial Study (SSS)” and were also included in the machine learning models. Other data types collected from SSS that were not part of CM were published in a separate data package (https://data.ess-dive.lbl.gov/view/doi:10.15485/1969566). This data package was originally published in February 2023. It was updated in June 2023 (v2; new and modified files); December 2023 (v3; new and modified files); June 2024 (v4; new and modified files); April 2024 (v5; new and modified files); and September 2025 (v6; modified files). See the change history section in the readme for more details. For details on how to navigate data packages generated by this project, see https://data.ess-dive.lbl.gov/portals/PNNLRiverCorridorSFA/About. This dataset is comprised of two folders of field photos and videos, one folder of raw Fourier transform ion cyclotron resonance mass spectrometry (FTICR-MS) data and one main data folder containing (1) file-level metadata; (2) data dictionary; (3) field metadata; (4) readme; (5) international generic sample number (IGSN) mapping file; (6) field protocols; (7) a subfolder with sample data; and (8) a subfolder with sensor data. The sample data subfolder contains (1) surface water and sediment dissolved organic carbon (DOC, measured as non-purgeable organic carbon, NPOC) data and averages; (2) surface water and sediment total nitrogen data and averages; (3) surface water major cations and anions and averages; (4) sediment grain size data; (5) sediment iron (II) data and averages; (6) wet sediment mass, dry sediment mass, water mass, and wet sediment volume in incubation and sediment ICR vials; (7) sediment incubation respiration rate data and averages; (8) normalized respiration rate data and averages; (9) methods codes; (10) sediment specific surface area; (11) sediment percent carbon and nitrogen; (12) sediment gravimetric moisture and averages; (15) sediment X-ray diffraction (XRD) data; (16) sediment adenosine triphosphate (ATP) and averages; (17) a subfolder with sediment incubation respiration data, scripts, and plots; (18) surface water and sediment FTICR methods; and (19) a subfolder of 9.4 Tesla (9.4T) FTICR-MS data. This folder contains five subfolders, one containing the sediment .xml data files, one containing the water .xml files, one containing the sediment CoreMS output files, one containing the water CoreMS output files, and the other containing instructions and scripts for processing the files in CoreMS (https://github.com/EMSL-Computing/CoreMS).The sensor data subfolder contains (1) a subfolder with miniDOT dissolved oxygen and temperature data and plots; (2) miniDOT dissolved oxygen and temperature summary data; and (3) miniDOT installation methods. All files are .csv, .pdf, .R, .xml, .d, .html, .Rmd, .py, .cal, .json, .jpg, .jpeg, .png, .mov, or .mp4. CORRECTION: Carbon and nitrogen content are reported as percentages. The current column headers "01395_C_percent_per_mg" and "01397_N_percent_per_mg" are incorrect. These should read "01395_C_percent" and "01397_N_percent" and will be corrected in the next version of this data package. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Machine learning model inputs, outputs, and scripts associated with “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions”

NOTE: The manuscript associated with this data package is currently in review. The data may be revised based on reviewer feedback. Upon manuscript acceptance, this data package will be updated with the final dataset and additional metadata. This data package is associated with the manuscript “Artificial intelligence-guided iterations between observations and modeling significantly improve environmental predictions” (Malhotra et al., in prep). This effort was designed following ICON (integrated, coordinated, open, and networked) principles to facilitate a model-experiment (ModEx) iteration approach, leveraging crowdsourced sampling across the contiguous United States (CONUS). New machine learning models were created every month to guide sampling locations. Data from the resulting samples were used to test and rebuild the machine learning models for the next round of sampling guidance. Associated sediment and water geochemistry and in situ sensor data can be found at https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1923689, https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1729719, and https://data.ess-dive.lbl.gov/datasets/doi:10.15485/1603775. This data package is associated with two GitHub repositories found at https://github.com/parallelworks/dynamic-learning-rivers and https://github.com/WHONDRS-Hub/ICON-ModEx_Open_Manuscript. In addition to this readme, this data package also includes two file-level metadata (FLMD) files that describes each file and two data dictionaries (DD) that describe all column/row headers and variable definitions. This data package consists of two main folders (1) dynamic-learning-rivers and (2) ICON-ModEx_Open_Manuscript which contain snapshots of the associated GitHub repositories. The input data, output data, and machine learning models used to guide sampling locations are within dynamic-learning-rivers. The folder is organized into five top-level directories: (1) “input_data” holds the training data for the ML models; (2) “ml_models” holds machine learning (ML) models trained on the data in “input_data”; (3) “examples” contains files for direct experimentation with the machine learning model, including scripts for setting up “hindcast” run; (4) “scripts” contains data preprocessing and postprocessing scripts and intermediate results specific to this data set that bookend the ML workflow; and (5) “output_data” holds the overall results of the ML model on that branch. Each trained ML model resides on its own branch in the repository; this means that inputs and outputs can be different branch-to-branch. There is also one hidden directory “.github/workflows”. This hidden directory contains information for how to run the ML workflow as an end-to-end automated GitHub Action but it is not needed for reusing the ML models archived here. Please see the top-level README.md in the GitHub repository for more details on the automation. The scripts and data used to create figures in the manuscript are within ICON-ModEx_Open_Manuscript. The folder is organized into four folders which contain the scripts, data, and pdf for each figure. Within the “fig-model-score-evolution” folder, there is a folder called “intermediate_branch_data” which contains some intermediate files pulled from dynamic-learning-rivers and reorganized to easily integrate into the workflows. NOTE: THIS FOLDER INCLUDES THE FILES AT THE POINT OF PAPER SUBMISSION. IT WILL BE UPDATED ONCE THE PAPER IS ACCEPTED WITH ANY REVISIONS AND WILL INCLUDE A DD/FLMD AT THAT POINT. We thank the United States Forest Service, Washington Department of Fish and Wildlife, Washington Department of Natural Resources, Cowiche Canyon Conservatory, Washington State Parks and Recreation Commission (Scientific Research Permit #210901), and the Confederated Tribes and Bands of the Yakama Nation for access to field locations where the samples labeled “SSS” were collected. We also thank the Yakama Nation Tribal Council and Yakama Nation Fisheries for working with us to facilitate sample collection and optimization of data usage according to their values and worldview. WHONDRS consortium members were asked to provide any acknowledgments for the collection of samples labeled “CM” and the following is a list of acknowledgments that were submitted with their corresponding Site IDs: (MART) Research activities were conducted in part on the Wind River Experimental Forest within the Gifford Pinchot National Forest; (MP- 100379) Philadelphia is part of Lenapehoking, the ancestral homelands of the Lenape peoples; (MP-102398) Land surveyed is the ancestral homelands of the Nookhose'iinenno (Arapaho), Tsis tsis'tas (Cheyenne), and Nuuchu (Ute); (MP-100749 and MP- 100747) Georgia Coastal Ecosystem LTER, OCE-1832178; (SP-70 and SP-72) Eastern Shoshone, Shoshone-Bannock; (MP- 102944) Funded by Oregon Watershed Enhancement Board. On the traditional lands of the Confederated Tribes of the Siletz, Confederated Tribes of the Grand Rhonde, and the Clatsop-Nehalem Confederated Tribe; (MP- 100607) Holiday Creek is located on the traditional territory of the Monacan Indian Nation; (SP-45) Lafayette Blue Springs State Park; (MP-102420) NSF DEB-2016749; (MP-100019) New Hampshire Agriculture Experiment Station; (SP-35) Rayonier (land owner; https://www.rayonier.com/); (MP- 101276) US Department of Energy, Office of Science, Biological and Environmental Research, Subsurface Biogeochemical Research, Watershed Dynamics and Evolution SFA at ORNL; (MP- 103224) Watershed Dynamics and Evolution SFA at ORNL; (MP- 101584) Traditional lands of the Oceti Sakowin (Dakota, Lakota, Nakoda) and Anishinaabe Peoples.

54 ENVIRONMENTAL SCIENCES↗

Camera settings and biome influence the accuracy of citizen science approaches to camera trap image classification

Scientists are increasingly using volunteer efforts of citizen scientists to classify images captured by motion-activated trail cameras. The rising popularity of citizen science reflects its potential to engage the public in conservation science and accelerate processing of the large volume of images generated by trail cameras. While image classification accuracy by citizen scientists can vary across species, the influence of other factors on accuracy is poorly understood. Inaccuracy diminishes the value of citizen science derived data and prompts the need for specific best-practice protocols to decrease error. We compare the accuracy between three programs that use crowdsourced citizen scientists to process images online: Snapshot Serengeti, Wildwatch Kenya, and AmazonCam Tambopata. We hypothesized that habitat type and camera settings would influence accuracy. To evaluate these factors, each photograph was circulated to multiple volunteers. All volunteer classifications were aggregated to a single best answer for each photograph using a plurality algorithm. Subsequently, a subset of these images underwent expert review and were compared to the citizen scientist results. Classification errors were categorized by the nature of the error (e.g., false species or false empty), and reason for the false classification (e.g., misidentification). Our results show that Snapshot Serengeti had the highest accuracy (97.9%), followed by AmazonCam Tambopata (93.5%), then Wildwatch Kenya (83.4%). Error type was influenced by habitat, with false empty images more prevalent in open-grassy habitat (27%) compared to woodlands (10%). For medium to large animal surveys across all habitat types, our results suggest that to significantly improve accuracy in crowdsourced projects, researchers should use a trail camera set up protocol with a burst of three consecutive photographs, a short field of view, and determine camera sensitivity settings based on in situ testing. Accuracy level comparisons such as this study can improve reliability of future citizen science projects, and subsequently encourage the increased use of such data.

54 ENVIRONMENTAL SCIENCES↗

Decision Making Under Uncertainty Human Subjects Data - Fire Evacuation Task

This dataset contains de-identified data from human subjects experiments, along with the images and code that were used to run the experiments (as a crowdsourced online study). In this study, participants were shown the probability of a house being in the burn zone of a wildfire. They were asked if they would stay in the house or evacuate in that scenario. The probability information was presented in different ways, including text and maps. The studies tested the impact of different visual cues on the participants' patterns of decisions.

Matzen, Laura E. [Sandia National Laboratories (SN↗

Assessing the Reliability of Relevant Tweets and Validation Using Manual and Automatic Approaches for Flood Risk Communication

While Twitter has been touted as a preeminent source of up-to-date information on hazard events, the reliability of tweets is still a concern. Our previous publication extracted relevant tweets containing information about the 2013 Colorado flood event and its impacts. Using the relevant tweets, this research further examined the reliability (accuracy and trueness) of the tweets by examining the text and image content and comparing them to other publicly available data sources. Both manual identification of text information and automated (Google Cloud Vision, application programming interface (API)) extraction of images were implemented to balance accurate information verification and efficient processing time. The results showed that both the text and images contained useful information about damaged/flooded roads/streets. This information will help emergency response coordination efforts and informed allocation of resources when enough tweets contain geocoordinates or location/venue names. This research will identify reliable crowdsourced risk information to facilitate near real-time emergency response through better use of crowdsourced risk communication platforms.

54 ENVIRONMENTAL SCIENCES↗

Galaxy Zoo: 3D – crowdsourced bar, spiral, and foreground star masks for MaNGA target galaxies

ABSTRACT The challenge of consistent identification of internal structure in galaxies – in particular disc galaxy components like spiral arms, bars, and bulges – has hindered our ability to study the physical impact of such structure across large samples. In this paper we present Galaxy Zoo: 3D (GZ:3D) a crowdsourcing project built on the Zooniverse platform that we used to create spatial pixel (spaxel) maps that identify galaxy centres, foreground stars, galactic bars, and spiral arms for 29 831 galaxies that were potential targets of the MaNGA survey (Mapping Nearby Galaxies at Apache Point Observatory, part of the fourth phase of the Sloan Digital Sky Surveys or SDSS-IV), including nearly all of the 10 010 galaxies ultimately observed. Our crowdsourced visual identification of asymmetric internal structures provides valuable insight on the evolutionary role of non-axisymmetric processes that is otherwise lost when MaNGA data cubes are azimuthally averaged. We present the publicly available GZ:3D catalogue alongside validation tests and example use cases. These data may in the future provide a useful training set for automated identification of spiral arm features. As an illustration, we use the spiral masks in a sample of 825 galaxies to measure the enhancement of star formation spatially linked to spiral arms, which we measure to be a factor of three over the background disc, and how this enhancement increases with radius.

Masters, Karen L. (ORCID:0000000308469578)↗

Laboratory earthquake forecasting: A machine learning competition

Earthquake prediction, the long-sought holy grail of earthquake science, continues to confound Earth scientists. Could we make advances by crowdsourcing, drawing from the vast knowledge and creativity of the machine learning (ML) community? We used Google’s ML competition platform, Kaggle, to engage the worldwide ML community with a competition to develop and improve data analysis approaches on a forecasting problem that uses laboratory earthquake data. The competitors were tasked with predicting the time remaining before the next earthquake of successive laboratory quake events, based on only a small portion of the laboratory seismic data. The more than 4,500 participating teams created and shared more than 400 computer programs in openly accessible notebooks. Complementing the now well-known features of seismic data that map to fault criticality in the laboratory, the winning teams employed unexpected strategies based on rescaling failure times as a fraction of the seismic cycle and comparing input distribution of training and testing data. In addition to yielding scientific insights into fault processes in the laboratory and their relation with the evolution of the statistical properties of the associated seismic data, the competition serves as a pedagogical tool for teaching ML in geophysics. The approach may provide a model for other competitions in geosciences or other domains of study to help engage the ML community on problems of significance.

58 GEOSCIENCES↗

Assessing residential PM 2.5 concentrations and infiltration factors with high spatiotemporal resolution using crowdsourced sensors

Building conditions, outdoor climate, and human behavior influence residential concentrations of fine particulate matter (PM 2.5 ). To study PM 2.5 spatiotemporal variability in residences, we acquired paired indoor and outdoor PM 2.5 measurements at 3,977 residences across the United States totaling >10,000 monitor-years of time-resolved data (10-min resolution) from the PurpleAir network. Time-series analysis and statistical modeling apportioned residential PM 2.5 concentrations to outdoor sources (median residential contribution = 52% of total, coefficient of variation = 69%), episodic indoor emission events such as cooking (28%, CV = 210%) and persistent indoor sources (20%, CV = 112%). Residences in the temperate marine climate zone experienced higher infiltration factors, consistent with expectations for more time with open windows in milder climates. Likewise, for all climate zones, infiltration factors were highest in summer and lowest in winter, decreasing by approximately half in most climate zones. Large outdoor–indoor temperature differences were associated with lower infiltration factors, suggesting particle losses from active filtration occurred during heating and cooling. Absolute contributions from both outdoor and indoor sources increased during wildfire events. Infiltration factors decreased during periods of high outdoor PM 2.5 , such as during wildfires, reducing potential exposures from outdoor-origin particles but increasing potential exposures to indoor-origin particles. Time-of-day analysis reveals that episodic emission events are most frequent during mealtimes as well as on holidays (Thanksgiving and Christmas), indicating that cooking-related activities are a strong episodic emission source of indoor PM 2.5 in monitored residences.

54 ENVIRONMENTAL SCIENCES↗

Crowdsourcing the Frontier: Advancing Hybrid Physics‐ML Climate Simulation via a $\$$50,000 Kaggle Competition

Subgrid machine-learning (machine learning [ML]) parameterizations have the potential to introduce a new generation of climate models that incorporate the effects of higher-resolution physics without incurring the prohibitive computational cost associated with more explicit physics-based simulations. However, important issues, ranging from online instability to inconsistent online performance, have limited their operational use for long-term climate projections. To more rapidly drive progress in solving these issues, domain scientists and ML researchers opened up the offline aspect of this problem to the broader ML and data science community with the release of ClimSim, a NeurIPS Data sets and Benchmarks publication, and an associated Kaggle competition. This paper reports on the downstream results of the Kaggle competition by coupling emulators inspired by the winning teams' architectures to an interactive climate model (including full cloud microphysics, a regime historically prone to online instability) and systematically evaluating their online performance. Our results demonstrate that online stability in the low-resolution real-geography setting is reproducible across multiple diverse architectures, which we consider a key milestone. All tested architectures exhibit strikingly similar offline and online biases, though their responses to architecture-agnostic design choices (e.g., expanding the list of input variables) can differ significantly. Multiple Kaggle-inspired architectures achieve state-of-the-art results on certain metrics such as zonal mean bias patterns and global Root Mean Squared Error, indicating that crowdsourcing the essence of the offline problem is one path to improving online performance in hybrid physics-AI climate simulation.

Environmental sciences↗

The CatWISE2020 Catalog

The CatWISE2020 Catalog consists of 1,890,715,640 sources over the entire sky selected from Wide-field Infrared Survey Explorer (WISE) and NEOWISE survey data at 3.4 and 4.6 μm (W1 and W2) collected from 2010 January 7 to 2018 December 13. This data set adds two years to that used for the CatWISE Preliminary Catalog, bringing the total to six times as many exposures spanning over 16 times as large a time baseline as the AllWISE catalog. The other major change from the CatWISE Preliminary Catalog is that the detection list for the CatWISE2020 Catalog was generated using crowdsource from Schlafly et al., while the CatWISE Preliminary Catalog used the detection software used for AllWISE. These two factors result in roughly twice as many sources in the CatWISE2020 Catalog. The scatter with respect to Spitzer photometry at faint magnitudes in the COSMOS field, which is out of the Galactic Plane and at low ecliptic latitude (corresponding to lower WISE coverage depth) is similar to that for the CatWISE Preliminary Catalog. The 90% completeness depth for the CatWISE2020 Catalog is at W1 = 17.7 mag and W2 = 17.5 mag, 1.7 mag deeper than in the CatWISE Preliminary Catalog. In comparison to Gaia, CatWISE2020 motions are accurate at the 20 mas yr-1 level for W1~15 mag sources and at the ~100 mas yr-1 level for W1~17 mag sources. This level of accuracy represents a 12 improvement over AllWISE. Finally, the CatWISE catalogs are available in the WISE/NEOWISE Enhanced and Contributed Products area of the NASA/IPAC Infrared Science Archive.

79 ASTRONOMY AND ASTROPHYSICS↗