Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Science Metadata”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Comparative analysis of nutrient concentrations in generalist and specialist tree species and soils, Manaus, Brazil

This dataset was collected near Manaus, Brazil, at ZF-2 site, inside the North-South transect plots from 20221011 to 20221020. Measurements were made on specialists and generalist tree species along topographic gradient (in upland high-clay content soils of plateaus and high sandy content and partially flooded soils of valleys). We selected nine species (with four replicates each, totaling 35 individuals) occurring in different topographic positions: three plateau specialists, three valley specialists, and three generalists, where leaf and trunk samples were collected from each individual, and soil samples for carbon and nutrient analysis and quantification. Three soil pits were opened around each sample tree, about one meter apart (total of 105 soil pits each 60-cm deep), where soil samples were collected at four depths: 0-5, 5-10, 10-30 and 30-50 cm. In each of the three pits around each tree, one single sample was taken at each depth and combined to obtain a composite sample per depth per individual tree (35 trees × 4 depths = 140 soil samples). The files “Plant_Nutrient_Concentrations_NS_Transect_Manaus.csv” and “Soil_Nutrient_Concentrations_NS_Transect_Manaus.csv” contain the nutrient concentration data from plant and soil material, respectively. Additionally, the file “Sample_Info.csv” contains details about each variable including units and data type. The file “Species_Info.csv” includes information about each sampled individual, such as species, family, diameter at the breast height (DBH), and more. The dataset is ready to be used in any programming language like python or R. This dataset was originally published on the NGEE Tropics Archive and is being mirrored on ESS-DIVE for long-term archival Acknowledgement: Funding for NGEE-Tropics data resources was provided by the U.S. Department of Energy Office of Science, Office of Biological and Environmental Research.

54 ENVIRONMENTAL SCIENCES↗

Genesis Mission Data cards

As data-intensive research and artificial intelligence become central to DOE mission science, the need for machine-actionable dataset documentation has grown accordingly. However, many DOE-aligned communities, including the Office of Science, NNSA, and cross-laboratory collaborations, have developed independent metadata practices. This fragmentation creates friction for discovery, federation, and reuse across programs. To address these challenges, this talk introduces the Genesis Data Card: a shared metadata artifact developed in collaboration with a broad DOE community (Jefferson Lab and the National Lab of the Rockies, Oak Ridge, Sandia, Idaho, Berkeley, and Los Alamos). The Genesis Data Card aims to standardize dataset documentation across DOE-aligned initiatives while remaining extensible to discipline-specific needs. This talk will describe the data card template and the supporting code to validate completed data cards, using a companion LinkML schema. I'll walk through the design decisions behind the template, its alignment with existing standards, its treatment of sensitivity and governance metadata, and the phased roadmap toward lifecycle-integrated "xCards" that support autonomous discovery and reuse. The talk closes with current gaps, ongoing work, and how others can contribute datasets and feedback to the shared repository.

McSpadden, Helen [Thomas Jefferson National Accele↗

Earthdata Search UX Lessons Learned

Crafting a great user experience is hard. Crafting a great user experience for Earth science applications is fraught with challenges. From the variability in metadata to the experience profile of various users the possible permutations of use cases introduce layer upon layer of complexities that must be designed against. In this session, the Earthdata Search team would like to highlight lessons learned over the lifespan of the application the good, the bad, and the ugly.

Earthdata Search↗

Revisiting the Solar Research Cyberinfrastructure Needs: A White Paper of Findings and Recommendations

Solar and Heliosphere physics are areas of remarkable data-driven discoveries. Recent advances in high cadence, high-resolution multiwavelength observations, growing amounts of data from realistic modeling, and operational needs for uninterrupted science-quality data coverage generate the demand for a solar metadata standardization and overall healthy data infrastructure. This white paper is prepared as an effort of the working group “Uniform Semantics and Syntax of Solar Observations and Events” created within the “Towards Integration of Heliophysics Data, Modeling, and Analysis Tools” EarthCube Research Coordination Network (@HDMIEC RCN), with primary objectives to discuss current advances and identify future needs for the solar research cyberinfrastructure. The white paper summarizes presentations and discussions held during the special working group session at the EarthCube Annual Meeting on June 19th, 2020, as well as community contribution gathered during a series of preceding workshops and subsequent RCN working group sessions. The authors provide examples of the current standing of the solar research cyberinfrastructure, and describe the problems related to current data handling approaches. The list of the top-level recommendations agreed by the authors of the current white paper is presented at the beginning of the paper.

SMD↗

Time-lapse imagery in 2017 and 2018 at the Lower Montane site in the East River Watershed, Colorado

Time-lapse imagery was collected using an automated RGB camera mounted on a pole at the base of the northeast-facing hillslope at the Lower Montane site in the East River Watershed, Colorado. The imagery was intended to support a better understanding of plant dynamics and their controls during the growing season. The dataset includes RGB images archived in four zip files (containing imagery in JPEG format), corresponding to photos taken from the hillslope and the adjacent floodplain during 2017 and 2018. A fifth zip file contains a few AVI movies that compare imagery between the two years. The AVI files can be read with most media players applications. The archive contains a total of five *.zip files and three csv metadata files (flmd.csv, dd.csv, and locations.csv).This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Hyporheic zone, river, and groundwater metagenome resolved genomes and rpS3 genes in East River Watershed, Colorado USA Summer 2020, 2021

Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from water filter collected across 8 locations along the East River Watershed, CO, and 1 nearby groundwater well. The purpose was to look for connectivity and similarities across the network and to see the impact of the groundwater. As a part of Lawrence Berkeley National Laboratory (LBNL) Watershed Science Focus Area (SFA), we assessed community composition and strain similarities between the sites and we also compared it to previous metagenomic studies within the watershed looking at floodplain (Matheus Carnevali et al. 2021) and hillslope (Lavy et al. 2019) microbiomes. Here we present metagenome assembled genomes (MAGs) for the bacterial and archaeal communities from filters across 8 locations during August 2020 and July 2021. This resulted in 32 samples. The groundwater sample was sequenced at UC Berkley's QB3. The other 31 samples were sequenced at University of Maryland. Metagenomes were assembled using four autobinners and the best bins were selected using dasTool. The genomes were dereplicated at 95% with dRep and the subset of winning genomes were manually curated based on visual inspection of taxonomic profile, GC content, coverage, and a set of 51 bacterial single copy genes (BSCG), and 38 archaeal signal copy genes (ASCG). The dataset includes a zip file of 311 genomes (HZ_River_SW_MAGS_Dereplicated_95.zip). The dataset additionally includes a zipped file of ribosomal protein small subunit 3 (rpS3) proteins from the hyporheic zone and river data (rpS3_Proteins_HZ_River.zip), a metadata file used to register associated samples with IGSNs (International Generic Sample Numbers) (samples.csv), a location metadata file (locations.csv). This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

DNA↗

The SPASE Data Model for Heliophysics Data: Is it Working?

The Space Physics Archive Search and Extract (SPASE) Data Model was developed to provide a metadata standard for describing Heliophysics (Space and Solar Physics) data within that science discipline. The SPASE Data Model has matured over the many years of its creation and is presently represented by Version 2.2.1. Information about SPASE can be obtained from the website group.org. The Data Model defines terms and values as well as the relationships between them in order to describe the data resources in the Heliophysics data environment. This data environment is quite complex, consisting of Virtual Observatories, Resident Archives, Data Providers, Partnering Data Centers, Services, Final Archives, and a Deep Archive. SPASE is the metadata language standard intended to permeate the complexity and provide a common method of obtaining and understanding data. Is it working in this capacity? SPASE has been used to describe a wide range of data. Examples range from ground-based magnetometer data to interplanetary satellite measurements to space weather model results. Has it achieved the goal of making the data easier to find and use? To find data of interest it is necessary that all the data of importance be described using the SPASE Data Model. Within the part of the data community associated with NASA (supported through NASA funding) there are obligations to use SPASE and (0 describe the old and new data using the SPASE XML schema. Although this pan of the community is not near 100% compliance with the mandate, there is good progress being made and the goal should be reachable in the future. Outside of the NASA data community there is still work to be done to convince the international community that SPASE descriptions are w011h the cost of their generation. Some of these groups such as Cluster, HELlO, GAIA, NOAA/NGDe. CSSDP, VSTO, SuperMAG, and IUGONET have agreed to use SPASE. but there are still other groups of importance that need (0 be reached. It is also assumed that the terminology is sufficiently broad and the descriptions are sufficiently complete that researchers needing data of a specific type or from a specific period can find and acquire what they need. A valid SPASE description can be very brief or very thorough depending on the willingness of the author to spend the time necessary to make the description useful. There is evidence that users are finding what they need through the SPASE descriptions, and this standard is a big step forward in Heliophysics data location. Does SPASE make it easier to use the data once they are found,) Thorough descriptions of data using SPASE can describe the data down to the level of individual parameters and exactly how the data are organized and stored. Should the SPASE data descriptions be written in such a way that they can be automatically ingested and understood by software tools'? Heliophysics instruments are becoming morc versatile all the time and the complexity of the data makes it tedious and time consuming to write SPASE descriptions with this level of sophistication even with the improvement of the tools used to generate the descriptions. Is it better to just write human-readable descriptions of the data at the parameter level or to refer to references that provide this information? This is a debate that is presently taking place and software is being developed to test what is possible.

Thieman, James↗

Online Metadata Directories: A way of preserving, sharing and discovering scientific information

The Global Change Master Directory (GCMD) assists the scientific community in the discovery of and linkage to Earth Science data and provides data holders a means to advertise their data to the community through its portals, i.e. online customized subset metadata directories. These directories are effectively serving communities like the Joint Committee on Antarctic Data Management (JCADM), the Global Observing System Information Center (GOSIC), and the Global Ocean Ecosystems Dynamic Program (GLOBEC) by increasing the visibility of their data holding. The purpose of the Gulf of Maine Ocean Data Partnership (GoMODP) is to "promote and coordinate the sharing, linking, electronic dissemination, and use of data on the Gulf of Maine region". The participants have decided that a "coordinated effort is needed to enable users throughout the Gulf of Maine region and beyond to discover and put to use the vast and growing quantities of data in their respective databases". GoMODP members have invited the GCMD to discuss potential collaborations associated with this effort. The presentation will focus on the use of the GCMD s metadata directory as a powerful tool for data discovery and sharing. An overview of the directory and its metadata authoring tools will be given.

Meaux, M.↗

Syntactic and Semantic Validation without a Metadata Management System

The ability to maintain quality information is essential to securing the confidence in any system for which the information serves as a data source. NASA's Global Change Master Directory (GCMD), an online Earth science data locator, holds over 9000 data set descriptions and is in a constant state of flux as metadata are created and updated on a daily basis. In such a system, the importance of maintaining the consistency and integrity of these-metadata is crucial. The GCMD has developed a metadata management system utilizing XML, controlled vocabulary, and Java technologies to ensure the metadata not only adhere to valid syntax, but also exhibit proper semantics.

Pollack, Janine↗

HDF-EOS Web Server

A shell script has been written as a means of automatically making HDF-EOS-formatted data sets available via the World Wide Web. ("HDF-EOS" and variants thereof are defined in the first of the two immediately preceding articles.) The shell script chains together some software tools developed by the Data Usability Group at Goddard Space Flight Center to perform the following actions: Extract metadata in Object Definition Language (ODL) from an HDF-EOS file, Convert the metadata from ODL to Extensible Markup Language (XML), Reformat the XML metadata into human-readable Hypertext Markup Language (HTML), Publish the HTML metadata and the original HDF-EOS file to a Web server and an Open-source Project for a Network Data Access Protocol (OPeN-DAP) server computer, and Reformat the XML metadata and submit the resulting file to the EOS Clearinghouse, which is a Web-based metadata clearinghouse that facilitates searching for, and exchange of, Earth-Science data.

Ullman, Richard↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, there-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA's Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomatic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related 'omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata 'omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data-use, resulting in 40 enabled publications by open data. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA "Open Science Data Repositories (OSDR)" and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Flourescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to "big data" from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology.

omics↗

Metagenome-assembled genomes measured at 3 depths during snowmelt period in East River, CO (March, May, and June, September 2017)

Snowmelt is a critical biogeochemical period that accounts for large nitrogen (N) export events from high-elevation watersheds. Soil microbial populations bloom and immobilize N during snowmelt, yet the population size crashes in spring, which releases a pulse of soil N. We sought to discover the N sources fueling this microbial bloom and determine the fate of N following microbial die-off. Here, focusing on the snowmelt period within a headwater catchment of the Upper Colorado River Basin (East River, CO), we deployed strain-resolved metagenomics to identify the metabolic pathways and processes that mobilize soil N during and after snowmelt. Soil metagenome samples were taken from 6 snowpits from 3 depths (0-5cm, 5-15cm, >15cm) at 4 time points during snowmelt period (March 2017, May 2017, and June 2017, September 2017) generating 48 metagenomes. We reconstructed 474 metagenome-assembled genomes (MAGs) across all metagenomes.All 48 metagenomes were sequenced at JGI and raw data can be found under JGI (Joint Genome Institute) GOLD Study Gs0135149. Metagenome assemblies from IMG under the same study were used for genome binning. This dataset (1) a zip file of 474 MAGs (as fasta files, Gs0135149_bins_tar.gz), (2) sample metadata file with sample IGSNs (International Generic Sample Numbers) (samples.csv), (3) bounding box coordinates for the sampled locations (Gs0135149.kml), (4) metagenome metadata file listing IMG/M (Integrated Microbial Genomes/Metagenomes) metagenome accessions linking samples to metagenomes (metagenomes.csv), (5) location metadata file (locations.csv), (6) file-level metadata file (flmd.csv) and (7) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

Open Science for Life in Space: Data Sharing and Tools for Knowledge Discovery

The next era in human space exploration is rapidly approaching and will require the use of countermeasures to deep space health hazards. The development of countermeasures (or, the re-purposing of existing agents) will be highly dependent on our understanding of basic biological responses to space stressors (e.g. ionizing radiation, altered gravitational fields, altered day-night cycles, confinement, isolation, hostile-closed environments, distance-duration from Earth, exposure to celestial regolith, etc.). The fast-growing array of space biological data, which in the past was simply archived after minimal analysis, holds great potential if it can be reorganized and formatted for Open Science. Organizing the data for such analysis is a challenge because of its diverse nature (molecular, cellular, tissue, imaging, whole organism and behavior). We will discuss here several strategies that NASA’s Biological and Physical Science Division has put in place to maximize the return on investment for spaceflight bioscience data. Open Science, as a scientific philosophy, is the concept that the more people who have access to the data, the more knowledge will be gained from it. This guiding principle led NASA to develop GeneLab in 2015. GeneLab houses spaceflight and relevant ground-based multi-omics data, and has grown to ~400 transcriptomic, proteomic, metabolomic and epigenomic datasets from plant, rodent, small animal, and microbial space experiments. GeneLab provides users with various tools for data analysis and a visualization portal that allows users to interact with gene expression data from space-related ‘omics experiments. Open Science is also about building scientific communities, and with this spirit in mind, GeneLab has spawned several Analysis Working Groups (AWGs), comprised of more than 200 volunteer scientists. The AWGs initially provided feedback on the processing pipeline and metadata ‘omics standards for GeneLab. Over the last few years, they have become a community-driven science enterprise, engaging in large meta-analysis of GeneLab datasets, resulting in 10 publications (beyond the originally submitted research). Overall, the Open Science nature of GeneLab has resulted in a high degree of data re-use, resulting in 38 additional publications derived from the original 67 publication over the past four years. The enormous success and knowledge gained from GeneLab has led to a collection of sister NASA “Open Science Data Repositories (OSDR)” and research support groups. These include the NASA Ames Life Sciences Data Archive (ALSDA), the NASA Biological Institutional Scientific Collection (NBISC), and the Biospecimen Sharing Program (BSP). All are adopting the GeneLab data architecture system to maximize open-access, find-ability, accessibility, interoperability, and reusability (FAIR). ALSDA collects and curates phenotypic-physiological bioimaging-behavioral data from space and space-relevant non-human experiments, oftentimes coming from the same omics-associated experimental datasets found in GeneLab. Since 2021, a community of ~100 researchers have rallied around ALSDA, to provide feedback in a new ALSDA AWG focused on phenotypic-physiological investigation-sample-assay metadata standards (e.g., Micro-Computed Tomography, Light/Fluorescence Microscopy, Western Blot, Flow Cytometry, Novel Object Recognition, Elevated Plus Maze, etc. of ~50 assays collected). These standards are part of a new single point-of-entry data submission portal for all non-human Space Biology and Human Research Program principal investigators, to submit, curate, and share their research data. With open-access space biological data now collected and curated together with rich metadata, and with the potential for linkage to “big data” from the international biological and medical communities (NIH, EBI, etc.), the artificial intelligence and machine learning (AI/ML) era has started for Space Biology. Several other talks will cover these topics in this conference.

life sciences↗

Geophysical survey associated with NEON AOP survey, East River, CO 2018

The package contains data layers developed and used in Falco et al. 2024: “EcoImaging: Advanced Sensing to Investigate Plant and Abiotic Hierarchical Spatial Patterns in Mountainous Watersheds". The package is part of the DOE Watershed Function Science Focus Area (SFA) project and includes geophysical measurements collected at the East River, Colorado, in conjunction with the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP) survey conducted in June 2018. This dataset provide soil geophysical information and were used to investigate soil-plant relationships. The dataset consists of: - NEON_2018_EMI_survey.zip: the electromagnetic induction (EMI) survey as shape-file; - NEON_plot_TDR.csv: plot‑level data from Time‑Domain Reflectometry (TDR) measurements, providing: * volumetric water content (VWC) in percent (%); * soil temperature in degrees Celsius (°C); - file level metadata (flmd.csv) - data dictionary (dd.csv) file This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Embracing Open Source for NASA's Earth Science Data Systems

The overarching purpose of NASAs Earth Science program is to develop a scientific understanding of Earth as a system. Scientific knowledge is most robust and actionable when resulting from transparent, traceable, and reproducible methods. Reproducibility includes open access to the data as well as the software used to arrive at results. Additionally, software that is custom-developed for NASA should be open to the greatest degree possible, to enable re-use across Federal agencies, reduce overall costs to the government, remove barriers to innovation, and promote consistency through the use of uniform standards. Finally, Open Source Software (OSS) practices facilitate collaboration between agencies and the private sector. To best meet these ends, NASAs Earth Science Division promotes the full and open sharing of not only all data, metadata, products, information, documentation, models, images, and research results but also the source code used to generate, manipulate and analyze them. This talk focuses on the challenges to open sourcing NASA developed software within ESD and the growing pains associated with establishing policies running the gamut of tracking issues, properly documenting build processes, engaging the open source community, maintaining internal compliance, and accepting contributions from external sources. This talk also covers the adoption of existing open source technologies and standards to enhance our custom solutions and our contributions back to the community. Finally, we will be introducing the most recent OSS contributions from NASA Earth Science program and promoting these projects for wider community review and adoption.

Earth Science↗

Groundwater elevation data for monitoring wells within the East and Taylor River basins, Colorado (USA)

This dataset is comprised of temporal variations in groundwater elevation data for the 24 monitoring wells located throughout the East River watershed. Seasonal to annual variations in groundwater elevations are a critical property of mountainous watersheds needed to understand both hydrological and below ground biogeochemical processes. Such data serve as a critical constraint for numerical models describing coupled groundwater-surface water behavior within the watershed. Additionally, the offset between the maximum and minimum groundwater elevations defines the extent of the bedrock weathering zone, with annual excursions in the groundwater hydrographic (i.e., the rising and falling hydrographic limbs) imposing primary controls on bedrock saturation state and redox conditions that govern biogeochemical reactions impacting nitrogen, carbon, and metals cycling. Manufacturer-specific software is used to download pressure data from each transducer, with broadly available spreadsheet software (e.g. Microsoft Excel) used to convert temporal variations in water pressure to elevations in units of meters above mean sea level. As additional monitoring wells are installed within the East River watershed and new groundwater monitoring wells are installed in the Taylor River watershed, temporal groundwater elevation data will be included as a part of this master dataset. Details regarding the metadata associated with each monitoring well location, including well depths, screened intervals, well location coordinates, and bedrock type, are included, as is a standard operating procedure for generating groundwater elevation data from water pressure values recorded by the pressure transducers. This dataset includes: (1) a zip file (East_River_Watershed_Compiled_Groundwater_Elevation_Data_Plots.zip), containing (a) PNG of groundwater hydrographs, (b) a CSV file with groundwater elevation data, and (c) CSV file containing metadata organized by location; (2) an Excel file (East_River_Watershed_Compiled_Groundwater_Elevation_Data_Plots.xlsx) with the groundwater elevation data, groundwater hydrographs, and metadata organized by location; (3) a Word file (Groundwater_elevation_data_protocols.docx) and a PDF file version (Groundwater_elevation_data_protocols.pdf) containing field protocols and methods; (4) a location metadata (locations.csv) file; (5) a file level metadata (flmd.csv); and (6) data dictionary (dd.csv) file.This work was supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

54 ENVIRONMENTAL SCIENCES↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

Semantic Web Data Discovery of Earth Science Data at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC)

Mirador is a web interface for searching Earth Science data archived at the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). Mirador provides keyword-based search and guided navigation for providing efficient search and access to Earth Science data. Mirador employs the power of Google's universal search technology for fast metadata keyword searches, augmented by additional capabilities such as event searches (e.g., hurricanes), searches based on location gazetteer, and data services like format converters and data sub-setters. The objective of guided data navigation is to present users with multiple guided navigation in Mirador is an ontology based on the Global Change Master directory (GCMD) Directory Interchange Format (DIF). Current implementation includes the project ontology covering various instruments and model data. Additional capabilities in the pipeline include Earth Science parameter and applications ontologies.

Hegde, Mahabaleshwara↗