Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km x 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

climate projection↗

Uncertainty Assessment of the NASA Earth Exchange Global Daily Downscaled Climate Projections (NEX-GDDP) Dataset

The NASA Earth Exchange Global Daily Downscaled Projections (NEX-GDDP) dataset is comprised of downscaled climate projections that are derived from 21 General Circulation Model (GCM) runs conducted under the Coupled Model Intercomparison Project Phase 5 (CMIP5) and across two of the four greenhouse gas emissions scenarios (RCP4.5 and RCP8.5). Each of the climate projections includes daily maximum temperature, minimum temperature, and precipitation for the periods from 1950 through 2100 and the spatial resolution is 0.25 degrees (approximately 25 km by 25 km). The GDDP dataset has received warm welcome from the science community in conducting studies of climate change impacts at local to regional scales, but a comprehensive evaluation of its uncertainties is still missing. In this study, we apply the Perfect Model Experiment framework (Dixon et al. 2016) to quantify the key sources of uncertainties from the observational baseline dataset, the downscaling algorithm, and some intrinsic assumptions (e.g., the stationary assumption) inherent to the statistical downscaling techniques. We developed a set of metrics to evaluate downscaling errors resulted from bias-correction ("quantile-mapping"), spatial disaggregation, as well as the temporal-spatial non-stationarity of climate variability. Our results highlight the spatial disaggregation (or interpolation) errors, which dominate the overall uncertainties of the GDDP dataset, especially over heterogeneous and complex terrains (e.g., mountains and coastal area). In comparison, the temporal errors in the GDDP dataset tend to be more constrained. Our results also indicate that the downscaled daily precipitation also has relatively larger uncertainties than the temperature fields, reflecting the rather stochastic nature of precipitation in space. Therefore, our results provide insights in improving statistical downscaling algorithms and products in the future.

general circulation model (GCM)↗

Internal Consistency of the NVAP Water Vapor Dataset

The NVAP (NASA Water Vapor Project) dataset is a global dataset at 1 x 1 degree spatial resolution consisting of daily, pentad, and monthly atmospheric precipitable water (PW) products. The analysis blends measurements from the Television and Infrared Operational Satellite (TIROS) Operational Vertical Sounder (TOVS), the Special Sensor Microwave/Imager (SSM/I), and radiosonde observations into a daily collage of PW. The original dataset consisted of five years of data from 1988 to 1992. Recent updates have added three additional years (1993-1995) and incorporated procedural and algorithm changes from the original methodology. Since each of the PW sources (TOVS, SSM/I, and radiosonde) do not provide global coverage, each of these sources compliment one another by providing spatial coverage over regions and during times where the other is not available. For this type of spatial and temporal blending to be successful, each of the source components should have similar or compatible accuracies. If this is not the case, regional and time varying biases may be manifested in the NVAP dataset. This study examines the consistency of the NVAP source data by comparing daily collocated TOVS and SSM/I PW retrievals with collocated radiosonde PW observations. The daily PW intercomparisons are performed over the time period of the dataset and for various regions.

Suggs, Ronnie J.↗

Highlights of the Version 8 SBUV and TOMS Datasets Released at this Symposium

Last October was the 25th anniversary of the launch of the SBUV and TOMS instruments on NASA's Nimbus-7 satellite. Total Ozone and ozone profile datasets produced by these and following instruments have produced a quarter century long record. Over time we have released several versions of these datasets to incorporate advances in UV radiative transfer, inverse modeling, and instrument characterization. In this meeting we are releasing datasets produced from the version 8 algorithms. They replace the previous versions (V6 SBUV, and V7 TOMS) released about a decade ago. About a dozen companion papers in this meeting provide details of the new algorithms and intercomparison of the new data with external data. In this paper we present key features of the new algorithm, and discuss how the new results differ from those released previously. We show that the new datasets have better internal consistency and also agree better with external datasets. A key feature of the V8 SBUV algorithm is that the climatology has no influence on inter-annual variability and trends; it only affects the mean values and, to a limited extent, the seasonal dependence. By contrast, climatology does have some influence on TOMS total O3 trends, particularly at large solar zenith angles. For this reason, and also because TOMS record has gaps, md EP/TOMS is suffering from data quality problems, we recommend using SBUV total ozone data for applications where the high spatial resolution of TOMS is not essential.

Bhartia, Pawan K.↗

Online Visualization and Analysis of Merged Global Geostationary Satellite Infrared Dataset

The NASA Goddard Earth Sciences Data Information Services Center (GES DISC) is home of Tropical Rainfall Measuring Mission (TRMM) data archive. The global merged IR product also known as the NCEP/CPC 4-km Global (60 degrees N - 60 degrees S) IR Dataset, is one of TRMM ancillary datasets. They are globally merged (60 degrees N - 60 degrees S) pixel-resolution (4 km) IR brightness temperature data (equivalent blackbody temperatures), merged from all available geostationary satellites (GOES-8/10, METEOSAT-7/5 and GMS). The availability of data from METEOSAT-5, which is located at 63E at the present time, yields a unique opportunity for total global (60 degrees N- 60 degrees S) coverage. The GES DISC has collected over 8 years of the data beginning from February of 2000. This high temporal resolution dataset can not only provide additional background information to TRMM and other satellite missions, but also allow observing a wide range of meteorological phenomena from space, such as, mesoscale convection systems, tropical cyclones, hurricanes, etc. The dataset can also be used to verify model simulations. Despite that the data can be downloaded via ftp, however, its large volume poses a challenge for many users. A single file occupies about 70 MB disk space and there is a total of approximately 73,000 files (approximately 4.5 TB) for the past 8 years. In order to facilitate data access, we have developed a web prototype to allow users to conduct online visualization and analysis of this dataset. With a web browser and few mouse clicks, users can have a full access to over 8 year and over 4.5 TB data and generate black and white IR imagery and animation without downloading any software and data. In short, you can make your own images! Basic functions include selection of area of interest, single imagery or animation, a time skip capability for different temporal resolution and image size. Users can save an animation as a file (animated gif) and import it in other presentation software, such as, Microsoft PowerPoint. The prototype will be integrated into GIOVANNI and existing GIOVANNI capabilities, such as, data download, Google Earth KMZ, etc will be available. Users will also be able to access other data products in the GIOVANNI family.

Liu, Zhong↗

Newly Released TRMM Version 7 Products, Other Precipitation Datasets and Data Services at NASA GES DISC

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) is home of global precipitation product archives, in particular, the Tropical Rainfall Measuring Mission (TRMM) products. TRMM is a joint U.S.-Japan satellite mission to monitor tropical and subtropical (40 S - 40 N) precipitation and to estimate its associated latent heating. The TRMM satellite provides the first detailed and comprehensive dataset on the four dimensional distribution of rainfall and latent heating over vastly undersampled tropical and subtropical oceans and continents. The TRMM satellite was launched on November 27, 1997. TRMM data products are archived at and distributed by GES DISC. The newly released TRMM Version 7 consists of several changes including new parameters, new products, meta data, data structures, etc. For example, hydrometeor profiles in 2A12 now have 28 layers (14 in V6). New parameters have been added to several popular Level-3 products, such as, 3B42, 3B43. Version 2.2 of the Global Precipitation Climatology Project (GPCP) dataset has been added to the TRMM Online Visualization and Analysis System (TOVAS; URL: http://disc2.nascom.nasa.gov/Giovanni/tovas/), allowing online analysis and visualization without downloading data and software. The GPCP dataset extends back to 1979. Version 3 of the Global Precipitation Climatology Centre (GPCC) monitoring product has been updated in TOVAS as well. The product provides global gauge-based monthly rainfall along with number of gauges per grid. The dataset begins in January 1986. To facilitate data and information access and support precipitation research and applications, we have developed a Precipitation Data and Information Services Center (PDISC; URL: http://disc.gsfc.nasa.gov/precipitation). In addition to TRMM, PDISC provides current and past observational precipitation data. Users can access precipitation data archives consisting of both remote sensing and in-situ observations. Users can use these data products to conduct a wide variety of activities, including case studies, model evaluation, uncertainty investigation, etc. To support Earth science applications, PDISC provides users near-real-time precipitation products over the Internet. At PDISC, users can access tools and software. Documentation, FAQ and assistance are also available. Other capabilities include: 1) Mirador (http://mirador.gsfc.nasa.gov/), a simplified interface for searching, browsing, and ordering Earth science data at NASA Goddard Earth Sciences Data and Information Services Center (GES DISC). Mirador is designed to be fast and easy to learn; 2)TOVAS; 3) NetCDF data download for the GIS community; 4) Data via OPeNDAP (http://disc.sci.gsfc.nasa.gov/services/opendap/). The OPeNDAP provides remote access to individual variables within datasets in a form usable by many tools, such as IDV, McIDAS-V, Panoply, Ferret and GrADS; 5) The Open Geospatial Consortium (OGC) Web Map Service (WMS) (http://disc.sci.gsfc.nasa.gov/services/wxs_ogc.shtml). The WMS is an interface that allows the use of data and enables clients to build customized maps with data coming from a different network.

Liu, Zhong↗

Comparison of Four Precipitation Forcing Datasets in Land Information System Simulations over the Continental U.S.

The NASA Short ]term Prediction Research and Transition (SPoRT) Center in Huntsville, AL is running a real ]time configuration of the NASA Land Information System (LIS) with the Noah land surface model (LSM). Output from the SPoRT ]LIS run is used to initialize land surface variables for local modeling applications at select National Weather Service (NWS) partner offices, and can be displayed in decision support systems for situational awareness and drought monitoring. The SPoRT ]LIS is run over a domain covering the southern and eastern United States, fully nested within the National Centers for Environmental Prediction Stage IV precipitation analysis grid, which provides precipitation forcing to the offline LIS ]Noah runs. The SPoRT Center seeks to expand the real ]time LIS domain to the entire Continental U.S. (CONUS); however, geographical limitations with the Stage IV analysis product have inhibited this expansion. Therefore, a goal of this study is to test alternative precipitation forcing datasets that can enable the LIS expansion by improving upon the current geographical limitations of the Stage IV product. The four precipitation forcing datasets that are inter ]compared on a 4 ]km resolution CONUS domain include the Stage IV, an experimental GOES quantitative precipitation estimate (QPE) from NESDIS/STAR, the National Mosaic and QPE (NMQ) product from the National Severe Storms Laboratory, and the North American Land Data Assimilation System phase 2 (NLDAS ]2) analyses. The NLDAS ]2 dataset is used as the control run, with each of the other three datasets considered experimental runs compared against the control. The regional strengths, weaknesses, and biases of each precipitation analysis are identified relative to the NLDAS ]2 control in terms of accumulated precipitation pattern and amount, and the impacts on the subsequent LSM spin ]up simulations. The ultimate goal is to identify an alternative precipitation forcing dataset that can best support an expansion of the real ]time SPoRT ]LIS to a domain covering the entire CONUS.

Case, Jonathan L.↗

Compiling a Comprehensive EVA Training Dataset for NASA Astronauts

Training for a spacewalk or extravehicular activity (EVA) is considered a hazardous duty for NASA astronauts. This places astronauts at risk for decompression sickness as well as various musculoskeletal disorders from working in the spacesuit. As a result, the operational and research communities over the years have requested access to EVA training data to supplement their studies. The purpose of this paper is to document the comprehensive EVA training data set that was compiled from multiple sources by the Lifetime Surveillance of Astronaut Health (LSAH) epidemiologists to investigate musculoskeletal injuries. The EVA training dataset does not contain any medical data, rather it only documents when EVA training was performed, by whom and other details about the session. The first activities practicing EVA maneuvers in water were performed at the Neutral Buoyancy Simulator (NBS) at the Marshall Spaceflight Center in Huntsville, Alabama. This facility opened in 1967 and was used for EVA training until the early Space Shuttle program days. Although several photographs show astronauts performing EVA training in the NBS, records detailing who performed the training and the frequency of training are unavailable. Paper training records were stored within the NBS after it was designated as a National Historic Landmark in 1985 and closed in 1997, but significant resources would be needed to identify and secure these records, and at this time LSAH has not pursued acquisition of these early training records. Training in the NBS decreased when the Johnson Space Center in Houston, Texas, opened the Weightless Environment Training Facility (WETF) in 1980. Early training records from the WETF consist of 11 hand-written dive logbooks compiled by individual workers that were digitized at the request of LSAH. The WETF was integral in the training for Space Shuttle EVAs until its closure in 1998. The Neutral Buoyancy Laboratory (NBL) at the Sonny Carter Training Facility near JSC opened in March 1997 and is the current site for US EVA training. Other space agencies also have used water to simulate weightlessness and train for EVAs. Russia has a training facility similar to the NBL named the Hydro Lab. The Hydro Lab began operations at the Gagarin Cosmonaut Training Center (GCTC) in 1980 and has been used extensively to the present. Although a majority of training in the Hydro Lab uses the Russian Orlan suit, a small number of sessions have been conducted using a NASA suit. The Japanese Weightlessness Environment Test System (WETS) went into service at the Tsukuba Space Center in 1997 but was closed in 2011 due to extensive earthquake damage. Several sessions were performed using a NASA suit, but these sessions were short and considered "development" runs. LSAH has assembled records from the WETF, NBL and Hydro Lab. Recording of the EVA training data has changed considerably from 1967 to present. The goal of early record keeping was to track use of hardware components, and the person involved was treated as a suited operator, not as a focus of interest. Records from the past two decades are fairly precise with the person, date, suit type and size noted. On occasion the length of the session was listed, but this data is not included on all records. Records were merged from data sources and extensive cleaning of the records was required since the multiple sources frequently overlapped and duplicated records. To date the LSAH EVA training dataset includes over 12,500 EVA training sessions performed by NASA astronauts since 1981. The following variables are included for most records: Name, Sex, Event date, Event name, HUT type, HUT size, Facility, and Estimated run time. For a smaller subset of records, the following variables are available: Actual run time, Time inverted, and the suit components Waist bearing type, Shoulder harness, Shoulder pads, and Teflon inserts. The LSAH dataset is currently the most complete resource for data regarding EVA training sessions performed by NASA astronauts. However, it is not 100 percent complete since the WETS (Japan) and NBS (Marshall) training facility data were not included. This dataset has been compiled by LSAH to study the relationship of EVA training to musculoskeletal injuries but has many other non-medical applications. This dataset can be provided to other groups in order to respond to program and research questions with appropriate board approvals.

Laughlin, M. S.↗

Comparison of Radiative Energy Flows in Observational Datasets and Climate Modeling

This study examines radiative flux distributions and local spread of values from three major observational datasets (CERES, ISCCP, and SRB) and compares them with results from climate modeling (CMIP3). Examinations of the spread and differences also differentiate among contributions from cloudy and clear-sky conditions. The spread among observational datasets is in large part caused by noncloud ancillary data. Average differences of at least 10Wm(exp -2) each for clear-sky downward solar, upward solar, and upward infrared fluxes at the surface demonstrate via spatial difference patterns major differences in assumptions for atmospheric aerosol, solar surface albedo and surface temperature, and/or emittance in observational datasets. At the top of the atmosphere (TOA), observational datasets are less influenced by the ancillary data errors than at the surface. Comparisons of spatial radiative flux distributions at the TOA between observations and climate modeling indicate large deficiencies in the strength and distribution of model-simulated cloud radiative effects. Differences are largest for lower-altitude clouds over low-latitude oceans. Global modeling simulates stronger cloud radiative effects (CRE) by +30Wmexp -2) over trade wind cumulus regions, yet smaller CRE by about -30Wm(exp -2) over (smaller in area) stratocumulus regions. At the surface, climate modeling simulates on average about 15Wm(exp -2) smaller radiative net flux imbalances, as if climate modeling underestimates latent heat release (and precipitation). Relative to observational datasets, simulated surface net fluxes are particularly lower over oceanic trade wind regions (where global modeling tends to overestimate the radiative impact of clouds). Still, with the uncertainty in noncloud ancillary data, observational data do not establish a reliable reference.

Raschke, Ehrhard↗

The Cumulus and Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD)

Low clouds continue to contribute greatly to the uncertainty in cloud feedback estimates. Depending on whether a region is dominated by cumulus (Cu) or stratocumulus (Sc) clouds, the interannual low-cloud feedback is somewhat different in both spaceborne and large-eddy simulation studies. Therefore, simulating the correct amount and variation of the Cu and Sc cloud distributions could be crucial to predict future cloud feedbacks. Here we document spatial distributions and profiles of Sc and Cu clouds derived from Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO) and CloudSat measurements. For this purpose, we create a new dataset called the Cumulus And Stratocumulus CloudSat-CALIPSO Dataset (CASCCAD), which identifies Sc, broken Sc, Cu under Sc, Cu with stratiform outflow and Cu. To separate the Cu from Sc, we design an original method based on the cloud height, horizontal extent, vertical variability and horizontal continuity, which is separately applied to both CALIPSO and combined CloudSat–CALIPSO observations. First, the choice of parameters used in the discrimination algorithm is investigated and validated in selected Cu, Sc and Sc–Cu transition case studies. Then, the global statistics are compared against those from existing passive- and active-sensor satellite observations. Our results indicate that the cloud optical thickness – as used in passive-sensor observations – is not a sufficient parameter to discriminate Cu from Sc clouds, in agreement with previous literature. Using clustering-derived datasets shows better results although one cannot completely separate cloud types with such an approach. On the contrary, classifying Cu and Sc clouds and the transition between them based on their geometrical shape and spatial heterogeneity leads to spatial distributions consistent with prior knowledge of these clouds, from ground-based, ship-based and field campaigns. Furthermore, we show that our method improves existing Sc–Cu classifications by using additional information on cloud height and vertical cloud fraction variation. Finally, the CASCCAD datasets provide a basis to evaluate shallow convection and stratocumulus clouds on a global scale in climate models and potentially improve our understanding of low-level cloud feedbacks. The CASCCAD dataset (Cesana, 2019, https://doi.org/10.5281/zenodo.2667637) is available on the Goddard Institute for Space Studies (GISS) website at https://data.giss.nasa.gov/clouds/casccad/ (last access: 5 November 2019) and on the zenodo website at https://zenodo.org/record/2667637 (last access: 5 November 2019).

Cesana, Gregory V.↗

Version 2 Ozone Monitoring Instrument SO2 Product (OMSO2 V2): New Anthropogenic SO2 Vertical Column Density Dataset

The Ozone Monitoring Instrument (OMI) has been providing global observations of SO2 pollution since 2004. Here we introduce the new anthropogenic SO2 vertical column density (VCD) dataset in the version 2 OMI SO2 product (OMSO2 V2). As with the previous version (OMSO2 V1.3), the new dataset is generated with an algorithm based on principal component analysis of OMI radiances, but features several updates. The most important among those is the use of expanded lookup tables and model a priori profiles to estimate SO2 Jacobians for individual OMI pixels, in order to better characterize pixel-to-pixel variations in SO2 sensitivity, including over snow and ice. Additionally, new data screening and spectral fitting schemes have been implemented to improve the quality of the spectral fit. As compared with the planetary boundary layer SO2 dataset in OMSO2 V1.3, the new dataset has substantially better data quality, especially over areas that are relatively clean or affected by the south Atlantic anomaly. The updated retrievals over snow/ice yield more realistic seasonal changes in SO2 at high latitudes and offer enhanced sensitivity to sources during wintertime. An error analysis has been conducted to assess uncertainties in SO2 VCDs from both the spectral fit and Jacobian calculations. The uncertainties from spectral fitting are reflected in SO2 slant column densities (SCDs) and largely depend on the signal-to-noise ratio of the measured radiances, as implied by the generally smaller SCD uncertainties over clouds or for smaller solar zenith angles. The SCD uncertainties for individual pixels are estimated to be~0.15-0.3 DU (Dobson Units) between ~40°S and ~40°N and to be~0.2-0.5 DU at higher latitudes. The uncertainties from the Jacobians are approximately ~50-100% over polluted areas, and primarily attributed to errors in SO2 a priori profiles and cloud pressures, as well as the lack of explicit treatment for aerosols. Finally, the daily mean and median SCDs over the presumably SO2-free equatorial East Pacific have increased by only~0.0035 DU and ~0.003 DU respectively over the entire 15-year OMI record; while the standard deviation of SCDs has grown by only~0.02 DU or ~10%. Such remarkable long-term stability makes the new dataset particularly suitable for detecting regional changes in SO2 pollution.

OMI, SO2, Remote Sensing↗

PERSIANN Dynamic Infrared–Rain Rate (PDIR-Now): A Near-Real-Time, Quasi-Global Satellite Precipitation Dataset

This study presents the Precipitation Estimation from Remotely Sensed Information Using Artificial Neural Networks–Dynamic Infrared Rain Rate (PDIR-Now) near-real-time precipitation dataset. This dataset provides hourly, quasi-global, infrared-based precipitation estimates at 0.04° × 0.04° spatial resolution with a short latency (15–60 min). It is intended to supersede the PERSIANN–Cloud Classification System (PERSIANN-CCS) dataset previously produced as the near-real-time product of the PERSIANN family. We first provide a brief description of the algorithm’s fundamentals and the input data used for deriving precipitation estimates. Second, we provide an extensive evaluation of the PDIR-Now dataset over annual, monthly, daily, and subdaily scales. Last, the article presents information on the dissemination of the dataset through the Center for Hydrometeorology and Remote Sensing (CHRS) web-based interfaces. The evaluation, conducted over the period 2017–18, demonstrates the utility of PDIR-Now and its improvement over PERSIANN-CCS at all temporal scales. Specifically, PDIR-Now improves the estimation of rain/no-rain days as demonstrated by a critical success index (CSI) of 0.53 compared to 0.47 of PERSIANN-CCS. In addition, PDIR-Now improves the estimation of seasonal and diurnal cycles of precipitation as well as regional precipitation patterns erroneously estimated by PERSIANN-CCS. Finally, an evaluation is carried out to examine the performance of PDIR-Now in capturing two extreme events, Hurricane Harvey and a cluster of summer thunderstorms that occurred over the Netherlands, where it is shown that PDIR-Now adequately represents spatial precipitation patterns as well as subdaily precipitation rates with a correlation coefficient (CORR) of 0.64 for Hurricane Harvey and 0.76 for the Netherlands thunderstorms.

Rainfall↗

An Automated Approach to Labelling Datasets in Earth Science Publications

NASA Data Active Archive Centers, orDAACs, ingest, store, and distribute dataacquired from satellites, ground systems as well asreanalysis models. Many authors use this datain their research. However, most of the datasets usedin Earth Science Publications are not citedcorrectly or not cited at all. Thus, there is no directlink between the datasets used and thescientific publications which reference them. Thisleads to issues with reproducibility of theresults, attribution of the research results, anddiscovery of new datasets. This project began byexploring various methods of automatically labellingGoddard Earth Sciences Data andInformation Services Center (GES DISC) datasets usingSupervised Machine Learning and EarthData Search Common Metadata Repository (CMR) queries.The ultimate goal was to create alibrary of citations that utilized automated citationlabeling to directly link the researchpublications to the data they use. Supervised MachineLearning approaches struggled due to thelimited amount of labelled training data to learnfrom. Increasing the volume of training data isdifficult as it requires subject matter experts todevote time to manually reviewing journalarticles and determining the datasets used. The CMRqueries were inconsistent because theunderlying metadata is continuously being updated.Thus, it is hard to generalize theeffectiveness of the CMR results as they are dependenton the internal state of CMR. Theseapproaches helped inform the decision to transitionthe project into using a Knowledge Graph.Another key aspect of this project focused on theautomated extraction of features (platform,instrument, variables, etc) and explicit citationsfrom within Earth Science Publications. Theseautomated extractions were used to classify researchpapers based on their platform/instrumentcouples. This information was input into the CitationManagement System for GES DISC. Theseplatform/instrument couples also provide an additionalfacet that can be searched on the GESDISC website.

Edward Jahoda↗

AssistTaxi: A Comprehensive Dataset for Taxiway Analysis and Autonomous Operations

The availability of high-quality datasets play a crucial role in advancing research and development especially, for safety critical and autonomous systems. This poster presents AssistTaxi, which is a comprehensive novel dataset which is a collection of images for runway and taxiway analysis. The dataset comprises of more than 300,000 frames of diverse and carefully collected data, gathered from Melbourne (MLB) and Grant-Valkaria (X59) general aviation airports. The importance of AssistTaxi lies in its potential to advance autonomous operations, enabling researchers and developers to train and evaluate algorithms for efficient and safe taxiing. Researchers can utilize AssistTaxi to benchmark their algorithms, assess performance, and explore novel approaches for runway and taxiway analysis. Additionally, the dataset serves as a valuable resource for validating and enhancing existing algorithms as well as facilitating innovation in autonomous operations for aviation. We also propose an initial approach to label the dataset using a contour based detection and line extraction technique.

Data Collection↗

Bridging the Gap: Enhancing Prominence and Provenance of NASA Datasets in Research Publications

Attribution of datasets that were used to generate research results described in peer-reviewed publications to the original source of these datasets (which are often archived at NASA Earth Science data centers) has been very challenging. Even though the data citation standard of citing datasets as research artifacts and citing them with Digital Object Identifiers (DOIs) was introduced over a decade ago, most authors do not properly reference the data used in their studies and merely mention them in the text. The lack of proper citations of datasets makes the peer-reviewed publication less transparent, imperils reproducibility, and impedes open science. We offer an open-source publication management methodology and a tool that can help to enhance usage-based data discovery, prominence, and provenance of the data; reproducibility of the research results; and potentially increase the return on investment on NASA-funded research.

open-source↗

A Large Dataset of Fluvial Hydraulic and Geometry Attributes Derived From USGS Field Measurement Records

Accurate representation of river channel geometry is important for hydrologic and hydraulic modeling of fluvial systems. Often, channel geometry is estimated using simple rating curves that can be applied across various spatial scales. However, such methods are limited to power law relations that do not employ many potentially relevant catchment and river attributes. This paper introduce a new dataset, IFMHA (Inventory of Field Measurement of Hydraulic Attributes), to enable research studies on channel geometry and streamflow characteristics. IFMHA is derived from the National Water Information System (NWIS) site inventory for surface water field measurements and stream attributes from the National Hydrography Dataset (NHD). IFMHA includes 2,802,532 records from 10,050 sites (NWIS streamgaging stations). The dataset utility is demonstrated here by presenting a series of conceptual models for estimating channel geometry parameters (i.e., channel mean depth, channel maximum depth, wetted perimeter, and roughness) based on the available field attributes within IFMHA. Such a dataset and attributed channel geometry parameters can enhance the performance of operational flood forecasting frameworks (e.g. National Water Model) by providing more accurate initial conditions used in hydrologic and hydraulic routing models.

Hydrology↗

Segmentation of Unstructured Datasets

Datasets generated by computer simulations and experiments in Computational Fluid Dynamics tend to be extremely large and complex. It is difficult to visualize these datasets using standard techniques like Volume Rendering and Ray Casting. Object Segmentation provides a technique to extract and quantify regions of interest within these massive datasets. This thesis explores basic algorithms to extract coherent amorphous regions from two-dimensional and three-dimensional scalar unstructured grids. The techniques are applied to datasets from Computational Fluid Dynamics and from Finite Element Analysis.

Bhat, Smitha↗

Backscatter Modeling at 2.1 Micron Wavelength for Space-Based and Airborne Lidars Using Aerosol Physico-Chemical and Lidar Datasets

Space-based and airborne coherent Doppler lidars designed for measuring global tropospheric wind profiles in cloud-free air rely on backscatter, beta from aerosols acting as passive wind tracers. Aerosol beta distribution in the vertical can vary over as much as 5-6 orders of magnitude. Thus, the design of a wave length-specific, space-borne or airborne lidar must account for the magnitude of 8 in the region or features of interest. The SPAce Readiness Coherent Lidar Experiment under development by the National Aeronautics and Space Administration (NASA) and scheduled for launch on the Space Shuttle in 2001, will demonstrate wind measurements from space using a solid-state 2 micrometer coherent Doppler lidar. Consequently, there is a critical need to understand variability of aerosol beta at 2.1 micrometers, to evaluate signal detection under varying aerosol loading conditions. Although few direct measurements of beta at 2.1 micrometers exist, extensive datasets, including climatologies in widely-separated locations, do exist for other wavelengths based on CO2 and Nd:YAG lidars. Datasets also exist for the associated microphysical and chemical properties. An example of a multi-parametric dataset is that of the NASA GLObal Backscatter Experiment (GLOBE) in 1990 in which aerosol chemistry and size distributions were measured concurrently with multi-wavelength lidar backscatter observations. More recently, continuous-wave (CW) lidar backscatter measurements at mid-infrared wavelengths have been made during the Multicenter Airborne Coherent Atmospheric Wind Sensor (MACAWS) experiment in 1995. Using Lorenz-Mie theory, these datasets have been used to develop a method to convert lidar backscatter to the 2.1 micrometer wavelength. This paper presents comparison of modeled backscatter at wavelengths for which backscatter measurements exist including converted beta (sub 2.1).

Srivastava, V.↗