Engineering PapersSearch

SEARCH · Engineering Papers

Results for “dataset”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Improved Estimates of Pentad Precipitation through the Merging of Independent Precipitation Datasets

Three independent, quasi-global, gridded datasets of precipitation (a rain gauge-based dataset, the satellite-only component of the NASA Integrated Multi-satellitE Retrievals for Global Precipitation Measurement mission [IMERG] Final Run precipitation product, and precipitation estimates derived from NASA Soil Moisture Active Passive [SMAP] soil moisture retrievals), are objectively combined into a single pentad precipitation dataset at 36-km resolution using a unique approach based on extended triple collocation. The quality of each of the four datasets is then evaluated against independent observations. When a global land surface model at 36-km resolution is integrated four times, once utilizing the merged precipitation forcing and once with each of the three contributing datasets, the near-surface soil moisture variations produced with the merged forcing validate best against independent satellite-based soil moisture fields. In addition, the merged dataset is found to be more consistent, relative to each contributor, with estimates of air temperature variations across the globe. The merged dataset thus appears to draw successfully on the complementary strengths of each contributor: the particularly high quality of the rain gauge-based dataset in areas of high gauge density, the more uniform accuracy across the globe of the IMERG data, and the moderate accuracy, particularly in semi-arid regions, of the soil moisture retrieval-based data. Plain Language Summary Obtaining measurements of precipitation across the globe can be challenging. Rain gauges in some ways provide the most accurate measurements, but gauges are absent in many parts of the world, and even where they exist, they only measure precipitation at the gauge itself and therefore may not provide an accurate large-scale average. Satellite-based estimates of precipitation largely overcome these problems, but such data have their own issues, notably a “snapshot” (rather than a time-average) character of the measurements and difficulty associated with interpreting the measured radiances in the presence of complex land surfaces. In the present paper, we use a novel approach to generate a “merged” dataset, one that optimally combines the gauge precipitation information and the satellite-based precipitation information with a third set of estimates derived from soil moisture retrievals. The merged precipitation dataset and each of the three contributors (aggregated here to 5-day averages at a spatial resolution of about 36-km) are then evaluated for consistency with independent geophysical fields. The merged dataset is found to perform best, a clear indication that it takes proper advantage of the complementary strengths of each contributor and, accordingly, that the presented approach for merging the different contributors is indeed viable.

Precipitation

Novel insights enabled by combining mouse muscle datasets from the Rodent Research-1 mission

Biological space experiments are often expensive and difficult to conduct. As such, it is critical to maximize the value of the data that is collected during these experiments. One way to do this is to combine multiple–previously separate–datasets. This can increase the number of replicates for the conditions of interest (and hence statistical power), allow new multi-factor questions to be asked, and potentially highlight new patterns that otherwise would not have been identified from single-dataset studies. However, the process of combining datasets introduces noise due to inherent technical variations between experiments. To better understand the insights that can be gained from multi-dataset analyses and the problems that may arise from joining multiple datasets, several mouse muscle RNA-Seq datasets from the Rodent Research-1 mission were first selected. Then, using the R package DESeq2, principal component analysis (PCA) plots and differentially expressed gene (DEG) lists between ground and flight muscle samples were generated for individual datasets and for different pairwise combinations of datasets. Several new DEGs were identified in the combined datasets, and patterns in the PCA plots were affected depending on which datasets were joined. Understanding the results of this work will be critical for future studies that seek to perform multi-dataset analyses.

spaceflight

Daily evaluation of 26 precipitation datasets using Stage-IV gauge-radar data for the CONUS

New precipitation (P) datasets are released regularly, following innovations in weather forecasting models, satellite retrieval methods, and multi-source merging techniques. Using the conterminous US as a case study, we evaluated the performance of 26 gridded (sub-)daily P datasets to obtain insight into the merit of these innovations. The evaluation was performed at a daily timescale for the period 2008–2017 using the Kling–Gupta efficiency (KGE), a performance metric combining correlation, bias, and variability. As a reference, we used the high-resolution (4 km) Stage-IV gauge-radar P dataset. Among the three KGE components, the P datasets performed worst overall in terms of correlation (related to event identification). In terms of improving KGE scores for these datasets, improved P totals (affecting the bias score) and improved distribution of P intensity (affecting the variability score) are of secondary importance. Among the 11 gauge-corrected P datasets, the best overall performance was obtained by MSWEP V2.2, underscoring the importance of applying daily gauge corrections and accounting for gauge reporting times. Several uncorrected P datasets outperformed gauge-corrected ones. Among the 15 uncorrected P datasets, the best performance was obtained by the ERA5-HRES fourth-generation reanalysis, reflecting the significant advances in earth system modeling during the last decade. The (re)analyses generally performed better in winter than in summer, while the opposite was the case for the satellite-based datasets. IMERGHH V05 performed substantially better than TMPA-3B42RT V7, attributable to the many improvements implemented in the IMERG satellite P retrieval algorithm. IMERGHH V05 outperformed ERA5-HRES in regions dominated by convective storms, while the opposite was observed in regions of complex terrain. The ERA5-EDA ensemble average exhibited higher correlations than the ERA5-HRES deterministic run, highlighting the value of ensemble modeling. The WRF regional convection-permitting climate model showed considerably more accurate P totals over the mountainous west and performed best among the uncorrected datasets in terms of variability, suggesting there is merit in using high-resolution models to obtain climatological P statistics. Our findings provide some guidance to choose the most suitable P dataset for a particular application.

Hylke E. Beck

Maximizing efficiency of dataset compression for machine learning potentials with information theory

Machine learning interatomic potentials (MLIPs) balance high accuracy and lower costs compared to density functional theory calculations, but their performance often depends on the size and diversity of training datasets. Large datasets improve model accuracy and generalization but are computationally expensive to produce and train on, while smaller datasets risk discarding rare but important atomic environments and compromising MLIP accuracy/reliability. Here, we develop an information-theoretical framework to quantify the efficiency of dataset compression methods and propose an algorithm that maximizes this efficiency. By framing atomistic dataset compression as an instance of the minimum set cover (MSC) problem over atom-centered environments, our method identifies the smallest subset of structures that contains as much information as possible from the original dataset while pruning redundant information. The approach is extensively demonstrated on the GAP-20 and TM23 datasets and validated on 64 varied datasets from the ColabFit repository. Across all cases, MSC consistently retains outliers, preserves dataset diversity, and reproduces the long-tail distributions of forces even at high compression rates, outperforming other subsampling methods. Furthermore, MLIPs trained on MSC-compressed datasets exhibit reduced error for out-of-distribution data even in low-data regimes. We explain these results using an outlier analysis and show that such quantitative conclusions could not be achieved with conventional dimensionality reduction methods. The algorithm is implemented in the open-source QUESTS package and can be used for several tasks in atomistic modeling, from data subsampling, outlier detection, and training improved MLIPs at a lower cost.

36 MATERIALS SCIENCE

Xanthos-Lake Dataset

The Xanthos-Lake v1.0 dataset provides the input data, trained machine-learning models, and simulation outputs needed to characterize lake water balance, snow and ice conditions, and mixing-layer temperature within the Xanthos global hydrological modeling framework. The dataset supports lake representation across a wide range of lake sizes and hydroclimatic conditions by combining xLSIM, a basin-specific machine-learning emulator of lake snow, ice, ice-cover fraction, and mixing-layer temperature, with the Xanthos-Lake water-balance model. The archive contains NetCDF datasets used to train and evaluate xLSIM, trained model weights, processed meteorological and lake-property inputs, and basin- and lake-category-specific simulation outputs. These materials are organized into four primary data groups, described below. Snowice_model_inputs: Contains the NetCDF input data used to train xLSIM. The xLSIM machine-learning framework uses three lake-based datasets. The meteorological forcing dataset provides monthly relative humidity, specific humidity, surface wind speed, maximum and minimum air temperature, downward longwave and shortwave radiation, snowfall, surface air pressure, and total precipitation. Lake surface area is included as an additional static predictor. The target-state dataset provides lake ice thickness, snow depth, snow cover, and lake mixing-layer temperature, while a companion lake-surface dataset provides the lake ice-cover fraction. Before training, ice thickness and snow depth are converted from meters to centimeters, mixing-layer temperature is converted from kelvin to degrees Celsius and constrained to nonnegative values, and ice-cover fraction is converted from a fraction to a percentage. The predictor variables are normalized using statistics calculated across the selected lakes and time steps. Snowice_model_outputs: Contains the NetCDF outputs generated by xLSIM. For each basin, xLSIM produces a file containing observed and predicted lake-state variables for the training, validation, and testing periods. The modeled variables include lake ice thickness, snow depth, snow cover, mixing-layer temperature, and lake ice-cover fraction. For basins without a sufficiently persistent snow-and-ice signal, the emulator predicts only mixing-layer temperature. The outputs also include training and validation loss histories, the selected model configuration, identifiers of the lakes used in training, and SHAP-based feature-importance information at the global, lake, and seasonal-regime levels. The trained machine-learning model weights are provided separately within the dataset archive. Together, these files support model evaluation and subsequent coupling with the Xanthos-Lake water-balance framework. XanthosLAKES: Contains the NetCDF input data used by the Xanthos-Lake framework. Monthly meteorological inputs include relative and specific humidity, downward shortwave and longwave radiation, mean, maximum, and minimum air temperature, wind speed, precipitation, snowfall, and surface air pressure. Static lake-property datasets provide lake identifiers, geographic locations, surface area, volume, mean depth, elevation, drainage area, fetch, outlet-routing information, and associated Xanthos grid-cell attributes. Separate bathymetric datasets provide the coefficients of the area–depth and volume–depth relationships for each aggregated lake unit. GLEV-based records provide observed lake surface area and evaporation data used to initialize lake states, define reference conditions, and calibrate and evaluate the model. Xanthos-Lake Outputs: Contains the basin- and lake-category-specific NetCDF outputs generated by Xanthos-Lake. Monthly variables include lake surface area, storage volume, outlet discharge, evaporation rate, evaporation volume, lake–groundwater exchange, lake inflow, ice thickness, snow depth, snow-cover fraction, ice-cover fraction, and mixing-layer temperature. The files also contain lake-specific calibration and validation statistics, including normalized root-mean-square error, mean absolute error, Nash–Sutcliffe efficiency, Kling–Gupta efficiency, and percent bias. Stored calibrated and derived parameters include the weir discharge coefficient, fractional freeboard, groundwater exchange coefficient, reference water level, corresponding reference surface area and storage volume, weir-width adjustment factor, and the fraction of routed inflow entering the lake. Basin identifiers, lake category, simulation period, calibration and validation periods, and parameter-schema information are retained as NetCDF metadata.

Abeshu, Guta [Pacific Northwest National Laborator

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets such as sex or age of the model organism used. In the present study, NASA GeneLab-hosted RNAseq datasets from rodent liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC, to determine statistical differences between datasets before and after correction, Principal Component Analysis, to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the standard approach. Thus, the most robust standard correction will be implemented in the GeneLab Visualization 2.0 platform when datasets are combined.

GeneLab, RNA-seq, Batch Correction

Evaluation of Correction Methods for NASA GeneLab Transcriptomic Datasets

Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. The results showed that the reference-based approach introduced several additional (and likely artificial) DEGs when compared with the respective standard approach. Of the methods tested, standard ComBat and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

GeneLab

Combining RNA-SEQ Datasets from NASA GENELAB: An Evaluation of Correction Methods

Background: Conducting space biology experiments aboard the International Space Station, particularly those utilizing complex model organisms like mice, is expensive and difficult due to limited crew availability, hardware, and space. As a result, sample numbers from these studies are low, reducing the statistical power of any one experiment. Aggregating spaceflight datasets serves as a method to increase sample numbers, allowing for novel insights through bioinformatic analysis of ‘omics data from merged datasets. However, aggregating datasets can introduce unwanted variation including 1) differences in sample handling, processing, and sequencing platforms between datasets (technical variation) as well as 2) differences in experimental design between datasets. Methods: In the present study, NASA GeneLab-hosted RNAseq datasets from mouse liver tissues were used to evaluate several statistical methods to correct for this unwanted variation through two approaches, reference-based and standard. The following correction algorithms were applied with (reference-based) and/or without (standard) considering Universal Mouse RNA Reference samples: ComBat and ComBat_seq from the SVA package, the median polish, empirical Bayes, and ANOVA-based algorithms from the MBatch package, and negative binomial regression normalization in the DESeq2 package. For each approach, after the correction algorithm was applied, differential gene expression (DGE) analysis of flight and ground control samples was performed with the combined data. The robustness of each tool was evaluated using BatchQC to determine statistical differences between datasets before and after correction, Principal Component Analysis to evaluate global gene expression in samples before and after correction, and by comparing DGE analysis of individual datasets and combined datasets before and after correction. Results: The results showed that the reference-based approach introduced several additional (and likely artificial) differentially expressed genes when compared with the respective standard approach. Conclusions: Of the methods tested, standard ComBat_seq and DESeq2 were identified as the most robust correction methods for combining spaceflight mouse liver RNAseq datasets hosted on GeneLab.

Finsam Samson

Harmonizing Multi-Disciplinary Data for Applied Research: A Comprehensive Analysis of GES DISC Datasets through NLP and TF-IDF Techniques

Data centers distribute data encompassing multiple disciplines, making it necessary to evaluate dataset applicability to applied research. Typically, these datasets are generated by specialized science teams and consist of single-discipline data, such as atmospheric temperature, pressure, and precipitation, leading to unique dataset formats and access services. However, in applied research, the utilization of datasets from multiple disciplines is commonly necessary. The increasing availability of research literature citing Earth Science datasets presents an opportunity to analyze the usage of datasets in multi-disciplinary research. This study proposes a novel approach wherein research publications citing datasets archived at the GES DISC (Goddard Earth Sciences Data and Information Services Center) are collected, and each publication is associated with specific research topics through the application of Natural Language Processing (NLP), using the Term Frequency Inverse Document Frequency (TF-IDF) technique on the publication titles and abstracts. Through this analysis, we gain insights into the distribution of dataset disciplines as they are being used in various applied research areas. This knowledge is essential for the development of dataset tools and services tailored to effectively support applied research studies, as it enhances a data center's comprehension of how datasets from multiple disciplines are integrated into research endeavors.

Infometrics

Enhancing Dataset Discovery and Usage Tracking in Earth Sciences: Integrating Knowledge Graphs and Large Language Models

NASA's Data Active Archive Centers (DAACs) have played a crucial role in supporting a wide range of applied research in Earth and Environmental sciences. To date, over 20,000 publications have been collected, citing more than 3,000 NASA Earth science datasets. We present an innovative approach that links datasets and collected publications through a knowledge graph (KG). This KG enables the tracking of dataset citations throughout the dataset's lifecycle, revealing patterns of dataset usage across various applied research areas. We fine-tuned the pre-trained NASA IMPACT INDUS-Base Retriever Large Language Model (LLM) using a set of labeled publication abstracts. Our results indicate that 87% of the publications were classified into one of twenty applied research areas, while the remaining 13% were categorized into non-applied research areas. The classified publications linked to datasets are used to discover datasets by users interested in specific applied research and by dataset providers to determine dataset usage for applications.

open-source

A globally sampled high-resolution hand-labeled validation dataset for evaluating surface water extent maps

Effective monitoring of global water resources is increasingly critical due to climate change and population growth. Advancements in remote sensing technology, specifically in spatial, spectral, and temporal resolutions, are revolutionizing water resource monitoring, leading to more frequent and high-quality surface water extent maps using various techniques such as traditional image processing and machine learning algorithms. However, satellite imagery datasets contain trade-offs that result in inconsistencies in performance, such as disparities in measurement principles between optical (e.g., Sentinel-2) and radar (e.g., Sentinel-1) sensors and differences in spatial and spectral resolutions among optical sensors. Therefore, developing accurate and robust surface water mapping solutions requires independent validations from multiple datasets to identify potential biases within the imagery and algorithms. However, high-quality validation datasets are expensive to build, and few contain information on water resources. For this purpose, we introduce a globally sampled, high-spatial-resolution dataset labeled using 3 m PlanetScope imagery. Our surface water extent dataset comprises 100 images, each with a size of 1024×1024 pixels, which were sampled using a stratified random sampling strategy covering all 14 biomes. We highlighted urban and rural regions, lakes, and rivers, including braided rivers and coastal regions. We evaluated two surface water extent mapping methods using our dataset – Dynamic World, based on Sentinel-2, and the NASA IMPACT model, based on Sentinel-1. Dynamic World achieved a mean intersection over union (IoU) of 72.16 % and F1 score of 79.70 %, while the NASA IMPACT model had a mean IoU of 57.61 % and F1 score of 65.79 %. Performance varied substantially across biomes, highlighting the importance of evaluating models on diverse landscapes to assess their generalizability and robustness. Our dataset can be used to analyze satellite products and methods, providing insights into their advantages and drawbacks. Our dataset offers a unique tool for analyzing satellite products, aiding the development of more accurate and robust surface water monitoring solutions. The dataset can be accessed via https://doi.org/10.25739/03nt-4f29.

54 ENVIRONMENTAL SCIENCES

Meta-analysis of North American Arctic and boreal aboveground biomass datasets: assessing accuracy, dynamics, and similarities

The North American arctic and boreal regions (ABRs) are rapidly warming and experiencing intensifying disturbances. Accurately quantifying aboveground biomass (AGB) is critical for understanding the impacts of these changes on the carbon cycle and for designing climate change mitigation strategies. Several AGB maps have been developed for the North American ABRs, including recent contributions from National Aeronautics and Space Administration’s Arctic-Boreal Vulnerability Experiment (ABoVE) campaign. However, these maps differ widely in training data, methodology, and resulting AGB density estimates. Presently, a comprehensive comparative evaluation is lacking, making it difficult for users to select datasets suited to their research or management needs. Here, in this study, we conducted a comparative analysis of nine AGB density datasets across North American ABRs, specifically for Alaska and Canada. We (1) summarized AGB by ecoregion and Canadian provinces, (2) evaluated their accuracy against field-based measurements, (3) analyzed spatial and temporal similarities among datasets, and (4) assessed their ability to capture disturbance (fire and harvest) impacts on AGB. We found substantial variation in regional and local AGB estimates across datasets, with overall accuracy ranging from R 2 = 0.25–0.62 and Bias% from −47.8% to 69.9% when validated against field plots. Despite these differences, most datasets have comparatively consistent spatial patterns in AGB (r > 0.8 for most cases). In contrast, agreement on the temporal patterns of AGB change is generally low. We found datasets with spatial resolutions ⩽300 m are capable of capturing disturbance impacts on AGB dynamics, though sensitivity varies across products. Our findings and dataset summary provide guidance for selecting appropriate AGB datasets for different applications within our study area. Our analysis also highlights the need to decrease map bias and increase capability to detect temporal change to decrease uncertainty of AGB datasets potentially by using training data which is representative of major plant functional types within the mapped area.

ABoVE

Observational ozone datasets over the global oceans and polar regions (version 2024)

Studying tropospheric ozone over the remote areas of the planet, such as the open oceans and the polar regions, is crucial to understand the role of ozone as a global climate forcer and regulator of atmospheric oxidative capacity. A focus on the pristine oceanic and polar regions complements the available land-based datasets and provides insights into key photochemical and depositional loss processes that control the concentrations and spatiotemporal variability in ozone as well as the physicochemical mechanisms driving these patterns. However, an assessment of the role of ozone over the oceanic and polar regions has been hampered by a lack of comprehensive observational datasets. Here, we present the first comprehensive collection of ozone data over the oceans and the polar regions. The overall dataset consists of 77 ship cruises/buoy-based observations and 48 aircraft-based campaigns. The dataset, consisting of more than 630 000 independent ozone measurement data points covering the period from 1977 to 2022 and an altitude range from the surface to 5000 m (with a focus on the lowest 2000 m), allows systematic analyses of the spatiotemporal distribution and long-term trends over the 11 defined ocean/polar regions. The datasets from ships, buoys, and aircraft are complemented by ozonesonde data from 29 launch sites or field campaigns and by 21 non-polar and 17 polar ground-based station datasets. The datasets contain information on how long the observed air masses were isolated from land, as estimated by backward trajectories from the individual observation points. To extract observations representative of oceanic conditions, we recommend using a subset of the data with an isolation time of 72 h or longer, from the analysis with coincident radon observations. These filtered oceanic and polar data showed typically flat diurnal cycles at high latitudes, whereas daytime decreases in ozone (11 %–16 %) were observed at lower latitudes. The ship/buoy- and aircraft-based datasets presented here will supplement the land-based ones in the TOAR-II (Tropospheric Ozone Assessment Report Phase II) database to provide a fully global assessment of tropospheric ozone. The described dataset is available at https://doi.org/10.17596/0004044 (Kanaya et al., 2025).

Kanaya, Yugo [Japan Agency for Marine-Earth Scienc

Performance of wind assessment datasets in United States coastal areas

The atmospheric dynamics that occur near the intersection of land and water offer exciting and challenging opportunities for wind energy deployment in coastal locations. New models and tools are continually being developed in support of wind resource assessment, and three recent products are explored in this work for their performance in representing characteristics of the wind resource at coastal locations: the Global Wind Atlas 3 (GWA3), the 2023 National Offshore Wind dataset (NOW-23), and the wind climate simulations that are a component of the Wind Integration National Dataset (WIND) Toolkit Long-Term Ensemble Dataset (WTK-LED Climate). These relatively new products are freely available and user-friendly so that anyone – from a utility-scale developer to a resident or business owner – can evaluate the potential for wind energy generation at their location of interest. The validations in this work provide guidance on the accuracy of wind resource assessments for coastal customers interested in installing small or midsize wind turbines (≤ 1 MW in capacity) to support energy needs at the residential, business, or community scale, such as the island and remotely located participants of the U.S. Department of Energy's Energy Transitions Initiative Partnership Project. At 23 coastal locations across the United States, dataset performance varies according to different evaluation metrics. All three recent datasets tend to overestimate the observed coastal wind resource. GWA3 produces the smallest annual average wind speed relative errors, whereas WTK-LED Climate is in best agreement in terms of representing diurnal wind speed cycles. NOW-23 is the highest performing of the datasets for representing seasonal and interannual trends in the coastal wind resource. While GWA3 and WTK-LED Climate are relatively insensitive to the dataset output heights selected for wind resource assessment at small and midsize wind turbine hub heights (20–60 m), significant variation in the NOW-23 representation of wind shear across the wind profile in the lowest 100 m of the atmosphere leads to notable differences in wind speed estimates according to the dataset output heights selected for evaluation. GWA3 exhibits challenges in the representation of observed wind speed diurnal cycles at small and midsize turbine hub heights, likely due to the dataset's consistent treatment of hourly wind speed trends regardless of altitude.

17 WIND ENERGY

Assessment of NASA's Physiographic and Meteorological Datasets as Input to HSPF and SWAT Hydrological Models

This paper documents the use of simulated Moderate Resolution Imaging Spectroradiometer land use/land cover (MODIS-LULC), NASA-LIS generated precipitation and evapo-transpiration (ET), and Shuttle Radar Topography Mission (SRTM) datasets (in conjunction with standard land use, topographical and meteorological datasets) as input to hydrological models routinely used by the watershed hydrology modeling community. The study is focused in coastal watersheds in the Mississippi Gulf Coast although one of the test cases focuses in an inland watershed located in northeastern State of Mississippi, USA. The decision support tools (DSTs) into which the NASA datasets were assimilated were the Soil Water & Assessment Tool (SWAT) and the Hydrological Simulation Program FORTRAN (HSPF). These DSTs are endorsed by several US government agencies (EPA, FEMA, USGS) for water resources management strategies. These models use physiographic and meteorological data extensively. Precipitation gages and USGS gage stations in the region were used to calibrate several HSPF and SWAT model applications. Land use and topographical datasets were swapped to assess model output sensitivities. NASA-LIS meteorological data were introduced in the calibrated model applications for simulation of watershed hydrology for a time period in which no weather data were available (1997-2006). The performance of the NASA datasets in the context of hydrological modeling was assessed through comparison of measured and model-simulated hydrographs. Overall, NASA datasets were as useful as standard land use, topographical , and meteorological datasets. Moreover, NASA datasets were used for performing analyses that the standard datasets could not made possible, e.g., introduction of land use dynamics into hydrological simulations

Alacron, Vladimir J.

Hydrology Research with the North American Land Data Assimilation System (NLDAS) Datasets at the NASA GES DISC Using Giovanni

The North American Land Data Assimilation System (NLDAS) is a collaboration project between NASA/GSFC, NOAA, Princeton Univ., and the Univ. of Washington. NLDAS has created a surface meteorology dataset using the best-available observations and reanalyses the backbone of this dataset is a gridded precipitation analysis from rain gauges. This dataset is used to drive four separate land-surface models (LSMs) to produce datasets of soil moisture, snow, runoff, and surface fluxes. NLDAS datasets are available hourly and extend from Jan 1979 to near real-time with a typical 4-day lag. The datasets are available at 1/8th-degree over CONUS and portions of Canada and Mexico from 25-53 North. The datasets have been extensively evaluated against observations, and are also used as part of a drought monitor. NLDAS datasets are available from the NASA GES DISC and can be accessed via ftp, GDS, Mirador, and Giovanni. GES DISC news articles were published showing figures from the heat wave of 2011, Hurricane Irene, Tropical Storm Lee, and the low-snow winter of 2011-2012. For this presentation, Giovanni-generated figures using NLDAS data from the derecho across the U.S. Midwest and Mid-Atlantic will be presented. Also, similar figures will be presented from the landfall of Hurricane Isaac and the before-and-after drought conditions of the path of the tropical moisture into the central states of the U.S. Updates on future products and datasets from the NLDAS project will also be introduced.

Mocko, David M.

Review and Analysis of Algorithmic Approaches Developed for Prognostics on CMAPSS Dataset

Benchmarking of prognostic algorithms has been challenging due to limited availability of common datasets suitable for prognostics. In an attempt to alleviate this problem several benchmarking datasets have been collected by NASA's prognostic center of excellence and made available to the Prognostics and Health Management (PHM) community to allow evaluation and comparison of prognostics algorithms. Among those datasets are five C-MAPSS datasets that have been extremely popular due to their unique characteristics making them suitable for prognostics. The C-MAPSS datasets pose several challenges that have been tackled by different methods in the PHM literature. In particular, management of high variability due to sensor noise, effects of operating conditions, and presence of multiple simultaneous fault modes are some factors that have great impact on the generalization capabilities of prognostics algorithms. More than 70 publications have used the C-MAPSS datasets for developing data-driven prognostic algorithms. The C-MAPSS datasets are also shown to be well-suited for development of new machine learning and pattern recognition tools for several key preprocessing steps such as feature extraction and selection, failure mode assessment, operating conditions assessment, health status estimation, uncertainty management, and prognostics performance evaluation. This paper summarizes a comprehensive literature review of publications using C-MAPSS datasets and provides guidelines and references to further usage of these datasets in a manner that allows clear and consistent comparison between different approaches.

Uncertainty

Total Ozone Trends from 1979 to 2016 Derived from Five Merged Observational Datasets - The Emergence into Ozone Recovery

We report on updated trends using different merged datasets from satellite and ground-based observations for the period from 1979 to 2016. Trends were determined by applying a multiple linear regression (MLR) to annual mean zonal mean data. Merged datasets used here include NASA MOD v8.6 and National Oceanic and Atmospheric Administration (NOAA) merge v8.6, both based on data from the series of Solar Backscatter UltraViolet (SBUV) and SBUV-2 satellite instruments (1978–present) as well as the Global Ozone Monitoring Experiment (GOME)-type Total Ozone (GTO) and GOME-SCIAMACHY-GOME-2 (GSG) merged datasets (1995-present), mainly comprising satellite data from GOME, the Scanning Imaging Absorption Spectrometer for Atmospheric Chartography (SCIAMACHY), and GOME-2A. The fifth dataset consists of the monthly mean zonal mean data from ground-based measurements collected at World Ozone and UV Data Center (WOUDC). The addition of four more years of data since the last World Meteorological Organization (WMO) ozone assessment (2013-2016) shows that for most datasets and regions the trends since the stratospheric halogen reached its maximum (approximately 1996 globally and approximately 2000 in polar regions) are mostly not significantly different from zero. However, for some latitudes, in particular the Southern Hemisphere extratropics and Northern Hemisphere subtropics, several datasets show small positive trends of slightly below +1 percent decade(exp. -1) that are barely statistically significant at the 2 Sigma uncertainty level. In the tropics, only two datasets show significant trends of +0.5 to +0.8 percent(exp.-1), while the others show near-zero trends. Positive trends since 2000 have been observed over Antarctica in September, but near-zero trends are found in October as well as in March over the Arctic. Uncertainties due to possible drifts between the datasets, from the merging procedure used to combine satellite datasets and related to the low sampling of ground-based data, are not accounted for in the trend analysis. Consequently, the retrieved trends can be only considered to be at the brink of becoming significant, but there are indications that we are about to emerge into the expected recovery phase. However, the recent trends are still considerably masked by the observed large year-to-year dynamical variability in total ozone.

column ozone trends