Engineering PapersSearch

SEARCH · Engineering Papers

Results for “data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

University Data Management Pilot Utilizing the Nuclear Research Data System

Background In 2022, the Office of Science and Technology Policy (OSTP) issued a memo that significantly reshaped the landscape of access to federally funded research. The memo mandated that all taxpayer-funded research be made available to the public without delay upon publication, without an embargo period, superseding the 2013 OSTP public access policy. This public access policy promotes transparency and the democratization of knowledge, ensuring that the fruits of scientific endeavors funded by federal agencies could be immediately accessed and built upon by scientists, educators, students, and the public at large. To implement the requirements of the OSTP guidance and DOE Public Access Plan, the Office of Nuclear Energy (NE) has implemented public access plan guidance and has identified several areas where better data management practices would further expand public access to important nuclear energy related scientific data, reports, and other technical products. Significant NE supported efforts are already underway for data management and public access to important nuclear energy related data.1 2 To address gaps in data management practices, and improve retention and accessibility of data, NE is actively exploring enhanced data management options utilizing its high-performance computing resources administered by its Nuclear Scientific User Facility Program. A newly piloted system, the Nuclear Research Data System (NRDS) acts as a portal for data collection and dissemination. Nuclear Energy University Program Research and Development Portfolio According to Web of Science, NEUP has produced 2,345 journal publication that have been cited more than 61,000 times3 and countless conference proceedings. These publications are publicly available through OSTI.gov and in the open literature. Additional scientific and technical products including project milestones that are not publications and NEUP project final reports are vetted through OSTI.gov and released once reviewed and approved by DOE. Since 2009, NEUP has awarded close to 1,000 different R&D projects in technical areas across the NE research programs. As of June 2023, 512 NEUP reports are publicly available on OSTI. The underlying data for projects is still held at universities, and data transfer, co-location, and dissemination has not occurred in a systematic way. NEUP data is currently accessible through myriad university-based data repositories, or through direct requests to PIs. The program identified this patchwork of repositories, or often lack of publicly available data, as a significant barrier to an organized, accessible, and comprehensive solution to sharing data with the larger nuclear energy community. Approach The goal of this pilot project is to establish a pathway to a consolidated long-term repository for NEUP project data. To accomplish this goal, the pilot strives to accomplish the following objectives: Establish data collection standards, including a standard set of required supplementary information to contextualize and support raw data files. Work with the HPC group collect and upload information and to modify the NRDS system, as needed, to support a standardized approach. Resolve potential barriers to successful roll out of an expanded data collection strategy, including modifying data management plan guidelines and establishing a document and data release process that accounts for potential intellectual property and/or export control concerns. Results Overall, the pilot was successful in collecting 8,982 raw and processes data files, 220 reports, 56 calibration files, and 5,931 other supplementary documents. Supplementary documents included experimental plans, methods, journal publications and conference proceedings, milestone reports, and final reports. Figure 2 shows the number of data sets and supplementary project information provided by each project. Projects has significantly different input, depending on experimental data produced and completeness of the datasets provided.

Data collection

Quantitative Highlights of 20 years Aqua Data Archive and Data Usage

NASA’s Aqua satellite carries six Earth-observing instruments Atmospheric Infrared Sounder (AIRS), Advanced Microwave Scanning Radiometer for EOS (AMSR-E), Advanced Microwave Sounding Unit (AMSU), Clouds and the Earth’s Radiant Energy System (CERES), Humidity Sounder for Brazil (HSB) and Moderate Resolution Imaging Spectroradiometer (MODIS). Currently only four of six instruments are collecting data, two instruments that stopped transmitting data are AMSR-E that suffered a major anomaly in October 2011 and was powered off in March 2016 while as HSB failed in February 2003. NASA’s Earth Science Data and Information System (ESDIS) Project makes these data, along with derived products, available to worldwide data users. Since the launch of Aqua on May 4, 2002, more than 10,000 data products have been archived and distributed by NASA-funded Distributed Active Archive Centers (DAACs) that are part of NASA’s Earth Observing System Data and Information System (EOSDIS). At the end of the 2021 Fiscal Year with over 100,000 orbits data, about 1,000 Aqua data products constituted almost 16.5 % of the entire EOSDIS data archive volume (8.6 PB out of approximately 55.2 PB), and 7.5 PB of Aqua data were distributed to over half-a-million public users worldwide. By categorizing the Aqua data products and their distribution, we can get a quantitative assessment of Aqua data usage. NASA’s ESDIS Project has collected archive, distribution, and user information from EOSDIS data users since February 2000. These metrics are available through the ESDIS Metrics System (EMS). EMS information is stored in a relational database from which quantitative metrics of Aqua data use can be retrieved and analyzed. The purposes of this study are to: 1) perform a comprehensive investigation of the 20-year trend in the archive and distribution of Aqua data products; 2) identify and characterize data product usage over the last 20 years; and 3) identify and characterize the global user community for these data. In addition to revealing how Aqua data use has evolved over time, the results of this study provide insights on identifying the various user communities for different kinds of Earth science data products. Also, because of the enormous quantity of data handled by EOSDIS DAACs, the study provides guidance of the requirements for future data systems that will be needed to effectively and efficiently handle the ever-increasing amounts of Earth science data produced by future (and ongoing) Earth science missions.

Lalit Wanchoo

DOE EV Data Collection - Vehicle Data

Vehicle data consist of electric vehicle performance data collected directly from the vehicle during standard operations. Data were collected using onboard data loggers that were either installed by the project team or preinstalled by the original equipment manufacturer. Data recorded by the data loggers were made accessible via an online web portal or an application programming interface. Different data loggers were used (HEM, ViriCiti, and Geotab), and the method for each vehicle is defined in the vehicle attributes file. Some systems collected data on a “trip-level” basis, in which each row of a table represents a single trip (the period between a key-on and key-off event), whereas other data were collected on a per-day basis, in which each row represents a single day of operation. Data were collected over a range of data collection periods, depending on the project. Data have been anonymized by removing information or decreasing information resolution as necessary so that fleets are not identifiable. Due to the wide range of vehicle types represented and variation in data collection, data parameters and frequencies differ between vehicles and fleets The **Performance Data Daily/Trip Data Dictionaries** contain definitions for each available parameter associated with a vehicle’s operations, aggregated at either a daily or trip level. The parameters available will vary from vehicle to vehicle, but every possible parameter will be defined. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. The **Vehicle Data** tables contain the data from each vehicle’s operations, aggregated at either a daily or trip level, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Increasing Discovery and Usability of Earth Science Satellite Data with My NASA Data

For 20 years, the My NASA Data project at NASA Langley Research Center has developed innovative approaches to increase the use of NASA’s satellite data by learners. My NASA Data offers a variety of authentic Earth Science datasets and a data visualization tool, eliminating the need for educators and/or learners to obtain specialized knowledge of GIS data formats and software to access and use authentic Earth Science data. While there is no shortage of available data, as federal government agencies such as NASA house petabytes of freely accessible Earth Science datasets, much of the data are only available for download and visualization in specialized formats and software, limiting their accessibility to educators and learners, especially those in primary and secondary school. Using the Google Earth Engine platform, the My NASA Data team has recently reinvented their data visualization tool, called the Earth System Data Explorer (ESDE). The ESDE gives users the capability to explore over 60 Earth Science satellite datasets in a multitude of formats such as maps, graphs, and data table Its new and improved user interface design was developed based on the preferences of educators, whom the My NASA Data project has over 20 years’ experience working with. Earth Science and GIS Subject Matter Experts (SMEs) structured the data in a professional and scientific manner. During Fiscal Year 2023, the My NASA Data website received over 1 million digital engagements, with over one-third being visitors to the data visualization tool. These metrics highlight the interest in a visualization tool that is simple and free to use with reliable and trusted datasets. The ESDE empowers users to readily relate and analyze NASA Earth Science data within their area of interest. The team used a user-centered design (UCD) framework to receive and incorporate feedback into the application’s design. Core requested features include the ability to create time series graphs, comparative analysis of maps, and download the data as CSV file. Responses indicate that advances in data visualization tools such as the ESDE make authentic Earth Science data more accessible. This presentation will cover how the My NASA Data project develops tools to enhance data discovery and accessibility, as well as how SME and user suggestions are incorporated.

Desiray Wilson

DOE EV Data Collection - Charging Data

Charging data are collected from one of three sources, each with varying levels of additional information. These sources, in approximate order from most to least additional information, are: • The electric vehicle supply equipment (charger) • Onboard the vehicle itself • From a utility submeter. Many chargers provide software that allows for the collection and reporting of charging session data. If unavailable, data may be recorded by the charging vehicle’s onboard systems. If neither of these options is available, data can be acquired from utility submeters that simply track the energy flowing to one or more chargers. Data collected directly from the electric vehicle supply equipment (EVSE) are typically the most accurate and highest frequency. However, it is not always possible to discern which exact vehicle is being charged during any one session. EVSE-side data can be identified where a single charger ID but a range of vehicle IDs are present (e.g., CH001, EV001-EV005). Data collected from the vehicle’s onboard systems usually does not provide information on which exact charger is being used. Vehicle-side data can be identified where a single Vehicle ID but a range of Charger IDs are present (e.g., EV001, CH001-CH005). Data collected from utility submeters provide no information on which specific vehicle is charging or which specific charger is in use. Submeter data can be identified where multiple Vehicle IDs and multiple Charger IDs are present, but only a single Fleet ID is present (e.g., EV001-EV005, CH001-CH005, Fleet01). The **Charge Data Daily/Session Dictionaries** contains definitions for each available parameter collected as part of an individual charging session, aggregated at either a daily or session level. The parameters available will vary between vehicles and chargers. The **Charger Attributes** table contains specific charger characteristics, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. The **Charger Attributes Data Dictionary** contains definitions for each available parameter collected on the physical and operational characteristics of the charging hardware itself. The **Vehicle Attributes Data Dictionary** contains definitions for each available parameter associated with a vehicle’s physical and functional attributes and fleet context. The **Vehicle Attributes** table contains specific vehicle characteristics, coded to an anonymous Vehicle ID. This Vehicle ID can be used as a key between vehicle data and vehicle attribute tables, and in cases where charging data are supplied, links a vehicle with the charger(s) that supplied it power. The **Charging Data** tables contain the data from each charger’s operations, coded to at least one anonymous Charger ID and linked to either a single or a range of Vehicle IDs. Vehicle ID can be used as a key between charging data and vehicle attribute tables. Data is being uploaded quarterly through 2023 and subject to change until the conclusion of the project.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

The Open Data Repositorys Data Publisher

Data management and data publication are becoming increasingly important components of researcher's workflows. The complexity of managing data, publishing data online, and archiving data has not decreased significantly even as computing access and power has greatly increased. The Open Data Repository's Data Publisher software strives to make data archiving, management, and publication a standard part of a researcher's workflow using simple, web-based tools and commodity server hardware. The publication engine allows for uploading, searching, and display of data with graphing capabilities and downloadable files. Access is controlled through a robust permissions system that can control publication at the field level and can be granted to the general public or protected so that only registered users at various permission levels receive access. Data Publisher also allows researchers to subscribe to meta-data standards through a plugin system, embargo data publication at their discretion, and collaborate with other researchers through various levels of data sharing. As the software matures, semantic data standards will be implemented to facilitate machine reading of data and each database will provide a REST application programming interface for programmatic access. Additionally, a citation system will allow snapshots of any data set to be archived and cited for publication while the data itself can remain living and continuously evolve beyond the snapshot date. The software runs on a traditional LAMP (Linux, Apache, MySQL, PHP) server and is available on GitHub (http://github.com/opendatarepository) under a GPLv2 open source license. The goal of the Open Data Repository is to lower the cost and training barrier to entry so that any researcher can easily publish their data and ensure it is archived for posterity.

Astrobiology data

A cost and community perspective on the barriers to microbiome data reuse

Microbiome research is becoming a mature field with a wealth of data amassed from diverse ecosystems, yet the ability to fully leverage multi-omics data for reuse remains challenging. To provide a view into researchers’ behavior and attitudes towards data reuse, we surveyed over 700 microbiome researchers to evaluate data sharing and reuse challenges. We found that many researchers are impeded by difficulties with metadata records, challenges with processing and bioinformatics, and problems with data repository submissions. We also explored the cost constraints of data reuse at each step of the data reuse process to better understand “pain points” and to provide a more quantitative perspective from sixteen active researchers. The bioinformatics and data processing step was estimated to be the most time consuming, which aligns with some of the most frequently reported challenges from the community survey. From these two approaches, we present evidence-based recommendations for how to address data sharing and reuse challenges with concrete actions for future work.

59 BASIC BIOLOGICAL SCIENCES

High-performance data format for scientific data storage and analysis

Here, in this article, we present the High-Performance Output (HiPO) data format developed at Jefferson Laboratory for storing and analyzing data from Nuclear Physics experiments. The format was designed to efficiently store large amounts of experimental data, utilizing modern fast compression algorithms. The purpose of this development was to provide organized data in the output, facilitating access to relevant information within the large data files. The HiPO data format has features that are suited for storing raw detector data, reconstruction data, and the final physics analysis data efficiently, eliminating the need to do data conversions through the lifecycle of experimental data. The HiPO data format is implemented in C++ and JAVA, and provides bindings to FORTRAN, Python, and Julia, providing users with the choice of data analysis frameworks to use. In this paper, we will present the general design and functionalities of the HiPO library and compare the performance of the library with more established data formats used in data analysis in High Energy and Nuclear Physics (such as ROOT and Parquete). In columnar data analysis, HiPO surpasses established data formats in performance and can be effectively applied to data analysis in other scientific fields.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Evaluate data lake design for the accelerator control system

Increasing precision in automation for modern particle accelerators not only creates a requirement to gather data from all devices but also demands scalable and high-performance data infrastructure with the capability of handling vast incoming device data. A well architected data lake is suitable for such a system which integrates real-time data acquisition, transient data caching, and long-term storage. This paper evaluates data lake architecture for an Accelerator Control System (ACS), focusing on two critical components of a data lake, data cache and long-term storage.

Jaikar, Amol [Fermilab]

Predicting Li-Ion Battery Capacity Fade Using Early-Life Data and a Hybrid Data-Driven Gaussian Process-Bayesian Regression Approach

Accurately predicting Li-ion battery capacity trajectories using early-life data can dramatically improve battery-life understandings and be used to rapidly evaluate design/cost/performance trade-offs when developing new battery materials. Accurate early-life predictions enable researchers to quickly iterate over cell designs and material precursor properties without consistently cycling cells to failure. To this end, we present a toolbox that uses a combined Gaussian Process and Bayesian regression approach that capitalizes on signals other than just capacity (e.g., dQ/dV, voltage drops) to rapidly predict capacity-fade trajectories. The prediction tool uses Bayesian regression to fit functional forms, e.g., power law, sigmoids, etc., to predict capacity-fade dynamics. By fitting functional forms, the capacity fade can be interrogated at any point in the future, allowing for early cell-failure prediction. Additionally, Bayesian regression allows for accurate uncertainty estimates that account for cell-to-cell variability (aleatoric uncertainty) and the lack of observation data (epistemic uncertainty). By only using early cycle data to predict the capacity fade trajectory, uncertainty bounds at end-of-life can be extremely large. The large uncertainty bounds are further exacerbated because there is no systematic way to define the prior distribution of the functional forms' parameters. We improve our the predicted trajectory confidence interval of our predicted trajectory using two methods. First, we shows that a small amount of held-out cycling data is sufficientuse some train cells, that have been cycled to failure to derive information regarding the appropriate prior distributions for the functional forms' parameters of the functional form, effectively leading to data-driven priors.. We propose constructing the data-driven priors by first running a Bayesian regression starting with uninformed priors to generate intermediate cell-specific posterior parameter distributions. These posterior distributions are combined using a Ggaussian mixture model for each parameter to create the data-driven priors. These mixture models serve as the data-driven prior distributions for the parameters for. Second, we derive multiple features, e.g., C_dchg 0.5 DoD 0.5, log (|mean(dQ/dV_(w_3-w_0 ) (V)|), etc., from the train cellsheld-out cycling data, identify which the features are that best predicting capacity at early/mid-life cycles, and then create Ggaussian process regression models that are used for predicting capacity at early/mid-life cycles for the test cells (see blue dots with error bars in Fig 1b). Finally, these predicted data-points are used in addition to the actual early cycle data capacity fade to construct the Bayesian regression trajectory for the test cell s. Notably. We note that these two methods are complementary and can be combined with each other. We evaluate the performance of our proposed method on an testing open-source dataset from Iowa State University and Iowa Lakes Community College (ISU-ILCC). This dataset comprises of 251 nickel-manganese-cobalt/graphite Lithium-ion cells that are cycled under 63 different conditions. We compute the mean average percentage error (MAPE) and negative log predictive density (NLPD) to quantify the efficacy of our method. Our initial findings suggest that, when only few observations are available, for test cells, when using only Bayesian regression with uninformed priors, a power law functional provides the most accurate predictions. with very few data points. However, asHowever, a the number of data points increases, a twin sigmoidal function becomes more accurate as the number of observations further increases. We also find that using as little as 10% of the data set towards generating data-driven priors can lead to significant improvement in prediction accuracy when using early cycle data. Lastly, we found that augmenting early-cycle data with Gaussian process-predicted capacity data for Bayesian regression greatly improves the prediction accuracy. We will present a comprehensive comparison of our methods to other methods available in the literature and apply this method to additional battery datasets.

42 ENGINEERING

BASIN-3D Data Integration for Selected ARM Data Field Campaign Report

The purpose of this data services request was to demonstrate integration of the Atmospheric Radiation Measurement (ARM) User Facility’s “met” datastreams with time series data from other earth science data sources using the BASIN-3D data synthesis software tool. BASIN-3D is an open-source Python library that enables researchers to integrate data across configured public and private data sources. It provides a common query language for researchers to request measurement locations and time series data based on specified locations, variables, time period, statistics, aggregation, and data quality. BASIN-3D acquires the data that match the query from each configured data source and translates the results into harmonized vocabularies, thus reducing researchers' data-wrangling effort. In addition, because the queries are executed on demand, researchers can easily regenerate their synthesized data sets as new data and/or data updates become available, eliminating one-off data products. BASIN-3D can output data using a variety of different data structures for end-user applications including Python pandas data frames and hdf5 output formats.

54 ENVIRONMENTAL SCIENCES

Tools and Data Services from the GSFC Earth Sciences DAAC for Aura Science Data Users

In these times of rapidly increasing amounts of archived data, tools and data services that manipulate data and uncover nuggets of information that potentially lead to scientific discovery are becoming more and more essential. The Goddard Space Flight Center (GSFC) Earth Sciences (GES) Distributed Active Archive Center (DAAC) has made great strides in facilitating science and applications research by, in consultation with its users, developing innovative tools and data services. That is, as data users become more sophisticated in their research and more savvy with information extraction methodologies, the GES DAAC has been responsive to this evolution. This presentation addresses the tools and data services available and under study at the GES DAAC, applied to the Earth sciences atmospheric data. Now, with the data from NASA's latest Atmospheric Chemistry mission, Aura, being readied for public release, GES DAAC tools, proven successful for past atmospheric science missions such as MODIS, AIRS, TRMM, TOMS, and UARS, provide an excellent basis for similar tools updated for the data from the Aura instruments. GES DAAC resident Aura data sets are from the Microwave Limb Sounder (MLS), Ozone Monitoring Instrument (OMI), and High Resolution Dynamics Limb Sounder (HIRDLS). Data obtained by these instruments afford researchers the opportunity to acquire accurate and continuous visualization and analysis, customized for Aura data, will facilitate the use and increase the usefulness of the new data. The Aura data, together with other heritage data at the GES DAAC, can potentially provide a long time series of data. GES DAAC tools will be discussed, as well as the GES DAAC Near Archive Data Mining (NADM) environment, the GIOVANNI on-line analysis tool, and rich data search and order services. Information can be found at: http://daac.gsfc.nasa.gov/upperatm/aura/. Additional information is contained in the original extended abstract.

Kempler, S.

Restoration of Apollo Data by the NSSDC and the PDS Lunar Data Node

The Lunar Data Node (LDN), under the auspices of the Geosciences Node of the Planetary Data System (PDS), is restoring Apollo data archived at the National Space Science Data Center. The Apollo data were arch ived on older media (7 -track tapes. microfilm, microfiche) and in ob solete digital formats, which limits use of the data. The LDN is maki ng these data accessible by restoring them to standard formats and archiving them through PDS. The restoration involves reading the older m edia, collecting supporting data (metadata), deciphering and understa nding the data, and organizing into a data set. The data undergo a pe er review before archive at PDS. We will give an update on last year' s work. We have scanned notebooks from Otto Berg, P.1. for the Lunar Ejecta and Meteorites Experiment. These notebooks contain information on the data and calibration coefficients which we hope to be able to use to restore the raw data into a usable archive. We have scanned Ap ollo 14 and 15 Dust Detector data from microfilm and are in the proce ss of archiving thc scans with PDS. We are also restoring raw dust de tector data from magnetic tape supplied by Yosio Nakamura (UT Austin) . Seiichi Nagihara (Texas Tech Univ.) and others in cooperation with NSSDC are recovering ARCSAV tapes (tapes containing raw data streams from all the ALSEP instruments). We will be preparing these data for archive with PDS. We are also in the process of recovering and archivi ng data not previously archived, from the Apollo 16 Gamma Ray Spectro meter and the Apollo 17 Infrared Spectrometer.

Williams, David R.

Famine Early Warning Systems Network (FEWS NET) Land Data Assimilation System (LDAS) and Other Assimilated Hydrological Data at NASA GES DISC

The NASA Goddard Earth Sciences Data and Information Services Center (GES DISC) provides science support for several data sets relevant to agriculture and food security, including the Famine Early Warning Systems Network (FEWS NET) Land Data Assimilation System (LDAS), or FLDAS data set. The GES DISC is one of twelve NASA Earth Observing System (EOS) data centers that process, archive, document, and distribute data from Earth science missions and related projects. The GES DISC hosts a wide range of remote sensing and model data, and provides reliable and robust data access and other services to users worldwide. Beyond data archive and access, the GES DISC offers many services to visualize and analyze the data. This presentation provides a summary of the hydrological data available at the GES DISC, along with an overview of related data services. Specifically, the FLDAS data set has been adapted to work with domains, data streams, and monitoring and forecast requirements associated with food security assessment in data-sparse, developing country settings. The FLDAS global monthly data have a 0.1 x 0.1 degree spatial resolution covering the period from January 1982 to present. Global FLDAS monthly anomaly and monthly climatology data are also available at the GES DISC to evaluate how current conditions compare to averages over the FLDAS 35-year period. Several case studies using the FLDAS soil moisture, evapotranspiration, rainfall, runoff, and surface temperature data will be presented.

Loeser, Carlee

GLDAS-2 Land Surface Model Data and Data Services at NASA GES DISC

The goal of the NASA Global Land Data Assimilation System (GLDAS, https://ldas.gsfc.nasa.gov/gldas(https://ldas.gsfc.nasa.gov/gldas)) is to generate optimal fields of land surface states and fluxes by ingesting satellite- and ground-based observational data products, using advanced land surface modeling and data assimilation techniques (Rodell et al., 2004).The GLDAS dataset currently archived at and distributed by the NASA Goddard Earth Sciences Data and Information Services Center (GES DISC, https://disc.gsfc.nasa.gov/ (https://disc.gsfc.nasa.gov/)) is GLDAS Version 2 (GLDAS-2). It contains a series of output fields from the upgraded Noah-3.6, Catchment-F2.5, and VIC-4.1.2 Land Surface Models (LSMs) in the Land Information System (LIS-V7, https://lis.gsfc.nasa.gov/ (https://lis.gsfc.nasa.gov/)). GLDAS-2 has three components:GLDAS-2.0, GLDAS-2.1, and GLDAS-2.2. GLDAS-2.0 is forced entirely with the upgraded Princeton Meteorological ForcingV2.2 Dataset and provides a temporally consistent series from 1948 through 2014. GLDAS-2.1 is forced with a combination of model and observation data, with data spanning from 2000 to the present. The GLDAS-2.2 product suite uses data assimilation(DA), whereas the GLDAS-2.0 and GLDAS-2.1 products are "open-loop" (i.e., no data assimilation). The choice of forcing data, as well as DA observation source, variable, and scheme, varies for different GLDAS-2.2 products. The currently availableGLDAS-2.2 data contain a daily 0.25-degree output from the Catchment-F2.5 LSM in LIS-V7. The data are forced with the meteorological analysis fields from the operational European Centre for Medium-Range Weather Forecasts Integrated Forecasting System (ECMWF-IFS) and assimilated with GRACE and GRACE-FO data, ranging from February 1, 2003 to the present. The current GLDAS-2.0 and 2.1 Noah LSM data were reprocessed in November 2019 and January 2020 respectively and their data from Catchment and VIC LSMs are new to the GLDAS-2 collection. This presentation provides a summary of theGLDAS-2 data products, their land surface fields, and their related data services at the GES DISC; and a description of the majorGLDAS-2 climatological characteristics as well as the intercomparison with the data of the previous version.

Hydrology

Recommendations for developing, documenting, and distributing data products derived from NEON data

The National Ecological Observatory Network (NEON) provides over 180 distinct data products from 81 sites (47 terrestrial and 34 freshwater aquatic sites) within the United States and Puerto Rico. These data products include both field and remote sensing data collected using standardized protocols and sampling schema, with centralized quality assurance and quality control (QA/QC) provided by NEON staff. Such breadth of data creates opportunities for the research community to extend basic and applied research while also extending the impact and reach of NEON data through the creation of derived data products—higher level data products derived by the user community from NEON data. Derived data products are curated, documented, reproducibly-generated datasets created by applying various processing steps to one or more lower level data products—including interpolation, extrapolation, integration, statistical analysis, modeling, or transformations. Derived data products directly benefit the research community and increase the impact of NEON data by broadening the size and diversity of the user base, decreasing the time and effort needed for working with NEON data, providing primary research foci through the development via the derivation process, and helping users address multidisciplinary questions. Creating derived data products also promotes personal career advancement to those involved through publications, citations, and future grant proposals. However, the creation of derived data products is a nontrivial task. Here we provide an overview of the process of creating derived data products while outlining the advantages, challenges, and major considerations.

54 ENVIRONMENTAL SCIENCES

Laying The Foundations for FAIR-ER Science: ISA And LSDA Data Submission Process in NASA’s Evolving Data Management Environment

The Life Sciences Data Archive (LSDA) archives data resulting from research on the effects of spaceflight on humans and the development of countermeasures to mitigate spaceflight hazards. Archivists work with researchers to ensure that unique and high value data products and their metadata are preserved and managed to support current and future research. Currently, LSDA is updating its procedures and data submission requirements in response to the evolving data preservation environment at NASA. LSDA is implementing best practices for research data management through the establishment of clear data submission guidelines, integration of the FAIR (Findability, Accessibility, Interoperability, Reusability) principles, and use of the ISA (Investigation, Study, Assay) research metadata framework for data discoverability and transparency into the data management processes. These changes directly impact LSDA’s requirements for research data submissions. The newly revised Research Data Submission Agreement (RDSA), formerly the Data Submission Agreement (DSA), introduces ISA-compatible metadata collection standards to LSDA’s process. Adherence to LSDA’s data submission guidelines enhances the FAIR-ness of the repository’s collections for future users. This presentation will discuss (1) how submission of research data and associated metadata are impacted by current data management policies, (2) benefits of the adoption of FAIR principles and the ISA metadata framework for retrospective studies utilizing existing LSDA datasets and historic data collections, and (3) the support LSDA will provide to researchers during this transition.

Data submission

Four-Dimensional Oceanic and Atmosperic Data Assimilation with Tropical Rainfall Measuring Mission Data

An oceanic data assimilation system which allows to utilize the forthcoming Tropical Rainfall Measuring Mission (TRMM) data has been developed and applied to the Pacific Ocean to produce the velocity field. The assimilated data will be indispensable to examine the effects of rainfall and its variability on the structure and circulation of the tropical oceans and to assess the impact of global warming due to the increase of carbon dioxide on the ocean circulation system and the marine pollution caused by oil spill and ocean damping of radionuclide. The data will also provide the verification for the oceanic and ocean-atmosphere coupled General Circulation Models (GCM's). The system consists of oceanic GCM, analysis scheme and data. In the system the flow field has been determined to be physically consistent with the observed density field and the sea surface winds derived from the Special Sensor Microwave Imagery (SSM/I) data which drive the ocean current. The time integration has been performed for five years until the flow field near the surface attained the steady state starting from the rest ocean with observed temperature and salinity fields, and the SSM/I surface wind velocity. The resultant flow field showed high producibility of the system. Especially the flow near the ocean surface agreed well with available observed data. The system, for the first time, succeeded to produce the eastward subtropical current which has been discovered in the joint investigation on Kuroshio current (CSK) in the 1960s. To verify the quality of the flow field a trajectory analysis has been carried out and compared with the Algos buoy data. BRIEF DESCRIPTION OF THE DATA ASSIMILATION SYSTEM ## Oceanic GCM and analysis scheme--The basic equations are much the same as used for the GCM's, except for the Newtonian damping terms introduced into the prediction equations for the potential temperature and salinity to maintain these fields as observed. The C grid of 2'lat. by 2'long. in horizontal and the 11 vertical levels are applied to the entire Pacific Ocean. At the east and west ocean boundaries the periodic boundary conditions are applied creating fictitious ocean there. The SMAC Method is used to increase the accuracy of mass conservation. * Data--The JODC temperature and salinity data obtained from 1906 to 1988 are used in the system between Long.100'E. and 60'W. The surface wind data are derived from the SSM/I data by Dr-R. Atlas of NASA/GSFC. The data set contains every 6 hours data from July 1987 to June 1989 on the grid of 2'lat. by 2.5'long. The averaged for the whole period and then interpolated into the 2'lat. by 2'long. grid data are used to force the system. The sea bottom topography data was based on the General Bathymetric Chart of the Ocean (GEBCO) supplied by the Canadian Hydrographic Service under contract with the International Hydrographic Organization and International Oceanographic Commission of UNESCO.

Takano, Kenji