Engineering PapersSearch

Engineering topics

Sean Leavor

Publications and source records attributed to Sean Leavor.

Spatially-Coordinated Airborne Data and Complementary Products for Aerosol, Gas, Cloud, and Meteorological Studies: the Nasa Activate Dataset

The NASA Aerosol Cloud meTeorology Interactions oVer the western ATlantic Experiment (ACTIVATE) produced a unique dataset for research into aerosol–cloud–meteorology interactions, with applications extending from process-based studies to multi-scale model intercomparison and improvement as well as to remote-sensing algorithm assessments and advancements. ACTIVATE used two NASA Langley Research Center aircraft, a HU-25 Falcon and King Air, to conduct systematic and spatially coordinated flights over the northwest Atlantic Ocean, resulting in 162 joint flights and 17 other single-aircraft flights between 2020 and 2022 across all seasons. Data cover 574 and 592 cumulative flights hours for the HU-25 Falcon and King Air, respectively. The HU-25 Falcon conducted profiling at different level legs below, in, and just above boundary layer clouds (< 3 km) and obtained in situ measurements of trace gases, aerosol particles, clouds, and atmospheric state parameters. Under cloud-free conditions, the HU-25 Falcon similarly conducted profiling at different level legs within and immediately above the boundary layer. The King Air (the high-flying aircraft) flew at approximately ∼ 9 km and conducted remote sensing with a lidar and polarimeter while also launching dropsondes (785 in total). Collectively, simultaneous data from both aircraft help to characterize the same vertical column of the atmosphere. In addition to individual instrument files, data from the HU-25 Falcon aircraft are combined into “merge files” on the publicly available data archive that are created at different time resolutions of interest (e.g., 1, 5, 10, 15, 30, 60 s, or matching an individual data product's start and stop times). This paper describes the ACTIVATE flight strategy, instrument and complementary dataset products, data access and usage details, and data application notes. The data are publicly accessible through https://doi.org/10.5067/SUBORBITAL/ACTIVATE/DATA001 (ACTIVATE Science Team, 2020).

Aerosol Cloud meTeorology Interactions oVer the we

ICARTT File Format Enhancements: Supporting FAIRness of Airborne and Field Campaign Data

The ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign in 2004. The ICARTT file format is text-based and composed of a header with important data description information and the data section. The ICARTT format, built on the NASA Ames and GTE data formats, was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to the success of the ICARTT campaign, the ICARTT file format was exposed to a broad range of airborne researchers and was adopted for use in many other field campaigns sponsored by NASA and other partner agencies. The ICARTT format standards became a NASA standard in 2010 and was amended in January 2017 providing many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. The ICARTT format can host metadata that is critical for proper use of the data, especially for in-situ measurements. However, the information that needs to be included is often in free text, meaning the information are human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary substantially between principal investigators and campaigns. To support interoperability and FAIR principles, further enhancements to the ICARTT standards are recommended. Possible recommendations include standardizing timestamps for easier data comparisons and analysis; potential use of controlled and consistent vocabulary for variable short name and certain common metadata elements; and providing guidance on variable measurement units and how they are reported.

Megan Buzanowicz

Improving “Domain-Relevant Metadata Requirements” for Supporting Open-Source Science Initiative

Implementation of the NASA Open-Source Science Initiative (OSSI) requires sharing of all relevant information to ensure “open reproducible science” [1]. However, there are several challenges in applying the OSSI to airborne field campaigns focused on atmospheric composition, which often involve a wide variety of in-situ measurements for trace gases, aerosol and cloud properties, meteorological parameters, and radiation fields. To ensure open reproducibility from airborne field campaigns, it is essential to obtain detailed measurement descriptions, which include the detection principle, sample procedure and treatment, and data processing and correction method. The challenge is that some information, e.g., sampling procedure and treatment, may be instrument-specific and campaign or platform-dependent. The data processing may also involve empirical corrections which may evolve over time. In addition, these details (especially operation- or campaign-specific ones) are often not given in journal publications. Given these issues, there is a need to leverage and improve the current “domain-relevant metadata requirements” to represent the measurement description in standardized metadata. These requirements can then facilitate systematic collection of measurement specific metadata and serve as a foundation to develop tools for making the information accessible and data more interoperable and usable or reusable. Here we show a review of existing metadata collections, use cases, and needs for new standards.

Sean Leavor

ICARTT File Format Enhancements: Supporting FAIRness and Data Discovery of Suborbital Campaign Data

Suborbital campaigns aim to accomplish a wide variety of goals and can include a variety of platforms, instruments, and parameters measured. In 2004, the ICARTT (International Consortium for Atmospheric Research on Transport and Transformation) standards were developed to fulfill data management needs for the ICARTT campaign. The ICARTT file format is text-based and composed of a header with important data description information and the data section. Built on the NASA Ames and GTE data formats, the ICARTT format was created to facilitate data exchange and promote collaborations among the science teams for achieving the ICARTT campaign goals. Due to its success and adaptation for use in many other field campaigns, the ICARTT file format became a NASA standard in 2010 and was amended in January 2017. These changes provided many enhancements, including the requirement for variable standard names. Primarily designed for airborne field studies, ICARTT has been further utilized for ground-based studies. NASA has made a commitment to build an inclusive open science community over the next decade. Open-source science strives to make publicly funded scientific research transparent, inclusive, accessible, and reproducible. The ICARTT format can host metadata that is critical for proper use of the data, particularly for in-situ measurements, and can enhance data discovery and accessibility. However, the required fields are often free text, meaning that the information is human readable, but not machine interpretable. Furthermore, the amount and type of information provided can vary significantly between principal investigators and campaigns. To support FAIR principles and interoperability, enhancements to the ICARTT standards are recommended. Possible recommendations include potential use of controlled and consistent vocabulary for variable standard name and certain common metadata elements; standardizing timestamps for easier data comparisons and analysis; and providing guidance on variable measurement units and how they are reported. Enhancing ICARTT metadata can further streamline the process to make suborbital data more readily available to the data user and improve variable-level metadata. Providing more variable-level metadata can enhance data searching and discovery, supporting NASA’s Open-Source Science Initiative (OSSI).

Megan Buzanowicz

Additional Metadata Guidelines to Improve the Structure and Usability of HDF and NetCDF Files

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global attributes or variables and are popular with the modeling community. HDF and NetCDF files are largely used with remote sensing data and have been used to support measurements from numerous campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. This metadata might be present in another file in the dataset containing the same data using the International Consortium for Atmospheric Research on Transport and Transformation (ICARTT) format. Recently, the Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a set of requirements for the HDF and NetCDF files for field studies, aiming to make the data products more interoperable and usable. Particularly, these requirements seek to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. These requirements would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. To be presented are the details of the MACIE requirements as well as examples of the implementation of these requirements for merge files and lidar observation data files.

Sean Leavor

Blast from the Past: ASDC Curation for NASA Suborbital Legacy Missions to Promote Data Discovery and Accessibility

NASA has an extensive history of conducting suborbital field campaigns to further advances in atmospheric sciences. Beginning with the Chemical Instrument Test and Evaluation (CITE) conducted in 1983-1984, NASA has completed many suborbital campaigns over the past three decades. Since the early 2010s, suborbital missions are typically assigned to a NASA Distributed Active Archive Center (DAAC) prior to the mission for long-term archival and distribution. Efforts are being made by NASA’s Earth Science Data and Information System (ESDIS) Project and the Airborne Data Management Group (ADMG) to assign legacy missions to DAACs for permanent archival and distribution, so that these valuable datasets remain to be available to the scientific community. NASA’s Atmospheric Science Data Center (ASDC) has been named the assigned DAAC for nearly 20 atmospheric composition legacy missions, including missions conducted as part of the Global Tropospheric Experiment (GTE) and expects to be named the assigned DAAC for more of these missions over the next few years. The primary goal of the ASDC is to provide access to the datasets as they are currently formatted to the broad user community and enhance their findability and accessibility. However, data reporting standards have evolved significantly since 1983 and the datasets span a wide variety of file formats, including text, Ames, GTE, and ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), and the amount of metadata and relevant information included in the files also varies greatly and can not be readily extracted without subject matter knowledge. This has caused challenges for the ASDC’s suborbital metadata extraction pipeline in ensuring that accurate and necessary metadata is being provided for the missions by all the ASDC’s existing search mechanisms. To make the data more findable and accessible, the ASDC has begun researching ways to further enhance the datasets, including distributing value-added products (i.e. consistent file format such as ICARTT or netCDF), adding standard names from the ESDIS Standards Coordination Office (ESCO)-approved Atmospheric Composition Variable Standard Names Convention (ACVSNC), and creating outreach materials such as ArcGIS StoryMaps, User Guides, and Micro Articles, providing overviews of the missions and what type of data was collected during the missions. These efforts also help support NASA’s Open-Source Science by enhancing the FAIRness of the legacy data products. This presentation will review the ASDC’s ongoing efforts, progress made, and future plans for legacy missions.

Megan Buzanowicz

Extending CF Conventions to Enhance Data FAIRness for Atmospheric Composition Observations

The Hierarchical Data Format (HDF) and Network Common Data Form (NetCDF) are data file formats created to aid users in the creation or use of scientific data. These file formats are useful for handling large data volumes and hosting extensive metadata as global, group, or variable attributes and are popular with the modeling community. HDF and NetCDF files are widely used with atmospheric remote sensing data and have been used to support measurements from numerous field campaigns, from satellite to aircraft or ground and mobile based measurements. The files from airborne field studies, however, vary greatly in terms of the file structure and the amount and content of their metadata. Information relevant to the file that can be useful to the user such as the data producer, location where data was taken, variable descriptions, or information about the instrument might not be included in the file. Recently, the Measurements of Aerosols, Clouds, and their Interactions for Earth System Models (MACIE) group started a grassroots effort to develop a CF-based template for the HDF and NetCDF files for field studies, with the aim of making the data products more interoperable and usable. This template seeks to make the files more compliant to Climate and Forecast (CF) metadata conventions and to standardize the file structure and the global and variable attributes. The template would help to ensure that HDF and NetCDF files contain adequate metadata to better support their use for research, e.g., the modeling community, and to enhance the usability and interoperability of data for research communities at large. The draft template has been applied to recent field studies for various instruments and their merge files in support of the Atmosphere Observing System (AOS) project. The details of the revised template are to be presented, as well as examples of the implementation of these requirements for merge files and lidar observation data files and issues revealed during the implementation process.

Sean Leavor

ASDC’s Python-Based Metadata Extraction Pipeline for Suborbital Campaigns

The FAIRness of data products, especially findability and accessibility depend on rich metadata which, when extracted, can allow for proper curation. Over the past few years, the Atmospheric Science Data Center (ASDC) suborbital science support team has developed a metadata extraction pipeline to ensure the required metadata can be retrieved systematically, effectively, and efficiently to ensure the data can be used by a broad community. The development of a pipeline has presented many, but necessary, challenges to support archival and distribution of ASDC’s 30+ suborbital missions. Though sufficient metadata is provided by instrument scientists, the metadata may not be readily machine actionable due to different formats and templates. Further complicating metadata extraction, our team has found that the nature of metadata can be quite diverse given the difference in measurement types, instruments, and measurement platforms. A metadata extraction pipeline has been developed to provide an efficient, plugin-in based, method for adding new parsers, a configuration system that lets non-developers customize how files are processed, and a system for identifying and logging metadata quality issues to ensure they are readily found and addressed. The metadata extraction pipeline identifies critical pieces of metadata that are needed to promote data FAIRness, including location, file revision, measurement start/end datetime and can be easily modified to extract further information (such as variables). Given the wide-ranging datasets, the pipeline has been modified to accommodate multiple file formats, including multiple versions of ICARTT (International Consortium for Atmospheric Research on Transport and Transformation), HDF (Hierarchical Data Format), netCDF (network Common Data Form), and multiple versions of the Ames File Format. The pipeline also supports building metadata for file formats that cannot have metadata easily extracted from them, such as PDF (Portable Document Format) and GIF (Graphics Interchange Format). The pipeline has allowed our team to maintain a consistent flow of data and metadata to archival and distribution services, ensuring the ASDC meets the needs of the suborbital science community. This presentation will highlight the ASDC’s suborbital metadata extraction pipeline, its development, how it’s been modified to support data FAIRness, and plans for maintaining the pipeline and adding new features.

Abraham Porter