Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Landsat 7 Science Data Processing: An Overview

The Landsat 7 Science Data Processing System, developed by NASA for the Landsat 7 Project, provides the science data handling infrastructure used at the Earth Resources Observation Systems (EROS) Data Center (EDC) Landsat Data Handling Facility (DHF) of the United States Department of Interior, United States Geological Survey (USGS) located in Sioux Falls, South Dakota. This paper presents an overview of the Landsat 7 Science Data Processing System and details of the design, architecture, concept of operation, and management aspects of systems used in the processing of the Landsat 7 Science Data.

Schweiss, Robert J.↗

Deriving Earth Science Data Analytics Requirements

Data Analytics applications have made successful strides in the business world where co-analyzing extremely large sets of independent variables have proven profitable. Today, most data analytics tools and techniques, sometimes applicable to Earth science, have targeted the business industry. In fact, the literature is nearly absent of discussion about Earth science data analytics. Earth science data analytics (ESDA) is the process of examining large amounts of data from a variety of sources to uncover hidden patterns, unknown correlations, and other useful information. ESDA is most often applied to data preparation, data reduction, and data analysis. Co-analysis of increasing number and volume of Earth science data has become more prevalent ushered by the plethora of Earth science data sources generated by US programs, international programs, field experiments, ground stations, and citizen scientists.Through work associated with the Earth Science Information Partners (ESIP) Federation, ESDA types have been defined in terms of data analytics end goals. Goals of which are very different than those in business, requiring different tools and techniques. A sampling of use cases have been collected and analyzed in terms of data analytics end goal types, volume, specialized processing, and other attributes. The goal of collecting these use cases is to be able to better understand and specify requirements for data analytics tools and techniques yet to be implemented. This presentation will describe the attributes and preliminary findings of ESDA use cases, as well as provide early analysis of data analytics toolstechniques requirements that would support specific ESDA type goals. Representative existing data analytics toolstechniques relevant to ESDA will also be addressed.

data analytics↗

Mathematics: The Tao of Data Science

The two pieces, "Ten Research Challenge Areas in Data Science" by Jeannette M. Wing and “Challenges and Opportunities in Statistics and Data Science: Ten Research Areas” by Xuming He and Xihong Lin, provide an impressively complete list of data science challenges from luminaries in the field of data science. They have done an extraordinary job, so this response offers a complementary viewpoint from a mathematical perspective and evangelizes advanced mathematics as a key tool for meeting the challenges they have laid out. Notably, we pick up the themes of scientific understanding of machine learning and deep learning, computational considerations such as cloud computing and scalability, balancing computational and statistical considerations, and inference with limited data. We propose that mathematics is an important key to establishing rigor in the field of data science and as such has an essential role to play in its future.

97 MATHEMATICS AND COMPUTING↗

A Relevancy Algorithm for Curating Earth Science Data Around Phenomenon

Earth science data are being collected for various science needs and applications, processed using different algorithms at multiple resolutions and coverages, and then archived at different archiving centers for distribution and stewardship causing difficulty in data discovery. Curation, which typically occurs in museums, art galleries, and libraries, is traditionally defined as the process of collecting and organizing information around a common subject matter or a topic of interest. Curating data sets around topics or areas of interest addresses some of the data discovery needs in the field of Earth science, especially for unanticipated users of data. This paper describes a methodology to automate search and selection of data around specific phenomena. Different components of the methodology including the assumptions, the process, and the relevancy ranking algorithm are described. The paper makes two unique contributions to improving data search and discovery capabilities. First, the paper describes a novel methodology developed for automatically curating data around a topic using Earthscience metadata records. Second, the methodology has been implemented as a standalone web service that is utilized to augment search and usability of data in a variety of tools.

earth science phenomena↗

1235 Preparing for TEMPO: A Review of Planned Metadata, Data Structure, and Distribution by NASA’s Atmospheric Science Data Center

The Atmospheric Science Data Center (ASDC) is in the Science Directorate located at the NASA Langley Research Center (LaRC), in Hampton, Virginia. The ASDC is one of NASA’s Distributed Active Archive Centers (DAAC) and supports over 60 projects and provides access to more than 1,000 archived collections. These datasets were created from satellite measurements, field experiments, and modeled data products. ASDC projects focus on the following Earth science disciplines: Radiation Budget, Clouds, Aerosols, and Tropospheric Composition. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the upcoming Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument.. The instrument will share a ride on a commercial satellite as a hosted payload and will be launched to an orbit about 22,000 miles above Earth's equator. The investigation will, for the first time, use a space-based instrument to make accurate observations of tropospheric pollution concentrations of ozone, nitrogen dioxide, formaldehyde, and aerosols with high resolution and frequency over the U.S, Canada, and Mexico.

Ashlee Autore↗

1235 Preparing for TEMPO: A Review of Planned Metadata, Data Structure, and Distribution by NASA’s Atmospheric Science Data Center

The Atmospheric Science Data Center (ASDC) is in the Science Directorate located at the NASA Langley Research Center (LaRC), in Hampton, Virginia. The ASDC is one of NASA’s Distributed Active Archive Centers (DAAC) and supports over 60 projects and provides access to more than 1,000 archived collections. These datasets were created from satellite measurements, field experiments, and modeled data products. ASDC projects focus on the following Earth science disciplines: Radiation Budget, Clouds, Aerosols, and Tropospheric Composition. The ASDC is the official Distributed Active Archive Center (DAAC) of record for the upcoming Tropospheric Emissions: Monitoring of Pollution (TEMPO) instrument.. The instrument will share a ride on a commercial satellite as a hosted payload and will be launched to an orbit about 22,000 miles above Earth's equator. The investigation will, for the first time, use a space-based instrument to make accurate observations of tropospheric pollution concentrations of ozone, nitrogen dioxide, formaldehyde, and aerosols with high resolution and frequency over the U.S, Canada, and Mexico.

Ashlee Autore↗

Lunar laser ranging data deposited in the National Space Science Data Center normal points, filtered observations, and unfiltered photon detections

The lunar laser ranging project at McDonald Observatory provides the unique opportunity to acquire successfully precise range data for the earth-moon system. From the experiment's inception, the obligation was recognized to make these data available to the general scientific community in a reasonably useable form and in a realistic time frame. The documentation to be used in conjunction with the 1979 April deposit into the National Space Science Data Center which contains normal points, filtered observations and unfiltered photon stops for the months July through December, 1978 are reported.

Shelus, P. J.↗

MISR - Science Data Validation Plan

This Science Data Validation Plan describes the plans for validating a subset of the Multi-angle Imaging SpectroRadiometer (MISR) Level 2 algorithms and data products and supplying top-of-atmosphere (TOA) radiances to the In-flight Radiometric Calibration and Characterization (IFRCC) subsystem for vicarious calibration.

MISR science data validation Plan↗

Analyzing the Impact of Canadian Wildfires on Air Quality in the U.S. Mid-Atlantic: with Data and Tools from NASA’s Atmospheric Sciences Data Center

Wildfires pose a growing concern in North America due to their harmful impacts on air quality and public health, with increased wildfire activity in recent years leading to widespread smoke plumes that can transcend borders. The exposure of New York City (NYC), the most populous city in North America, to Canadian wildfire smoke highlights the substantial implications for public health and urban environments. To better understand the impact of Canadian wildfires on air quality in NYC, satellite data from the NASA Atmospheric Science Data Center (ASDC) at Langley Research Center, along with ground-based measurements and atmospheric modeling results, are analyzed. We examine concentrations of atmospheric aerosols—particularly PM2.5 particulate matter originating from Canadian wildfires—their dispersion patterns, and the duration and intensity of smoke events impacting NYC. Data from multiple satellites, such as those from the Earth Polychromatic Imaging Camera (EPIC), are synergistically used to identify regions affected by wildfires and estimate aerosol loading. Ground-based measurements, including data from air quality monitoring stations, provide localized information for validation and calibration purposes. The findings of this study contribute to our understanding of the impact of Canadian wildfires on NYC's air quality and emphasize the importance of monitoring and prediction of transboundary smoke events using data synthesized from multiple sources, such as those provided by the ASDC. This information is crucial for policymakers, public health officials, and residents in affected areas to develop effective strategies for mitigating the health risks associated with wildfire smoke and improving air quality during wildfire seasons. The utilization of ASDC data in this research highlights the critical role of atmospheric remote sensing in addressing the challenges posed by wildfires and their consequences on regional scales.

Ingrid Garcia-Solera↗

NASA's astrophysics archives at the National Space Science Data Center

NASA maintains an archive facility for Astronomical Science data collected from NASA's missions at the National Space Science Data Center (NSSDC) at Goddard Space Flight Center. This archive was created to insure the science data collected by NASA would be preserved and useable in the future by the science community. Through 25 years of operation there are many lessons learned, from data collection procedures, archive preservation methods, and distribution to the community. This document presents some of these more important lessons, for example: KISS (Keep It Simple, Stupid) in system development. Also addressed are some of the myths of archiving, such as 'scientists always know everything about everything', or 'it cannot possibly be that hard, after all simple data tech's do it'. There are indeed good reasons that a proper archive capability is needed by the astronomical community, the important question is how to use the existing expertise as well as the new innovative ideas to do the best job archiving this valuable science data.

Vansteenberg, M. E.↗

Data Science in Chemical Engineering: Applications to Molecular Science

Chemical engineering is being rapidly transformed by the tools of data science. On the horizon, artificial intelligence (AI) applications will impact a huge swath of our work, ranging from the discovery and design of new molecules to operations and manufacturing and many areas in between. Early adoption of data science, machine learning, and early examples of AI in chemical engineering has been rich with examples of molecular data science—the application tools for molecular discovery and property optimization at the atomic scale. Here, we summarize key advances in this nascent subfield while introducing molecular data science for a broad chemical engineering readership. We introduce the field through the concept of a molecular data science life cycle and discuss relevant aspects of five distinct phases of this process: creation of curated data sets, molecular representations, data-driven property prediction, generation of new molecules, and feasibility and synthesizability considerations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Combined Industry, Space and Earth Science Data Compression Workshop

The sixth annual Space and Earth Science Data Compression Workshop and the third annual Data Compression Industry Workshop were held as a single combined workshop. The workshop was held April 4, 1996 in Snowbird, Utah in conjunction with the 1996 IEEE Data Compression Conference, which was held at the same location March 31 - April 3, 1996. The Space and Earth Science Data Compression sessions seek to explore opportunities for data compression to enhance the collection, analysis, and retrieval of space and earth science data. Of particular interest is data compression research that is integrated into, or has the potential to be integrated into, a particular space or earth science data information system. Preference is given to data compression research that takes into account the scien- tist's data requirements, and the constraints imposed by the data collection, transmission, distribution and archival systems.

Kiely, Aaron B.↗

Framework for Processing Citizens Science Data for Applications to NASA Earth Science Missions

Citizen science (or crowdsourcing) has drawn much high-level recent and ongoing interest and support. It is poised to be applied, beyond the by-now fairly familiar use of, e.g., Twitter for natural hazards monitoring, to science research, such as augmenting the validation of NASA earth science mission data. This interest and support is seen in the 2014 National Plan for Civil Earth Observations, the 2015 White House forum on citizen science and crowdsourcing, the ongoing Senate Bill 2013 (Crowdsourcing and Citizen Science Act of 2015), the recent (August 2016) Open Geospatial Consortium (OGC) call for public participation in its newly-established Citizen Science Domain Working Group, and NASA's initiation of a new Citizen Science for Earth Systems Program (along with its first citizen science-focused solicitation for proposals). Over the past several years, we have been exploring the feasibility of extracting from the Twitter data stream useful information for application to NASA precipitation research, with both "passive" and "active" participation by the twitterers. The Twitter database, which recently passed its tenth anniversary, is potentially a rich source of real-time and historical global information for science applications. The time-varying set of "precipitation" tweets can be thought of as an organic network of rain gauges, potentially providing a widespread view of precipitation occurrence. The validation of satellite precipitation estimates is challenging, because many regions lack data or access to data, especially outside of the U.S. and in remote and developing areas. Mining the Twitter stream could augment these validation programs and, potentially, help tune existing algorithms. Our ongoing work, though exploratory, has resulted in key components for processing and managing tweets, including the capabilities to filter the Twitter stream in real time, to extract location information, to filter for exact phrases, and to plot tweet distributions. The key step is to process the "precipitation" tweets to be compatible with satellite-retrieved precipitation data. These key components for processing and managing "precipitation" tweets (and additional ones to be developed) are not limited to precipitation, nor are they limited to the Twitter social medium. Indeed, to maximize the value of our work for NASA earth science programs, these components should be generalized and be part of an overall framework for processing citizen science data for science research. In this paper, we outline such a framework.

earth science satellite data↗

The European HST Science Data Archive

The paper describes the European HST Science Data Archive. Particular attention is given to the flow from the HST spacecraft to the Science Data Archive at the Space Telescope European Coordinating Facility (ST-ECF); the archiving system at the ST-ECF, including the hardware and software system structure; the operations at the ST-ECF and differences with the Data Management Facility; and the current developments. A diagram of the logical structure and data flow of the system managing the European HST Science Data Archive is included.

Pasian, F.↗

NASA ESDS Citizen Science Data Working Group

This document provides guidelines for legal, policy, and ethical issues; standards for citizen science data collection and management; information on ensuring usability of citizen science data and communication regarding its use; and best practices for long-term archival of citizen science data.Section1 contains a detailed discussion of policy, ethical, and legal considerations influencing citizen science data collection. Section 2 considers standards for documentation, including documentation of instrumentation, procedures, and the data itself. It concludes with a discussion of how citizen science data should be attributed. Section 3 provides guidance about how to ensure citizen science data are collected and stored in a useable way. It also considers how NASA and data producers should notify the scientific community, including citizen scientists and the public, about citizen science datasets and the scientific conclusions reached using them. Finally, Section 4 provides detailed information regarding what should be archived from projects using a citizen science approach, including data and code. It provides guidance about archive location, process, and timeframe, as well as information about data access and distribution services provided by NASA that may be relevant to data producers working with citizen scientists.

Citizen Science↗

Landsat 7 Science Data Processing: A System's Overview

The Landsat Science Data Processing System, developed by NASA for the Landsat 7 Project provides science data handling infrastructure used at the EROS Data Center Landsat 7 Data Handling Facility of the USGS Department of Interior. This paper presents an overview the designs, architectures, and details of the various systems used in the processing of the Landsat 7 Science Data.

Schweiss, Robert↗