Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “data and data science”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

DataHub - Science data management in support of interactive exploratory analysis

DataHub addresses four areas of significant need: scientific visualization and analysis; science data management; interactions in a distributed, heterogeneous environment; and knowledge-based assistance for these functions. The fundamental innovation embedded within the DataHub is the integration of three technologies, viz. knowledge-based expert systems, science visualization, and science data management. This integration is based on a concept called the DataHub. With the DataHub concept, science investigators are able to apply a more complete solution to all nodes of a distributed system. Both computational nodes and interactive nodes are able to effectively and efficiently use the data services (access, retrieval, update, etc.) in a distributed, interdisciplinary information system in a uniform and standard way. This allows the science investigators to concentrate on their scientific endeavors, rather than to involve themselves in the intricate technical details of the systems and tools required to accomplish their work. Thus, science investigators need not be programmers. The emphasis is on the definition and prototyping of system elements with sufficient detail to enable data analysis and interpretation leading to information. The DataHub includes all the required end-to-end components and interfaces to demonstrate the complete concept.

Handley, Thomas H., Jr.↗

Space and Earth Science Data Compression Workshop

The workshop explored opportunities for data compression to enhance the collection and analysis of space and Earth science data. The focus was on scientists' data requirements, as well as constraints imposed by the data collection, transmission, distribution, and archival systems. The workshop consisted of several invited papers; two described information systems for space and Earth science data, four depicted analysis scenarios for extracting information of scientific interest from data collected by Earth orbiting and deep space platforms, and a final one was a general tutorial on image data compression.

Tilton, James C.↗

HypsIRI On-Board Science Data Processing

Topics include On-board science data processing, on-board image processing, software upset mitigation, on-board data reduction, on-board 'VSWIR" products, HyspIRI demonstration testbed, and processor comparison.

Flatley, Tom↗

A Concept of Operations for Earth Science Data Archive and Distribution in the Cloud

Science data systems can enable more comprehensive Earth system research by evolving to take advantage of advances in commercial computer technology services. Since their inception twenty five years ago, NASA's Earth Observing System Data and Information System (EOSDIS) Distributed Active Archive Centers (DAACs) have periodically evolved to utilize new technology and expand research using the exponential growth and diversity of Earth observations. Recently, with the advent of a maturing commercial compute services industry and upcoming high data volume missions such as the Surface Water and Ocean Topography (SWOT) mission and the NASA-Indian Space Research Organization Synthetic Aperture Radar (NISAR) mission, options were explored and a decision made to utilize commercial compute and storage services. This paper presents an overview of the concept of operations under development for the DAACs in the Cloud. We highlight the goals and expected advantages of utilizing Cloud services. We outline EOSDIS operations tenets and driving principles. A high-level view of EOSDIS system of systems target architecture serves as context for describing principle interactions. Concepts for key DAAC system and EOSDIS enterprise functions characterize automated end-to-end operations but mark nominal check and recovery points. Concepts are presented for managing Cloud resources, including organizational roles and responsibilities of the NASA project and DAAC personnel. Scenarios we use to further distinguish between what the system will do and what configuration and controls operators will have. Examples include interactions with data providers and data consumers with both in-cloud and on-premise facilities.

Moses, John F.↗

A distributed component framework for science data product interoperability

Correlation of science results from multi-disciplinary communities is a difficult task. Traditionally data from science missions is archived in proprietary data systems that are not interoperable. The Object Oriented Data Technology (OODT) task at the Jet Propulsion Laboratory is working on building a distributed product server as part of a distributed component framework to allow heterogeneous data systems to communicate and share scientific results.

XML distributed framework science interprobability↗

Science Data Preservation: Implementation and Why It Is Important

Remote Sensing data generation by NASA to study Earth s geophysical processes was initiated in 1960 with the launch of the first Television Infrared Observation Satellite Program (TIROS), to develop a meteorological satellite information system. What would be deemed as a primitive data set by today s standards, early Earth science missions were the foundation upon which today s remote sensing instruments have built their scientific success, and tomorrow s instruments will yield science not yet imagined. NASA Scientific Data Stewardship requirements have been documented to ensure the long term preservation and usability of remote sensing science data. In recent years, the Federation of Earth Science Information Partners and NASA s Earth Science Data System Working Groups have organized committees that specifically examine standards, processes, and ontologies that can best be employed for the preservation of remote sensing data, supporting documentation, and data provenance information. This presentation describes the activities, issues, and implementations, guided by the NASA Earth Science Data Preservation Content Specification (423-SPEC-001), for preserving instrument characteristics, and data processing and science information generated for 20 Earth science instruments, spanning 40 years of geophysical measurements, at the NASA s Goddard Earth Sciences Data and Information Services Center (GES DISC). In addition, unanticipated preservation/implementation questions and issues in the implementation process are presented.

Kempler, Steven J.↗

Photometer Performance Assessment in Kepler Science Data Processing

This paper describes the algorithms of the Photometer Performance Assessment (PPA) software component in the science data processing pipeline of the Kepler mission. The PPA performs two tasks: One is to analyze the health and performance of the Kepler photometer based on the long cadence science data down-linked via Ka band approximately every 30 days. The second is to determine the attitude of the Kepler spacecraft with high precision at each long cadence. The PPA component is demonstrated to work effectively with the Kepler flight data.

Li, Jie↗

DataHub: Science data management in support of interactive exploratory analysis

The DataHub addresses four areas of significant needs: scientific visualization and analysis; science data management; interactions in a distributed, heterogeneous environment; and knowledge-based assistance for these functions. The fundamental innovation embedded within the DataHub is the integration of three technologies, viz. knowledge-based expert systems, science visualization, and science data management. This integration is based on a concept called the DataHub. With the DataHub concept, science investigators are able to apply a more complete solution to all nodes of a distributed system. Both computational nodes and interactives nodes are able to effectively and efficiently use the data services (access, retrieval, update, etc), in a distributed, interdisciplinary information system in a uniform and standard way. This allows the science investigators to concentrate on their scientific endeavors, rather than to involve themselves in the intricate technical details of the systems and tools required to accomplish their work. Thus, science investigators need not be programmers. The emphasis on the definition and prototyping of system elements with sufficient detail to enable data analysis and interpretation leading to information. The DataHub includes all the required end-to-end components and interfaces to demonstrate the complete concept.

Handley, Thomas H., Jr.↗

Open Science for Plants in Space: Improvements in NASA's Open Science Data Repository

Upcoming deep space missions will rely on plants for crew and ecosystem health. Open access space biology data enables scientists to examine the biological responses of plants to ionizing radiation, altered gravity, elevated CO2, and many other abiotic stressors. NASA has declared 2023 as the ‘Year of Open Science’ and created a 5-year Transform to Open Science (TOPS) initiative designed to rapidly transform the agency toward an inclusive culture of open science. NASA’s Open Science Data Repository (OSDR) combines two databases, GeneLab and Ames Life Sciences Data Archive (ALSDA) to maximize access to standardized ‘omics (e.g., transcriptomics, proteomics) and phenotypic data (e.g., microscopy, biomass), respectively. Current OSDR standards include the ISA (Investigation-Study-Assay) experiment model, assay metadata configurations, and standardized terminology and ontologies. In 2024 OSDR will include a new suite of features for improved FAIR compliance including downloadable plant metadata templates, data submission tools and overall improved AI-readiness of plant datasets. AI/ML methods can be helpful tools to overcome the inherent challenges of space biology research (small sample size, sparse and heterogeneous data etc.). However these methods are built on an assumption of normalized and well-curated data. OSDR’s new curation tools will improve users ability to leverage ML and AI methods to model space biology data and better understand the complex effects of spaceflight on living systems across hierarchical biological levels. We look forward to sharing our advances with the spaceflight community.

FAIR↗

EOS MLS Science Data Processing System: A Description of Architecture and Capabilities

This paper describes the architecture and capabilities of the Science Data Processing System (SDPS) for the EOS MLS. The SDPS consists of two major components--the Science Computing Facility and the Science Investigator-led Processing System. The Science Computing Facility provides the facilities for the EOS MLS Science Team to perform the functions of scientific algorithm development, processing software development, quality control of data products, and scientific analyses. The Science Investigator-led Processing System processes and reprocesses the science data for the entire mission and delivers the data products to the Science Computing Facility and to the Goddard Space Flight Center Earth Science Distributed Active Archive Center, which archives and distributes the standard science products.

Microwave Limb Sounder (MLS)↗

Data Science Shows that Entropy Correlates with Accelerated Zeolite Crystallization in Monte Carlo Simulations

We have performed a data science study of Monte Carlo simulation trajectories to understand factors that can accelerate formation of zeolite nanoporous crystals, a process that can take days or even weeks. In previous work, Monte Carlo simulations predicted and experiments confirmed that using a secondary organic structure-directing agent (OSDA) accelerates crystallization of all-silica LTA zeolite, with experiments finding a three-fold speedup [PCCP 24, 142-148 (2022)]. However, it remains unclear what physical factors cause the speed-up. Here, we apply data science to analyze the simulation trajectories to discover what drives accelerated zeolite crystallization in Monte Carlo going from a one-OSDA synthesis (1OSDA) to a two-OSDA version (2OSDA). We encoded simulation snapshots using the Smooth Overlap of Atomic Positions approach, which represents all 2- and 3-body correlations within a given cutoff distance. Principal component analyses failed to discriminate datasets of structures from 1OSDA and 2OSDA simulations, while the Support Vector Machine (SVM) approach succeeded at classifying such structures with an area-under-curve (AUC) score of 0.99 (where AUC = 1 is a perfect classification) with all 3-body correlations, and as high as 0.94 with only 2-body correlations. SVM decision functions reveal relatively broad / narrow histograms for 1OSDA / 2OSDA datasets, suggesting that the two simulations differ strongly in information heterogeneity. Informed by these results, we performed pair (2-body) entropy calculations during crystallization, resulting in entropy differences that semi-quantitatively account for the speedup observed in the previous Monte Carlo simulations. We conclude that altering synthesis conditions in ways that substantially changes the entropy of labile silica networks may accelerate zeolite crystallization, and we discuss possible approaches for achieving such acceleration.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

A Science Data System Approach for the SMAP Mission

Though Science Data System (SDS) development has not traditionally been part of the mission concept phase, lessons learned and study of past Earth science missions indicate that SDS functionality can greatly benefit algorithm developers in all mission phases. We have proposed a SDS approach for the SMAP Mission that incorporates early support for an algorithm testbed, allowing scientists to develop codes and seamlessly integrate them into the operational SDS. This approach will greatly reduce both the costs and risks involved in algorithm transitioning and SDS development.

PCS↗

James Webb Space Telescope - L2 Communications for Science Data Processing

JWST is the first NASA mission at the second Lagrange point (L2) to identify the need for data rates higher than 10 megabits per second (Mbps). JWST will produce approximately 235 Gigabits of science data every day that will be downlinked to the Deep Space Network (DSN). To get the data rates desired required moving away from X-band frequencies to Ka-band frequencies. To accomplish this transition, the DSN is upgrading its infrastructure. This new range of frequencies are becoming the new standard for high data rate science missions at L2. With the new frequency range, the issues of alternatives antenna deployment, off nominal scenarios, NASA implementation of the Ka-band 26 GHz, and navigation requirements will be discussed in this paper. JWST is also using Consultative Committee for Space Data Systems (CCSDS) standard process for reliable file transfer using CCSDS File Delivery Protocol (CFDP). For JWST the use of the CFDP protocol provides level zero processing at the DSN site. This paper will address NASA implementations of Ground Stations in support of Ka-band 26 GHz and lesson learned from implementing a file base (CFDP) protocol operational system.

Johns, Alan↗

NASA Earth Sciences Data Support System and Services for the Northern Eurasia Earth Science Partnership Initiative

The presentation describes data management of NASA remote sensing data for Northern Eurasia Earth Science Partnership Initiative (NEESPI). Many types of ground and integrative (e.g., satellite, GIs) data will be needed and many models must be applied, adapted or developed for properly understanding the functioning of Northern Eurasia cold and diverse regional system. Mechanisms for obtaining the requisite data sets and models and sharing them among the participating scientists are essential. The proposed project targets integration of remote sensing data from AVHRR, MODIS, and other NASA instruments on board US- satellites (with potential expansion to data from non-US satellites), customized data products from climatology data sets (e.g., ISCCP, ISLSCP) and model data (e.g., NCEPNCAR) into a single, well-architected data management system. It will utilize two existing components developed by the Goddard Earth Sciences Data & Information Services Center (GES DISC) at the NASA Goddard Space Flight Center: (1) online archiving and distribution system, that allows collection, processing and ingest of data from various sources into the online archive, and (2) user-friendly intelligent web-based online visualization and analysis system, also known as Giovanni. The former includes various kinds of data preparation for seamless interoperability between measurements by different instruments. The latter provides convenient access to various geophysical parameters measured in the Northern Eurasia region without any need to learn complicated remote sensing data formats, or retrieve and process large volumes of NASA data. Initial implementation of this data management system will concentrate on atmospheric data and surface data aggregated to coarse resolution to support collaborative environment and climate change studies and modeling, while at later stages, data from NASA and non-NASA satellites at higher resolution will be integrated into the system.

Leptoukh, Gregory↗