Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Exploring Autoencoder-based Error-bounded Compression for Scientific Data

Error-bounded lossy compression is becoming an indispensable technique for the success of today's scientific projects with vast volumes of data produced during the simulations or instrument data acquisitions. Not only can it significantly reduce data size, but it also can control the compression errors based on user-specified error bounds. Autoencoder (AE) models have been widely used in image compression, but few AE-based compression approaches support error-bounding features, which are highly required by scientific applications. To address this issue, we explore using convolutional autoencoders to improve error-bounded lossy compression for scientific data, with the following three key contributions. (1) We provide an in-depth investigation of the characteristics of various autoencoder models and develop an error-bounded autoencoder-based framework in terms of the SZ model. (2) We optimize the compression quality for main stages in our designed AE-based error-bounded compression framework, fine-tuning the block sizes and latent sizes and also optimizing the compression efficiency of latent vectors. (3) We evaluate our proposed solution using five real-world scientific datasets and comparing them with six other related works. Experiments show that our solution exhibits a very competitive compression quality from among all the compressors in our tests. In absolute terms, it can obtain a much better compression quality (100%similar to 800% improvement in compression ratio with the same data distortion) compared with SZ2.1 and ZFP in cases with a high compression ratio.

Liu, Jinyang↗

Optimizing Error-Bounded Lossy Compression for Scientific Data With Diverse Constraints

Vast volumes of data are produced by today's scientific simulations and advanced instruments. These data cannot be stored and transferred efficiently because of limited I/O bandwidth, network speed, and storage capacity. Error-bounded lossy compression can be an effective method for addressing these issues: not only can it significantly reduce data size, but it can also control the data distortion based on user-defined error bounds. In practice, many scientific applications have specific requirements or constraints for lossy compression, in order to guarantee that the reconstructed data are valid for post hoc analysis. For example, some datasets contain irrelevant data that should be isolated in particular and users often have intuition regarding value ranges, geospatial regions, and other data subsets that are crucial for subsequent analysis. Existing state-of-the-art error-bounded lossy compressors, however, do not consider these constraints during compression, resulting in inferior compression ratios with respect to user's post hoc analysis, due to the fact that the data itself provides little or no value for post hoc analysis. In this work we address this issue by proposing an optimized framework that can preserve diverse constraints during the error-bounded lossy compression, e.g., cleaning the irrelevant data, efficiently preserving different precision for multiple value intervals, and allowing users to set diverse precision over both regular and irregular regions. We perform our evaluation on a supercomputer with up to 2,100 cores. Experiments with six real-world applications show that our proposed diverse constraints based error-bounded lossy compressor can obtain a higher visual quality or data fidelity on reconstructed data with the same or even higher compression ratios compared with the traditional state-of-the-art compressor SZ. Furthermore, our experiments also demonstrate very good scalability in compression performance compared with the I/O throughput of the parallel file system.

97 MATHEMATICS AND COMPUTING↗

User Scientific Data Systems: Experience Report

This paper presents an abbreviated history of NASA science data management system development over the past ten years by selecting two case studies, each representative of a distinct era of science data management systems.

Scientific Data↗

A General Framework for Error-controlled Unstructured Scientific Data Compression

Data compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the −6 to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets.

Gong, Qian↗

Scientific data from precipitation driver response model intercomparison project

This data descriptor reports the main scientific values from General Circulation Models (GCMs) in the Precipitation Driver and Response Model Intercomparison Project (PDRMIP). The purpose of the GCM simulations has been to enhance the scientific understanding of how changes in greenhouse gases, aerosols, and incoming solar radiation perturb the Earth’s radiation balance and its climate response in terms of changes in temperature and precipitation. Here we provide global and annual mean results for a large set of coupled atmospheric-ocean GCM simulations and a description of how to easily extract files from the dataset. The simulations consist of single idealized perturbations to the climate system and have been shown to achieve important insight in complex climate simulations. We therefore expect this data set to be valuable and highly used to understand simulations from complex GCMs and Earth System Models for various phases of the Coupled Model Intercomparison Project.

54 ENVIRONMENTAL SCIENCES↗

Scientific Data From Precipitation Driver Response Model Intercomparison Project

This data descriptor reports the main scientific values from General Circulation Models (GCMs) in the Precipitation Driver and Response Model Intercomparison Project (PDRMIP). The purpose of the GCM simulations has been to enhance the scientific understanding of how changes in greenhouse gases, aerosols, and incoming solar radiation perturb the Earth’s radiation balance and its climate response in terms of changes in temperature and precipitation. Here we provide global and annual mean results for a large set of coupled atmospheric-ocean GCM simulations and a description of how to easily extract files from the dataset. The simulations consist of single idealized perturbations to the climate system and have been shown to achieve important insight in complex climate simulations. We therefore expect this data set to be valuable and highly used to understand simulations from complex GCMs and Earth System Models for various phases of the Coupled Model Intercomparison Project

Precipitation Driver Response Model Intercompariso↗

A Language for Probabilistic Modeling of Scientific Data

We describe a portable container which allows specification of the probabilistic relations between several variables. The applications motivating this work are mainly scientific inference problems.

probabilistic modeling scientific data↗

SDMS: A scientific data management system

SDMS is a data base management system developed specifically to support scientific programming applications. It consists of a data definition program to define the forms of data bases, and FORTRAN-compatible subroutine calls to create and access data within them. Each SDMS data base contains one or more data sets. A data set has the form of a relation. Each column of a data set is defined to be either a key or data element. Key elements must be scalar. Data elements may also be vectors or matrices. The data elements in each row of the relation form an element set. SDMS permits direct storage and retrieval of an element set by specifying the corresponding key element values. To support the scientific environment, SDMS allows the dynamic creation of data bases via subroutine calls. It also allows intermediate or scratch data to be stored in temporary data bases which vanish at job end.

Massena, W. A.↗

In situ compression artifact removal in scientific data using deep transfer learning and experience replay

The massive amount of data produced during simulation on high-performance computers has grown exponentially over the past decade, exacerbating the need for streaming compression and decompression methods for efficient storage and transfer of this data---key to realizing the full potential of large-scale computational science. Lossy compression approaches such as JPEG when applied to scientific simulation data realized as a stream of images can achieve good compression rates but at the cost of introducing compression artifacts and loss of information. This paper develops a unified framework for in situ compression artifact removal in which the fully convolutional neural network architectures are combined with scalable training, transfer learning, and experience replay to achieve superior accuracy and efficiency while significantly decreasing the storage footprint as compared with the traditional optimization-based approaches. We demonstrate the proposed approach and compare it with compressed sensing postprocessing and other baseline deep learning models using climate simulations and nuclear reactor simulations, both of which are driven by hyperbolic partial differential equations. Our approach when applied to remove the compression artifacts on the JPEG-compressed nuclear reactor simulation data (using a transfer-trained model that was pretrained on the climate simulation data and updated incrementally as the nuclear reactor simulation progressed), achieved a significant improvement---mean peak signal-to-noise ratio of 42.438 as compared with 27.725 obtained with the compressed sensing approach.

97 MATHEMATICS AND COMPUTING↗

Researcher-driven Campaigns Engage Nature's Notebook Participants in Scientific Data Collection

One of the many benefits of citizen science projects is the capacity they hold for facilitating data collection on a grand scale and thereby enabling scientists to answer questions they would otherwise not been able to address. Nature's Notebook, the plant and animal phenology observing program of the USA National Phenology Network (USA-NPN) suitable for scientists and non-scientists alike, offers scientifically-vetted data collection protocols and infrastructure and mechanisms to quickly reach out to hundreds to thousands of potential contributors. The USA-NPN has recently partnered with several research teams to engage participants in contributing to specific studies. In one example, a team of scientists from NASA, the New Mexico Department of Health, and universities in Arizona, New Mexico, Oklahoma, and California are using juniper phenology observations submitted by Nature's Notebookparticipants to improve predictions of pollen release and inform asthma and allergy alerts. In a second effort, researchers from the University of Maryland Center for Environmental Science are engaging Nature's Notebookparticipants in tracking leafing phenophases of poplars across the U.S. These observations will be compared to information acquired via satellite imagery and used to determine geographic areas where the tree species are most and least adapted to predicted climate change. Researchers in these partnerships receive benefits primarily in the form of ground observations. Launched in 2010, the juniper pollen effort has engaged participants in several western states and has yielded thousands of observations that can play a role in model ground validation. Periodic evaluation of these observations has prompted the team to improve and enhance the materials that participants receive, in an effort to boost data quality. The poplar project is formally launching in spring of 2013 and will run for three years; preliminary findings from 2013 will be presented. Participants in these special campaigns benefit through direct engagement in science. This form of researcher partnership has now been successfully pilot-tested and implemented in several instances, and provides a template for future research project campaigns.

Crimmins, Theresa M.↗

Researcher-driven Campaigns Engage Nature's Notebook Participants in Scientific Data Collection

One of the many benefits of citizen science projects is the capacity they hold for facilitating data collection on a grand scale and thereby enabling scientists to answer questions they would otherwise not been able to address. Nature's Notebook, the plant and animal phenology observing program of the USA National Phenology Network (USA-NPN) suitable for scientists and non-scientists alike, offers scientifically-vetted data collection protocols and infrastructure and mechanisms to quickly reach out to hundreds to thousands of potential contributors. The USA-NPN has recently partnered with several research teams to engage participants in contributing to specific studies. In one example, a team of scientists from NASA, the New Mexico Department of Health, and universities in Arizona, New Mexico, Oklahoma, and California are using juniper phenology observations submitted by Nature's Notebookparticipants to improve predictions of pollen release and inform asthma and allergy alerts. In a second effort, researchers from the University of Maryland Center for Environmental Science are engaging Nature's Notebookparticipants in tracking leafing phenophases of poplars across the U.S. These observations will be compared to information acquired via satellite imagery and used to determine geographic areas where the tree species are most and least adapted to predicted climate change. Results/Conclusions Researchers in these partnerships receive benefits primarily in the form of ground observations. Launched in 2010, the juniper pollen effort has engaged participants in several western states and has yielded thousands of observations that can play a role in model ground validation. Periodic evaluation of these observations has prompted the team to improve and enhance the materials that participants receive, in an effort to boost data quality. The poplar project is formally launching in spring of 2013 and will run for three years; preliminary findings from 2013 will be presented. Participants in these special campaigns benefit through direct engagement in science. This form of researcher partnership has now been successfully pilot-tested and implemented in several instances, and provides a template for future research project campaigns.

Crimmins, Theresa M.↗

Scientific Data Collection/Analysis: 1994-2004

This custom bibliography from the NASA Scientific and Technical Information Program lists a sampling of records found in the NASA Aeronautics and Space Database. The scope of this topic includes technologies for lightweight, temperature-tolerant, radiation-hard sensors. This area of focus is one of the enabling technologies as defined by NASA s Report of the President s Commission on Implementation of United States Space Exploration Policy, published in June 2004.

Source record↗

Standardizing Algorithm Documentation For Improved Scientific Data Understanding: The Algorithm Publication Tool Prototype

Algorithm Theoretical Basis Documents (ATBDs) are documents which accompany Earth observation data products generated from algorithms. While ATBDs are essential to scientific reproducibility, these key documents are not standardized and are often difficult to find. In this paper, we present the prototype Algorithm Publication Tool (APT), a cloud-based ATBD authoring and editing tool for NASA’s Earth science data systems. A standardized ATBD information model is also described as well as lessons learned from developing the prototype tool.

Kaylin Bugbee↗

Scientific data reduction and analysis plan: PI services

This plan comprises two parts. The first concerns the real-time data display to be provided by MSC during the mission. The prime goal is to assess the operation of the UVS and to identify any problem areas that could be corrected during the mission. It is desirable to identify any possible observations of unusual scientific interest in order to repeat these observations at a later point in the mission, or to modify the time line with respect to the operating modes of the UVS. The second part of the plan discusses the more extensive postflight analysis of the data in terms of the scientific objectives of this experiment.

Feldman, P. D.↗