Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization

In scientific visualization, despite the significant advances of deep learning for data generation, researchers have not thoroughly investigated the issue of data translation. We present a new deep learning approach called generalized multivariate translation (GMT) for multivariate time-varying data analysis and visualization. Like V2V, GMT assumes a preprocessing step that selects suitable variables for translation. However, unlike V2V, which only handles one-to-one variable translation during training and inference, GMT enables one-to-many and many-to-many variable translation in the same framework. We leverage the recent StarGAN design from multi-domain image-to-image translation to achieve this generalization capability. We experiment with different loss functions and injection strategies to explore the best choices and leverage pre-training for performance improvement. We compare GMT with other state-of-the-art methods (i.e., Pix2Pix, V2V, StarGAN). Furthermore, the results demonstrate the overall advantage of GMT in translation quality and generalization ability.

97 MATHEMATICS AND COMPUTING↗

Top Research Challenges and Opportunities for Near Real-Time Extreme-Scale Visualization of Scientific Data

The rapid advancement in scientific simulations and experimental facilities has resulted in the generation of vast amounts of data at unprecedented scales. The analysis and visualization of large amounts of data is a challenge in and of itself, but the requirements for timeliness significantly magnify these difficulties. Near real-time visualization is critical to monitor and analyze the data produced by these large facilities, but current production tools are not well-suited to these requirements. In this position paper, we share our perspective on some of the challenges, and thus, opportunities for research that stand in the way of near-real-time visualization of large scientific data.

Pugmire, Dave↗

Black-box statistical prediction of lossy compression ratios for scientific data

Lossy compressors are increasingly adopted in scientific research, tackling volumes of data from experiments or parallel numerical simulations and facilitating data storage and movement. In contrast with the notion of entropy in lossless compression, no theoretical or data-based quantification of lossy compressibility exists for scientific data. Users rely on trial and error to assess lossy compression performance. As a strong data-driven effort toward quantifying lossy compressibility of scientific datasets, we provide a statistical framework to predict compression ratios of lossy compressors. Our method is a two-step framework where (i) compressor-agnostic predictors are computed and (ii) statistical prediction models relying on these predictors are trained on observed compression ratios. Proposed predictors exploit spatial correlations and notions of entropy and lossyness via the quantized entropy. We study 8+ compressors on 6 scientific datasets and achieve a median percentage prediction error less than 12%, which is substantially smaller than that of other methods while achieving at least a 8.8× speedup for searching for a specific compression ratio and 7.8× speedup for determining the best compressor out of a collection.

97 MATHEMATICS AND COMPUTING↗

Error-controlled Progressive Retrieval of Scientific Data under Derivable Quantities of Interest

The unprecedented amount of scientific data has introduced heavy pressure on the current data storage and transmission systems. Progressive compression has been proposed to mitigate this problem, which offers data access with on-demand precision. However, existing approaches only consider precision control on primary data, leaving uncertainties on the quantities of interest (QoIs) derived from it. In this work, we present a progressive data retrieval framework with guaranteed error control on derivable QoIs. Our contributions are three-fold. (1) We carefully derive the theories to strictly control QoI errors during progressive retrieval. Our theory is generic and can be applied to any QoIs that can be composited by the basis of derivable QoIs proved in the paper. (2) We design and develop a generic progressive retrieval framework based on the proposed theories, and optimize it by exploring feasible progressive representations. (3) We evaluate our framework using five real-world datasets with a diverse set of QoIs. Experiments demonstrate that our framework can faithfully respect any user-specified QoI error bounds in the evaluated applications. This leads to over 2.02× performance gain in data transfer tasks compared to transferring the primary data while guaranteeing a QoI error that is less than 1E-5.

Wu, Xuan↗

A Framework for Compressing Unstructured Scientific Data via Serialization

We present a general framework for compressing unstructured scientific data with known local connectivity. A common application is simulation data defined on arbitrary finite element meshes. The framework employs a greedy topology preserving reordering of original nodes which allows for seamless integration into existing data processing pipelines. This reordering process depends solely on mesh connectivity and can be performed offline for optimal efficiency. However, the algorithm’s greedy nature also supports on-the-fly implementation. The proposed method is compatible with any compression algorithm that leverages spatial correlations within the data. The effectiveness of this approach is demonstrated on a large-scale real dataset using several compression methods, including MGARD, SZ, and ZFP.

Reshniak, Viktor [ORNL] (ORCID:0000000315454462)↗

MR-CDF: Managing multi-resolution scientific data

MR-CDF is a system for managing multi-resolution scientific data sets. It is an extension of the popular CDF (Common Data Format) system. MR-CDF provides a simple functional interface to client programs for storage and retrieval of data. Data is stored so that low resolution versions of the data can be provided quickly. Higher resolutions are also available, but not as quickly. By managing data with MR-CDF, an application can be relieved of the low-level details of data management, and can easily trade data resolution for improved access time.

Salem, Kenneth↗

Investigating the Future of Scientific Data Search [Slides]

Searching for usable, actionable, data in a trustworthy manner is a challenge across scientific communities. Artificial Intelligence (AI) and Machine Learning (ML) techniques may be leveraged to increase the utility of scientific data by: Demystify unstructured data to aid curation & sharing Surfacing hard to find datasets. User Experience (UX) Research can help uncover scientists needs & challenges finding data and using AI/ML enabled tools.

97 MATHEMATICS AND COMPUTING↗

Integrating HPC, AI, and Workflows for Scientific Data Analysis: Report from Dagstuhl Seminar 23352

The Dagstuhl Seminar 23352, titled “Integrating HPC, AI, and Workflows for Scientific Data Analysis,” held from August 27 to September 1, 2023, was a significant event focusing on the synergy between High-Performance Computing (HPC), Artificial Intelligence (AI), and scientific workflow technologies. The seminar recognized that modern Big Data analysis in science rests on three pillars: workflow technologies for reproducibility and steering, AI and Machine Learning (ML) for versatile analysis, and HPC for handling large data sets. These elements, while crucial, have traditionally been researched separately, leading to gaps in their integration. The seminar aimed to bridge these gaps, acknowledging the challenges and opportunities at the intersection of these technologies. The event highlighted the complex interplay between HPC, workflows, and ML, noting how ML has increasingly been integrated into scientific workflows, thereby enhancing resource demands and bringing new requirements to HPC architectures, like support for GPUs and iterative computations. The seminar also addressed the challenges in adapting HPC for large-scale ML tasks, including in areas like deep learning, and the need for workflow systems to evolve to leverage ML in data analysis fully. Moreover, the seminar explored how ML could optimize scientific workflow systems and HPC operations, such as through improved scheduling and fault tolerance. A key focus was on identifying prestigious use cases of ML in HPC and understanding their unique, unmet requirements. The stochastic nature of ML and its impact on the reproducibility of data analysis on HPC systems was also a topic of discussion.

97 MATHEMATICS AND COMPUTING↗

The NASA Scientific Data Purchase: Summary and Evaluation

This report summarizes the results of the NASA Scientific Data Purchase (SDP) program implemented by the Stennis Space Center Earth Science Application Directorate in fiscal years 1998-2002. The SDP was conducted in support of NASA's Mission to Planet Earth (MTPE) program (currently known as NASA's Earth Science Enterprise). This Earth Science Enterprise (ESE) provides major observational capabilities for NASA's Earth system science research and the U.S. Global Change Research Program. Observations supported through the SDP were selected based upon their application to the five science themes of the MTPE/ESE: 1) Land cover and land use change research; 2) Seasonal to interannual climate variability and prediction; 3) Natural hazards research and applications; 4) Long-term climate: Natural variability and change research; 5) Atmospheric ozone research. The MTPE/ESE science themes, although they have evolved over time, are driven by a set of key science questions that help focus the research program on characterizing the Earth system.

Source record↗

The Infrared Astronomical Satellite /IRAS/ Scientific Data Analysis System /SDAS/ sky flux subsystem

The sky flux subsystem of the Infrared Astronomical Satellite Scientific Data Analysis System is described. Its major output capabilities are (1) the all-sky lune maps (8-arcminute pixel size), (2) galactic plane maps (2-arcminute pixel size) and (3) regional maps of small areas such as extended sources greater than 1-degree in extent. The major processing functions are to (1) merge the CRDD and pointing data, (2) phase the detector streams, (3) compress the detector streams in the in-scan and cross-scan directions, and (4) extract data. Functional diagrams of the various capabilities of the subsystem are given. Although this device is inherently nonimaging, various calibrated and geometrically controlled imaging products are created, suitable for quantitative and qualitative scientific interpretation.

Stagner, J. R.↗

Domain-Specific Type-Safe APIs for Hierarchical Scientific Data with Modern C++

General-purpose library application programming interfaces (APIs) for self-describing hierarchical scientific data storage, such as the HDF5 and NetCDF libraries, are traditionally of runtime nature. Runtime errors for entry existence and data types are typically caught later in the development process of higher-level application-specific APIs. In this paper, we propose exploiting modern C++ metaprogramming features to add compile-time type-safety to improve the interaction with a well-defined metadata-rich scientific schema in domain-specific hierarchical datasets. We tackle two aspects of common use: (i) direct data access, (ii) flexible “in-memory” index models for efficient search and data processing. The proposed APIs use C++17’s template type auto deduction features, C++11’s enum class for type-safety and C-style preprocessor macros for generative templated code. We showcase the pros and cons of our initial work on the standard NeXus schema used for annotating and storing experimental neutron scattering data at several facilities around the world on top of HDF5. Extendable compile-time type-safe APIs are a desirable feature that could be indexed by any modern integrated development environment (IDE). Hence, such APIs can help ease the learning curve for domain scientists using a less error-prone software interaction to enhance the findability of their data without resorting to a domain-specific language (DSL).

Godoy, William↗

Earth Science Enterprise Scientific Data Purchase Project: Verification and Validation

This paper presents viewgraphs on the Earth Science Enterprise Scientific Data Purchase Project's verification,and validation process. The topics include: 1) What is Verification and Validation? 2) Why Verification and Validation? 3) Background; 4) ESE Data Purchas Validation Process; 5) Data Validation System and Ingest Queue; 6) Shipment Verification; 7) Tracking and Metrics; 8) Validation of Contract Specifications; 9) Earth Watch Data Validation; 10) Validation of Vertical Accuracy; and 11) Results of Vertical Accuracy Assessment.

Jenner, Jeff↗

Representation of Serendipitous Scientific Data

A computer program defines and implements an innovative kind of data structure than can be used for representing information derived from serendipitous discoveries made via collection of scientific data on long exploratory spacecraft missions. Data structures capable of collecting any kind of data can easily be implemented in advance, but the task of designing a fixed and efficient data structure suitable for processing raw data into useful information and taking advantage of serendipitous scientific discovery is becoming increasingly difficult as missions go deeper into space. The present software eases the task by enabling definition of arbitrarily complex data structures that can adapt at run time as raw data are transformed into other types of information. This software runs on a variety of computers, and can be distributed in either source code or binary code form. It must be run in conjunction with any one of a number of Lisp compilers that are available commercially or as shareware. It has no specific memory requirements and depends upon the other software with which it is used. This program is implemented as a library that is called by, and becomes folded into, the other software with which it is used.

James, Mark↗

AI Curation Methods for NASA Scientific Data

The NASA Open Science Data Repository (OSDR) serves as a central hub for sharing and accessing NASA's vast collection of scientific data, supporting researchers across diverse fields. To enhance the efficiency, accuracy, and accessibility of this data, we are leveraging advanced artificial intelligence (AI) techniques as part of the AI for Curation project. By integrating large language models (LLMs) into our data curation workflow, we aim to streamline the entire process—from data submission to user interaction. This initiative focuses on improving key areas, including data ingestion, curation, and user engagement with curated datasets, impacting multiple domains and a wide user base. First, we are developing tools that can automatically parse data in various formats, using LLMs to convert unstructured data into structured, standardized formats. This reduces the manual effort required for curation, allowing curators to focus on more critical scientific analyses. Additionally, AI and machine learning (ML) models are being implemented to automate data validation and verification, ensuring the highest standards of data quality and reliability. Finally, we are creating a conversational AI agent to interact with the curated scientific studies in OSDR, helping users easily navigate the repository and access relevant data. By enhancing data discoverability and accessibility, these advancements will foster new research opportunities and promote the principles of open science.

Walter Alvarado↗

STZ: A High Quality and High Speed Streaming Lossy Compression Framework for Scientific Data

Error-bounded lossy compression is one of the most efficient solutions to reduce the volume of scientific data. For lossy compression, progressive decompression and random-access decompression are critical features that enable on-demand data access and flexible analysis workflows. However, these features can severely degrade compression quality and speed. To address these limitations, we propose a novel streaming compression framework that supports both progressive decompression and random-access decompression while maintaining high compression quality and speed. Our contributions are three-fold: (1) we design the first compression framework that simultaneously enables both progressive decompression and random-access decompression; (2) we introduce a hierarchical partitioning strategy to enable both streaming features, along with a hierarchical prediction mechanism that mitigates the impact of partitioning and achieves high compression quality—even comparable to state-of-the-art (SOTA) non-streaming compressor SZ3; and (3) our framework delivers high compression and decompression speed, up to 6.7 × faster than SZ3.

Wang, Daoce [University of Nebraska, Omaha]↗

Improving Progressive Retrieval for HPC Scientific Data using Deep Neural Network

As the disparity between compute and I/O on high-performance computing systems has continued to widen, it has become increasingly difficult to perform post-hoc data analytics on full-resolution scientific simulation data due to the high I/O cost. Error-bounded data decomposition and progressive data retrieval framework has recently been developed to address such a challenge by performing data decomposition before storage and reading only part of the decomposed data when necessary. However, the performance of the progressive retrieval framework has been suffering from the over-pessimistic error control theory, such that the achieved maximum error of recomposed data is significantly lower than the required error. Therefore, more data than required is fetched for recomposition, incurring additional I/O overhead. In order to tackle this issue, we propose a DNN-based progressive retrieval framework that can better identify the minimum amount of data to be retrieved. Our contributions are as follows: 1) We provide an in-depth investigation of the recently developed progressive retrieval framework; 2) We propose two designs of prediction models (named D-MGARD and E-MGARD) to estimate the amount of retrieved data size based on error bounds. 3) We evaluate our proposed solutions using scientific datasets generated by real-world simulations from two domains. Evaluation results demonstrate the effectiveness of our solution in accurately predicting the amount of retrieval data size, as well as the advantages of our solution over the traditional approach to reducing the I/O overhead. Based on our evaluation, our solution is shown to read significantly less data (5% - 40% with D-MGARD, 20% - 80% with E-MGARD).

Wang, Jinzhen↗