Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Scientific data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Position Papers for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Report for the ASCR Workshop on the Management and Storage of Scientific Data

The purpose of this workshop is to identify priority research directions in the area of data management for high-performance and scientific computing above and beyond HPC’s traditional "the parallel file system is the data-management system" model. Supporting the breadth of the DOE mission, including the explosion of AI uses and the growing needs of experimental and observational science, motivates revisiting our assumptions about data management. There are many facets of this topic to explore including: (1) Interfaces for accessing data that resides on traditional persistent storage as well as memory devices; (2) Storage-system architecture design that supports scientific workflows on varied hierarchical storage and networking devices; (3) Devising metadata management infrastructure to support FAIR principles (Findability, Accessibility, Interoperability, and Reusability); (4) Capturing provenance information about scientific data; (5) Utilizing AI to learn I/O patterns of emerging workloads for efficient data management; (6) Providing data management support for AI and complex workflows; and (7) Understanding the overlap between traditional storage systems and I/O (SSIO) efforts and data management. While the program committee has identified these topics as important areas for discussion, we welcome position papers from the community that propose additional topics of interest for discussion at the workshop. The workshop agenda will include breakout sessions for discussing these and selected topic areas to inform priority research directions for data management for high-performance and scientific computing.

97 MATHEMATICS AND COMPUTING↗

Asynchronous and Load-Balanced Union-Find for Distributed and Parallel Scientific Data Visualization and Analysis

We present a novel distributed union-find algorithm that features asynchronous parallelism and k-d tree based load balancing for scalable visualization and analysis of scientific data. Applications of union-find include level set extraction and critical point tracking, but distributed union-find can suffer from high synchronization costs and imbalanced workloads across parallel processes. In this study, we prove that global synchronizations in existing distributed union-find can be eliminated without changing final results, allowing overlapped communications and computations for scalable processing. We also use a k-d tree decomposition to redistribute inputs, in order to improve workload balancing. We benchmark the scalability of our algorithm with up to 1,024 processes using both synthetic and application data. Here, we demonstrate the use of our algorithm in critical point tracking and super-level set extraction with high-speed imaging experiments and fusion plasma simulations, respectively.

97 MATHEMATICS AND COMPUTING↗

RAPIDS: Reconciling Availability, Accuracy, and Performance in Managing Geo-Distributed Scientific Data

In modern science, big data plays an increasingly important role. Many scientific applications, such as running simulations on supercomputers or conducting experiments on advanced instruments, produce huge amount of data at unprecedented speed. Analyzing and understanding such big data is the key for scientists to make scientific breakthroughs. However, data might become unavailable for scientists to access when outages or maintenance of the storage system occur, which severely hinders scientific discovery. To improve the data availability, data duplication and erasure coding (EC) are often used. But as the scientific data gets larger, using these two methods can cause considerable storage and network overhead.In this paper, we propose RAPIDS, a hybrid approach that combines the multigrid-based error-bounded lossy compression with erasure coding, to significantly reduce the storage and network overhead required for maintaining high data availability. Our experiments show that RAPIDS reduces the storage overhead by up to 7.5x and network overhead by up to 3x to achieve the same level of availability compared to the regular EC method. We improve RAPIDS by building two models to optimize the fault tolerance configurations and data gathering strategy. We demonstrate that RAPIDS significantly improves performance when running on many CPU cores in parallel or on GPUs.

Wan, Lipeng↗

Space and Earth Sciences, Computer Systems, and Scientific Data Analysis Support, Volume 1

This Final Progress Report covers the specific technical activities of Hughes STX Corporation for the last contract triannual period of 1 June through 30 Sep. 1993, in support of assigned task activities at Goddard Space Flight Center (GSFC). It also provides a brief summary of work throughout the contract period of performance on each active task. Technical activity is presented in Volume 1, while financial and level-of-effort data is presented in Volume 2. Technical support was provided to all Division and Laboratories of Goddard's Space Sciences and Earth Sciences Directorates. Types of support include: scientific programming, systems programming, computer management, mission planning, scientific investigation, data analysis, data processing, data base creation and maintenance, instrumentation development, and management services. Mission and instruments supported include: ROSAT, Astro-D, BBXRT, XTE, AXAF, GRO, COBE, WIND, UIT, SMM, STIS, HEIDI, DE, URAP, CRRES, Voyagers, ISEE, San Marco, LAGEOS, TOPEX/Poseidon, Pioneer-Venus, Galileo, Cassini, Nimbus-7/TOMS, Meteor-3/TOMS, FIFE, BOREAS, TRMM, AVHRR, and Landsat. Accomplishments include: development of computing programs for mission science and data analysis, supercomputer applications support, computer network support, computational upgrades for data archival and analysis centers, end-to-end management for mission data flow, scientific modeling and results in the fields of space and Earth physics, planning and design of GSFC VO DAAC and VO IMS, fabrication, assembly, and testing of mission instrumentation, and design of mission operations center.

Estes, Ronald H.↗

Prioritizing Scientific Data for Transmission

A software system has been developed for prioritizing newly acquired geological data onboard a planetary rover. The system has been designed to enable efficient use of limited communication resources by transmitting the data likely to have the most scientific value. This software operates onboard a rover by analyzing collected data, identifying potential scientific targets, and then using that information to prioritize data for transmission to Earth. Currently, the system is focused on the analysis of acquired images, although the general techniques are applicable to a wide range of data modalities. Image prioritization is performed using two main steps. In the first step, the software detects features of interest from each image. In its current application, the system is focused on visual properties of rocks. Thus, rocks are located in each image and rock properties, such as shape, texture, and albedo, are extracted from the identified rocks. In the second step, the features extracted from a group of images are used to prioritize the images using three different methods: (1) identification of key target signature (finding specific rock features the scientist has identified as important), (2) novelty detection (finding rocks we haven t seen before), and (3) representative rock sampling (finding the most average sample of each rock type). These methods use techniques such as K-means unsupervised clustering and a discrimination-based kernel classifier to rank images based on their interest level.

Castano, Rebecca↗

Autonomic and Apoptotic, Aeronautical and Aerospace Systems, and Controlling Scientific Data Generated Therefrom

A self-managing system that uses autonomy and autonomicity is provided with the self-* property of autopoiesis (self-creation). In the event of an agent in the system self-destructing, autopoiesis auto-generates a replacement. A self-esteem reward scheme is also provided and can be used for autonomic agents, based on their performance and trust. Art agent with greater self-esteem may clone at a greater rate compared to the rate of an agent with lower self-esteem. A self-managing system is provided for a high volume of distributed autonomic/self-managing mobile agents, and autonomic adhesion is used to attract similar agents together or to repel dissimilar agents from an event horizon. An apoptotic system is also provided that accords an "expiry date" to data and digital objects, for example, that are available on the internet, which finds usefulness not only in general but also for controlling the loaning and use of space scientific data.

Sterritt, Roy↗

A Data-scientific Noise-removal Method for Efficient Submillimeter Spectroscopy With Single-dish Telescopes

For submillimeter spectroscopy with ground-based single-dish telescopes, removing the noise contribution from the Earth’s atmosphere and the instrument is essential. For this purpose, here we propose a new method based on a data-scientific approach. The key technique is statistical matrix decomposition that automatically separates the signals of astronomical emission lines from the drift noise components in the fast-sampled (1–10 Hz) time-series spectra obtained by a position-switching (PSW) observation. Because the proposed method does not apply subtraction between two sets of noisy data (i.e., on-source and off-source spectra), it improves the observation sensitivity by a factor of √2. It also reduces artificial signals such as baseline ripples on a spectrum, which may also help to improve the effective sensitivity. We demonstrate this improvement by using the spectroscopic data of emission lines toward a high-redshift galaxy observed with a 2 mm receiver on the 50 m Large Millimeter Telescope. Since the proposed method is carried out offline and no additional measurements are required, it offers an instant improvement on the spectra reduced so far with the conventional method. It also enables efficient deep spectroscopy driven by the future 50 m class large submillimeter single-dish telescopes, where fast PSW observations by mechanical antenna or mirror drive are difficult to achieve.

47 OTHER INSTRUMENTATION↗

A classification and evaluation of data movement technologies for the delivery of highly voluminous scientific data products

In this paper, we present a preliminary study of several different electronic data movement technologies. We detail our approach to classifying the technologies included in our study and present the preliminary results of some initial performance benchmarking. Our studies suggest that highly parallel TCP/IP streaming technologies, such as GridFTP and bbFTP, outperform commercial and open-source UDP-bursting technologies in several of the key data movement dimensions that we studied.

technologies↗

NASA Scientific Data Purchase Laboratory Characterization of Positive Systems ADAR 5500 Sensor

This report summarizes the laboratory characterization of the Positive Systems ADAR 5500 SN4 multi-spectral imager measured dirijg the period of April 10-11, 2000 in the NASA CRSP Commercial Instrumentation Validation Laboratory (CIVL). This characterization is one component of the NASA Scientific Data Purchase (SDP) validation and verification process. A complete characterization will require analysis of airborne data. The spectral, radiometric, and spatial characteristics of this system were measured and compared with the product data specifications as defined in the SDP contract. These results are summarized in the specification compliance matrices in this summary. A complete description of the system, the SDP product specification and the sensor laboratory characterization is provided in the body of the main report.

Ryan, James↗

Performance of customized DCT quantization tables on scientific data

We show that it is desirable to use data-specific or customized quantization tables for scaling the spatial frequency coefficients obtained using the Discrete Cosine Transform (DCT). DCT is widely used for image and video compression (MP89, PM93) but applications typically use default quantization matrices. Using actual scientific data gathered from divers sources such as spacecrafts and electron-microscopes, we show that the default compression/quality tradeoffs can be significantly improved upon by using customized tables. We also show that significant improvements are possible for the standard test images Lena and Baboon. This work is part of an effort to develop a practical scheme for optimizing quantization matrices for any given image or video stream, under any given quality or compression constraints.

Ratnakar, Viresh↗

Infrared Astronomical Satellite /IRAS/ Scientific Data Analysis System

The Infrared Astronomical Satellite (IRAS), to be launched in 1982, is discussed. It will systematically survey the entire sky over a large percentage of the infrared spectrum, in the wavelength region of 8 to 120 microns, at sensitivities a hundred times greater than previously achieved from high-altitude observatories, aircraft, balloons or sounding rockets. The Scientific Data Analysis System (SDAS), an off-line data processing facility, is examined. Its primary function is to produce a catalog of inertially fixed infrared-emitting point sources (mainly stars and galaxies) observed during the IRAS survey. Details for source detection and confirmation are given. It is estimated that the catalog will contain approximately a million objects having a brightness of 10 amtowatts per square centimeter or greater; 125,000 SDAS detections, if spurious events of signal-to-noise ratios greater than 2.5 are included, will be made every day.

Duxbury, J. H.↗

Understanding and Estimating Error Propagation in Neural Networks for Scientific Data Analysis

Neural networks are increasingly integrated into scientific discovery, where input data reduction and model quantization play a key role in accelerating inference. However, understanding and mitigating the impact of these techniques on output error is critical for ensuring reliable results, particularly in tasks demanding high numerical precision. This paper introduces a comprehensive framework for optimizing neural network inference in scientific computing by combining data reduction and weight quantization while maintaining error-controlled outcomes. We develop theoretical analyses to bound error propagation under these reductions and propose a framework that balances computational performance with error constraints. Evaluation on real-world learning-based combustion simulations and satellite image classification demonstrates that our derived error bounds accurately predict observed errors while enabling significant computational speedup under our framework. This work highlights the potential for further leveraging advancements in modern lossy compression algorithms and hardware accelerators that support lower-precision formats.

He, Weiming [New Jersey Institute of Technology]↗

Error-Bounded Learned Scientific Data Compression with Preservation of Derived Quantities

Scientific applications continue to grow and produce extremely large amounts of data, which require efficient compression algorithms for long-term storage. Compression errors in scientific applications can have a deleterious impact on downstream processing. Thus, it is crucial to preserve all the “known” Quantities of Interest (QoI) during compression. To address this issue, most existing approaches guarantee the reconstruction error of the original data or primary data (PD), but cannot directly control the problem of preserving the QoI. In this work, we propose a physics-informed compression technique that is composed of two parts: (i) reduction of the PD with bounded errors and (ii) preservation of the QoI. In the first step, we combine tensor decompositions, autoencoders, product quantizers, and error-bounded lossy compressors to bound the reconstruction error at high levels of compression. In the second step, we use constraint satisfaction post-processing followed by quantization to preserve the QoI. To illustrate the challenges of reducing the reconstruction errors of the PD and QoI, we focus on simulation data generated by a large-scale fusion code, XGC, which can produce tens of petabytes in a single day. The results show that our approach can achieve a high compression amount while accurately preserving the QoI within scientifically acceptable bounds.

97 MATHEMATICS AND COMPUTING↗

Uncovering interpretable relationships in high-dimensional scientific data through function preserving projections

Abstract In many fields of science and engineering, we frequently encounter experiments or simulations datasets that describe the behavior of complex systems and uncovering human interpretable patterns between their inputs and outputs via exploratory data analysis is essential for building intuition and facilitating discovery. Often, we resort to 2D embeddings for examining these high-dimensional relationships (e.g. dimensionality reduction). However, most existing embedding methods treat the dimensions as coordinates for samples in a high-dimensional space, which fail to capture the potential functional relationships, and the few methods that do take function into consideration either only focus on linear patterns or produce non-linear embeddings that are hard to interpret. To address these challenges, we proposed function preserving projections (FPP), which construct 2D linear embeddings optimized to reveal interpretable yet potentially non-linear patterns between the domain and the range of a high-dimensional function. The intuition here is that humans are good at understanding potentially non-linear patterns in 2D but unable to interpret non-linear mapping from high-dimensional space to 2D. Therefore, we should restrict the projection to linear but not the pattern we are seeking. Using FPP on real-world datasets, one can obtain fundamentally new insights about high-dimensional relationships in extremely large datasets that could not be processed with existing dimension reduction methods.

97 MATHEMATICS AND COMPUTING↗

Pilot climate data system: A state-of-the-art capability in scientific data management

The Pilot Climate Data System (PCDS) was developed by the Information Management Branch of NASA's Goddard Space Flight Center to manage a large collection of climate-related data of interest to the research community. The PCDS now provides uniform data catalogs, inventories, access methods, graphical displays and statistical calculations for selected NASA and non-NASA data sets. Data manipulation capabilities were developed to permit researchers to easily combine or compare data. The current capabilities of the PCDS include many tools for the statistical survey of climate data. A climate researcher can examine any data set of interest via flexible utilities to create a variety of two- and three-dimensional displays, including vector plots, scatter diagrams, histograms, contour plots, surface diagrams and pseudo-color images. The graphics and statistics subsystems employ an intermediate data storage format which is data-set independent. Outside of the graphics system there exist other utilities to select, filter, list, compress, and calculate time-averages and variances for any data of interest. The PCDS now fully supports approximately twenty different data sets and is being used on a trial basis by several different in-house research grounds.

Smith, P. H.↗