Engineering Papers⌕ Search

Engineering topics

Koziol, Quincey

Publications and source records attributed to Koziol, Quincey.

LBNL Superfacility Project Report

The Superfacility model is designed to leverage HPC for experimental science. It is more than simply a model of connected experiment, network, and HPC facilities; it encompasses the full ecosystem of infrastructure, software, tools, and expertise needed to make connected facilities easy to use. The three-year Lawrence Berkeley National Laboratory (LBNL) Superfacility project was initiated in 2019 to coordinate work being performed at LBNL to support this model, and to provide a coherent and comprehensive set of science requirements to drive existing and new work. A key component of the project was the in-depth engagements with eight science teams that represent challenging use cases across the DOE Office of Science.

97 MATHEMATICS AND COMPUTING↗

A case study on parallel HDF5 dataset concatenation for high energy physics data analysis

In High Energy Physics (HEP), experimentalists generate large volumes of data that, when analyzed, helps us better understand the fundamental particles and their interactions. This data is often captured in many files of small size, creating a data management challenge for scientists. In order to better facilitate data management, transfer, and analysis on large scale platforms, it is advantageous to aggregate data further into a smaller number of larger files. However, this translation process can consume significant time and resources, and if performed incorrectly the resulting aggregated files can be inefficient for highly parallel access during analysis on large scale platforms. In this paper, we present our case study on parallel I/O strategies and HDF5 features for reducing data aggregation time, making effective use of compression, and ensuring efficient access to the resulting data during analysis at scale. We focus on NOvA detector data in this case study, a large-scale HEP experiment generating many terabytes of data. Here, the lessons learned from our case study inform the handling of similar datasets, thus expanding community knowledge related to this common data management task.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Characterizing Impacts of Storage Faults on HPC Applications: A methodology and insights

In recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to equip with more advanced storage infrastructures such as solid-state disks (SSDs). While SSDs offer high-performance I/O, it remains unclear about the reliability challenges faced by the HPC applications under the SSD-related failures, in particular, failures resulting in data corruptions. The goal of this paper is to understand the impact of SSD-related data corruptions on the behaviors of complex HPC applications. To this end, we propose FFIS, a FUSE-based fault injection framework that systematically introduces storage faults into the application layer to model the errors originated from SSDs. FFIS is able to plant different I/O related faults into the data returned from underlying file systems, which also enables the investigation on the error resilience characteristics of the scientific file format for the first time. We demonstrate the use of FFIS with three representative real HPC applications, show how each application reacts to the data corruptions, and provide insights on the error resilience of the widely-adopted HDF5 file format for the HPC applications.

Fang, Bo↗

Alternate physical formats for storing data in HDF

Since its inception HDF has evolved to meet new demands by the scientific community to support new kinds of data and data structures, larger data sets, and larger numbers of data sets. The first generation of HDF supported simple objects and simple storage schemes. These objects were used to build more complex objects such as raster images and scientific data sets. The second generation of HDF provided alternate methods of storing data elements, making it possible to do such things as store extendible objects with in HDF, to store data externally from HDF files, and support data compression effectively. As we look to the next generation of HDF, we are considering fundamental changes to HDF, including a redefinition of the basic HDF object from a simple object to a more general, higher-level scientific data object that has certain characteristics, such as dimensionality, a more general atomic number type, and attributes. These changes suggest corresponding changes to the HDF file format itself.

Folk, Mike↗