DOE OSTI · 1736017
SIRIUS: Science-Driven Data Management for Multi-Tiered Storage
Abstract
The data sets being generated by large applications on very large-scale systems are increasing in both size and complexity. At the same time, there are new ways available to store and access these data sets. The goal in this project is to develop software that applications can use to make use of new and existing storage technologies in more sophisticated ways. One challenge in scientific data management is handling ‘hot’ vs ‘cold’ data. Data that is hot is data that is needed (or will be needed soon) in order for the program to continue progressing, while cold data is either output (and so will not be need further during the life of the program) or will not be needed until significantly later in the program’s run. Hot data should be stored in a way that allows fast access. On most systems, economic factors lead to an inverse relationship between storage performance and storage capacity and so fast access storage is limited. This makes it important to correctly place hot and cold data and avoid cold data unnecessarily consuming precious resources. In this reporting period, we addressed this challenge in various ways and at various levels. Data management frameworks offer only limited control to applications in how data is stored. We have added software capabilities for seamlessly moving data between layers of the storage technology using promote and demote functions to existing software frameworks. This gives direct control to applications in deciding what priority data receives. Additionally, we integrated different storage layer management frameworks in order to allow data to be exchanged and moved between storage layers in a consistent way across the application. Further, applications are not always able to directly decide what storage level makes sense for a given piece of data without an understanding of the underlying storage technologies. Data storage frameworks are often positioned to make these sorts of decisions in service of the application. We have added machine-learning based capabilities to data staging frameworks in order to make intelligent decisions about where data should be stored given learning about patterns in previous usage of similar data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Parashar, Manish. 2021-01-19. SIRIUS: Science-Driven Data Management for Multi-Tiered Storage. https://doi.org/10.2172/1736017
Cite the original work for its findings. Save a collection to share your selection of sources.