Engineering PapersSearch

DOE OSTI · 3363926

Data Readiness for Scientific AI at Scale

Abstract

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Brewer, Wes [ORNL] (ORCID:0000000236393956), Widener, Patrick [ORNL] (ORCID:0000000258820816), Anantharaj, Valentine [ORNL] (ORCID:0000000293561311), Wang, Feiyi [ORNL] (ORCID:0000000200991559), Beck, Tom [ORNL] (ORCID:0000000189737145), Shankar, Mallikarjun (Arjun) [ORNL] (ORCID:0000000152897460), Oral, Sarp [ORNL] (ORCID:0000000187457078). 2025-12-01. Data Readiness for Scientific AI at Scale. https://doi.org/10.1145/3750720.3757282

Cite the original work for its findings. Save a collection to share your selection of sources.