Usable Data Abstractions for Next-Generation Scientific Workflows
Data- and computationally-intensive scientific research, such as numerical simulations and inversions or the training of large neural networks in machine learning applications, that are well suited for HPC environments also often require expert insight and evaluation throughout the computation which can be greatly facilitated with the use of interactive computing tools, such as those in the Jupyter ecosystem. HPC workflows and interactive workflows are typically treated as orthogonal, however, the next generation of research will require both. The first challenge we face in this project is thus designing the right level of abstractions to allow interactive capabilities in the JupyterLab environment to allow the working scientist to flexibly explore and query their data at multiple levels, with a minimal amount of customization required of the underlying optimized codes. In addition to these questions regarding the high-level representation of data for interactive use in HPC, we tackled two additional issues that are part of the entire lifecycle of research and that become particularly acute in HPC contexts: how to improve the experience of interfacing with the HPC system's scheduling environment for a scientist focused on exploratory questions, and how can that scientist then best share the results of their work with others in a self-contained, reproducible manner.