How data science methods can improve the quality and efficiency of ICF and HEDP research
Data Science methods (many that are Bayesian based) are widely used in the physical sciences to estimate model parameters from experimental data, synthesize heterogeneous data, calibrate models, design experiments, and determine statistical significance of data. These methods provide a wealth of advantages over traditional analysis techniques because: 1) uncertainties are rigorously defined and propagated naturally through complex systems including covariance, 2) prior information is captured within the analysis framework (including rad-MHD and rad-hydro simulations), 3) competing models can be selected and/or ruled out using quantitative criteria, and 4) complex, heterogeneous data can be incorporated simultaneously. While these methods have been widely adopted as the gold standard in fields such as particle physics, astronomy, and biology, they have been slow to catch on in Inertial Confinement Fusion (ICF) and High Energy Density Physics (HEDP) research. Recently, several teams at LLNL, SNL, LANL, and the LLE have been exploring the use of these tools in their research and have found success. Here we propose that a concerted effort to consolidate these independent research efforts by developing and deploying common tools for use across the complex can revolutionize the way we approach data analysis, assimilation of theory and experiment, and decision making. The Bayesian formalism provides a means to accomplish this, but we are lacking certain infrastructure to make it happen on a large scale. Furthermore, once adopted, these techniques can be used to develop standards by which discoveries can be judged, similar to the so-called 5σ rule in high energy particle physics. Such standards may be used in the future to address the issue of unknown reproducibility in ICF and HED experiments caused by low shot rate and high cost per experiment. Our goals as a group are to advance the state of the art in HED measurement science by enabling: 1) better inferences from data with well-defined uncertainties, 2) better use of the data we have and continue to collect, 3) intelligent synthesis of data, 4) evaluation of the statistical significance of our data, and 5) informed decision making regarding the design of new experiments and instruments.