Engineering Papers⌕ Search

Engineering topics

Spears, Brian

Publications and source records attributed to Spears, Brian.

Advanced Research Directions on AI for Science, Energy, and Security: Report on Summer 2022 Workshops

Over the past decade, fundamental changes in artificial intelligence (AI)—from foundational to applied—have delivered dramatic insights across a wide breadth of U.S. Department of Energy (DOE) mission space. AI is helping to augment and improve scientific and engineering workflows (e.g., for control, design, and dramatic performance gains through surrogate models) in national security, the Office of Science, and DOE’s applied energy programs. The progress and potential for AI in DOE science was captured in the 2020 “AI for Science” report from the DOE laboratory community in collaboration with academia and industry. Specific scientific areas ready to further leverage the power of AI ranged from the scale and performance of computational models to data analysis to creating new classes of observations using computer vision. Since that report, the scale and scope of scientific AI have accelerated, revealing new, emergent properties that yield insights that go beyond enabling opportunities to being potentially transformative in the way that scientific problems are posed and solved. Thus, under the guidance of both the Office of Science (SC) and the National Nuclear Security Administration (NNSA), the DOE national laboratories organized a series of workshops in 2022 to gather input on new and rapidly emerging opportunities and challenges of scientific AI. This 2023 report is a synthesis of those workshops. The scientific community believes AI can have a foundational impact on a broad range of DOE missions, including science, energy, and national security. Further, DOE has unique capabilities that enable the community to drive progress in scientific use of AI, building on long-standing DOE strengths and investments in computation, data, and communications infrastructure, spanning the Energy Sciences Network (ESnet), the Exascale Computing Project (ECP), and integrative programs such as the NNSA Office of Defense Programs Advanced Simulation and Computing (ASC) and the SC Scientific Discovery through Advanced Computing (SciDAC) programs.

97 MATHEMATICS AND COMPUTING↗

Iterative sampling of expensive simulations for faster deep surrogate training

Deep neural network (DNN) surrogates of expensive physics simulations are enabling a rapid change in the way that common experimental design and analysis tasks are approached. Surrogate models allow simulations to be performed in parallel and separately from downstream tasks, thereby enabling analyses that would be impossible with the simulation in-the-loop; surrogates based on DNNs can effectively emulate diverse non-scalar data of the types collected in fusion and laboratory-astrophysics experiments. The challenge is in training the surrogate model, for which large ensembles of physics simulations must be run, preferably without wasting computational effort on uninteresting simulations. Here, in this paper, we present an iterative sampling scheme that can preferentially propose simulations in interesting regions of parameter space without neglecting unexplored regions, allowing high-quality and wide-ranging surrogate models to be trained using 2–3 times fewer simulations compare to space-filling designs. Our approach uses an explicit importance function defined on the simulation output space, balanced against a measure of simulation density which serves as a proxy for surrogate accuracy. It is easy to implement and can be tuned to find interesting simulations early in the study, allowing surrogates to be trained quickly and refined as new simulations become available; this represents an important step towards the routine generation of deep surrogate models quickly enough to be truly relevant to experimental work.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Enabling machine learning-ready HPC ensembles with Merlin

With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.

97 MATHEMATICS AND COMPUTING↗