Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Data Readiness for Scientific AI at Scale

This paper examines how Data Readiness for AI (DRAI) principles apply to leadership-scale scientific datasets used to train foundation models. We analyze archetypal workflows across four representative domains—climate, nuclear fusion, bio/health, and materials—to identify common preprocessing patterns and domain-specific constraints. We introduce a two-dimensional readiness framework that combines canonical preprocessing patterns with a five-level operational readiness scale, both tailored to high-performance computing (HPC) environments. This framework helps outline key challenges in transforming large-scale scientific data into formats suitable for scalable AI training. Together, these dimensions form a conceptual maturity matrix that characterizes scientific data readiness and guides infrastructure development toward standardized, cross-domain support for scalable and reproducible AI for science.

Brewer, Wes [ORNL] (ORCID:0000000236393956)↗

High Performance Computing Facility Operational Assessment 2022: Oak Ridge Leadership Computing Facility

The Oak Ridge Leadership Computing Facility (OLCF) was established to accelerate scientific discovery by providing world-leading computational performance and advanced data infrastructure. As a US Department of Energy (DOE) Office of Science user facility, the OLCF has managed the successful deployment and operation of a succession of leadership-class resources dedicated to open science. In addition to these resources, the OLCF staff continually strive to develop innovative processes and technologies, improve security, and empower users through allocation management and comprehensive user support and training. These efforts support the advancement of science by the OLCF users and benefit high-performance computing (HPC) facilities around the world. In calendar year (CY) 2022, the OLCF supported 1,681 users and 570 projects and exceeded all targets for user satisfaction. The facility received an average satisfaction score of 4.6 out of 5 on the annual user survey, and 96% of respondents reported a high satisfaction rate with the OLCF overall. Of the 3,212 user tickets submitted in CY 2022, OLCF staff resolved 97% within 3 business days. The facility also introduced several new services for users this year, including weekly virtual office hours with subject matter experts from ORNL and vendor partners; new views in MyOLCF that allow users to analyze allocation and compute usage for a project; the ability to build and run containers on Summit; and improved data visualization support and training resources.

97 MATHEMATICS AND COMPUTING↗

Artificial Intelligence for Data Center Operations (AI Ops)

HPC data centers such as the one at NREL's ESIF will increasingly need to rely on automation to keep pace with exascale growth in compute capability and to manage and optimize the data center environment and facility resources. Artificial intelligence (AI) and machine learning (ML) approaches provide the means to improve HPC data center efficiency (energy, operational, and managerial efficiency) and resiliency by learning historical trends and training models to operate on real-time data collected from both IT and facilities sources. The goal of coupled improvement of data center resiliency and energy efficiency through automated data collection and AI has led to a multi-year, multi-staged collaboration between NREL and Hewlett-Packard Enterprise's Advanced Technology Group, referred to as Artificial Intelligence for Data Center Operations (AIOps). The extended efforts within the AIOps project include a common goal of building capabilities for an advanced smart facility and demonstration of data collection and AI modeling techniques in the ESIF data center.

97 MATHEMATICS AND COMPUTING↗

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

Artificial Intelligence/Deep Learning FRNN Software for Prediction & Real-Time Control of DIII-D Plasma Control System (PCS)

This collaborative project integrated an improved version of the Artificial Intelligence/Deep Learning FRNN prediction and control software into the real-time DIII-D PCS (plasma control system). A key AI/DL software challenge is to build a modern high-performance computing (HPC) enabled “synthetic plasma simulator” capable of carrying out HPC-driven real-time plasma control applications. This involves development of a deep learning framework to train the surrogate model for a first-principles-based instability analysis simulator (“SGTC”) derived from the global gyrokinetic code GTC. The role of SGTC is to provide accurate and detailed plasma instability information from a real-time AI-based simulator capability to complement the deep learning prediction and control from experimentally-measured signals, such as ECE Imaging, supplemented by synthetic SGTC-ECEI.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Comparative Study of Large Language Model Architectures on Frontier

Large language models (LLMs) have garnered significant attention in both the AI community and beyond. Among these, the Generative Pre-trained Transformer (GPT) has emerged as the dominant architecture, spawning numerous variants. However, these variants have undergone pre-training under diverse conditions, including variations in input data, data preprocessing, and training methodologies, resulting in a lack of controlled comparative studies. Here we meticulously examine two prominent open-sourced GPT architectures, GPT-NeoX and LLaMA, leveraging the computational power of Frontier, the world’s first Exascale supercomputer. Employing the same materials science text corpus and a comprehensive end-to-end pipeline, we conduct a comparative analysis of their training and downstream performance. Our efforts culminate in achieving state-of-the-art performance on a challenging materials science benchmark. Furthermore, we investigate the computation and energy efficiency, and propose a computationally efficient method for architecture design. To our knowledge, these pre-trained models represent the largest available for materials science. Our findings provide practical guidance for building LLMs on HPC platforms.

Yin, Junqi↗

US Department of Energy, Office of Science, High-Performance Computing Facility: 2023 Operational Assessment Oak Ridge Leadership Computing Facility

The Oak Ridge Leadership Computing Facility (OLCF) was established to accelerate scientific discovery by providing world-leading computational performance and advanced data infrastructure. As a US Department of Energy (DOE) Office of Science user facility, the OLCF has managed the successful deployment and operation of a succession of leadership-class resources dedicated to open science. In addition to these resources, the OLCF staff continually strive to develop innovative processes and technologies, improve security, and empower users through effective allocation management and comprehensive user support and training. These efforts support the advancement of science by the OLCF users and benefit high-performance computing (HPC) facilities around the world. In calendar year (CY) 2023, the OLCF supported 1,676 users and 598 projects and exceeded all targets for user satisfaction. The facility received an average satisfaction score of 4.52 out of 5 on the annual user survey, and 94% of respondents reported a high satisfaction rate with the OLCF overall. Of the 3,619 user tickets submitted in CY 2023, OLCF staff resolved 97% within 3 business days. The facility opened Frontier to full scientific operations this year. Two projects conducted on Frontier received the Association for Computing Machinery (ACM) Gordon Bell Prize and the Gordon Bell Special Prize for Climate Modeling, and a third earned a nomination as a Gordon Bell Prize finalist. The facility’s previous flagship machine, Summit, gained new life and was extended through 2024 in part to help provide resources to the Integrated Research Infrastructure (IRI) projects and the National Artificial Intelligence Research Resource (NAIRR) pilot program. The facility instantiated an Advanced Computing Ecosystem testbed in part to support IRI workflows. OLCF made interactivity easier and more accessible to users than ever through tools like Jupyter notebooks and workflows.

97 MATHEMATICS AND COMPUTING↗

US Department of Energy, Office of Science, High-Performance Computing Facility 2024 Operational Assessment Oak Ridge Leadership Computing Facility

The Oak Ridge Leadership Computing Facility (OLCF) was established to accelerate scientific discovery by providing world-leading computational performance and advanced data infrastructure to the US Department of Energy (DOE) computing community. As a DOE Office of Science user facility, the OLCF has managed the successful deployment and operation of a succession of leadership-class resources dedicated to open science. In addition to these resources, the OLCF staff continually strive to develop innovative processes and technologies, improve security, and empower users through effective allocation management and comprehensive user support and training. These efforts support the advancement of science by the OLCF users and benefit high-performance computing (HPC) facilities around the world.

97 MATHEMATICS AND COMPUTING↗

Unified Language Frontend for Physic-Informed AI/ML

Artificial intelligence and machine learning (AI/ML) are becoming important tools for scientific modeling and simulation as in several other fields such as image analysis and natural language processing. ML techniques can leverage the computing power available in modern systems and reduce the human effort needed to configure experiments, interpret and visualize results, draw conclusions from huge quantities of raw data, and build surrogates for physics based models. Domain scientists in fields like fluid dynamics, microelectronics and chemistry can automate many of their most difficult and repetitive tasks or improve the design times by use of the faster ML-surrogates. However, modern ML and traditional scientific highperformance computing (HPC) tend to use completely different software ecosystems. While ML frameworks like PyTorch and TensorFlow provide Python APIs, most HPC applications and libraries are written in C++. Direct interoperability between the two languages is possible but is tedious and error-prone. In this work, we show that a compiler-based approach can bridge the gap between ML frameworks and scientific software with less developer effort and better efficiency. We use the MLIR (multi-level intermediate representation) ecosystem to compile a pre-trained convolutional neural network (CNN) in PyTorch to freestanding C++ source code in the Kokkos programming model. Kokkos is a programming model widely used in HPC to write portable, shared-memory parallel code that can natively target a variety of CPU and GPU architectures. Our compiler-generated source code can be directly integrated into any Kokkosbased application with no dependencies on Python or cross-language interfaces.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Machine Learning Assisted HPC Workload Trace Generation for Leadership Scale Storage Systems

Monitoring and analyzing a wide range of I/O activities in an HPC cluster is important in maintaining mission-critical performance in a large-scale, multi-user, parallel storage system. Center-wide I/O traces can provide high-level information and fine-grained activities per application or per user running in the system. Studying such large-scale traces can provide helpful insights into the system. It can be used to develop predictive methods for making predictive decisions, adjusting scheduling policies, or providing decisions for the design of next-generation systems. However, sharing real-world I/O traces to expedite such research efforts leaves a few concerns; i) the cost of sharing the large traces is expensive due to this large size, and ii) privacy concern is an issue.We address such issues by building an end-to-end machine learn- ing (ML) workflow that can generate I/O traces for large-scale HPC applications. We leverage ML based feature selection and gener- ative models for I/O trace generation. The generative models are trained on I/O traces collected by the darshan I/O characterization tool over a period of one year. We present a two-step generation process consisting of two deep-learning models, called the feature generator and the trace generator. The combination of two-step generative models provides robustness by reducing the bias of the model and accounting for the stochastic nature of the I/O traces across different runs of an application. We evaluate the performance of the generative models and show that the two-step model can generate time-series I/O traces with less than 20% root mean square error.

Paul, Arnab↗

Data Center Immersion Cooling: A Case Study and Summary of High-Performance Computing Cooling Technologies

The future of High Performance Computing (HPC) is carved out for all high-performance facilities. As computer processing increases exponentially, especially with the demand for AI training, more speed will equate to more computing density and ultimately more power draw and resource demand. Sandia National Laboratories has not seen or known of anything that would lead us to think this might change in the next 5 to 10 years.

97 MATHEMATICS AND COMPUTING↗

Large-Scale Materials Modeling at Quantum Accuracy: Ab Initio Simulations of Quasicrystals and Interacting Extended Defects in Metallic Alloys

Ab initio electronic-structure has remained dichotomous between achievable accuracy and length-scale. Quantum many-body (QMB) methods realize quantum accuracy but fail to scale. Density functional theory (DFT) scales favorably but remains far from quantum accuracy. We present a framework that breaks this dichotomy by use of three interconnected modules: (i) invDFT: a methodological advance in inverse DFT linking QMB methods to DFT; (ii) MLXC: a machine-learned density functional trained with invDFT data, commensurate with quantum accuracy; (iii) DFT-FE-MLXC: an adaptive higher-order spectral finite-element (FE) based DFT implementation that integrates MLXC with efficient solver strategies and HPC innovations in FE-specific dense linear algebra, mixed-precision algorithms, and asynchronous compute-communication. Furthermore, we demonstrate a paradigm shift in DFT that not only provides an accuracy commensurate with QMB methods in ground-state energies, but also attains an unprecedented performance of 659.7 PFLOPS (43.1% peak FP64 performance) on 619,124 electrons using 8,000 GPU nodes of Frontier supercomputer.

density functional theory↗

Scaling Ensembles of Data-Intensive Quantum Chemical Calculations for Millions of Molecules

Deep learning models are efficient computational tools that can accelerate the inverse design of molecules with desired functional properties by generating predictions at a fraction of the time required by traditional quantum chemical approaches. To ensure that a model maintains accuracy and transferability across broad regions of the chemical space explored during the inverse design, it must be trained on massively large volumes of simulation data. This requires running large-scale ensemble quantum chemical calculations on high-performance computing (HPC) systems for data collection. However, the efficient execution of such large ensemble calculations and the management of large volumes of output data require tools that can judiciously utilize computational resources and manage metadata overhead on the file system. Therefore, we present a high-performance, scalable, ensemble management framework for performing data-intensive quantum chemical electronic structure calculations for organic molecules. This framework provides abstractions to plug different ab initio, first principles, and first principles-based semi-empirical methods and executes them efficiently at large scale on HPC systems. It dynamically distributes tasks to resources and uses tiered storage for managing large collections of files. We employed this framework to process over ten million organic molecules and generate open-source datasets that provide UV-vis absorption spectra by running time-dependent density-functional tight-binding calculations. It is the largest database containing molecular optical spectra that were simulated with quantum chemical methods in a consistent manner.

Mehta, Kshitij↗

Efficient distributed continual learning for steering experiments in real-time

Deep learning has emerged as a powerful method for extracting valuable information from large volumes of data. However, when new training data arrives continuously (i.e., is not fully available from the beginning), incremental training suffers from catastrophic forgetting (i.e., new patterns are reinforced at the expense of previously acquired knowledge). Training from scratch each time new training data becomes available would result in extremely long training times and massive data accumulation. Rehearsal-based continual learning has shown promise for addressing the catastrophic forgetting challenge, but research to date has not addressed performance and scalability. To fill this gap, we propose an approach based on a distributed rehearsal buffer that efficiently complements data-parallel training on multiple GPUs to achieve high accuracy, short runtime, and scalability. It leverages a set of buffers (local to each GPU) and uses several asynchronous techniques for updating these local buffers in an embarrassingly parallel fashion, all while handling the communication overheads necessary to augment input minibatches using unbiased, global sampling. We further propose a generalization of rehearsal buffers to support both classification and generative learning tasks, as well as more advanced rehearsal strategies (notably Dark Experience Replay, leveraging knowledge distillation). We illustrate this approach with a real-life HPC streaming application from the domain of ptychographic image reconstruction. Furthermore, we run extensive experiments on up to 128 GPUs of the ThetaGPU supercomputer to compare our approach with baselines representative of training-from-scratch (the upper bound in terms of accuracy) and incremental training (the lower bound). Results show that rehearsal-based continual learning achieves a top-5 validation accuracy close to the upper bound, while simultaneously exhibiting a runtime close to the lower bound.

Asynchronous data management↗

S4PST: Sustainability for Programming Systems and Tools: May Workshop Report

The US Department of Energy (DOE) Exascale Computing Project (ECP) has fostered and strengthened the use of modern software engineering practices for developing applications and libraries, and this effort has resulted in the coordinated and interoperable E4S1 and xSDK2 ecosystems. Although this approach is cost-effective, it relies on robust programming systems and tools (PST) as the underlying foundation for our HPC software. At present, our primary PST stack consists of traditional high-performance computing (HPC) languages, namely Fortran, C, C++, and the popular Python language for data analysis and AI workflows. These languages support various programming frameworks and run-time abstractions that enable parallelism and concurrency across multiple node architectures and thousands of nodes through a variety of interconnect systems. However, to accommodate users’ diverse needs, certain aspects of the HPC ecosystem are delegated to vendor-specific or third-party implementations that extend beyond a particular scientific domain. This broader scope results in a multitude of specifications and variations, which leads to a complex orchestration of many-ecosystems. Unfortunately, this complexity in the ecosystem imposes additional overhead costs on consumers during the latter stages of the development cycle. In addition to the software ecosystem challenge, the upcoming conclusion of the ECP by December 2023 has raised significant concerns within the HPC programming systems community, from both the economic and social perspectives. The ECP has implemented a management structure for software development and funding decisions across all ECP participants by following a conventional hierarchical and centralized approach. However, this structure has prompted certain considerations within the community, particularly in anticipation of the Software Sustainability initiative by the DOE’s Advanced Scientific Computing Research Program (ASCR). For the success of this new initiative, it is of utmost importance to secure consistent funding and foster close engagement with researchers and core developers of existing programming-system products. This collaboration is vital to maintaining the critical capabilities of the current software during the transition phase while proactively adapting to future technology and workforce trends. The community recognizes the significance of adapting to emerging trends and is aware of the inherent fragility of the HPC software ecosystem, particularly in relation to programming systems that cater to all users. The ability to adapt and evolve is essential to staying relevant and effectively addressing these technical, economic, and social challenges. The S4PST team, which represents one of the six ASCR Software Sustainability seedling projects, is dedicated to tackling these challenges through community-based approaches that go beyond the scope of the DOE. This involves collaboration between national laboratories with academia, non-DOE institutions, hardware and system vendors, and international partners. By fostering these partnerships, we aim to create a robust and sustainable HPC software ecosystem that can effectively meet the needs of the community. This new community effort, driven by the eight DOE labs, will take on the responsibility of guiding funding decisions for programming-systems development and maintenance with transparency and consistency across all decisions. Additionally, the team will offer common technical services to the programming systems community, irrespective of their funding situations, and facilitate community-wide incubation to proactively nurture the software ecosystem. By actively engaging with stakeholders and employing a collaborative approach, we can collectively shape the future of programming systems and ensure a robust and thriving HPC software landscape. On May 11–12, 2023, the S4PST team conducted its inaugural kick-off workshop at the Innovative Computing Laboratory (ICL) in the University of Tennessee, Knoxville, hosted by Hartwig Anzt. The workshop encompassed various sessions dedicated to presentations and discussions, with the aim of comprehending the team members’ perspectives on the vision of software sustainability. Additionally, the workshop aimed to identify the technical, economic, and social requirements for sustaining the programming-systems community in the field of HPC. This report provides a summary of the S4PST effort by highlighting five major thrust areas discussed during the workshop: (i) community, (ii) technical support, (iii) training and diversity, (iv) verification, validation and correctness, and (v) emerging technologies. It also encompasses an overview of the presentations and discussions held throughout the event, our views and potential synergies with other seedling efforts, along with the outcomes and key takeaways from our initial discussions.

97 MATHEMATICS AND COMPUTING↗

The Challenge of Disproportionate Importance of Temporal Features in Predicting HPC Power Consumption

In this work, we demonstrate the challenges in predicting HPC cluster power consumption in the face of significant temporal skew in power consumption behavioral patterns. Predicting large power swings that extend several megawatts has significant operational value for HPC centers, however, prediction is challenging due to the relative rarity of such events and also due to the abrupt or disjoint deviation from the average power consumption levels. To study the impact of this challenge, we have trained a recurrent neural network (RNN) as a reasonably sophisticated model to predict power consumption of the one-year worth of node power consumption data from the Summit supercomputer located in the Oak Ridge Leadership Computing Facility. By studying the prediction results, we have found that although simple usage of RNN models can provide good results on average power consumption levels, it would fail at predicting the power swings that have more operational value. With such results, we discuss potential next steps in addressing such issues aiming towards a robust usage of power prediction techniques in HPC operations.

Li, Chengcheng↗

Towards Low-Overhead Resilience for Data Parallel Deep Learning

Data parallel techniques have been widely adopted both in academia and industry as a tool to enable scalable training of deep learning models. At scale, DL training jobs can fail due to software or hardware bugs, may need to be preempted or terminated due to unexpected events, or may perform suboptimally because they were misconfigured. Under such circumstances, there is a need to recover and/or reconfigure data-parallel DL training jobs on-the-fly, while minimizing the impact on the accuracy of the DNN model and the runtime overhead. In this regard, state-of-art techniques adopted by the HPC community mostly rely on checkpoint-restart, which inevitably leads to loss of progress, thus increasing the runtime overhead. In this paper we explore alternative techniques that exploit the properties of modern deep learning frameworks (overlapping of gradient averaging and weight updates with local gradient computations through pipeline parallelism) to reduce the overhead of resilience/elasticity. To this end we introduce a failure simulation framework and two resilience strategies (immediate mini-batch rollback and lossy forward recovery), which we study compared with checkpoint-restart approaches in a variety of settings in order to understand the trade-offs between the accuracy loss of the DNN model and the runtime overhead.

data-parallel training↗

Can the United States Maintain Its Leadership in High-Performance Computing? - A report from the ASCAC Subcommittee on American Competitiveness and Innovation to the ASCR Office

The United States (U.S.) is no longer the unambiguous leader in the vitally important field of high performance computing (HPC). Japan, the European Union (EU), and China have fielded systems that are on par with our fastest supercomputers. The supply chain for everything from semiconductors to scientific software is globally distributed. Yet our economic future and security depend critically on our ability to innovate faster than our competitors, and the speed of innovation depends increasingly on large-scale computational science and engineering and thus HPC. How should the United States respond to this challenge? This report seeks to initiate a new and potentially transformative national discussion on this vital question. The Department of Energy’s (DOE) Advanced Scientific Computing Research (ASCR) program is well-positioned to make informed, targeted decisions about where the United States should cooperate and where it should compete in the global market for scientific exploration and discovery. By setting its sights on problems critical to our nation and the world, by establishing productive new collaborations, and by making strategic investments, ASCR can restore and maintain U.S. scientific leadership in the critical areas described in this report while strengthening our research infrastructure and training a large, diverse cohort of scientists. In doing so, ASCR and its scientists will pave the way for a secure and prosperous future for America. For more than 30 years, the ASCR program has provided the HPC and networking capabilities and expertise needed to support DOE’s mission to advance the national, economic, and energy security of the United States. The program now faces the challenge of developing and deploying the next generation of HPC systems and technologies, as well as supporting the application of HPC and artificial intelligence (AI) technologies to a wide range of scientific and engineering research problems. Through its research and development efforts, the ASCR program must also advance the state of the art in HPC and accelerate the pace of scientific discovery and technological innovation. Fulfilling this promise will require significantly increased investments, as well as innovative policies and programs. This subcommittee is aware that we are making recommendations and calls for action at a time when federal resources are limited. We understand that a wide range of competing priorities must be balanced by the nation’s leaders and that there is a need to leverage resources in new ways and seek efficiencies in facilities and operations. However, we must not let these realities limit our imagination or silence our advocacy. The ASCR program is a key part of the U.S. research infrastructure and an important component of economic growth and U.S. competitiveness. ASCR has a responsibility to pursue its mission, including advanced scientific computing, applications of AI technologies, and the required advanced research facilities, with determination and enthusiasm. To fulfill the scientific enterprise’s responsibility to the nation, the ASCR program must not only develop and publish a clear vision with an associated list of goals, priorities, and recommendations but also demonstrate scientific leadership by consistently securing long-term funding. This will allow the program to build on its achievements to date, to realize its ambitious vision, and to make lasting contributions to the field.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗