Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC training”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

HPC-FAIR: A Framework Managing Data and AI Models for Analyzing and Optimizing Scientific Applications

The increasing reliance on machine learning (ML) to analyze and optimize large-scale scientific applications on supercomputers faces a significant bottleneck: the lack of readily available, high-quality training datasets and the difficulty in reusing existing AI models. This project was motivated by the urgent need to address the “FAIR” principles (Findability, Accessibility, Interoperability, Reusability) for both training datasets and AI models in the high-performance computing (HPC) domain. The project developed HPC-FAIR, a high-performance computing data management framework designed to centralize HPC-related datasets and AI models within a unified hub. To ensure interoperability, the framework established a standardized representation and vocabulary (ontology) for both data and models. HPC-FAIR also implemented automated workflows to streamline data processing, model access, and benchmarking. Additionally, the project focused on optimizing data harnessing efficiency through advanced techniques like deep reuse and compression-based analytics.

97 MATHEMATICS AND COMPUTING↗

Interpreting Write Performance of Supercomputer I/O Systems with Regression Models

This work seeks to advance the state of the art in HPC I/O performance analysis and interpretation. In particular, we demonstrate effective techniques to: (1) model output performance in the presence of I/O interference from production loads; (2) build features from write patterns and key parameters of the system architecture and configurations; (3) employ suitable machine learning algorithms to improve model accuracy. We train models with five popular regression algorithms and conduct experiments on two distinct production HPC platforms. We find that the lasso and random forest models predict output performance with high accuracy on both of the target systems. We also explore use of the models to guide adaptation in I/O middleware systems, and show potential for improvements of at least 15% from model-guided adaptation on 70% of samples, and improvements up to 10× on some samples for both of the target systems.

Xie, Bing↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre- defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self- learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compare to state-of-the- art approaches.

Baheri, Betis↗

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗

Addressing GPU memory limitations for Graph Neural Networks in High-Energy Physics applications

Introduction Reconstructing low-level particle tracks in neutrino physics can address some of the most fundamental questions about the universe. However, processing petabytes of raw data using deep learning techniques poses a challenging problem in the field of High Energy Physics (HEP). In the Exa.TrkX Project, an illustrative HEP application, preprocessed simulation data is fed into a state-of-art Graph Neural Network (GNN) model, accelerated by GPUs. However, limited GPU memory often leads to Out-of-Memory (OOM) exceptions during training, due to the large size of models and datasets. This problem is exacerbated when deploying models on High-Performance Computing (HPC) systems designed for large-scale applications. Methods We observe a high workload imbalance issue during GNN model training caused by the irregular sizes of input graph samples in HEP datasets, contributing to OOM exceptions. We aim to scale GNNs on HPC systems, by prioritizing workload balance in graph inputs while maintaining model accuracy. Our paper introduces diverse balancing strategies aimed at decreasing the maximum GPU memory footprint and avoiding the OOM exception, across various datasets. Results Our experiments showcase memory reduction of up to 32.14% compared to the baseline. We also demonstrate the proposed strategies can avoid OOM in application. Additionally, we create a distributed multi-GPU implementation using these samplers to demonstrate the scalability of these techniques on the HEP dataset. Discussion By assessing the performance of these strategies as data loading samplers across multiple datasets, we can gauge their effectiveness in both single-GPU and distributed environments. Our experiments, conducted on datasets of varying sizes and across multiple GPUs, broaden the applicability of our work to various GNN applications that handle input datasets with irregular graph sizes.

Lee, Claire Songhyun↗

Access Patterns and Performance Behaviors of Multi-layer Supercomputer I/O Subsystems under Production Load

Scientific computing workloads at HPC facilities have been shifting from traditional numerical simulations to AI/ML applications for training and inference while processing and producing ever-increasing amounts of scientific data. To address the growing need for increased storage capacity, lower access latency, and higher bandwidth, emerging technologies such as non-volatile memory are integrated into supercomputer I/O subsystems. With these emerging trends, we need a better understanding of the multilayer supercomputer I/O systems and ways to use these subsystems efficiently. In this work, we study the I/O access patterns and performance characteristics of two representative supercomputer I/O subsystems. Through an extensive analysis of year-long I/O logs on each system, we report new observations in I/O reads and writes, unbalanced use of storage system layers, and new trends in user behaviors at the HPC I/O middleware stack.

Bez, JL↗

RingX: Scalable Parallel Attention for Long-Context Learning on HPC

The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.

Yin, Junqi [ORNL] (ORCID:0000000338435520)↗

Nuclear Science User Facilities High Performance Computing: Provide a Science Gateway for HPC Users

Idaho National Laboratory (INL), supported by the Department of Energy Office of Nuclear Energy (DOE-NE) through the Nuclear Science User Facilities (NSUF), provides direct access to the Barracuda Virtual Reactor and 18 Multiphysics Object-Oriented Simulation Environment (MOOSE) applications via a web-based science gateway developed using Open Ondemand on the INL high performance computing (HPC) systems. This gateway features the computational tools of the Nuclear Computational Resource Center (NCRC) and the computing resources of the INL high performance computing systems. These computational tools are a key foundation of collaboration and innovation in nuclear energy systems research. High performance computing resources and INL staff directly support the mission and objectives of DOE-NE. The Barracuda Virtual Reactor was the first science gateway deployed in Jan 2021 to support NSUF users. In July 2021, the science gateway was expanded to support access to NCRC codes for use across all supported INL HPC systems. The HPC science gateway currently supports 20 total applications. The gateway also includes access to training resources specific to NCRC tools.

99 GENERAL AND MISCELLANEOUS↗

Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads

Scientific workflows increasingly involve both HPC and machine-learning tasks, combining MPI-based simulations, training, and inference in a single execution. Launchers such as Slurm’s srun constrain concurrency and throughput, making them unsuitable for dynamic and heterogeneous workloads. We present a performance study of RADICAL-Pilot (RP) integrated with Flux and Dragon, two complementary runtime systems that enable hierarchical resource management and high-throughput function execution. Using synthetic and production-scale workloads on Frontier, we characterize the task execution properties of RP across runtime configurations. RP+Flux sustains up to 930 tasks/s, and RP+Flux+Dragon exceeds 1,500 tasks/s with over 99.6% utilization. In contrast, srun peaks at 152 tasks/s and degrades with scale, with utilization below 50%. For IMPECCABLE.v2 drug discovery campaign, RP+Flux reduces makespan by 30–60% relative to srun/Slurm and increases throughput more than four times on up to 1,024. These results demonstrate hybrid runtime integration in RP as a scalable approach for hybrid AI-HPC workloads.

HPC-AI↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

A Scalable Pipeline for Gigapixel Whole Slide Imaging Analysis on Leadership Class HPC Systems

Whole Slide Imaging (WSI) captures microscopic details of a patient's histopathological features at multiple res-olutions organized across different levels. Images produced by WSI are gigapixel-sized, and saving a single image in memory requires a few gigabytes which is scarce since a complicated model occupies tens of gigabytes. Performing a simple met-ric operation on these large images is also expensive. High-performance computing (HPC) can help us quickly analyze such large images using distributed training of complex deep learning models. One popular approach in analyzing these images is to divide a WSI image into smaller tiles (patches) and then train a simpler model with these reduced-sized but large numbers of patches. However, we need to solve three pre-processing challenges efficiently for pursuing this patch-based approach. 1) Creating small patches from a high-resolution image can result in a high number (hundreds of thousands per image) of patches. Storing and processing these images can be challenging due to a large number of I/O and arithmetic operations. To reduce I/Oand memory accesses, an optimal balance between the size and number of patches must exist to reduce I/O and memory accesses. 2) WSI images may have tiny annotated regions for cancer tissue and a significant portion with normal and fatty tissues; correct patch sampling should avoid dataset imbalance. 3) storing and retrieving many patches to and from disk storage might incur I/O latency while training a deep learning model. An efficient distributed data loader should reduce I/O latency during the training and inference steps. This paper explores these three challenges and provides empirical and algorithmic solutions deployed on the Summit supercomputer hosted at the Oak Ridge Leadership Computing Facility.

Dash, Sajal↗

Report of the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science

This report summarizes insights from the 2025 Workshop on Next-Generation Ecosystems for Scientific Computing: Harnessing Community, Software, and AI for Cross-Disciplinary Team Science, which convened more than 40 experts from national laboratories, academia, industry, and community organizations to chart a path toward more powerful, sustainable, and collaborative scientific software ecosystems. To address urgent challenges at the intersection of high-performance computing (HPC), AI, and scientific software, participants envisioned agile, robust ecosystems built through socio-technical co-design—the intentional integration of social and technical components as interdependent parts of a unified strategy. This approach combines advances in AI, HPC, and software with new models for cross-disciplinary collaboration, training, and workforce development. Key recommendations include building modular, trustworthy AI-enabled scientific software systems; enabling scientific teams to integrate AI systems into their workflows while preserving human creativity, trust, and scientific rigor; and creating innovative training pipelines that keep pace with rapid technological change. Pilot projects were identified as near-term catalysts, with initial priorities focused on hybrid AI/HPC infrastructure, cross-disciplinary collaboration and pedagogy, responsible AI guidelines, and prototyping of public-private partnerships. This report presents a vision of next-generation ecosystems for scientific computing where AI, software, hardware, and human expertise are interwoven to drive discovery, expand access, strengthen the workforce, and accelerate scientific progress.

97 MATHEMATICS AND COMPUTING↗

Active Learning for Metamaterial Optimization on HPC and QC Integrated Systems

Active learning algorithms, integrating machine learning, quantum computing and optics simulation in an iterative loop, offer a promising approach to optimizing metamaterials. However, these algorithms can face difficulties in optimizing highly complex structures due to computational limitations. High-performance computing (HPC) and quantum computing (QC) integrated systems can address these issues by enabling parallel computing. In this study, we develop an active learning algorithm working on HPC-QC integrated systems. We evaluate the performance of optimization processes within active learning (i.e., training a machine learning model, problem-solving with quantum computing, and evaluating optical properties through wave-optics simulation) for highly complex metamaterial cases. Our results showcase that utilizing multiple cores on the integrated system can significantly reduce computational time, thereby enhancing the efficiency of optimization processes. Therefore, we expect that leveraging HPC-QC integrated systems helps effectively tackle large-scale optimization challenges in general.

Kim, Seongmin↗

Dilute Combustion Control Using Spiking Neural Networks

Dilute combustion with exhaust gas recirculation (EGR) in spark-ignition engines presents a cost-effective method for achieving higher levels of engine efficiency. At high levels of EGR, however, cycle-to-cycle variability (CCV) of the combustion process is exacerbated by sporadic occurrences of misfires and partial burns. Previous studies have shown that temporal deterministic patterns emerge at such conditions and certain combustion cycles have a significant influence over future events. Due to the complexity of the combustion process and the nature of CCV, harnessing all the deterministic information for control purposes has remained challenging even with physics based 0-D, 1-D, and high-fidelity computational fluid dynamics (CFD) models. In this study, we present a data-driven approach to optimize the combustion process by controlling CCV adjusting the cycle-to-cycle fuel injection quantity. Readily available data from in-cylinder pressure was used to train a spiking neural network (SNN) which learns the optimal way to manage fuel injection in order to reduce CCV while maintaining acceptable levels of fuel consumption. SNNs are particularly well suited for powertrain control applications due to their ability to be deployed on FPGA-based neuromorphic hardware which are small, inexpensive, and have a low power demand. The high-performance computing (HPC) resources of Oak Ridge National Laboratory were used to run an evolutionary-based training approach for choosing the best SNN configuration that minimizes the size of the network while achieving the desired goal. The neuromorphic hardware with the optimized SNN deployed was connected to the rapid prototyping engine control system for real-time control implementation and tested on a single cylinder version of a GM LNF 4-cylinder engine. The results show a significant reduction of CCV with a small percentage of additional fuel used to stabilize the charge.

33 ADVANCED PROPULSION SYSTEMS↗

Exercise Versus +Gz Acceleration Training

Decreased working capacity and "orthostatic" intolerance are two major problems for astronauts during and after landing from spaceflight in a return vehicle. The purpose was to test the hypotheses that (1) supine-passive-acceleration training, supine-interval-exercise plus acceleration training, and supine exercise plus acceleration training will improve orthostatic tolerance (OT) in ambulatory men; and that (2) addition of aerobic exercise conditioning will not influence this enhanced OT from that of passive-acceleration training. Seven untrained men (24-38 yr) underwent 3 training regimens (30 min/d x 5d/wk x 3wk on the human-powered centrifuge - HPC): (a) Passive acceleration (alternating +1.0 Gz to 50% Gzmax); (b) Exercise acceleration (alternating 40% - 90% V02max leg cycle exercise plus 50% of HPCmax acceleration); and (c) Combined intermittent exercise-acceleration at 40% to 90% HPCmax. Maximal supine exercise workloads increased (P < 0.05) by 8.3% with Passive, by 12.6% with Exercise, and by 15.4% with Combined; but maximal V02 and HR were unchanged in all groups. Maximal endurance (time to cessation) was unchanged with Passive, but increased (P < 0.05) with Exercise and Combined. Resting pre-tilt HR was elevated by 12.9% (P < 0.05) only after Passive training, suggesting that exercise training attenuated this HR response. All resting pre-tilt blood pressures (SBP, DBP, MAP) were not different pre- vs. post-training. Post-training tilt-tolerance time and HR were increased (P < 0.05) only with Passive training by 37.8% and by 29.1%, respectively. Thus, addition of exercise training attenuated the increased Passive tilt tolerance. Resting (pre-tilt) and post-tilt cardiac R-R interval, stroke volume, end-diastolic volume, and cardiac output were all uniformly reduced (P < 0.05) while peripheral resistance was uniformly increased (P < 0.05) pre-and post-training for the three regimens indicating no effect of any training regimen on those cardiovascular variables. Plasma volume (% delta) was uniformly decreased by 8% to 14% (P < 0.05) at tilt-tolerance pre- vs. post-training for all regimens indicating no effect of these training regimens on the level of vascular fluid shifts.

Greenleaf, John E.↗

Educating HPC Users in the use of advanced computing technology

We examine a multi-modal approach to educating and training users of an advanced computing technology testbed at the Institute for Advanced Computational Science at Stony Brook University. Ookami provides researchers worldwide with access to 176 Fujitsu A64FX compute nodes, this being the same processor technology powering the Japanese Fugaku supercomputer, the fastest computer in the world since June 2020. However, achieving high-performance on this Arm-based, leadership computing technology requires that users be familiar with details of computer architecture, performance analysis and modeling, and high-performance programming models that are commonly omitted in introductory programming courses. Indeed, regardless of their seniority, many of the testbed users are surprisingly unfamiliar with basic concepts such as vectorization, pipelining, latency/bandwidth, roofline models, computing energy/power, threads, and non-uniform memory access. These same concepts also pervade mainstream x86 technologies, so this is of widespread concern. Due to the national/global nature of our user community that is also very diverse in both discipline and experience, the inability to offer formal classes, and our experience that most people do not tend to read online documentation or training materials in sufficient depth, we have consciously employed multiple approaches that heavily emphasize (online) personal interactions and transfer of skills. Online documentation has been organized around best-practices and FAQs; twice-weekly hackathons and office hours via Zoom enable deep dives by both the team and the user community with multiple broad benefits; a Slack channel provides both real time and archived answers and discussions; and workshops, training and webinars target community needs as they arise. Furthermore, the perspective that these tools are being used in an educational setting rather than just for project communication makes them more effective and contributes to community success.

A64FX↗