LLNL/cyme-launcher
Utility to help run CYME in an HPC environment. Ensures sufficient licenses are available for running CYME in as part of a batch job, and helps start CYME running under WINE on Linux.
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Utility to help run CYME in an HPC environment. Ensures sufficient licenses are available for running CYME in as part of a batch job, and helps start CYME running under WINE on Linux.
Network switches, such as those from Arista and Mellanox, often have underutilized computational resources in the form of built-in processors and memory. By leveraging these untapped resources, we can optimize functionality and efficiency of computational cluster networks. Our research focuses on deploying containers directly onto these switches to execute various auxiliary tasks ranging from metric logging to system-wide management via post-boot configuration. By doing so, we can significantly enchance the capabilities of the cluster without the need for additional dedicated hardware. Our research involved five distinct scenarios where switch utilization could have a profound impact on HPC Clusters: run cloud-init services via link-local connection; configuring a Telegraf container to export metrics; deploying a caching proxy; creating a reconfigurable IPv6 DHCP/DNS provider for VLAN; and implementing a client detection with Magellan discovery. These scenarios were containerized with podman and docker, and tested both physically on the switch virtually on a QEMU VM both running SONiC OS. Testing and findings indicate that network switches can indeed be used for these scenarios. They offer a wide range of possibilities beyond these applications. They run as expected as containers on the switches, and although there were some minor issues, work-arounds were implemented. Overall, this is a positive result that can be further explored with more scenarios.
Abstract. Some programming languages are easy to develop at the cost of slow execution, while others are fast at runtime but much more difficult to write. Julia is a programming language that aims to be the best of both worlds – a development and production language at the same time. To test Julia's utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, the Model for Prediction Across Scales–Ocean (MPAS-Ocean), as well as a Python shallow water code. Three versions of the Julia shallow water code were created: for single-core CPU, graphics processing unit (GPU), and Message Passing Interface (MPI) CPU clusters. Comparing identical simulations revealed that our first version of the Julia model was 13 times faster than Python using NumPy, where both used an unthreaded single-core CPU. Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20× speed-up of the single-core CPU Julia model. The GPU-accelerated Julia code was almost identical in terms of performance to the MPI parallelized code on 64 processes, an unexpected result for such different architectures. Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts and ranges from 2× faster to 2× slower for higher processor counts. Our experience is that Julia development is fast and convenient for prototyping but that Julia requires further investment and expertise to be competitive with compiled codes. We provide advice on Julia code optimization for HPC systems.
A scalable density functional electronic code with Gaussian basis set, called UTEP-NRLMOL, is developed to perform simulations of molecular systems in the presence of the environment with particular attention to the memory requirements. In the electronic structure calculations, the memory and computation time are proportional to the number of atoms. Memory requirements for density functional calculations scale as N*N, where N is the number of atoms. While the recent advances in HPC offer platforms with large numbers of cores, the limited amount of memory available on a given node and poor scalability of the electronic structure codes hinder their efficient usage of these platforms. We have introduced new scaling and parallelization paradigms using MPI-3 shared-memory functionality combined with usage of sparse algebra and storage of matrices in sparse format. This extends the range of applicability of the UTEP-NRLMOL code to large systems over 10,000 atoms, or using up to 67,000 basis functions, and making use of HPC architectures using over 6,000 processors utilizing all available cores. We have also interfaced code with effective fragment potential and polarizable continuum model libraries. The code was used in simulations of several applications which are published in reputed scientific journals.
With the rising demand for high performance computing (HPC) and artificial intelligence (AI) systems, maintaining a stable and efficient power supply is increasingly critical. The HPC team at Idaho National Laboratory is spearheading efforts to seamlessly integrate HPC systems with nuclear reactors. This lightning talk explores one early strategy for managing power fluctuations using software-defined controls. To effectively harness nuclear reactors for power generation, control mechanisms are essential to address the slow load-following capabilities of reactors, which are typically around 5% per minute. While this rate is sufficient for many uses, large HPC systems can experience rapid power consumption changes by tens of megawatts when jobs start or stop running. A reactor could overproduce power and match the peak power rating for the HPC system, however when the system is not running a job or a job unexpectedly stops, the load-following of the system would be affected leading to power being wasted and the likelihood of power transient occurrences increases. Controlling the increase or decrease of power consumption on these systems at the same rate as the load-following of reactors is one piece of the puzzle to properly utilizing nuclear reactors as a power source for HPC systems.
Software developers must adapt to keep up with the changing capabilities of platforms so that they can utilize the power of High-Performance Computers (HPC), including exascale systems. OpenMP, a directive-based parallel programming model, allows developers to include directives to existing C, C++, or Fortran code to allow node level parallelism without compromising performance. This paper describes our CI/CD efforts to provide easy evaluation of the support of OpenMP across different compilers using existing testsuites and benchmark suites on HPC platforms. Our main contributions include (1) the set of a Continuous Integration (CI) and Continuous Development (CD) workflow that captures bugs and provides faster feedback to compiler developers, (2) an evaluation of OpenMP (offloading) implementations supported by AMD, HPE, GNU, LLVM, and Intel, and (3) evaluation of the quality of compilers across different heterogeneous HPC platforms. With the comprehensive testing through the CI/CD workflow, we aim to provide a comprehensive understanding of the current state of OpenMP (offloading) support in different compilers and heterogeneous platforms consisting of CPUs and GPUs from NVIDIA, AMD, and Intel.
HPC and AI are businesses necessitating growth for providing performance improvement with increased power and cooling. As HPC and AI computers and their supporting facilities approach utility scale with a frequency of technology innovation outpacing utility and construction timelines, understanding and designing for this trend has become critical. This paper will provide historical trend data for facilities and compute racks, relate power trend data to cooling technology capabilities, and reason through constraints impacting anticipated future power densities to aid the reader in futureproofing a facility’s power and cooling systems through the mid-2030’s.
We have developed some scripts and tools that compliment an existing open-source product (Podman, https://podman.io/). This enables podman to function as a high-performance container solution. This capability is similar to Shifter and Singularity (both developed at LBNL). The advantage of leveraging Podman is it has greater support in the broader community outside of HPC and has capabilities missing from many of the HPC-only solutions. The specific enhancements we wish to share our how to enable Podman to scale to very large HPC jobs and our integrations we have done to utilize high-performance interconnects and GPUs.
This User Guide is intended to provide new and current users the necessary information to effectively and efficiently login and utilize their Idaho National Laboratory (INL) High Performance Computing (HPC) account, as well as inform users of available HPC resources
Writing large amounts of data concurrently to stable storage is a typical I/O pattern of many HPC workflows. This pattern introduces high I/O overheads and results in increased storage space utilization especially for workflows that need to capture the evolution of data structures with high frequency as checkpoints. In this context, many applications, such as graph pattern matching, perform sparse updates to large data structures between checkpoints. For these applications, incremental checkpointing techniques that save only the differences from one checkpoint to another can dramatically reduce the checkpoint sizes, I/O bottlenecks, and storage space utilization. However, such techniques are not without challenges: it is non-trivial to transparently determine what data has changed since a previous checkpoint and assemble the differences in a compact fashion that does not result in excessive metadata. State-of-art data reduction techniques (e.g., compression and de-duplication) have significant limitations when applied to modern HPC applications that leverage GPUs: slow at detecting the differences, generate a large amount of metadata to keep track of the differences, and ignore crucial spatiotemporal checkpoint data redundancy. This paper addresses these challenges by proposing a Merkle tree-based incremental checkpointing method to exploit GPUs' high memory bandwidth and massive parallelism. Experimental results at scale show a significant reduction of the I/O overhead and space utilization of checkpointing compared with state-of-the-art incremental checkpointing and compression techniques.
High-performance computing systems rely upon scheduling algorithms to achieve high utilization. These schedulers rely upon user estimates of job resource requirements, such as runtime, to determine optimal scheduling of incoming jobs. These user estimates, however, are prone to error. To mitigate this error, significant research has been directed at providing better estimates of job runtime, usually employing machine learning techniques. These techniques are dependent upon the input features selected. Among the possible features is the primary application used by the job. In a survey of more than 20 papers directed at improving runtime prediction, only four included primary application as an input feature. We focus this investigation specifically on the value of adding primary application as an input feature, and find that it does improve model performance, especially for jobs with longer runtimes, though this improvement varies based on the application used. We recommend further research to determine the cause of this variability as well as an optimal strategy for employing a mixture of models both including and not including primary application as a feature.
High-performance computing systems rely upon scheduling algorithms to achieve high utilization. These schedulers rely upon user estimates of job resource requirements, such as runtime, to determine optimal scheduling of incoming jobs. These user estimates, however, are prone to error. To mitigate this error, significant research has been directed at providing better estimates of job runtime, usually employing machine learning techniques. These techniques are dependent upon the input features selected. Among the possible features is the primary application used by the job. In a survey of more than 20 papers directed at improving runtime prediction, only four included primary application as an input feature. We focus this investigation specifically on the value of adding primary application as an input feature, and find that it does improve model performance, especially for jobs with longer runtimes, though this improvement varies based on the application used. We recommend further research to determine the cause of this variability as well as an optimal strategy for employing a mixture of models both including and not including primary application as a feature.
Accurate degradation modeling is essential for predicting photovoltaic (PV) module performance, estimating longevity and informing design decisions. With degradation rates varying significantly by location, geospatial analysis is critical for PV and broader applications, such as agrivoltaics, weathering and environmental data analysis. This work presents PVDeg, an open-source tool designed for geospatial degradation analysis. PVDeg integrates meteorological data from global sources, including the National Solar Radiation Database (NSRDB) and Photovoltaic Geographical Information System (PVGIS), with degradation models. The toolkit enables users to customize geospatial workflows by integrating weather data, material parameters, and user-defined Python functions. It facilitates accelerated downloads of NSRDB and PVGIS datasets and optimizes geospatial point selection to preserve data density in regions of interest. Additionally, PVDeg provides a local database for storage and spatial queries, supporting large-scale analyses without the need for high-performance computing (HPC) resources. PVDeg provides a foundational workflow that extends its utility beyond PV applications, enabling researchers to analyze geospatial processes across discipline.
With the growing complexity of computational and experimental facilities, many scientific researchers are turning to machine learning (ML) techniques to analyze large scale ensemble data. With complexities such as multi-component workflows, heterogeneous machine architectures, parallel file systems, and batch scheduling, care must be taken to facilitate this analysis in a high performance computing (HPC) environment. Here, we present Merlin, a workflow framework to enable large ML-friendly ensembles of scientific HPC simulations. By augmenting traditional HPC with distributed compute technologies, Merlin aims to lower the barrier for scientific subject matter experts to incorporate ML into their analysis. As a producer–consumer workflow model, Merlin enables multi-machine, cross-batch job, dynamically allocated yet persistent workflows capable of utilizing surge-compute resources. Key features of Merlin are a flexible HPC-centric interface, low per-task overhead, multi-tiered fault recovery, and a hierarchical sampling algorithm that allows for $\mathscr{O}$(N) task execution and $\mathscr{O}$(N ln N) task queuing to ensembles of millions of tasks. In addition to Merlin’s design, we test the algorithm’s performance in an HPC center and demonstrate the ability to enqueue 40 million simulations in 100 s, with a 30 millisecond per-task overhead that is independent of ensemble size. Finally, we describe some example applications that Merlin has enabled on leadership-class HPC resources, such as the ML-augmented optimization of nuclear fusion experiments and the calibration of infectious disease models to study the progression of and possible mitigation strategies for COVID-19.
Managed under U.S. Department of Energy (DOE)-funded EVs@Scale Consortium, High-Power Electric Vehicle Charging Hub Integration Platform (eCHIP) project aims to design and develop a high-power, interoperable charging experimental platform to research, develop, and demonstrate the integration and control approaches for a DC distribution-based high-power charging (HPC) system. The eCHIP project addresses the crucial need to design and validate efficient, low-cost, reliable, and interoperable solutions for DC-coupled charging hub ('DC hub' for short). This report explains the design, development, and implementation process of the experimental platform for the DC hub. The utilization of DC distribution holds significant potential for enhancing the operation of an HPC station architecture. However, there are challenges establishing the DC hub, including interoperability, commoditization, distributed energy resource integration, stability, DC protection, and lack of common system level controllers. To address these challenges, a testing setup is required that accommodates commercial off-the-shelf (COTS) products as well as novel, in-house designed solutions to evaluate different use cases at rated power and voltage levels.
Mixed-precision algorithms have been proposed as a way for scientific computing to benefit from some of the gains seen for AI on recent high performance computing (HPC) platforms. A few applications dominated by dense matrix operations have seen substantial speedups by utilizing low precision formats such as FP16. However, a majority of scientific simulation applications are memory bandwidth limited. Beyond preliminary studies, the practical gain from using mixed-precision algorithms on a given high-performance computing (HPC) system is largely unclear. The High Performance GMRES Mixed Precision (HPG-MxP) benchmark has been proposed to measure the useful performance of a HPC system on sparse matrix-based mixed-precision applications. In this work, we present an implementation of the HPG-MxP benchmark for an exascale system and describe our algorithm enhancements. We show for the first time a speedup of 1.6x using a combination of double- and single-precision keeping the same residual level on modern GPU-based supercomputers.
Productivity Frameworks for HPC will include container recipes, build recipes, continuous integration scripts, and other software aimed at testing the portability of containerized HPC software across platforms and interconnects. In particular, it tests the utility of bind-mounting at the MPI layer (rather than the underlying fabric layer) to leverage a standardized protocol and avoid various technical debt and vendor lock-in. Since MPIs are often ABI-incompatible, trampolines such as the open-source Wi4MPI will be tested when such cases arise.
The attention mechanism has become foundational for remarkable AI breakthroughs since the introduction of the Transformer, driving the demand for increasingly longer context to power frontier models such as large-scale reasoning language models and high-resolution image/video generators. However, its quadratic computational and memory complexities present substantial challenges. Current state-of-the-art parallel attention methods, such as ring attention, are widely adopted for long-context training but utilize a point-to-point communication strategy that fails to fully exploit the capabilities of modern HPC network architectures. In this work, we propose ringX, a scalable family of parallel attention methods optimized explicitly for HPC systems. By enhancing workload partitioning, refining communication patterns, and improving load balancing, ringX achieves up to 3.4 × speedup compared to conventional ring attention on the Frontier supercomputer. Optimized for both bi-directional and causal attention mechanisms, ringX demonstrates its effectiveness through training benchmarks of a Vision Transformer (ViT) on a climate dataset and a Generative Pre-Trained Transformer (GPT) model, Llama3 8B. Our method attains an end-to-end training speedup of approximately 1.5 × in both scenarios. To our knowledge, the achieved 38% model FLOPs utilization (MFU) for training Llama3 8B with a 1M-token sequence length on 4,096 GPUs represents one of the highest training efficiencies reported for long-context learning on HPC systems. Our code implementation is available at https://github.com/jqyin/ringX-attention.