Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “HPC Access”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

UnifyFS: A User-level Shared File System for Unified Access to Distributed Local Storage

We introduce UnifyFS, a user-level file system that aggregates node-local storage tiers available on high performance computing (HPC) systems and makes them available to HPC applications under a unified namespace. UnifyFS employs transparent I/O interception, so it does not require changes to application code and is compatible with commonly used HPC I/O libraries. The design of UnifyFS supports the predominant HPC I/O workloads and is optimized for bulk-synchronous I/O patterns. Furthermore, UnifyFS provides customizable file system semantics to flexibly adapt its behavior for diverse I/O workloads and storage devices. In this paper, we discuss the unique design goals and architecture of UnifyFS and evaluate its performance on a leadership-class HPC system. In our experimental results, we demonstrate that UnifyFS exhibits excellent scaling performance for write operations and can improve the performance of application checkpoint operations by as much as 3× versus a tuned configuration.

Brim, Michael↗

Enabling the Write-Once Read-Many (WORM) Dashboard and Web Interface for HPC Data Storage

Idaho National Laboratory (INL), supported by the Department of Energy Office of Nuclear Energy (DOE-NE) through the Nuclear Science User Facilities, provides access to supercomputer systems and data storage along with support staff for system management, software installation, cybersecurity, and user support to the broader DOE-NE user community. A key component of this is providing long-term or perpetual data storage as part of efforts to improve reproducibility in modeling and simulation in addition to supporting simulations with regulatory compliance requirements. This milestone details the deployment of a new type of storage option for users requiring tamper resistant perpetual storage of data and the associated support for data control and release via the data storage web interface. The perpetual data storage system is integrated within the Open OnDemand framework deployed at INL and is called the ?WORM? in recognition of its write-once read-many characteristics.

99 GENERAL AND MISCELLANEOUS↗

The Efficiency and the Scalability of an Explicit Operator on an IBM POWER4 System

We present an evaluation of the efficiency and the scalability of an explicit CFD operator on an IBM POWER4 system. The POWER4 architecture exhibits a common trend in HPC architectures: boosting CPU processing power by increasing the number of functional units, while hiding the latency of memory access by increasing the depth of the memory hierarchy. The overall machine performance depends on the ability of the caches-buses-fabric-memory to feed the functional units with the data to be processed. In this study we evaluate the efficiency and scalability of one explicit CFD operator on an IBM POWER4. This operator performs computations at the points of a Cartesian grid and involves a few dozen floating point numbers and on the order of 100 floating point operations per grid point. The computations in all grid points are independent. Specifically, we estimate the efficiency of the RHS operator (SP of NPB) on a single processor as the observed/peak performance ratio. Then we estimate the scalability of the operator on a single chip (2 CPUs), a single MCM (8 CPUs), 16 CPUs, and the whole machine (32 CPUs). Then we perform the same measurements for a chache-optimized version of the RHS operator. For our measurements we use the HPM (Hardware Performance Monitor) counters available on the POWER4. These counters allow us to analyze the obtained performance results.

Frumkin, Michael↗

SOMA: Observability, monitoring, and in situ analytics for exascale applications

With the rise of exascale systems and large, data-centric workflows, the need to observe and analyze high performance computing (HPC) applications during their execution is becoming increasingly important. HPC applications are typically not designed with online monitoring in mind, therefore, the observability challenge lies in being able to access and analyze interesting events with low overhead while seamlessly integrating such capabilities into existing and new applications. We explore how our service-based observation, monitoring, and analytics (SOMA) approach to collecting and aggregating both application-specific diagnostic data and performance data addresses these needs. Furthermore, we present our SOMA framework and demonstrate its viability with LULESH, a hydrodynamics proxy application. Then we focus on Astaroth, a multi-GPU library for stencil computations, highlighting the integration of the TAU and APEX performance tools and SOMA for application and performance data monitoring.

97 MATHEMATICS AND COMPUTING↗

xGFabric: Coupling Sensor Networks and HPC Facilities with Private 5G Wireless Networks for Real-Time Digital Agriculture

Advanced scientific applications require coupling distributed sensor networks with centralized high-performance computing facilities. Citrus Under Protective Screening (CUPS) exemplifies this need in digital agriculture, where citrus research facilities are instrumented with numerous sensors monitoring environmental conditions and detecting protective screening damage. CUPS demands access to computational fluid dynamics codes for modeling environmental conditions and guiding real-time interventions like water application or robotic repairs. These computing domains have contrasting properties: sensor networks provide low-performance, limited-capacity, unreliable data access, while high-performance facilities offer enormous computing power through high-latency batch processing. Private 5G networks present novel capabilities addressing this challenge by providing low latency, high throughput, and reliability necessary for near-real-time coupling of edge sensor networks with HPC simulations. This work presents xGFabric, an end-to-end system coupling sensor networks with HPC facilities through Private 5G networks. The prototype connects remote sensors via 5G network slicing to HPC systems, enabling real-time digital agriculture simulation.

Digital Agriculture↗

MDLoader: A Hybrid Model-Driven Data Loader for Distributed Graph Neural Network Training

Scalable data management is essential for processing large scientific dataset on HPC platforms for distributed deep learning. In-memory distributed storage is preferred for its speed, enabling rapid, random, and frequent data access required by stochastic optimizers. Processes use one-sided or collective communication to fetch remote data, with optimal performance depending on (i) dataset characteristics, (ii) training scale, and (iii) interconnection network. Empirical analysis shows collective communication excels with larger mini-batch sizes and/or fewer processes, whereas one-sided communication outperforms at larger scales. We propose MDLoader, a hybrid in-memory data loader for distributed graph neural network training. MDLoader features a model-driven performance estimator that dynamically selects between one-sided and collective communication at the beginning of training using Tree of Parzen Estimators (TPE). Evaluations on NERSC Perlmutter and OLCF Summit show MDLoader outperforms single-backend loaders by up to 2.83 × and predicts the suitable communication method with 96.3% (Perlmutter) and 94.3% (Summit) success rate.

Bae, Jonghyun↗

NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

28 NREL Stratus - Enabling Workflows to Fuse Data Streams, Modeling, Simulation, and Machine Learning: Preprint

Integrating cloud services into advanced computing facilities provides significant new capabilities over focusing solely on traditional high performance computing (HPC) workloads. This brings complementary capabilities as well as enabling new focused roles for HPC. They are especially potent for workflows that fuse data streams, modeling and simulation ('modsim') and machine learning. A key challenge to adopting a hybrid edge-cloud-HPC model is to align optimal capability, data, and user intent on the right resources for each step in a workflow.?The NREL Stratus service provides a basis for this: Stratus layers capabilities needed to make?cloud services accessible to a lab-based scientific community on commercial offerings, and; currently supports upwards of 200 projects ranging from IOT integration to traditional modeling and simulation. This provides a real-world inventory of scientific workflow elements. A growing knowledge base enables placing these elements appropriately between the edge, cloud, and traditional HPC. This paper outlines a vision via reference architecture and the application of that architecture in a typical workflow highlighting multiple components: sensor data intake, cleaning and transforming (edge/cloud suitable); generation of synthetic data through modsim, computationally heavy ML training and hyperparameter optimization (HPC suitable), and; inference and deployment (cloud ideal). Every step in such a workflow involves a cost-benefit analysis regarding the data movement, computational efficiency, availability, latency, and resource capabilities. The reference architecture and examples outlined allow for understanding new opportunities in the context of emerging workflows that combine IOT, cloud, and HPC to bolster scientific productivity.

AI↗

Lustre Unveiled: Evolution, Design, Advancements, and Current Trends

The Lustre filesystem serves as a vital element in high-performance parallel storage, meeting the rising demands of scientific, research, and enterprise environments. Widely deployed across HPC environments, ranging from small-scale applications in AI/ML, to domains like oil and gas, drug discovery, and meteorology, and manufacturing, Lustre addresses the universal challenge of efficiently accessing vast and ever-increasing volumes of data. Lustre is the filesystem of choice on six out of the top 10 fastest supercomputers in the world today, over 65% of the top 100, and also for over 60% of the top 500. Despite its widespread popularity, there is a lack of a complete and up-to-date reference, covering Lustre’s evolution, design, and various advancements made over the years. In this journal, we aim to fill this gap by providing a comprehensive journey of Lustre, including its history with significant contributions to HPC, detailed architecture and design elements, exploration of advancements added through its evolution, and future directions. Additionally, we present a comparison of Lustre with other prominent storage technologies of the era. To illustrate the current state of Lustre, we analyze several filesystem trends, including utilization, performance, and usage patterns on Orion, the Lustre filesystem on the first exascale supercomputer Frontier. We hope that this journal serves as a comprehensive educational reference for the current and future generations interested in HPC filesystem storage aspects.

97 MATHEMATICS AND COMPUTING↗

Evaluating the potential of disaggregated memory systems for HPC applications

Summary Disaggregated memory is a promising approach that addresses the limitations of traditional memory architectures by enabling memory to be decoupled from compute nodes and shared across a data center. Cloud platforms have deployed such systems to improve overall system memory utilization, but performance can vary across workloads. High‐performance computing (HPC) is crucial in scientific and engineering applications, where HPC machines also face the issue of underutilized memory. As a result, improving system memory utilization while understanding workload performance is essential for HPC operators. Therefore, learning the potential of a disaggregated memory system before deployment is a critical step. This paper proposes a methodology for exploring the design space of a disaggregated memory system. It incorporates key metrics that affect performance on disaggregated memory systems: memory capacity, local and remote memory access ratio, injection bandwidth, and bisection bandwidth, providing an intuitive approach to guide machine configurations based on technology trends and workload characteristics. We apply our methodology to analyze thirteen diverse workloads, including AI training, data analysis, genomics, protein, fusion, atomic nuclei, and traditional HPC bookends. Our methodology demonstrates the ability to comprehend the potential and pitfalls of a disaggregated memory system and provides motivation for machine configurations. Our results show that eleven of our thirteen applications can leverage injection bandwidth disaggregated memory without affecting performance, while one pays a rack bisection bandwidth penalty and two pay the system‐wide bisection bandwidth penalty. In addition, we also show that intra‐rack memory disaggregation would meet the application's memory requirement and provide enough remote memory bandwidth.

Ding, Nan↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

An Efficient Storage-Driven Machine Learning Model for Performance in the Era of Multimodal Scientific Data

Scientific workflows are increasingly relying on machine learning (ML), simulation, and hybrid techniques to predict, understand, and optimize the behavior of complex experiments. High-performance computing has greatly improved researchers’ ability to acquire diverse data modalities in these workflows. Recent studies suggest that the performance of machine learning models can be improved by integrating data from various sources. Unfortunately, these workloads pose unprecedent pressure on the network storage to meet the demands associated with accessing these multimodal data. To mitigate the impact of intensive IO, we propose a solution that utilizes a multi-tier High-Performance Computing (HPC) distributed storage and data processing framework, placing computation where the data resides for better performance. By adopting this project, the scientific community will gain new opportunities to explore multimodal storage-driven possibilities, integrating multiple scientific data sources with advanced streaming frameworks. Additionally, our framework effectively utilizes computing resources and bridges the gaps identified by HPC experts. Our proposed approach tackles scalability and persistence challenges by leveraging native persistency, which has posed difficulties in traditional approaches. Furthermore, we seek to enhance fault-tolerance and load-balance of computations by leveraging real-time streaming in diverse scientific computing environments, thereby propelling advanced scientific computing research into the next generation.

97 MATHEMATICS AND COMPUTING↗

Enabling discovery data science through cross-facility workflows

Experimental and observational instruments for scientific research (such as light sources, genome sequencers, accelerators, telescopes and electron microscopes) increasingly require High Performance Computing (HPC) scale capabilities for data analysis and workflow processing. Next-generation instruments are being deployed with higher resolutions and faster data capture rates, creating a big data crunch that cannot be handled by modest institutional computing resources. Often these big data analysis pipelines also require near real-time computing and have higher resilience requirements than the simulation and modeling workloads more traditionally seen at HPC centers. While some facilities have enabled workflows to run at a single HPC facility, there is a growing need to integrate capabilities across HPC facilities to enable cross-facility workflows, either to provide resilience to an experiment, increase analysis throughput capabilities, or to better match a workflow to a particular architecture. In this paper we describe the barriers to executing complex data analysis workflows across HPC facilities and propose an architectural design pattern for enabling scientific discovery using cross-facility workflows that includes orchestration services, application programming interfaces (APIs), data access and co-scheduling.

Antypas, Katerina B.↗

HAM: Hotspot-Aware Manager for Improving Communications with 3D-Stacked Memory

merging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. %and an associated hotspot-aware page policy. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51\% and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81\% in average (up to 34.28\%) and power savings of 35.07\% over a standard 3D-stacked memory.

Wang, Xi↗

Optimization and Portability of a Fusion OpenACC-based FORTRAN HPC Code from NVIDIA to AMD GPUs

NVIDIA has been the main provider of GPU hardware in HPC systems for over a decade. Most applications that benefit from GPUs have thus been developed and optimized for the NVIDIA software stack. Recent exascale HPC systems are, however, introducing GPUs from other vendors, e.g. with the AMD GPU-based OLCF Frontier system just becoming available. AMD GPUs cannot be directly accessed using the NVIDIA software stack, and require a porting effort by the application developers. This paper provides an overview of our experience porting and optimizing the CGYRO code, a widely-used fusion simulation tool based on FORTRAN with OpenACC-based GPU acceleration. While the porting from the NVIDIA compilers was relatively straightforward using the CRAY compilers on the AMD systems, the performance optimization required more fine-tuning. In the optimization effort, we uncovered code sections that had performed well on NVIDIA GPUs, but were unexpectedly slow on AMD GPUs. After AMD-targeted code optimizations, performance on AMD GPUs has increased to meet our expectations. Modest speed improvements were also seen on NVIDIA GPUs, which was an unexpected benefit of this exercise.

Sfiligoi, Igor↗

Evaluating HPC Kernels for Processing in Memory

Memory subsystems contribute significantly to the performance and energy efficiency of high-performance computing (HPC) applications. Traditional memory technologies with conventional organization (e.g., DRAM) are struggling to keep up with the increasing memory requirements of modern applications. Techniques such as multilayer cache hierarchy and out-of-order execution are still falling short of mitigating the penalty incurred by memory accesses. Processing-in-memory (PIM), which involves moving memory-intensive kernels to memory for execution instead of bringing the data to the processing unit, is emerging as a promising technique. PIM has recently received traction among computer architecture researchers, and the increasing research activity surrounding this technique indicates its potential to alleviate main memory performance bottlenecks. In this paper, we characterize and identify memory-intensive HPC kernels, perform a first-order evaluation of the PIM technique for selected HPC kernels, quantify performance deviation, and analyze the key factors that affect PIM efficiency.

Asifuzzaman, Kazi↗

Proxy Applications for Converged Workloads: DMC LDRD Initiative

Modern scientific applications are complicated and require coordination of several components. Proxy application driven software-hardware co-design plays a vital role in driving innovation among the developments of applications, software infrastructure and hardware architecture. Proxy applications are self-contained and simplified codes that are intended to model the performance-critical computations within applications. Applications executing on modern High Performance Computing (HPC) systems are susceptible to network congestion, insufficient memory bandwidth within and across compute nodes, and inadvertent loss of performance due to bugs and unoptimized programming models. Modern numerical simulations and machine learning models play a critical role in studying physical phenomenon under myriad uncertainties. Such applications often exhibit irregular computation and memory accesses at specific regions of the application code, which can contribute to various performance bottlenecks at scale. To mitigate such issues and prepare the next generation hardware for a variety of computation and data movement contingencies, a well-known practice is to consider "proxy" applications as representative motifs for various classes of scientific applications. While there is disagreement in the HPC community on the mechanisms of construction of the proxy applications, there is a strong consensus on their positive impact in co-design. Proxy Applications for Converged Workloads (PACER) is about facilitating software-hardware co-design through proxy applications with the goal of improving the performance of converged science workflows on heterogeneous systems.

97 MATHEMATICS AND COMPUTING↗