Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “LDMS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

LDMS-GPU: Lightweight Distributed Metric Service (LDMS) for NVIDIA GPGPUs

GPUs are now a fundamental accelerator for many high-performance computing applications. They are viewed by many as a technology facilitator for the surge in fields like machine learning and Convolutional Neural Networks. To deliver the best performance on a GPU, we need to create monitoring tools to ensure that we optimize the code to get the most performance and efficiency out of a GPU. Since NVIDIA GPUs are currently the most commonly implemented in HPC applications and systems, NVIDIA tools are the solution for performance monitoring. The Light-Weight Distributed Metric System (LDMS) at Sandia is an infrastructure widely adopted for large-scale systems and application monitoring. Sandia has developed CPU application monitoring capability within LDMS. Therefore, we chose to develop a GPU monitoring capability within the same framework. In this report, we discuss the current limitations in the NVIDIA monitoring tools, how we overcame such limitations, and present an overview of the tool we built to monitor GPU performance in LDMS and its capabilities. Also, we discuss our current validation results. Most of the performance counter results are the same in both vendor tools and our tool when using LDMS to collect these results. Furthermore, our tool provides these statistics during the entire runtime of the tool as a time series and not just aggregate statistics at the end of the application run. This allows the user to see the progress of the behavior of the applications during their lifetime.

97 MATHEMATICS AND COMPUTING↗

NERSC_Lightweight Distributed Metric Service (NERSC_LDMS) v4.4.2

Miscellany This LDMS Loftsman/Helm Chart horizontally scales LDMS daemons in order to achieve a 1Hz sample rate from over 5,000 nodes, collecting 38k metrics per minute on Perlmutter. This LMDS Configuration relies on already running `ldmsd` producers running on nodes, which produce metrics via sampler plugins. The Helm chart distributes the collection of metrics from producer acrross many aggregator and storage `ldmsd` daemons, ensuring no damon is overloaded and data loss is avoided.

Stile, John [Lawrence Berkeley National Laboratory↗

LDMS New Features for Deployment in Advanced Environments and Feedback for Operations

We describe how LDMS is being used to collect application data concurrent with system data and how the low-latency availability of this data for analysis can be used for real-time data analysis and feedback in order to support efficient, resilient, and reliable system operations. Finally, we will describe current related research areas.

Brandt, James Michael [Sandia National Laboratorie↗

LDMS Job Summary Pipeline

Explore the source record for details and available documents.

Schwaller, Benjamin [Sandia National Laboratories ↗