Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interactive HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Enabling Seamless Transitions from Experimental to Production HPC for Interactive Workflows

The evolving landscape of scientific computing requires seamless transitions from experimental to production HPC environments for interactive workflows. This paper presents a structured transition pathway developed at OLCF that bridges the gap between development testbeds and production systems. We address both technological and policy challenges, introducing frameworks for data streaming architectures, secure service interfaces, and adaptive resource scheduling for time-sensitive workloads and improved HPC interactivity. Our approach transforms traditional batch-oriented HPC into a more dynamic ecosystem capable of supporting modern scientific workflows that require near real-time data analysis, experimental steering, and cross-facility integration.

Etz, Brian [ORNL] (ORCID:0000000208554863)↗

Towards Interactive, Reproducible Analytics at Scale on HPC Systems

The growth in scientific data volumes has resulted in a need to scale up processing and analysis pipelines using High Performance Computing (HPC) systems. These workflows need interactive, reproducible analytics at scale. The Jupyter platform provides core capabilities for interactivity but was not designed for HPC systems. In this paper, we outline our efforts that bring together core technologies based on the Jupyter Platform to create interactive, reproducible analytics at scale on HPC systems. Our work is grounded in a real world science use case-applying geophysical simulations and inversions for imaging the subsurface. Our core platform addresses three key areas of the scientific analysis workflow-reproducibility, scalability, and interactivity. We describe our implemention of a system, using Binder, Science Capsule, and Dask software. We demonstrate the use of this software to run our use case and interactively visualize real-Time streams of HDF5 data.

containers↗

Clippy

Clippy (CLI + PYthon) is a Python language interface to HPC resources. Precompiled binaries that execute on HPC systems are exposed as methods to a dynamically-created Clippy Python object, where they present a familiar interface to researchers, data scientists, and others. Clippy allows these users to interact with HPC resources in an easy, straightforward environment - at the REPL, for example, or within a notebook - without the need to learn complex HPC behavior and arcane job submission commands.

Bromberger, SethA.↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

In situ feature analysis for large-scale multiphase flow simulations

The study of multiphase flow is essential for designing chemical reactors such as fluidized bed reactors (FBR), as a detailed understanding of hydrodynamics is critical for optimizing reactor performance and stability. An FBR allows scientists to conduct different types of chemical reactions involving multiphase materials, especially interaction between gas and solids. During such complex chemical processes, the formation of void regions in the reactor, generally termed as bubbles, is an important phenomenon. The study of these bubbles has a deep implication in predicting the reactor’s overall efficiency. But physical experiments needed to understand bubble dynamics are costly and non-trivial due to the technical difficulties involved and harsh working conditions of the reactors. Therefore, to study such chemical processes and bubble dynamics, a state-of-the-art computational simulation MFIX-Exa is being developed. Despite the proven accuracy of MFIX-Exa in modeling bubbling phenomena, the large-scale output data prohibits the use of traditional post hoc analysis capabilities in both storage and I/O time. Herein, to address these issues and allow the application scientists to explore the bubble dynamics in an efficient and timely manner, we have developed an end-to-end analytics pipeline that enables in situ detection of bubbles, followed by a flexible post hoc visual exploration methodology of bubble dynamics. The proposed method enables interactive analysis of bubbles, along with quantification of several bubble characteristics, enabling experts to understand the bubble interactions in detail. Positive feedback from the experts has indicated the efficacy of the proposed approach for exploring bubble dynamics in very-large-scale multiphase flow simulations.

97 MATHEMATICS AND COMPUTING↗

Usage Pattern Analysis for the Summit Login Nodes

High performance computing (HPC) users interact with Summit through dedicated gateways, also known as login nodes. The performance and stability of these login nodes can have a significant impact on the user experience. In this study, the performance and stability of Summit’s five login nodes are evaluated by analyzing the log data from 2020 and 2021. The analysis focuses on the computing capability (CPU average load, users and tasks) and the storage performance, along with the associated job scheduler activity. The outcome of this study can serve as the foundation of a predictive modeling framework that enables the system admin of an HPC system to preemptively deploy countermeasures before the onset of a system failure.

Eiffert, Brett↗

Community Requirements Meta-Analysis: Characterizing Needs and Opportunities for HPDF

This High Performance Data Facility (HPDF) Project is creating a new scientific user facility to provide advanced infrastructure for data-intensive science, supporting the DOE’s Office of Science (SC) community. HPDF’s mission is to enable and accelerate scientific discovery by delivering state-of-the-art data management infrastructure, capabilities, and tools. This meta-analysis examines the needs of the breadth of the SC community, captured in publicly available community reports or mission documents. The meta-analysis identifies and provides initial characterization of fifteen core requirements for the HPDF Project team to consider during the conceptual design phase. The fifteen requirements illustrate how scientific work among SC communities requires modern, seamless user experiences across the ASCR Ecosystem to advance the use of large volumes of heterogeneous data. The scientific community requires support for the missing middle of compute between local and HPC to interactively and collaboratively use growing datasets. Data producers and end users will benefit from enhanced data catalogs and portals that improve data access through advanced search of well curated data. The fifteen requirements are examined here organized across five themes for discussion. Examples in each theme illustrate the array of scientific needs that convey the important role that the fully realized and operational High Performance Data Facility will be able to play as an integral part of the evolving ASCR Ecosystem. Our amalgamated data tables from ESnet reports demonstrate ranges to the volumes of data HPDF must be concerned with, but limitations are inherent to this meta-analysis (see Key Challenges & Limitations). Feedback and validation of these requirements along with additional details and emergent community requirements will be gathered through user research and design activities.

97 MATHEMATICS AND COMPUTING↗

Virtual Engineering Software Framework for Integrated Biomass Conversion Modeling

This presentation covers the design and implementation of a software tool to systematically connect computational models of unit operations to simulate an integrated process of low-temperature conversion of biomass to fuel. This virtual engineering (VE) software was designed with the overarching goal of connecting unit models written in various programming languages and requiring different computational resources within a single, flexible framework. The models and features currently considered for the VE library include mechanistic models for pretreatment, enzymatic hydrolysis, and aerobic bioreaction; high-fidelity computational fluid dynamics (CFD) simulations for enzymatic hydrolysis and aerobic bioreaction; and the capability to perform techno-economic analyses (TEA) using Aspen Plus, a commercial software package. The CFD models require access to high-performance computing (HPC) resources, so in addition to handling multiple programming languages and interfaces, the VE software must also be capable of interacting with an HPC scheduler to submit, run, and post-process jobs. Using the Python programming language, a new VE software package has been developed that contains functionality to manage the input-output communication between various unit models, schedule simulations to run on NREL's HPC and analyze those results, and interface with existing TEA software workflows. A Jupyter-notebook GUI was also created to solicit user input and provide documentation. In cases where multiple models for a particular unit-operation exist, selection between models is accomplished through a simple checkbox, with the appropriate inputs and outputs being parsed and converted seamlessly in the background. Each operation makes use of a different programming language, but the flow of information from pretreatment to enzymatic hydrolysis to bioreaction is managed with an intuitive, centralized file-communication strategy. In this talk, the programming approach and implementation details of the notebook are presented for multiple possibilities of the conversion process, including a demonstration of the ability to manage HPC resources. Additionally, an example of a sensitivity study of treatment parameters governing the overall conversion outcome is shown which highlights the ease of defining new problems using the VE Notebook workflow and leads into a discussion of ongoing work to enable outer-loop optimization studies.

biofuel↗

ChatBLAS: The First AI-Generated and Portable BLAS Library

We present ChatBLAS, the first AI-generated and portable Basic Linear Algebra Subprograms (BLAS) library on different CPU/GPU configurations. The purpose of this study is (i) to evaluate the capabilities of current large language models (LLMs) to generate a portable and HPC library for BLAS operations and (ii) to define the fundamental practices and criteria to interact with LLMs for HPC targets to elevate the trustworthiness and performance levels of the AI-generated HPC codes. The generated C/C++ codes must be highly optimized using device-specific solutions to reach high levels of performance. Additionally, these codes are very algorithm-dependent, thereby adding an extra dimension of complexity to this study. We used OpenAI’s LLM ChatGPT and focused on vector-vector BLAS level-1 operations. ChatBLAS can generate functional and correct codes, achieving high-trustworthiness levels, and can compete or even provide better performance against vendor libraries.

Valero Lara, Pedro↗

A Visual Comparison of Silent Error Propagation

High-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. Here, to accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution.

97 MATHEMATICS AND COMPUTING↗

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization↗

DXT Explorer v0.1

DTX Explorer is a tool to generate interactive data visualizations of Darshan I/O traces collected from HPC applications. Its goal is to provide an easy and interactive way for researchers and developers to explore their application's I/O behavior and detect possible I/O bottlenecks that are impacting performance.

Bez, Jean Luca↗

DXT Explorer v2.0

DTX Explorer is a tool to generate interactive data visualizations of Darshan I/O traces collected from HPC applications. Its goal is to provide an easy and interactive way for researchers and developers to explore their application's I/O behavior and detect possible I/O bottlenecks that are impacting performance.

Bez, JeanLuca↗

Integrating Artificial Intelligence into Science Gateways

Science gateways are altering the manner in which people interact with high performance computing (HPC) by providing a web browser based interface to advanced computing platforms. In particular, science gateways lower the barrier to using HPC by simplifying the process of submitting workloads to such systems and by offloading the efforts required to use HPC to the maintainers of the system. While science gateways decrease the time-to-science that comes with using such advanced systems, progress can still be made in improving the user's experience. In this paper we explore two strategies for integrating artificial intelligence tools commonly found in non-HPC service workflows: voice activated assistants and chatbots. Since August 2021, the HPC group at Idaho National Laboratory answers an average of 581 support tickets per month of which a large percentage could be addressed via these two strategies. This work defines the key capabilities that an HPC voice activated assistant and chatbot would need to address for a userbase consisting of largely non-expert users as well as a design for integration into the Open OnDemand science gateway.

97 MATHEMATICS AND COMPUTING↗

Understanding Interactive and Reproducible Computing With Jupyter Tools at Facilities

Increasingly Jupyter tools are being adopted and incorporated into High Performance Computing (HPC) and scientific user facilities. Adopting Jupyter tools enables more interactive and reproducible computational work at facilities across data life cycles. As the volume, variety, and scope of data grow, scientists need to be able to analyze and share results in user friendly ways. Human-centered research highlights design challenges around computational notebooks, and our qualitative user study shifts focus to better characterize how Jupyter tools are being used in HPC and science user facilities today. We conducted twenty-nine interviews, and obtained 103 survey responses from NERSC Jupyter users, to better understand the increasing role of interactive computing tools in DOE sponsored scientific work. We examine a range of issues that emerge using and supporting Jupyter in HPC ecosystems, including: how Jupyter is being used by scientists in HPC and user facility ecosystems; how facilities are purposefully supporting Jupyter in their ecosystems; feedback NERSC users have about the facility’s deployment, and, discuss features NERSC indicated would be helpful. We offer a variety of takeaways for staff supporting Jupyter at facilities, Project Jupyter and related open source communities, and funding agencies supporting interactive computing work.

97 MATHEMATICS AND COMPUTING↗

VISILIENCE: An Interactive Visualization Framework for Resilience Analysis using Control-Flow Graph

Soft errors have become one of the major concerns for the error resilience of HPC applications, as those errors can cause HPC applications to generate serious outcomes such as Silent Data Corruptions (SDCs). A large body of approaches has been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of the analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to conduct the resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of details, which may create hurdles to compare and explore the resilience of HPC program execution. To this end, we present VISILIENCE, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, VISILIENCE leverages an effective visualization approach Control Flow Graph (CFG) to present a function execution. In addition, three widely-used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly embedded into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework VISILIENCE.

Jiang, Hailong↗

A Novel use of Direct Simulation Monte-Carlo to Model Dynamics of COVID-19 Pandemic Spread

In this report, we evaluate a novel method for modeling the spread of COVID-19 pandemic. In this new approach we leverage methods and algorithms developed for fully-kinetic plasma physics simulations using Particle-In-Cell (PIC) Direct Simulation Monte-Carlo (DSMC) models. This approach then leverages Sandia-unique simulation capabilities, and High-Performance Computer (HPC) resources and expertise in particle-particle interactions using stochastic processes. Our hypothesis is that this approach would provide a more efficient platform with assumptions based on physical data that would then enable the user to assess the impact of mitigation strategies and forecast different phases of infection. This work addresses key scientific questions related to the assumptions this new approach must make to model the interactions of people using algorithms typically used for modeling particle interactions in physics codes (kinetic plasma, gas dynamics). The model developed uses rational/physical inputs while also providing critical insight; the results could serve as inputs to, or alternatives for, existing models. The model work presented was developed over a four-week time frame, thus far showing promising results and many ways in which this model/approach could be improved. This work is aimed at providing a proof-of-concept for this new pandemic modeling approach, which could have an immediate impact on the COVID-19 pandemic modeling, while laying a basis to model future pandemic scenarios in a manner that is timely and efficient. Additionally, this new approach provides new visualization tools to help epidemiologists comprehend and articulate the spread of this and other pandemics as well as a more general tool to determine key parameters needed in order to better predict pandemic modeling in the future. In the report we describe our model for pandemic modeling, apply this model to COVID-19 data for New York City (NYC), assess model sensitivities to different inputs and parameters and , finally, propagate the model forward under different conditions to assess the effects of mitigation and associated timing. Finally, our approach will help understand the role of asymptomatic cases, and could be extended to elucidate the role of recovered individuals in the second round of the infection, which is currently being ignored.

59 BASIC BIOLOGICAL SCIENCES↗