Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interactive HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Novel use of Direct Simulation Monte-Carlo to Model Dynamics of COVID-19 Pandemic Spread

In this report, we evaluate a novel method for modeling the spread of COVID-19 pandemic. In this new approach we leverage methods and algorithms developed for fully-kinetic plasma physics simulations using Particle-In-Cell (PIC) Direct Simulation Monte-Carlo (DSMC) models. This approach then leverages Sandia-unique simulation capabilities, and High-Performance Computer (HPC) resources and expertise in particle-particle interactions using stochastic processes. Our hypothesis is that this approach would provide a more efficient platform with assumptions based on physical data that would then enable the user to assess the impact of mitigation strategies and forecast different phases of infection. This work addresses key scientific questions related to the assumptions this new approach must make to model the interactions of people using algorithms typically used for modeling particle interactions in physics codes (kinetic plasma, gas dynamics). The model developed uses rational/physical inputs while also providing critical insight; the results could serve as inputs to, or alternatives for, existing models. The model work presented was developed over a four-week time frame, thus far showing promising results and many ways in which this model/approach could be improved. This work is aimed at providing a proof-of-concept for this new pandemic modeling approach, which could have an immediate impact on the COVID-19 pandemic modeling, while laying a basis to model future pandemic scenarios in a manner that is timely and efficient. Additionally, this new approach provides new visualization tools to help epidemiologists comprehend and articulate the spread of this and other pandemics as well as a more general tool to determine key parameters needed in order to better predict pandemic modeling in the future. In the report we describe our model for pandemic modeling, apply this model to COVID-19 data for New York City (NYC), assess model sensitivities to different inputs and parameters and , finally, propagate the model forward under different conditions to assess the effects of mitigation and associated timing. Finally, our approach will help understand the role of asymptomatic cases, and could be extended to elucidate the role of recovered individuals in the second round of the infection, which is currently being ignored.

59 BASIC BIOLOGICAL SCIENCES↗

Charliecloud 101

This workshop will provide participants with background and hands-on experience to use basic containers for HPC applications. We will discuss what containers are, why they matter for HPC, and how they work. We’ll give an overview of Charliecloud, the unprivileged container solution from HPC Division, and walk participants through installing it on their own compute resource. Participants will build toy containers and a real HPC application, and then run them in parallel on an HPC Division cluster. This will be a highly interactive workshop with lots of Q&A.

97 MATHEMATICS AND COMPUTING↗

Investigating User Experiences with Data Abstractions on High Performance Computing Systems

Scientific exploration generates expanding volumes of data that commonly require High Performance Computing (HPC) systems to facilitate research. HPC systems are complex ecosystems of hardware and software that frequently are not user friendly. The Usable Data Abstractions (UDA) project set out to build usable software for scientific workflows in HPC environments by undertaking multiple rounds of qualitative user research. Qualitative research investigates how individuals accomplish their work and our interview-based study surfaced a variety of insights about the experiences of working in and with HPC ecosystems. This report examines multiple facets to the experiences of scientists and developers using and supporting HPC systems. We discuss how stakeholders grasp the design and configuration of these systems, the impacts of abstraction layers on their ability to successfully do work, and the varied perceptions of time that shape this work. Examining the adoption of the Cori HPC at NERSC we explore the anticipations and lived experiences of users interacting with this system’s novel storage feature, the Burst Buffer. We present lessons learned from across these insights to illustrate just some of the challenges HPC facilities and their stakeholders need to account for when procuring and supporting these essential scientific resources to ensure their usability and utility to a variety of scientific practices.

97 MATHEMATICS AND COMPUTING↗

Performance Debugging and Tuning of Flash-X with Data Analysis Tools

State-of-the-art multiphysics simulations running on large scale leadership computing platforms have many variables contributing to their performance and scaling behavior. We recently encountered an interesting performance anomaly in Flash-X, a multiphysics multicomponent simulation software, when characterizing its performance behavior on several large-scale HPC platforms. The anomaly was tracked down to the interaction between the use of dynamic allocation of scratch data and data locality in the cache hierarchy. In this paper we present the details of unexpected performance variability of Flash-X, its extensive analysis using the performance measurement tool TAU to collect the data and Python data analysis libraries to explore the data, and our insights from this experience. In this process, we discovered and removed or mitigated two additional performance limiting bottlenecks for performance tuning.

Huck, Kevin↗

LCIO

Performance of file systems shift during their life cycles. Evaluating this performance change over time is not trivial. Complexity arises in the interplay between external (i.e. application I/O workloads) and internal (i.e. the filesystem state) factors. Many benchmarks can test how a filesystem performs at the current snapshot state, but to observe the change over time necessitates that the filesystem state mutate (age) between benchmark runs. For a large-scale HPC parallel filesystem, the sheer scale and amount of interacting components during I/O operations magnify these challenges. LCIO addresses the question - how will the filesystem perform at different stages of its life cycle? LCIO is a synthetic benchmark, which provides the file system aging process to increase the amount of information that existing benchmarks like IOR and MDTest yield, as well as provide additional points of data that will be useful to system architects and engineers.

Bachstein, Matthew↗

High throughput, accurate gene annotation through AI and HPC-enabled structural analysis

With the advances in next generation sequencing technologies, the number of sequenced genomes is growing exponentially, resulting in a technology bottleneck for the translation of sequence information into usable hypotheses about the function of each gene. We have proposed leveraging our leadership high-performance computing (HPC) resources to help break this annotation bottleneck. Here we design an HPC-based framework to infer gene function from gene sequence by incorporating information about protein structure and interactions predicted by deep learning approaches. Accurate functional prediction and gene annotation using computational methods will facilitate breakthroughs in the genomic sciences essential to understanding and harnessing life processes in bacteria, fungi and plants. The development and applications of the state-of-the-art deep neural networks to protein structural modeling, interaction prediction, sequence comparison, and quality assessment of protein structural models will be made possible by leadership computational resources. These HPC-enabled bioinformatics and molecular modeling tools will provide powerful insights into molecular functions of genes.

59 BASIC BIOLOGICAL SCIENCES↗

HPC4Mfg with ACS

Distillation in the chemical industry accounts for roughly 10% of energy use in the U.S. While porous mass separating agents (MSAs) appear capable of achieving the same separations for a fraction of the energy, the fundamental lack of a well-understood relationship between the behavior of fluid mixtures confined in MSA pores and the selectivity of MSA-based processes presents a major barrier to their widespread industrial application. This work is the first study to systematically and self-consistently explore a range of parameters describing various molecule-material interactions in MSA-based separations. Through high performance computing (HPC), this study has yielded a fundamental understanding of the influence of confined fluid behavior on selectivity. The results also serve as a knowledge base for subsequent investigations needed to transform the framework for rational design of MSA-based separation processes. This will significantly reduce the energy required for separations central to chemical manufacturing.

36 MATERIALS SCIENCE↗

Livermore Computing User and System Scripts

LCUSS is a collection of scripts used to improve productivity on HPC systems for both administrators and general users. It will include general scripts for user management, scripts for helping users interact with LC resource management software (e.g. SLURM and Flux), and scripts to automate common user command-line tasks on LC and other HPC machines. These scripts are intended to be made available to all LC users. Hosting them on GitHub will allow LC staff, users, and collaborators to work on them together.

Long, Jeffery↗

PLEXUS: A Pattern-Oriented Runtime System Architecture for Resilient Extreme-Scale High-Performance Computing Systems

For high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed.In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application.

Hukerikar, Saurabh↗

Interactive Supercomputing With Jupyter

Rich user interfaces like Jupyter have the potential to make interacting with a supercomputer easier and more productive, consequently attracting new kinds of users and helping to expand the application of supercomputing to new science domains. For the scientist-user, the ideal rich user interface delivers a familiar, responsive, introspective, modular, and customizable platform upon which to build, run, capture, document, re-run, and share analysis workflows. From the provider or system administrator perspective, such a platform would also be easy to configure, deploy securely, update, customize, and support. Jupyter checks most if not all of these boxes. But from the perspective of leadership computing organizations that provide supercomputing power to users, such a platform should also make the unique features of a supercomputer center more accessible to users and more composable with high performance computing (HPC) workflows. Project Jupyter’s core design philosophy of extensibility, abstraction, and agnostic deployment, has allowed HPC centers like NERSC to bring in advanced supercomputing capabilities that can extend the interactive notebook environment. This has enabled a rich scientific discovery platform, particularly for experimental facility data analysis and machine learning problems.

97 MATHEMATICS AND COMPUTING↗

DRAS: Deep Reinforcement Learning for Cluster Scheduling in High Performance Computing

Cluster schedulers are crucial in high-performance computing (HPC). They determine when and which user jobs should be allocated to available system resources. Existing cluster scheduling heuristics are developed by human experts based on their experience with specific HPC systems and workloads. However, the increasing complexity of computing systems and the highly dynamic nature of application workloads have placed tremendous burden on manually designed and tuned scheduling heuristics. More aggressive optimization and automation are needed for cluster scheduling in HPC. In this work, we present an automated HPC scheduling agent named DRAS (Deep Reinforcement Agent for Scheduling) by leveraging deep reinforcement learning. DRAS is built on a hierarchical neural network incorporating special HPC scheduling features such as resource reservation and backfilling. An efficient training strategy is presented to enable DRAS to rapidly learn the target environment. Once being provided a specific scheduling objective given by the system manager, DRAS automatically learns to improve its policy through interaction with the scheduling environment and dynamically adjusts its policy as workload changes. We implement DRAS into a HPC scheduling platform called CQGym. CQGym provides a common platform allowing users to flexibly evaluate DRAS and other scheduling methods such as heuristic and optimization methods. Furthermore, the experiments using CQGym with different production workloads demonstrate that DRAS outperforms the existing heuristic and optimization approaches by up to 50%.

97 MATHEMATICS AND COMPUTING↗

An Integrated ML/AI Framework for Digitizing, Structuring and Searching DOE U-TRU-Fuels Data with Gap Analysis of Non-DOE Records

The U.S. Department of Energy (DOE) Advanced Fuels Campaign (AFC) is advancing transmutation fuel technologies to reduce long-lived radioactive waste by converting minor actinides into shorter-lived or stable elements through irradiation in sodium-cooled fast reactors. Key experiments such as AFC-1, AFC-2, FUels for the transmutation of Trans-URanium elements In phéniX (FUTURIX)-Fortes Teneurs en Actinides (FTA), and Experimental Breeder Reactor-II (EBR-II) X501 have provided fuel fabrication, irradiation, and performance data on various transuranic-bearing fuel forms. This report documents the creation of an artificial-intelligence assisted database, which has consolidated all DOE-owned data related to Transuranic (TRU)-bearing fuel experiments and stored across it across both the Idaho National Laboratory (INL) Nuclear Data Management and Analysis System and the INL high performance computing (HPC) infrastructure. A dedicated webpage, hosted on the INL HPC system, has been developed to support role-based access and data interaction. The database architecture allows researchers to navigate large, heterogeneous archives with far greater speed and accuracy than manual search and lays the foundation for future expansion into multimodal nuclear materials analysis environments. The database represents a major step towards a nationally integrated fuels database utilizing artificial intelligence tools.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

An HPC benchmark survey and taxonomy for characterization

The field of High-Performance Computing (HPC) is defined by providing computing devices with highest performance for a variety of demanding scientific users. The tight co-design relationship between HPC providers and users propels the field forward, paired with technological improvements, achieving continuously higher performance and resource utilization. A key device for system architects, architecture researchers, and scientific users are benchmarks, allowing for well-defined assessment of hardware, software, and algorithms. Many benchmarks exist in the community, from individual niche benchmarks testing specific features, to large-scale benchmark suites for whole procurements. We survey the available HPC benchmarks, summarizing them in table form with key details and concise categorization, also through an interactive website. For categorization, we present a benchmark taxonomy for well-defined characterization of benchmarks.

Benchmarking↗

Evaluation of OpenAI Codex for HPC Parallel Programming Models Kernel Generation

We evaluate AI-assisted generative capabilities on fundamental numerical kernels in high-performance computing (HPC), including AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG. We test the generated kernel codes for a variety of language-supported programming models, including (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). We use the GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code as of April 2023 to generate a vast amount of implementations given simple + + prompt variants. To quantify and compare the results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. Results suggest that the OpenAI Codex outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general purpose Python can benefit from adding code keywords, while Julia prompts perform acceptably well for its mature programming models (e.g., Threads and CUDA.jl). We expect for these benchmarks to provide a point of reference for each programming model's community. Overall, understanding the convergence of large language models, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

Godoy, William↗

HPC-Driven Modeling with ML-Based Surrogates for Magnon-Photon Dynamics in Hybrid Quantum System

Here, we introduce a hybrid computational framework that merges HPC-based numerical solvers with physics-informed ML surrogates for efficient modeling of magnon-photon interactions. By running short-duration, high-fidelity Maxwell-LLG simulations and feeding their results into an ML model, we substantially cut simulation time while achieving accurate predictions across larger spatiotemporal domains.

Accuracy↗

Execute BEE workflows on private cloud infrastructure-2.3.6.01 - LANL ATDM ST / STNS01-22 Milestone Completion Documentation (BEE-FY21 P6-2) [Slides]

This work involves the creation of the Cloud Launcher, a new subcomponent of BEE, and the extension of the BEETaskManager to run on Cloud systems. BEE will be able to interact with the Google Compute Engine and OpenStack cloud APIs to set up simple Cloud clusters for launching HPC job scripts. BEE will use existing functionality to launch jobs that previously could only be launched on HPC systems. The BEETaskManager will handle launching tasks on the Cloud cluster.

97 MATHEMATICS AND COMPUTING↗

Frontier Job-Centric Telemetry Dataset

Comprehensive analysis of high-performance computing (HPC) systems requires linking workload execution to system behavior. This kind of analysis is vital for diagnosing performance issues, managing capacity, detecting anomalous workloads, and understanding how applications interact with system hardware. This job-centric telemetry dataset unifies scheduler job records with node-level measurements, enabling direct association between workloads and their corresponding power, thermal, and performance characteristics. It contains sanitized, scheduler related metadata for 152,400 individual jobs that ran on the Frontier supercomputer and ended on selected days throughout 2024 and 2025, a subpopulation of ~6.8% of the total number of allocated jobs with non-zero run time on the system over that same period. Each is linked with files that contain telemetry time series records of the power utilization and temperature behavior of its allocated nodes and their processors during the run time of the job. Where available, a portion of the job files also contain network performance time series. Jobs are sampled from select days that reflect normal levels of user activity and possess job size distributions with large numbers of leadership class jobs (>20% of Frontier nodes). Jobs in this dataset attempt to best represent successful user workflows.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Computational Fluid Dynamics Simulations to Support Efficiency Improvements in Aluminum Smelting Process

Smelting is broadly described as the extraction of a metal from its ore. In the United States, aluminum is commonly produced by smelting alumina in bauxite using the Hall-Héroult process. Optimization of equipment and processes in conventional smelting is crucial to enhancing process efficiency and productivity, is necessary for improving the techno-economic feasibility, which directly manifests as the growth of the American economy. To achieve optima, insightful data on the multiphysics phenomena that are inherent to the process must be obtained through physical investigation or high-fidelity numerical simulations. The resolution of relevant scales in time and space for smelting operations requires intensive, high-performance computing (HPC) simulations. Hostile operating conditions limit physical data acquisition to specific techniques; therefore, these data do not describe the multiscale interaction of simultaneous effects. Fortunately, in recent decades, significant advancements in computing hardware and computational methods have made the numerical resolution of such a complex process possible. In this study, a high-fidelity simulation of aluminum smelting was performed using an open-source tool, OpenFOAM, which analyzed many parameters characteristic to underlying phenomena. A multiphysics model based on the Eulerian-Eulerian multifluid approach was adopted. This model can resolve critical issues in the electrolytic smelting of aluminum, such as bubbling of carbon dioxide from the anode(s), magnetohydrodynamics from electromagnetic effects, ionic dissolution of the alumina in the electrolyte, and the evolution of thermal profiles. This study provides valuable connectivity for characteristic data that can direct the future designs of efficient smelters. A basic framework to model and simulate the smelting process using OpenFOAM is presented for user modification in keeping with process development. Of relevance to the flow field, a detailed investigation of vortices produced by bubble motion and electromagnetics is discussed, along with their impact on the evolution of thermal profiles. The predictions show small-scale vortices in the clearance between the anode and cathode caused by magnetic forces. Predictions also indicate relatively large-scale vortices in the inter-anode space resulting from carbon dioxide rising through the electrolytic flow field. The formation of vortices at the edges of anodes was shown to direct alumina charged by the feeder to the bottom of the anodes, thus preventing the entrapment of gas bubbles in the periphery of the bottom of the anode. Symmetry was observed in the location of cold spots in the electrolytic mixture in the vicinity of the feeder. Cold spots were also observed in the clearance between the anode and cathode due to the flow’s transmission of unconverted alumina to this region.

36 MATERIALS SCIENCE↗