Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “interactive HPC”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ATHENA: Analytical Tool for Heterogeneous Neuromorphic Architectures

The ASC program seeks to use machine learning to improve efficiencies in its stockpile stewardship mission. Moreover, there is a growing market for technologies dedicated to accelerating AI workloads. Many of these emerging architectures promise to provide savings in energy efficiency, area, and latency when compared to traditional CPUs for these types of applications — neuromorphic analog and digital technologies provide both low-power and configurable acceleration of challenging artificial intelligence (AI) algorithms. If designed into a heterogeneous system with other accelerators and conventional compute nodes, these technologies have the potential to augment the capabilities of traditional High Performance Computing (HPC) platforms [5]. This expanded computation space requires not only a new approach to physics simulation, but the ability to evaluate and analyze next-generation architectures specialized for AI/ML workloads in both traditional HPC and embedded ND applications. Developing this capability will enable ASC to understand how this hardware performs in both HPC and ND environments, improve our ability to port our applications, guide the development of computing hardware, and inform vendor interactions, leading them toward solutions that address ASC’s unique requirements.

97 MATHEMATICS AND COMPUTING↗

Assessment of ESM Readiness Level for Exascale HPC

Advancement of Earth System Models (ESMs) is becoming increasingly challenging due to a confluence of factors including increasing model complexity – to more fully represent the earth system, increasing spatial resolution - to achieve higher accuracy by resolving fine-scale dynamical to physical, biological, and chemical processes and their interaction, increasing ensemble size - to more accurately represent predictive uncertainty, and increased computing requirements – to enable more accurate and timely weather predictions and climate projections for societal benefit. The belief by many that computing will take care of itself is no longer valid given the disruptive changes in HPC that are driving up the cost of computing, increasing the difficulty of using emerging HPC effectively, and exposing limits in parallelism, portability and scalability of the ESM applications themselves.

54 ENVIRONMENTAL SCIENCES↗

QFw: A Quantum Framework for Large-scale HPC Ecosystems

This work extends Quantum Framework (QFw) by integrating it with Northwest Quantum Simulator (NWQ-Sim) and by introducing a lightweight python library that allows multiple frontends (e.g., Qiskit) to interact with QFw. This extension enables QFw to flexibly decouple frontends from backends (e.g., NWQ-Sim). We demonstrate this capability by executing a Greenberger-Horne-Zeilinger (GHZ) circuit using Qiskit and Pennylane with NWQ-Sim and Tensor-Network Quantum Virtual-Machine (TN-QVM). QFw enables easy scaling to multiple nodes. We showcase this with scaling tests using GHZ with up to 32 qubits for different number of nodes on the Frontier supercomputer. And, to demonstrate the use of QFw for real world problems, we solve a metamaterial optimization problem, using a Quantum Approximate Optimization Algorithm (QAOA). We observe that QFw over NWQ-Sim marginally improves Qiskit-aer’s accuracy in reaching the lowest energy state. These additions to QFw prepare it to run hybrid applications in a hybrid resource environment since it treats actual quantum hardware and simulators alike.

Chundury, Srikar↗

Scaling SQL to the Supercomputer for Interactive Analysis of Simulation Data

AI and simulation workloads consume and generate large amounts of data that need to be searched, transformed and merged with other data. With the goal of treating data as a first-class citizen inside a traditionally compute-centric HPC environment, we explore how the use of accelerators and high-speed interconnects can speed up tasks which otherwise constitute bottlenecks in computational discovery workflows. BlazingSQL is SQL engine that runs natively on NVIDIA GPUs and supports internode communication for fast analytics on terabyte-scale tabular data sets. We show how a fast interconnect improves query performance if leveraged through the Unified Communication X (UCX) middleware. We envision that future computing platforms will integrate accelerated database query capabilities for immediate and interactive analysis of large simulation data.

Glaser, Jens↗

Eureka: Enabling Fine-Grained Access and Range Queries on Compressed Scientific Data via Data-Index Co-Compression

Handling large-scale scientific data in high-performance computing (HPC) environments poses significant challenges, including excessive I/O, high storage costs, and slow query performance. Traditional approaches often require full data decompression and scans, making them impractical for real-time or interactive analysis. To address these limitations, we introduce Eureka, a unified data-index co-compression framework that enables fine-grained access and efficient range queries on compressed scientific datasets. Eureka integrates spatial domain decomposition with block-wise error-bounded lossy compression to support selective decompression. It constructs a hierarchical AVL-tree index during compression to capture block-level value ranges, enabling fast pruning during query execution. To reduce metadata overhead, the index itself is also compressed while ensuring recall-preserving results. Experiments on six diverse HPC simulation datasets show that Eureka achieves up to 25x data compression and over 300x index compression, surpassing state-of-the-art compressors such as SZ3 and ZFP in rate-distortion performance. Additionally, Eureka delivers over 30x speedup for low-selectivity range queries, making it a scalable and efficient solution for modern scientific data analysis.

Yan, Ning↗

High-Accuracy Simulations to Model Pyrometallurgical Processes in a Secondary Lead Reverberatory Furnace

The US manufacturing industry produces about 1.3 million tons of refined lead each year using secondary sources consisting mainly of lead batteries. ORNL is partnering with Gopher resource, the second largest lead recycling company in the United States, and GTI, to develop a high-fidelity CFD model of a directly fired, reverberatory-style, secondary lead furnace. These High Performance Computing (HPC) simulations are aimed to use first principles modeling for combustion and melting processes of the secondary lead feed while accounting for complex interphase interactions between the gas, solid charge (lead) material, slag, and metal phases. Through validation against operating plant data, this effort will enable significant improvements in design, operational parameters, and energy efficiency, thus improving productivity and refractory lifetime of secondary lead melting furnaces. Estimated savings/reduction of, at least, 1 trillion BTU, 1 million ton/year of greenhouse gas emissions, and $\$50$ million/year to the US lead industry can be expected. ORNL resources and expertise in high-performance computing and multicomponent, multiphase flows were utilized to realize this goal while advancing the understanding of the smelting and melting processes occurring within the furnace.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Vidyut3d: A Non-Equilibrium Plasma Modeling Tool [SWR-24-101]

Vidyut3d is a massively-parallel plasma-fluid solver for low-temperature plasmas (LTPs) that supports both local field (LFA) and local mean energy (LMEA) approximations, as well as complex gas and surface-phase chemistry. The solver supports 2D and 3D domains, and uses AMReX's adaptive mesh refinement capabilities to increase the grid resolution around complex structures (e.g. streamer heads and sheaths) while maintaining a tractable problem size. Vidyut specializes in simulating various types of gas-phase discharges, as well as plasma/surface interactions and surface chemistry (e.g. for plasma-mediated catalysis applications). The solver also supports hybrid CPU/GPU parallelization strategies, and has demonstrated excellent scaling on various HPC architectures for problem sizes consisting of O(100 M) control volumes.

Sitaraman, Hariswaran↗

Scalable Data-Intensive Geocomputation: A Design for Real-Time Continental Flood Inundation Mapping

The convergence of data-intensive and extreme-scale computing enables an integrated software and data ecosystem for scientific discovery. Developments in this realm will fuel transformative research in data-driven interdisciplinary domains. Geocomputation provides computing paradigms in Geographic Information Systems (GIS) for interactive computing of geographic data, processes, models, and maps. Because GIS is data-driven, the computational scalability of a geocomputation workflow is directly related to the scale of the GIS data layers, their resolution and extent, as well as the velocity of the geo-located data streams to be processed. Unique in high user interactivity and low end-to-end latency requirements, geocomputation applications will dramatically benefit from the convergence of high-end data analytics (HDA) and high-performance computing (HPC). The application level challenge, however, is to identify and eliminate computational bottlenecks that arise along a geocomputation workflow. Indeed, poor scalability at any of the workflow components is detrimental to the entire end-to-end pipeline. Here, we study a large geocomputation use case in flood inundation mapping that handles multiple national-scale geospatial datasets and targets low end-to-end latency. We discuss benefits and challenges for harnessing both HDA and HPC for data-intensive geospatial data processing and intensive numerical modeling of geographic processes. We propose an HDA+HPC geocomputation architecture design that couples HDA (e.g., Spark)-based spatial data handling and HPC-based parallel data modeling. Key techniques for coupling HDA and HPC to bridge the two different software stacks are reviewed and discussed.

Liu, Yan↗

Integration of Kokkos into MonteRay [Slides]

Monte Carlo Neutron Transport codes are C++-based and simulate interaction of nuclear particles with materials. MonteRay is a library for accelerating Monte Carlo ray-casting tallies with GPUs. Kokkos is commonly used to write performance portable applications for HPC platforms. A goal is to implement Kokkos into the Expected Path Length file that uses Cuda as a backend. The methodology and results are summarized.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Large language model evaluation for high–performance computing software development

We apply AI-assisted large language model (LLM) capabilities of GPT-3 targeting high-performance computing (HPC) kernels for (i) code generation, and (ii) auto-parallelization of serial code in C ++, Fortran, Python and Julia. Our scope includes the following fundamental numerical kernels: AXPY, GEMV, GEMM, SpMV, Jacobi Stencil, and CG, and language/programming models: (1) C++ (e.g., OpenMP [including offload], OpenACC, Kokkos, SyCL, CUDA, and HIP), (2) Fortran (e.g., OpenMP [including offload] and OpenACC), (3) Python (e.g., numpy, Numba, cuPy, and pyCUDA), and (4) Julia (e.g., Threads, CUDA.jl, AMDGPU.jl, and KernelAbstractions.jl). Kernel implementations are generated using GitHub Copilot capabilities powered by the GPT-based OpenAI Codex available in Visual Studio Code given simple + + prompt variants. To quantify and compare the generated results, we propose a proficiency metric around the initial 10 suggestions given for each prompt. For auto-parallelization, we use ChatGPT interactively giving simple prompts as in a dialogue with another human including simple “prompt engineering” follow ups. Results suggest that correct outputs for C++ correlate with the adoption and maturity of programming models. For example, OpenMP and CUDA score really high, whereas HIP is still lacking. We found that prompts from either a targeted language such as Fortran or the more general-purpose Python can benefit from adding language keywords, while Julia prompts perform acceptably well for its Threads and CUDA.jl programming models. Finally, we expect to provide an initial quantifiable point of reference for code generation in each programming model using a state-of-the-art LLM. Overall, understanding the convergence of LLMs, AI, and HPC is crucial due to its rapidly evolving nature and how it is redefining human-computer interactions.

97 MATHEMATICS AND COMPUTING↗

PeleLMeX [SWR-22-48]

PeleLMeX is a solver for high fidelity reactive flow simulations, namely direct numerical simulation (DNS) and large eddy simulation (LES). The solver combines a low Mach number approach, adaptive mesh refinement (AMR), embedded boundary (EB) geometry treatment and high performance computing (HPC) to provide a flexible tool to address research questions on platforms ranging from small workstations to the world's largest GPU-accelerated supercomputers. PeleLMeX has been used to study complex flame/turbulence interactions in RCCI engines and hydrogen combustion or the effect of sustainable aviation fuel on gas turbine combustion. PeleLMeX is part of the Pele combustion Suite (https://amrex-combustion.github.io/)

Day, Marcus↗

Developing a Vorticity-Velocity-Based Off-Body Solver to Perform Multifidelity Simulations of Wind Farms

Wind power has become a key player in satisfying the global energy needs. With increased market penetration, unanticipated unsteady loading induced failures, installation related reductions in power generation, and significant maintenance costs have underscored the need to predict the unsteady fluid-structure interactions related to turbine layout and off-design wind conditions. Contemporary turbine design tools are incapable of accounting for such loadings. As a result, researchers have started utilizing high-Performance-Computing (HPC) based Computational Fluid Dynamics (CFD) solvers, such as the U.S. Department of Energy sponsored ExaWind software package, to investigate these phenomena. Unfortunately, such HPC tools are computationally expensive for routine industrial use, often because of the sheer number of cells required to resolve the wake flowfield. This paper describes a preliminary effort to address this issue by developing a vorticity-velocity based CFD off-body solver, VorTran-M2-AMReX, that integrates directly with DOE's ExaWind wind turbine analysis system to perform accurate and reliable simulations of wind turbine/farm at a lower computational cost than ExaWind alone. This article summarizes work undertaken to date concerning the assembly of the proposed analysis tool, and provides preliminary validation and verification of the VorTran-M2-AMReX off-body solver.

adaptive mesh refinement↗

S&TR September 2025: Computing Grand Challenge Turns 20

Livermore’s Computing Grand Challenge Program enters its 20th year with more unclassified high-performance computing (HPC) power than ever before. This unique, peer-reviewed competition awards HPC allocations on top supercomputers to multidisciplinary teams with high-impact projects. The Grand Challenge encourages researchers to innovate, pushes scientific discovery to new heights, improves the Laboratory’s HPC capabilities, and extends HPC accessibility to collaborators. Awardees must adapt to successive generations of HPC hardware and learn to run simulations at scale. The feature article spotlights three Grand Challenge teams whose research broke new ground in key scientific pursuits—the essence of dark matter, explosion-generated seismic waves, and protein interactions linked to cancer—while underscoring the importance of academic partnerships and considering the program’s future.

07 ISOTOPE AND RADIATION SOURCES↗

2020 Exascale Computing Project Annual Meeting (Executive Summary Report)

The Exascale Computing Project (ECP) delivers specific applications, software products, and outcomes on DOE computing facilities. Integration across these elements for specific hardware technologies for exascale system instantiations is fundamental to ECP success. The outcome of the ECP is the delivery of a capable exascale computing ecosystem to provide breakthrough solutions addressing our most critical challenges in scientific discovery, energy assurance, economic competitiveness, and national security. This outcome is not a matter of ensuring more powerful computing systems. The ECP is designed to create more valuable and rapid insights from a wide variety of applications (“capable”), which requires a much higher level of inherent efficacy in all methods, software tools, and ECP-enabled computing technologies to be acquired by DOE laboratories (“ecosystem”). The ECP annual meeting provides a unique opportunity for the core technical expertise in the United States focused on achieving this next plateau of computational science and computing performance to engage in direct discussions on project execution. Face-to-face gatherings in technical communities like this are common and needed for the exchange of scientific ideas and technical performance. The ECP annual meeting stands apart from other technical conferences and meetings in the computing community as it is uniquely and solely focused on the execution of the ECP and the integration of technical activities leading to the creation of the exascale computing ecosystem for the future. The direct interaction of key critical technical staff, who are leaders in their respective fields, and the resulting give-and-take between software, applications, and hardware and the technical co-design therein, is unique and essential to the effective execution of the ECP. The first annual meeting was held in Knoxville, Tennessee, January 31 – February 2, 2017 and brought together, for the first time, a diverse collection of researchers from 16 DOE national laboratories as well as university computer and computational science researchers to discuss shared problems and joint solutions for the development of a capable exascale computing ecosystem. These interactions resulted in focused technical plans and an energized community centered on advances for ECP. The second annual meeting was held in Knoxville, Tennessee, February 5–9, 2018. It included 643 individual thought leaders and performers in application development, software research and deployment, and hardware research and integrators, all of whom are part of the multifaceted, billion dollar HPC community. This meeting provided a platform to discuss and disseminate numerous examples where researchers with common goals and synergistic solutions came together for the first time to deliver tangible results. Additionally, at the 2018 meeting, ECP researchers had the opportunity to digest all US HPC vendor R&D product roadmaps pointing to exascale – not only to learn how their research can play a role, but, more importantly, to influence those roadmaps to ensure successful delivery on DOE applications that will contribute to (if not solve) problems of national interest in national security, science, energy, and health, as well as growing security threats. The third annual meeting was held in Houston, Texas, January 14–17, 2019. With a 19% increase in the number of registrations (768 people), and the change in location, the third annual meeting was considered the most impactful of the three at the time. The new website provided a better platform for the dissemination of the content, the new venue as a meeting hotel instead of a conference center facilitated interactions and discussions after event hours, and the addition of an award-winning mobile event conference app (Whova) transformed dramatically the attendee experience at the event. This fourth annual meeting was held in Houston, Texas, February 3-7, 2020. This meeting had an increase in the number of attendees for a total of 824 people registered (782 attendees) and included numerous enhancements based on feedback and lessons learned from previous meetings, some of which are listed here: Improved quality of the sessions, their material and the whole program.; Had more industry participation and addition of external collaborators from overseas.; Published the full agenda earlier to better accommodate attendance and travel plans based on schedule.; Centralized all sessions in one venue.; Provided additional hotels and room blocks for the attendees.; Improved communication with the audience (links, material, directions, notifications, etc.) to go paperless.; Enhanced side meeting scheduling, management and user experience.; Made available additional space and tables for impromptu meetings and side discussions.; Improved IT and A/V solutions for speakers. In addition, our final survey captured the following points as opportunities for improvement in future meetings: consider a different meeting location that is more pedestrian friendly, reduce talks during working meals to allow more collaboration and informal time, adapt the agenda to acknowledge attendees from different timezones, consider recording some of the tutorials and/or sessions to share broadly with the HPC community, have a larger poster room, provide additional power strips, and improve the WiFi.

97 MATHEMATICS AND COMPUTING↗

Quantum/AI Topology-Aware Latency-Adaptive HPC Workflow Scheduling Optimization

The growing demand for more powerful high-performance computing (HPC) systems has led to a steady rise in energy consumption by supercomputing worldwide. This study is focused on comparing our Application-Topology Mapper (ATMapper) to the popular Simple Linux Utility for Resource Management (SLURM) for the purpose of exploring methods that can further optimize job-scheduling within HPC systems. ATMapper is an Artificial-Intelligence based approach to job-scheduling that is currently being enhanced with quantum annealing (QA) to generate optimal schedules faster. We are applying QA to speedup our ATMapper process to achieve higher computing efficiency, thereby reducing HPC energy consumption. Here, we examine how four job-scheduling approaches perform in processor node assignment when using an example network architecture of 4 interconnected nodes. Using a specialized script, we are assessing the schedule of a computation flow with 11 interdependent tasks. The data movements among nodes were tracked to count for the number of interactions (network hops) between nodes needed to complete the tasks. The total number of hops and the job completion time were then used to quantify the efficiency of the different mapping approaches. In addition to SLURM, we also compare our ATMapper to the QA-enabled LBNL TIGER and the D-Wave Distributed Computing processor assignment approaches. The preliminary results showed that our topology-aware, latency-adaptive ATMapper is significantly more efficient when compared to the other scheduling approaches due to its load-imbalance network allocation. The scheduler displayed a computing efficiency of 53% by performing significantly fewer network hops than its alternatives. By reducing the number of hops, ATMapper was able to perform all 11 tasks by using only 3 nodes out of given 4. This research indicates the potential to use QA/AI for HPC job-scheduling. Later, we will test a SLURM simulator program to draw further comparisons on the effectiveness of ATMapper's scheduling approach. The results of this comparison will serve as a baseline for later improving SLURM's performance using a QA-enhanced ATMapper approach.

Caraveo, Braulio [University of Huston - Clear Lak↗

sfapi_client v1.0

This software is a client designed to interact with the Superfacility API developed at NERSC. It allows users to easily access the api in python and programmatically interact with the compute resources available at NERSC. Other implementations are ad-hoc and made by our user base, the goal of the project is to encourage more users to adopt the API by making it easier to start building complex HPC workflows.

Tyler, Nicholas↗

RLC4CLR (Reinforcement Learning Controller for Critical Load Restoration Problems)

RLC4CLR demonstrates using a reinforcement learning controller (RLC) to solve a critical load restoration (CLR) problem, which improves the grid resilience after a substation outage event. RLC4CLR consists of two parts. (1) RL environment: This environment encapsulates the CLR problem to be solved and provides interfacing functions to follow the standard OpenAI Gym format. A power system simulator, i.e., OpenDSS, is included to provide the power flow solution. Controller inputs and outputs (RL state and action) as well as the reward are defined in this environment as well. In summary, the RL environment is the problem formulation from which the RL agent can learn. (2) RL training script: The training script enables the RL agent to learn its control policy by interacting with the RL environment. For RL training, an open-sourced RL library, i.e., RLlib, is leveraged which is based on a distributed computing framework (Ray). The training script is designed to be able to be run on both local machine or the NREL HPC system. Other components of RLC4CLR include input data, e.g., grid model (standard IEEE test feeders), and other files used for results analysis.

Zhang, Xiangyu↗

Study of interconnect errors, network congestion, and applications characteristics for throttle prediction on a large scale HPC system

Today’s High Performance Computing (HPC) systems contain thousand of nodes which work together to provide performance in the order of petaflops. The performance of these systems depends on various components like processors, memory, and interconnect. Among all, interconnect plays a major role as it glues together all the hardware components in an HPC system. A slow interconnect can impact a scientific application running on multiple processes severely as they rely on fast network messages to communicate and synchronize frequently. Unfortunately, the HPC community lacks a study that explores different interconnect errors, congestion events and applications characteristics on a large-scale HPC system. In our previous work, we process and analyze interconnect data of the Titan supercomputer to develop a thorough understanding of interconnects faults, errors, and congestion events. In this work, we first show how congestion events can impact application performance. We then investigate application characteristics interaction with interconnect errors and network congestion to predict applications encountering congestion with more than 90% accuracy.

97 MATHEMATICS AND COMPUTING↗