Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Practical Implementation of GPU-based Computing at the Grid Edge for Resilience Scenarios

This paper presents a practical implementation of GPU-accelerated computing at the grid edge to enhance power system resilience through next-generation smart meters. Advanced Metering Infrastructure (AMI) systems rely predominantly on centralized processing architectures, which limit real-time response capabilities during grid disturbances. This work proposes the integration of GPU-enabled computational platforms directly within smart meter to enable local execution support for power system analytics, fault detection algorithms, and optimization routines. The proposed framework uses the Julia programming language to leverage highperformance parallel computing capabilities while maintaining code portability and development efficiency. We use two experimental scenarios to benchmark the computational feasibility of this approach: sparse linear system solutions representative of power flow analyses, and multi-stage production cost simulations incorporating unit commitment and economic dispatch operations. Results demonstrate that computationally intensive power system algorithms, such as those supporting resilience scenario calculations, can be effectively executed at the distribution edge using commercially available embedded GPU hardware. Keywords—GPU acceleration, edge computing, smart meters, grid resilience, AMI, resilience.

De Souza, Reubun [School of Electrical Engineering↗

TTDFT: A GPU accelerated Tucker tensor DFT code for large-scale Kohn-Sham DFT calculations

We present the Tucker tensor DFT (TTDFT) code which uses a tensor-structured algorithm with graphic processing unit (GPU) acceleration for conducting ground-state DFT calculations on large-scale systems. The Tucker tensor DFT algorithm uses a localized Tucker tensor basis computed from an additive separable approximation to the Kohn-Sham Hamiltonian. The discrete Kohn-Sham problem is solved using Chebyshev filtered subspace iteration method that relies on matrix-matrix multiplications of a sparse symmetric Hamiltonian matrix and a dense wavefunction matrix, expressed in the localized Tucker tensor basis. These matrix-matrix multiplication operations, which constitute the most computationally intensive step of the solution procedure, are GPU accelerated providing ~8-fold GPU-CPU speedup for these operations on the largest systems studied. In conclusion, the computational performance of the TTDFT code is presented using benchmark studies on aluminum nano-particles and silicon quantum dots with system sizes ranging up to ~7,000 atoms.

97 MATHEMATICS AND COMPUTING↗

Practical Scalability of LuGo: Benchmarking the HHL Algorithm Using an Enhanced QPE Algorithm

The HHL algorithm is a prominent quantum algorithm that offers exponential speedup over its classical counterparts for solving a system of linear equations. However, synthesizing and executing HHL circuits demand significant computational resources from both classical and quantum systems. In this paper, we benchmark the HHL algorithm using the optimized Quantum Phase Estimation (QPE) generation algorithm, LuGo \cite{lu2025lugo}, to enhance its scalability and efficiency. We leverage the National Energy Research Scientific Computing Center's (NERSC) Perlmutter supercomputer to evaluate the scalability of generating HHL circuits and to measure the time to simulate the generated circuits. Additionally, we provide a comprehensive analysis of the algorithm's performance on various state-of-the-art superconducting and trapped-ion quantum devices, including studies on qubit connectivity, fidelity comparisons, and hardware compatibility and robustness. Our results offer preliminary insights into potential practical applications of the HHL algorithm enabled by LuGo and the performance of various types of quantum hardware.

Lu, Chao [ORNL] (ORCID:0000000179346933)↗

SCALE HTR-PROTEUS Benchmark Model

This dataset contains input and result files of computational simulations of HTR-PROTEUS benchmark with the latest version of SCALE code system. The simulations cover criticality control rod worth calculations as well as sensitivity analysis and uncertainty quantification. Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

High-Resolution Simulations of Geological CO 2 Injection: Application to the SPE11 Benchmark

Geological carbon sequestration (GCS) will play a critical role in decarbonization and in facilitating the transition to clean energy systems. Because CO 2 is highly mobile, ensuring its safe and permanent injection into subsurface geological formations involves monitoring over larger spatial domains and longer time periods than is typical for hydrocarbon reservoirs. This can benefit from simulation tools capable of modeling key CO 2 trapping mechanisms, particularly those optimized for speed and scalability on high-performance computing systems. Using isothermal versions of the SPE11B and SPE11C benchmark cases, we conduct a mesh refinement study simulating CO 2 injection into kilometer-scale rock formations at centimeter resolution with the GEOS open-source simulation framework. We focus on how mesh refinement improves the accuracy of convective mixing in both 2D and 3D simulations. The computational costs associated with achieving a converged solution highlight the need for predictive upscaling techniques. A systematic performance scaling analysis—including both central processing unit (CPU) and graphics processing unit (GPU) architectures—complements the “Results” section.

Geosciences↗

Effect of Inlet Throttling on Thermohydraulic Instability in a Large Scale Waterbased RCCS: A System-Level Analysis with RELAP5-3D

This paper presents results from system -level modeling of a water -based reactor cavity cooling system using RELAP5-3D. The computational model is benchmarked with experimental data from a half -scale RCCS test facility at Argonne National Laboratory. The model prediction is first compared with a two-phase oscillatory baseline experimental case where mixed accuracy is obtained. The model shows reasonable prediction of mass flow rate, pressure, and temperature but significant overprediction of void fraction. The model prediction is then compared with a fault case where the inlet of the risers is gradually reduced using a throttling valve. As the valve is closed, the model is able to predict some major flow phenomena observed in the experiment such as the dampening of oscillations, the reintroduction of oscillations, as well as boiling, flashing, and geysering in the risers. However, the timeline of these events are not well captured by the model. The model is also used to investigate the evolution of flow regime in the chimney. This work highlights that the semiempirical constitutive relations used in RELAP-3D could have a strong influence on the accuracy of the model in two-phase oscillatory flows.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Increased Fidelity and Associated Computational cost of Detailed Integral Experiment Benchmarks [Slides]

It does not seem like the system is significantly more sensitive to diameters of components near the center of the core. Intuitively it is, but was not detectable with simulations run to a Monte Carlo k eff uncertainty of 0.00002. The system is more sensitive to heights of components near the center of the core. Most (if not all) Zeus style benchmarks have perturbed core component heights individually.

42 ENGINEERING↗

Verifying MCNP Models of the TEX High 240 Plutonium Benchmark

Computational modeling programs are invaluable tools that allow us to understand systems, safely develop new processes, and make reliable predictions about future designs. However, the effectiveness of these codes is limited by the degree to which their parameters match the real world. In the field of nuclear engineering, cross section data is one of these vital parameters. Accurate cross section data on important fissile and fissionable isotopes promotes the design of safer and more efficient fabrication, transportation, storage, and stockpiling of nuclear fuel. Unfortunately, there are knowledge gaps in data on key isotopes. In 2011, a multinational meeting hosted by the US Department of Energy Nuclear Criticality Safety Program ranked the priority of certain cross section data needs. In response, Lawrence Livermore National Lab (LLNL) designed the Thermal and Epithermal eXperiment (TEX) series of benchmark experiments. Benchmark experiments are used to validate current cross section data. They validate data by comparing the results of an actual experiment to the predicted results from a computational model. The data a benchmark applies to depends on the isotope and energy range the experiment’s neutron multiplication factor ( k eff ) is most sensitive to. The development and testing of the TEX High 240 Plutonium Benchmark will help validate 240 Pu cross section data. The configuration and materials of this benchmark are designed to be most sensitive to 240 Pu's intermediate energy range (from 0.625 ev to 100 keV ). MCNP® models of the assembly have been developed by LLNL and the results have been written in the final design report. In order for the discrepancies between benchmark models and experiments to be attributed to cross section inaccuracies, the accuracy of the models needs to be verified. The goal of this project is to verify of the results of LLNL's modeling by creating a new set of MCNP models and comparing the results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Recent Advances of PyROS: A Pyomo Solver for Nonconvex Two-Stage Robust Optimization in Process Systems Engineering

This poster highlights uncertainty and technical risk reduction capabilities in CCSI2, with a focus on robust optimization. It presents recent advances of the two-stage robust optimization (RO) solver PyROS and applications to advanced energy systems optimization. To demonstrate the computational performance and reliability of PyROS, a benchmarking study on a library of over 8,500 small-scale RO problems is presented. Further, PyROS is used to obtain robust system designs of a MEA-based CO2 absorber under uncertainty in the thermodynamic property models for a variety of CO2 capture rate threshold requirements. Overall, the results demonstrate that the PyROS solver, including recent extensions to multi-stage RO settings, provides a reliable avenue to optimize the design and operation of advanced energy systems subject to various sources of parametric uncertainty.

Sherman, Jason↗

Recent Advances of PyROS: A Pyomo Solver for Nonconvex Two-Stage Robust Optimization in Process Systems Engineering

This poster highlights uncertainty and technical risk reduction capabilities in CCSI2, with a focus on robust optimization. It presents recent advances of the two-stage robust optimization (RO) solver PyROS and applications to advanced energy systems optimization. To demonstrate the computational performance and reliability of PyROS, a benchmarking study on a library of over 8,500 small-scale RO problems is presented. Further, PyROS is used to obtain robust system designs of a MEA-based CO2 absorber under uncertainty in the thermodynamic property models for a variety of CO2 capture rate threshold requirements. Overall, the results demonstrate that the PyROS solver, including recent extensions to multi-stage RO settings, provides a reliable avenue to optimize the design and operation of advanced energy systems subject to various sources of parametric uncertainty.

Sherman, Jason↗

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)↗

Evaluating Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this work, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, shared local memory accesses, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

97 MATHEMATICS AND COMPUTING↗

Unlocking hidden information in sparse small-angle neutron scattering measurements

Hypothesis Small-Angle Neutron Scattering (SANS) is a powerful technique for studying soft matter systems such as colloids, polymers, and lyotropic phases, providing nanoscale structural insights. However, its effectiveness is limited by low neutron flux, leading to long acquisition times and noisy data. Here, we hypothesize that Bayesian statistical inference using Gaussian Process Regression (GPR) can reconstruct high-fidelity scattering data from sparse measurements by leveraging intensity smoothness and continuity. Experiments and Simulations The method was benchmarked computationally and validated through SANS experiments on various soft matter systems, including wormlike micelles, colloidal suspensions, polymeric structures, and lyotropic phases. GPR-based inference was applied to both experimental and synthetic data to evaluate its effectiveness in noise reduction and intensity reconstruction. Findings GPR significantly enhances SANS data quality and therefore reducing measurement times by up to two orders of magnitude. This cost-effective approach maximizes experimental efficiency, enabling high-throughput studies and real-time monitoring of dynamic systems. It is particularly beneficial for weakly scattering and time-sensitive studies. Beyond SANS, this framework applies to other low-SNR techniques, including laboratory-based small-angle X-ray scattering and various dynamical scattering methods. Furthermore, it offers transformative potential for compact neutron sources, enhancing their viability for structural analysis in resource-limited settings.

Small angle neutron scattering↗

A Benchmark Case for the Grid Survivability Analysis

Among current priorities of the power system analysis is the development of metrics and computational tools for the resilience analysis during catastrophic events. New methods and tools are required for such an analysis and they have to be validated prior application to real systems. However, benchmark problems are not readily available due to the analysis novelty. The current paper presents a case based on the IEEE 14-bus system for this purpose. The grid is simplified to a graph with nodes representing generators, loads, and buses. Power inputs are imported from real-time simulations of the IEEE 14-bus system. Outcomes of all possible combinations of failed elements are presented in terms of probabilities for the grid to survive, partially survive, or fail. Only the power grid's ability to withstand adverse events (survivability) is analyzed. The grid's recoverability, the other part of the resilience analysis, is not considered.

power system↗

GPU Acceleration of Large-Scale Full-Frequency GW Calculations

Many-body perturbation theory is a powerful method to simulate electronic excitations in molecules and materials starting from the output of density functional theory calculations. By implementing the theory efficiently so as to run at scale on the latest leadership high-performance computing systems it is possible to extend the scope of GW calculations. Here, we present a GPU acceleration study of the full-frequency GW method as implemented in the WEST code. Excellent performance is achieved through the use of (i) optimized GPU libraries, e.g., cuFFT and cuBLAS, (ii) a hierarchical parallelization strategy that minimizes CPU-CPU, CPU-GPU, and GPU-GPU data transfer operations, (iii) nonblocking MPI communications that overlap with GPU computations, and (iv) mixed precision in selected portions of the code. A series of performance benchmarks has been carried out on leadership high-performance computing systems, showing a substantial speedup of the GPU-accelerated version of WEST with respect to its CPU version. Good strong and weak scaling is demonstrated using up to 25 920 GPUs. Finally, we showcase the capability of the GPU version of WEST for large-scale, full-frequency GW calculations of realistic systems, e.g., a nanostructure, an interface, and a defect, comprising up to 10 368 valence electrons.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Full-Induction Magnetohydrodynamics Solver for Liquid Metal Fusion Blankets in Vertex-CFD

Multiphysics modeling of liquid metal fusion blankets, which produce tritium and convert energy of neutrons created via fusion reactions into heat, is crucial for predicting performance, ensuring structural integrity, and optimizing energy production. While traditional blanket modeling of liquid metal flows during normal steady operating conditions commonly employs the inductionless approximation of the magnetohydrodynamics (MHD) equations, transient scenarios, when the plasma-confining magnetic field varies on millisecond time scales, require a full-induction MHD approach that dynamically evolves the magnetic field via the time-dependent induction equation. This paper presents the formulation, implementation, and initial verification of a full-induction MHD solver integrated within the open-source Vertex-CFD framework, which aims to achieve tight multiphysics coupling, a flexible software design enabling easy extension and addition of physics models, and performance portability across computing platforms. The solver utilizes finite element spatial discretization, implicit Runge–Kutta time integration, and an inexact Newton method to solve the resulting discrete nonlinear system, leveraging Trilinos packages for efficient computation. Verification against selected benchmark problems demonstrates accuracy and robustness of the solver. Furthermore, when the solver is applied to an idealized blanket model in 2.5D and full 3D, results obtained with Vertex-CFD are in good agreement with recently published quasi-2D simulations. These findings establish a computational foundation for future simulations of transient MHD phenomena in liquid metal blankets with Vertex-CFD, and open avenues for future extensions and performance optimizations.

Endeve, Eirik [ORNL] (ORCID:0000000312519507)↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

Benchmarking quantum computers

The rapid pace of development in quantum computing technology has sparked a proliferation of benchmarks to assess the performance of quantum computing hardware and software. However, not all benchmarks are of equal merit. Good ones empower scientists, engineers, programmers and users to understand the power of a computing system, whereas bad ones can misdirect research and inhibit progress. In this Perspective, we survey the science of quantum computer benchmarking. Here, we discuss the role of benchmarks and benchmarking and how good benchmarks can drive and measure progress towards the long-term goal of useful quantum computations, known as quantum utility. We explain how different kinds of benchmark quantify the performance of different parts of a quantum computer, discuss existing benchmarks, examine recent trends in benchmarking, and highlight important open research questions in this field.

Proctor, Timothy James [Sandia National Laboratori↗