Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “computer system benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Assessing and advancing the potential of quantum computing: A NASA case study

Quantum computing is one of the most enticing computational paradigms with the potential to revolutionize diverse areas of future-generation computational systems. While quantum computing hardware has advanced rapidly, from tiny laboratory experiments to quantum chips that can outperform even the largest supercomputers on specialized computational tasks, these noisy-intermediate scale quantum (NISQ) processors are still too small and non-robust to be directly useful for any real-world applications. In this paper, we describe NASA’s work in assessing and advancing the potential of quantum computing. We discuss advances in algorithms, both near- and longer-term, and the results of our explorations on current hardware as well as with simulations, including illustrating the benefits of algorithm-hardware co-design in the NISQ era. This work also includes physics-inspired classical algorithms that can be used at application scale today. We discuss innovative tools supporting the assessment and advancement of quantum computing and describe improved methods for simulating quantum systems of various types on high-performance computing systems that incorporate realistic error models. We provide an overview of recent methods for benchmarking, evaluating, and characterizing quantum hardware for error mitigation, as well as insights into fundamental quantum physics that can be harnessed for computational purposes.

Rieffel, Eleanor G.↗

In situ temperature measurements in sooting methane/air flames using synchrotron x-ray fluorescence of seeded krypton atoms

Synchrotron x-ray fluorescence has been used to measure temperatures in optically dense gases where traditional methods would fail. These data provide a benchmark for stringent tests of computational fluid dynamics models for complex systems where physical and chemical processes are intimately linked. The experiments measured krypton number densities in a sooting, atmospheric pressure, nonpremixed coflow flame that is widely used in combustion research. The experiments not only form targets for the models, but the simulations also identify potential sources of uncertainties in the measurements, allowing for future improvements.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Insights from Optimizing HPL Performance on Exascale Systems: A Comparative Analysis of Panel Factorization

High performance LINPACK (HPL) remains the primary benchmark for evaluating supercomputing performance. It includes many parts with substantial internal complexity, and its performance is affected by a large number of parameters that interact in ways that are difficult to predict on large-scale heterogeneous supercomputer systems. We present a comprehensive performance analysis of HPL on Frontier, the world’s first exascale supercomputer, which achieved HPL performance of 1.35 exaflops. Through empirical parameter tuning, detailed modeling, and comparative evaluation, we uncover critical performance insights, share lessons learned, and outline best practices for effective parameter tuning on exascale systems. We introduce and evaluate two novel PDFACT strategies: a dedicated-thread (DT) variant and a GPU-based variant (GPUPDFACT) implementation using HIP cooperative groups, demonstrating that GPU-based factorization outperforms conventional CPU-based PDFACT on Frontier’s architecture. Our findings establish key performance factors for HPL on exascale systems and offer valuable guidance for future high-performance computing and benchmarking efforts.

Lu, Hao [ORNL] (ORCID:000000018941870X)↗

NCCS High Performance GMRES Mixed Precision

HPG-MxP is a software package that performs a fixed number of multigrid preconditioned (using a Gauss-Seidel smoother) Generalized minimal residual (PGMRES) iterations in order to solve a possibly nonsymmetric large sparse linear system of equations. It is designed to be a benchmark to measure a computer's performance for sparse linear algebra workloads typical in scientific computing while allowing the use of mixed precision methods. The solution is required to have convergence characteristics and accuracy similar to double precision GMRES. It is based on the High Performance Conjugate Gradient Benchmark (HPCG) which restricts all implementations to use only the IEEE double precision format (FP64). The original implementation (https://github.com/hpg-mxp/hpg-mxp) was written by Ichitaro Yamazaki, Jennifer Loe, Christian Glusa, Sivasankaran Rajamanickam, Piotr Luszczek, and Jack Dongarra. Please refer to that repository for documentation on the original implementation. This version is maintained by the National Center for Computational Sciences at Oak Ridge National Laboratory. It is highly scalable and optimized for Oak Ridge Leadership Computing Facility (OLCF) systems, particularly Frontier.

Kashi, Aditya [Oak Ridge National Laboratory (ORNL↗

Unimolecular Dynamics of Partially Deuterated Hydroperoxyalkyl Intermediates (•QOOD) in Cyclohexane and Cyclopentane Oxidation

The hydroperoxyalkyl radicals (•QOOH) formed in cyclohexane and cyclopentane oxidation with carbon radical center at the β site relative to the hydroperoxy group have recently been observed through their infrared fingerprint and time- and energy-resolved unimolecular dissociation dynamics to hydroxyl (OH) radical and cyclic ether products. Partial deuteration shifts the IR transitions associated with OH/OD motion to lower frequencies, providing a new window on the IR spectroscopy and unimolecular dynamics of these transient intermediates. Specifically, the overtone OD stretch (2ν OD ) has been identified at 5215.0 cm –1 and 5216.0 cm –1 for the β-QOOD intermediates in cyclohexane and cyclopentane oxidation, respectively. Unimolecular decay rates of the •QOOD intermediates at these energies have been obtained from the time-resolved appearance of OD products. The experimental rates are compared with statistical microcanonical rates evaluated using RRKM theory, including heavy-atom tunneling associated with simultaneous O–O bond elongation and C–O–O angle contraction along the reaction pathway. The experimental rates provide a further test of the previously determined transition state (TS) barriers, which were computed utilizing a benchmarking approach that builds on higher-level reference calculations for the smaller ethane oxidation system. Here, the experimental rate measurements for both •QOOD intermediates agree well with the computed rates after accounting for small changes in zero-point energy and a minor empirical adjustment of the TS barriers in both systems.

Ethers↗

Enhancing f -Element Separations with ADAAM-EH: The Impact of Phase Modifiers and a DGA Aqueous Complexant

Recent investigations have used a 2-ethylhexyl diamide amine (ADAAM-EH) for Am/Cm separations in combination with N,N,N ',N '-tetraethyldiglycolamide as an aqueous complexant to achieve an unprecedented separation factor of 41. The aim of this research effort is to understand the speciation of trivalent lanthanide (Ln) and actinide (An) ions in the organic phase of an ADAAM-EH extraction system, both with and without phase modifiers (PM) (1-octanol and tri-n-butyl phosphate (TBP)). Leveraging spectroscopic techniques in combination with distribution ratio measurements provides an understanding of organic phase f-element ligand complexation. In the absence of PM, Ln is extracted in a stoichiometric 1:1 [M(ADAAM-EH) 1 (NO 3 ) x (H 2 O) 1 ](NO 3 ) 3-x complex. The addition of 1-octanol at 20 vol % results in multiple species present. One of the species is the same as the no PM case, and the other species results in an increased -OH coordination to the inner sphere, potentially displacing some NO 3 . In the case of TBP, increasing concentration results in additional red-shifted bands in the UV-visible spectra, suggesting the complexation of additional ligands of either ADAAM-EH or TBP. Finally, the new system knowledge obtained by and spectroscopic experiments will provide benchmarking information for computational studies of the inner- and outer-sphere coordination environments of f-element cations and insights into ADAAM-EH adduct formation with PM, like 1-octanol and TBP.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electronic Structure and Bonding of US, SUO, and US 2

Anion photoelectron spectra of US – and US 2 – were recorded using the third (355 nm) and fourth (266 nm) harmonics of an Nd:YAG laser, which yielded vertical detachment energies (VDEs) of 1.71 and 2.02 eV, respectively. The experimental results are supported by extensive relativistic ab initio calculations, primarily at the coupled cluster level of theory, with systematic sequences of correlation consistent basis sets. Calculations include the closely related SUO and SUO – molecules, as well as the oxide congeners UO/UO – and UO 2 /UO 2 – which are well-known experimentally and provide benchmark systems for the sulfide calculations. Adiabatic electron detachment energies (ADEs) are computed for UO – , UO 2 – , US – , SUO – , and US 2 – using the Feller–Peterson–Dixon (FPD) composite approach. Additionally, ADEs are determined for UO – and US – using a spinor-based coupled cluster approach where spin–orbit coupling is included at the orbital level. VDEs are derived from the ab initio results from Franck–Condon simulations of the photoelectron spectra.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

The accuracies of effective interactions in downfolding coupled-cluster approaches for small-dimensionality active spaces

Here, this paper evaluates the accuracy of the Hermitian form of the downfolding procedure using the double unitary coupled cluster (DUCC) ansatz on the benchmark systems of linear chains of hydrogen atoms, H6 and H8. The computational infrastructure employs the occupation-number-representation codes to construct the matrix representation of arbitrary second-quantized operators, allowing for the exact representation of exponentials of various operators. The tests demonstrate that external amplitudes from standard single-reference coupled cluster methods that sufficiently describe external (out-of-active-space) correlations reliably parameterize the Hermitian downfolded effective Hamiltonians in the DUCC formalism. The results show that this approach can overcome the problems associated with losing the variational character of corresponding energies in the corresponding SR-CC theories.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C↗

How Accurate Are Approximate Density Functionals for Noncovalent Interaction of Very Large Molecular Systems?

Noncovalent intermolecular interactions are very important in many research areas. Therefore, it is vital to understand the extent to which approximate density functionals give a proper description of noncovalent interactions. Previous research has demonstrated that some approximate density functionals can predict usefully accurate interaction energies for many noncovalent systems; however, most of that work is limited to small and moderate-sized molecules. Very recently though, accurate benchmarks have become available for some very large molecules. Here, the present work applies 21 approximate density functionals to compute the binding energies of seven large molecular systems that have a number of atoms ranging from 200 to 910. The results are judged by comparison to the recently published CIM-DLPNO-CCSD(T) results, which are assumed to provide a reliable benchmark. The five most accurate methods among those tested are found to be PW6B95-D4, PW6B95-D3(BJ), revM11, M06-L, and MN15.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Open‐Source Anaerobic Digestion Modeling Platform, Anaerobic Digestion Model No. 1 Fast (ADM1F)

An open‐source modeling platform, called Anaerobic Digestion Model No. 1 Fast (ADM1F), is introduced to achieve fast and numerically stable simulations of anaerobic digestion processes. ADM1F is compatible with an iPython interface to facilitate model configuration, simulation, data analysis, and visualization. Faster simulations and more stable results are accomplished by implementing an advanced open‐source library of numerical methods called Portable Extensive Toolkit for Scientific Computation (PETSc) to solve the ADM1 system of equations. Leveraging PETSc, ADM1F can consistently complete a steady‐state simulation under 0.2 s, over 99% faster than a benchmark ADM1 model implemented with MATLAB while achieving agreement of model outputs within 1% of those obtained with the benchmark model. For dynamic simulations, however, ADM1F has a computational speed advantage only when the influent characteristics update more frequently than every 4 h. The ability of ADM1F to be useful as a tool to study anaerobic digestion systems is demonstrated through two example implementations of ADM1F: (1) a two‐phase co‐digestion scenario evaluating the impact of the organic loading rate and the substrate composition on reactor performance and stability, and (2) a conventional digester scenario assessing the effectiveness of recovery strategies after disruptions that led to instability. These examples demonstrate how the high simulation speed and the convenience of the iPython interface allow ADM1F to complete complex analyses within minutes, much faster than computational strategies currently reported in the literature.

anaerobic co-digestion↗

Review of Experimental Data for Validating Computer Codes Used in Shielding Calculations for Spent Fuel Storage and Transportation Systems

This report presents a review of available radiochemical assay data and shielding benchmarks applicable to spent nuclear fuel (SNF) shielding calculations. The relevant information reviewed herein includes the Spent Fuel Composition (SFCOMPO) database, the Shielding Integral Benchmark Archive and Database (SINBAD), the International Handbook of Evaluated Criticality Safety Benchmark Experiments, and published measurements of external dose rates of casks loaded with SNF. The relevant experimental data identified in this report may be used to support verification and validation of computer codes used in SNF cask/transport shielding applications, as well as development of calculation uncertainties. It should be noted that a relatively small subset of the identified experimental data (e.g., criticality alarm experiments) is available in a standard format established by the international community participating in experimental isotopic and shielding data evaluations. An effort of the SFCOMPO Technical Review Group (TRG) is underway to publish first isotopic evaluations of individual assay data using a standard data evaluation format. The SINBAD TRG has recently initiated benchmark evaluations and modernization of the database. Therefore, more relevant information is expected in the future that will enable users to select quality experimental data in depletion code and shielding code validations for SNF applications.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CommBench: Micro-Benchmarking Hierarchical Networks with Multi-GPU, Multi-NIC Nodes

Modern high-performance computing systems have multiple GPUs and network interface cards (NICs) per node. The resulting network architectures have multilevel hierarchies of subnetworks with different interconnect and software technologies. These systems offer multiple vendor-provided communication capabilities and library implementations (IPC, MPI, NCCL, RCCL, OneCCL) with APIs providing varying levels of performance across the different levels. Understanding this performance is currently difficult because of the wide range of architectures and programming models (CUDA, HIP, OneAPI). We present CommBench, a library with cross-system portability and a high-level API that enables developers to easily build microbenchmarks relevant to their use cases and gain insight into the performance (bandwidth & latency) of multiple implementation libraries on different networks. We demonstrate CommBench with three sets of microbenchmarks that profile the performance of six systems. Our experimental results reveal the effect of multiple NICs on optimizing the bandwidth across nodes and also present the performance characteristics of four available communication libraries within and across nodes of NVIDIA, AMD, and Intel GPU networks.

Hidayetoglu, Mert↗

Benchmark Testing on the IBM-Q Network

The goal of this project is to evaluate a proposed set of candidate benchmarks being developed by the Standards and Performance Metrics Technical Advisory Group (TAG) of the Quantum Economic Development Consortium (QED-C), of which LANL is a member. The QED-C Standards and Performance Metrics TAG has developed implementations of several benchmark codes with the hope and expectation that these codes will be helpful to QED-C members interested in investigating various quantum computer platforms. The benchmark set will be run on LANL’s access to the IBM-Q system with the intent to evaluate the set for scalability, correct execution, and coverage of the application space. The goal is for the QED-C Standards and Performance Metrics TAC to be able to produce a coherent, consistent, scalable set of benchmarks that will run on multiple quantum computing platforms and will be available to members of QED-C, including LANL.

97 MATHEMATICS AND COMPUTING↗

Qutrit Randomized Benchmarking

Ternary quantum processors offer significant potential computational advantages over conventional qubit technologies, leveraging the encoding and processing of quantum information in qutrits (three-level systems). Therefore, to evaluate and compare the performance of such emerging quantum hardware it is essential to have robust benchmarking methods suitable for a higher-dimensional Hilbert space. We demonstrate extensions of industry standard randomized benchmarking (RB) protocols, developed and used extensively for qubits, suitable for ternary quantum logic. Using a superconducting five-qutrit processor, we find an average single-qutrit process infidelity of 3.8×10 -3 . Through interleaved RB, we characterize a few relevant gates, and employ simultaneous RB to fully characterize crosstalk errors. Finally, we apply cycle benchmarking to a two-qutrit CSUM gate and obtain a two-qutrit process fidelity of 0.85. Our results present and demonstrate RB-based tools to characterize the performance of a qutrit processor, and a general approach to diagnose control errors in future qudit hardware.

97 MATHEMATICS AND COMPUTING↗

SMR safety through HTTF modeling and benchmark efforts for code validation for gas-cooled reactor applications

Accurate modeling and simulation tools for thermal-hydraulics calculations are a key element needed to design and license new advanced reactors including Small Modular Reactors (SMR) and Microreactors. Uncertainties in modeling and simulation can have significant safety and economic implications. The High Temperature Test Facility (HTTF) at Oregon State University (OSU) is a scaled integral effects experiment designed to investigate transient behavior in high-temperature gas-cooled prismatic-block nuclear reactors. High-quality measurement data is available from the HTTF that is suitable for a thermal-hydraulics code validation benchmark for gas-cooled reactor simulations. Here, this paper summarizes individual HTTF modeling efforts to date for tool validation at Idaho National Laboratory (INL), Argonne National Laboratory (ANL), Oregon State University (OSU) and Canadian Nuclear Laboratories (CNL) using system thermal-hydraulics codes, Computational Fluid Dynamics (CFD) codes and system-CFD code couplings. Also, the paper introduces the ongoing OECD Nuclear Energy Agency (NEA) High Temperature Gas Reactor Thermal-Hydraulics (HTGR T/H) benchmark that allows for better comparisons of results between different international modeling teams. The benchmark provides well defined computational problems that include code-to-code comparisons and comparisons to measured data. These problems provide an avenue for quantifying accuracy and identifying sources of uncertainty in thermal-hydraulics calculations, including in measured thermophysical properties, as part of validation for gas-cooled reactor simulation tools.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

Entanglement Benchmarking in Quantum Simulations of Spin Systems

We simulate quantum spin systems and measure entanglement using circuits tailored for near-term quantum computers. Traditional tools like entanglement entropy are limited to pure states and require full state tomography, making them impractical on current hardware. Instead, we employ the novel approach, Positive Partial Transpose (PPT) criterion to efficiently detect pairwise entanglement from two-spin reduced density matrices, applicable to both pure and mixed states. This method enables scalable entanglement detection, providing a practical route to study quantum correlations, phase transitions, and benchmark quantum devices.

Baul, Anshumitra [ORNL] (ORCID:0000000268947191)↗

Validation of time-dependent shift using the pulsed sphere benchmarks

The detailed behavior of neutrons in a rapidly changing time-dependent physical system is a challenging computational physics problem, particularly when using Monte Carlo methods on heterogeneous high-performance computing architectures. A small number of algorithms and code implementations have been shown to be performant for time-independent (fixed source and k-eigenvalue) Monte Carlo, and there are existing simulation tools that successfully solve the time-dependent Monte Carlo problem on smaller computing platforms. To bridge this gap, a time-dependent version of ORNL’s Shift code has been recently developed. Shift’s history-based algorithm on CPUs, and its event-based algorithm on GPUs, have both been observed to scale well to very large numbers of processors, which motivated the extension of this code to solve time-dependent problems. The validation of this new capability requires a comparison with time-dependent neutron experiments. Lawrence Livermore National Laboratory’s (LLNL) pulsed sphere benchmark experiments were simulated in Shift to validate both the time-independent as well as new time-dependent features recently incorporated into Shift. A suite of pulsed-sphere models was simulated using Shift and compared to the available experimental data and simulations with MCNP. Overall results indicate that Shift accurately simulates the pulsed sphere benchmarks, and that the new time-dependent modifications of Shift are working as intended. Validated exascale neutron transport codes are essential for a wide variety of future multiphysics applications.

Palmer, Camille J.↗