Engineering PapersSearch

SEARCH · Engineering Papers

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Optimization Algorithms as Quantum Performance Benchmarks

Combinatorial optimization is anticipated to be one of the primary use cases for quantum computation in the coming years. The Quantum Approximate Optimization Algorithm (QAOA) and Quantum Annealing (QA) have the potential to demonstrate significant run-time performance benefits over current state-of-the-art solutions. Using existing methods for characterizing classical optimization algorithms, we analyze solution quality obtained by solving Max-Cut problems using a quantum annealing device and gate-model quantum simulators and devices. This is used to guide the development of an advanced benchmarking framework for quantum computers designed to evaluate the trade-off between run-time execution performance and the solution quality for iterative hybrid quantum-classical applications. The framework generates performance profiles through effective visualizations that show performance progression as a function of time for various problem sizes and illustrates algorithm limitations uncovered by the benchmarking approach. The framework is an enhancement to the existing open-source QED-C Application-Oriented Benchmark suite and can connect to the open-source analysis libraries. The suite can be executed on various quantum simulators and quantum hardware systems.

benchmarking

The HTR-Proteus Benchmark: Analysis and Use as a Verification and Validation Case

This presentation outlines the evaluation and application of the HTR-Proteus benchmark as a verification and validation (V&V) case for advanced reactor modeling tools. The work supports the U.S. Department of Energy’s HALEU Availability Program (HAP) and the joint DOE/NRC DNCSH project, which aims to reduce criticality safety uncertainties in commercial-scale HALEU fuel cycle and transportation systems. The HTR-Proteus experiments, conducted at the Paul Scherrer Institute, provide high-fidelity data for TRISO-fueled, graphite-moderated pebble bed reactors with high neutron leakage—conditions relevant to HALEU transport scenarios. This study focuses on Cores 4.2 and 4.3 of the HTR-Proteus benchmark, analyzing key sources of uncertainty including pebble packing, TRISO particle positioning, and core height. Using Project Chrono for realistic pebble geometries and Serpent, SHIFT, and MCNP for neutronics simulations, the study quantifies the impact of these uncertainties on the effective multiplication factor (keff). Results show that a sample size of 110 pebble configurations is sufficient to converge keff, with ±30 pcm uncertainty due to packing randomness. TRISO positioning and core height variations also significantly influence keff, highlighting the importance of detailed modeling in V&V efforts. The benchmark serves as a valuable test case for validating the Griffin reactor physics code and improving confidence in HALEU system simulations.

73 - NUCLEAR PHYSICS AND RADIATION PHYSICS

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)

SMR safety through HTTF modeling and benchmark efforts for code validation for gas-cooled reactor applications

Accurate modeling and simulation tools for thermal-hydraulics calculations are a key element needed to design and license new advanced reactors including Small Modular Reactors (SMR) and Microreactors. Uncertainties in modeling and simulation can have significant safety and economic implications. The High Temperature Test Facility (HTTF) at Oregon State University (OSU) is a scaled integral effects experiment designed to investigate transient behavior in high-temperature gas-cooled prismatic-block nuclear reactors. High-quality measurement data is available from the HTTF that is suitable for a thermal-hydraulics code validation benchmark for gas-cooled reactor simulations. Here, this paper summarizes individual HTTF modeling efforts to date for tool validation at Idaho National Laboratory (INL), Argonne National Laboratory (ANL), Oregon State University (OSU) and Canadian Nuclear Laboratories (CNL) using system thermal-hydraulics codes, Computational Fluid Dynamics (CFD) codes and system-CFD code couplings. Also, the paper introduces the ongoing OECD Nuclear Energy Agency (NEA) High Temperature Gas Reactor Thermal-Hydraulics (HTGR T/H) benchmark that allows for better comparisons of results between different international modeling teams. The benchmark provides well defined computational problems that include code-to-code comparisons and comparisons to measured data. These problems provide an avenue for quantifying accuracy and identifying sources of uncertainty in thermal-hydraulics calculations, including in measured thermophysical properties, as part of validation for gas-cooled reactor simulation tools.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Comparison of URANS and LES predictions for the open phase of the OECD NEA CSNI fluid structure interaction CFD benchmark

The OECD NEA CSNI WGAMA CFD Task Group ran a benchmark in 2020 and 2021 to assess the predictive capabilities of coupled fluid structure interaction (FSI) CFD analysis methods. This paper presents the predictions made for the open phase of the benchmark using URANS and LES turbulence modelling approaches, and a comparison of the results to the experimental data. The benchmark comprised a channel containing two inline cylinders in cross-flow. The cylinders were fixed at one end, free at the other, and had measured resonant frequencies and damping properties. The URANS modelling used ANSYS Fluent 2-way coupled to ANSYS Mechanical. The LES modelling used Nek5000, 1-way coupled to Diablo. Comparisons with cross-channel velocity profiles are presented, both for the mean flow and its RMS. Comparisons are also made to the frequency spectra for point measurements of fluid velocity and pressure, and for the accelerations of the free end of each cylinder. URANS predicts the average velocity profiles relatively well, and is able to predict the velocity and acceleration spectra at the shedding frequency. However, the frequency content at the 4th harmonic of the shedding frequency is low in the URANS flow fields, and so does not excite accelerations at the resonant frequency of the cylinders. LES makes better predictions of the average profiles, and the velocity spectra agree well at both the shedding frequency and at higher frequencies. In conclusion, the 1-way coupled LES results show good agreement for acceleration spectra.

22 GENERAL STUDIES OF NUCLEAR REACTORS

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

Many-Body Benchmark of Electronic Charge and Spin Densities for Li 1–x NiO 2

Accurate benchmarks are particularly important for highly correlated oxides as mean-field approximations often fail to describe the subtle balance of charge transfer and magnetism in these materials with an accuracy comparable to experimental needs. Here we present accurate diffusion Monte Carlo (DMC) results of the electronic charge and spin densities for the tunable highly correlated oxide Li 1–x NiO 2 for x = 0, 1/2, and 1. To enable quantitative comparisons, we introduce a robust density-partitioning scheme, extending Voronoi analysis to assign atomic charges from spatially noisy DMC densities. We then benchmark common approximations used in density functional theory (DFT). Comparison against DMC shows that r 2 SCAN delivers the most balanced performance across charge, spin, and radial density descriptors, nearly reproducing DMC results for LiNiO 2 and apical Ni sites in Li 0.5 NiO 2 . Hybrid functionals (PBE0, SCAN0) perform unexpectedly poorly, and PBE + U + V yields inconsistent trends between charge and spin densities. Therefore, the r 2 SCAN functional minimizes errors relative to DMC while capturing the variable valence of the Ni ion and also retaining the computational efficiency of DFT for large-scale simulations of the tunable structural and electronic phases of Li1−xNiO2. Our study highlights the importance of accurate benchmarking of the fundamental quantities involved in DFT to select appropriate DFT approximations in order to advance the predictive modeling of charge-transfer-driven phenomena in correlated electron systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

The neurobench framework for benchmarking neuromorphic computing algorithms and systems

Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. This article presents NeuroBench, a benchmark framework for neuromorphic algorithms and systems, which is collaboratively designed from an open community of researchers across industry and academia. NeuroBench introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent and hardware-dependent settings. For latest project updates, visit the project website (neurobench.ai).

Yik, Jason [Harvard Univ., Cambridge, MA (United S

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark—A Bayesian Inverse UQ-Based Approach for Data Assimilation

The Organisation for Economic Co-operation and Development Working Party on Nuclear Criticality Safety has proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian inverse uncertainty quantification (IUQ) employing scientific machine learning surrogate models as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of generalized linear least squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. Here, when comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that the GLLS predictions failed to replicate the computed response distributions for nonlinear applications, while MOCABA showed near agreement, and IUQ used the computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

Bayesian calibration

Benchmarking the performance of a high-Q cavity qudit using random unitaries

High-coherence cavity resonators are excellent resources for encoding quantum information in higher-dimensional Hilbert spaces, moving beyond traditional qubit-based platforms. A natural strategy is to use the Fock basis to encode information in qudits. One can perform quantum operations on the cavity mode qudit by coupling the system to a non-linear ancillary transmon qubit. However, the performance of the cavity-transmon device is limited by the noisy transmons. It is, therefore, important to develop practical benchmarking tools for these qudit systems in an algorithm-agnostic manner. We gauge the performance of these qudit platforms using sampling tests such as the heavy output generation test as well as the linear cross-entropy benchmark, by way of simulations of such a system subject to realistic dominant noise channels. We use selective number-dependent arbitrary phase and unconditional displacement gates as our universal gateset. Our results show that contemporary transmons comfortably enable controlling a few tens of Fock levels of a cavity mode. This framework allows benchmarking even higher dimensional qudits as those become accessible with improved transmons.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Benchmarking Variables for Checkpointing in HPC Applications

Checkpoint/Restart (C/R) is a widely used fault tolerance mechanism in converged systems of cloud, edge, and HPC. However, users often rely on their experience to determine which variables to checkpoint, as there is currently no benchmark that can provide a reference. This can result in checkpointing redundant or even incorrect variables. To address this issue, we propose a benchmark suite that includes critical variables for checkpointing, which have been manually identified, and a method for identifying those critical variables, with 20 representative HPC applications. Our method involves analyzing data dependency between variables to identify critical variables analytically. We verify the identified variables' correctness with a widely used C/R library FTI by an ablation study. With our benchmark suite and data dependency analysis, HPC practitioners now have a reference for identifying checkpointing variables and better knowledge of what kind of variables to checkpoint.

Fu, Xiang

LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.

Chitty-Venkata, Krishna Teja

Benchmarking large language models for materials synthesis: The case of atomic layer deposition

In this work, we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and, in particular, in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from the graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI’s GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1–5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, the specificity of the question, and the accuracy of the response as graded by the human experts. Furthermore, this emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Artificial intelligence

Binder-benchmarking

SAND2025-07593O Binder-benchmarking evaluates the speed and memory impacts of C++, Python, and Matlab code binders. As a repository, it provides a way to locally run computation-based and memory-based benchmark suites on pybind11 and nanobind-based code in a Docker image. The software runs simple-speed and memory benchmarks on primitive navigation and integration exemplar algorithms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),

bmdrc: Python package for quantifying phenotypes from chemical exposures with benchmark dose modeling

Though chemical exposures are known to potentially have negative impacts on health, including contributing to chronic diseases such as cancer, the quantitative contribution of risk is not fully understood for every chemical. A commonly used approach to quantify levels of risk is to measure the proportion of organisms (such as a total number of zebrafish on a plate or mice in a cage) with abnormal behavioral responses or morphology at increasing concentrations of chemical exposure. A particular challenge with processing the proportional data from these assays is the appropriate estimation of chemical concentration levels that result in malformations or acute toxicity, as these values typically vary between experimental measurements. The recommended approach by the Environmental Protection Agency (EPA) is to fit benchmark dose curves with specific filters and model fitting steps, which are crucial to properly processing the proportional data. Several tools exist for the fitting of benchmark dose response curves, but none are standalone Python libraries built to process both morphological and behavioral data as proportions with all the EPA recommended filters, filter parameters, models, and model parameters. Thus, here we present the benchmark dose response curve (bmdrc) Python library, which was built to closely follow these EPA guidelines with helpful visualizations of filters and fitted model curves, and reports for reproducibility purposes. bmdrc is open-source and has demonstrated utility as a support package to an existing web portal for information on chemicals (https://srp.pnnl.gov). Our package will support any toxicology analysis where the response is a proportional value at increasing levels of a concentration of a chemical or chemical mixture.

Superfund

Shielding Benchmark Comparison - MCNP6.2

This report documents the calculations performed for a shielding code comparison between various Department of Energy (DOE) sites. These shielding calculations were performed using an experiment drawn from the International Handbook of Evaluated Criticality Safety Benchmark Experiments, published by the Organisation for Economic Cooperation and Development/Nuclear Energy Agency (OECD/NEA). The benchmark selected for comparison is ALARM-CF-AIR-LAB-001 (“Neutron Fields in the Three-Section Concrete Labyrinth from Cf-252 Source, Benchmark ALARM CF AIR LAB-001”). The Y-12 submission for this code comparison was performed with Monte Carlo N Particle (MCNP) Transport Code System, Version 6.2 and Automated Variance Reduction Generator (ADVANTG).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS