Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges, requiring efficient hardware acceleration. Benchmarking the performance of LLMs across diverse hardware platforms is crucial to understanding their scalability and throughput characteristics. We introduce LLM-Inference-Bench, a comprehensive benchmarking suite to evaluate the hardware inference performance of LLMs. We thoroughly analyze diverse hardware platforms, including GPUs from Nvidia and AMD and specialized AI accelerators, Intel Habana and SambaNova. Our evaluation includes several LLM inference frameworks and models from LLaMA, Mistral, and Qwen families with 7B and 70B parameters. Our benchmarking results reveal the strengths and limitations of various models, hardware platforms, and inference frameworks. We provide an interactive dashboard to help identify configurations for optimal performance for a given hardware platform.

Chitty-Venkata, Krishna Teja

Benchmarking large language models for materials synthesis: The case of atomic layer deposition

In this work, we introduce an open-ended question benchmark, ALDbench, to evaluate the performance of large language models (LLMs) in materials synthesis, and, in particular, in the field of atomic layer deposition, a thin film growth technique used in energy applications and microelectronics. Our benchmark comprises questions with a level of difficulty ranging from the graduate level to domain expert current with the state of the art in the field. Human experts reviewed the questions along the criteria of difficulty and specificity, and the model responses along four different criteria: overall quality, specificity, relevance, and accuracy. We ran this benchmark on an instance of OpenAI’s GPT-4o. The responses from the model received a composite quality score of 3.7 on a 1–5 scale, consistent with a passing grade. However, 36% of the questions received at least one below average score. An in-depth analysis of the responses identified at least five instances of suspected hallucination. Finally, we observed statistically significant correlations between the difficulty of the question and the quality of the response, the difficulty of the question and the relevance of the response, the specificity of the question, and the accuracy of the response as graded by the human experts. Furthermore, this emphasizes the need to evaluate LLMs across multiple criteria beyond difficulty or accuracy.

Artificial intelligence

Binder-benchmarking

SAND2025-07593O Binder-benchmarking evaluates the speed and memory impacts of C++, Python, and Matlab code binders. As a repository, it provides a way to locally run computation-based and memory-based benchmark suites on pybind11 and nanobind-based code in a Docker image. The software runs simple-speed and memory benchmarks on primitive navigation and integration exemplar algorithms. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Walker II, Michael [Sandia National Lab. (SNL-CA),

bmdrc: Python package for quantifying phenotypes from chemical exposures with benchmark dose modeling

Though chemical exposures are known to potentially have negative impacts on health, including contributing to chronic diseases such as cancer, the quantitative contribution of risk is not fully understood for every chemical. A commonly used approach to quantify levels of risk is to measure the proportion of organisms (such as a total number of zebrafish on a plate or mice in a cage) with abnormal behavioral responses or morphology at increasing concentrations of chemical exposure. A particular challenge with processing the proportional data from these assays is the appropriate estimation of chemical concentration levels that result in malformations or acute toxicity, as these values typically vary between experimental measurements. The recommended approach by the Environmental Protection Agency (EPA) is to fit benchmark dose curves with specific filters and model fitting steps, which are crucial to properly processing the proportional data. Several tools exist for the fitting of benchmark dose response curves, but none are standalone Python libraries built to process both morphological and behavioral data as proportions with all the EPA recommended filters, filter parameters, models, and model parameters. Thus, here we present the benchmark dose response curve (bmdrc) Python library, which was built to closely follow these EPA guidelines with helpful visualizations of filters and fitted model curves, and reports for reproducibility purposes. bmdrc is open-source and has demonstrated utility as a support package to an existing web portal for information on chemicals (https://srp.pnnl.gov). Our package will support any toxicology analysis where the response is a proportional value at increasing levels of a concentration of a chemical or chemical mixture.

Superfund

Shielding Benchmark Comparison - MCNP6.2

This report documents the calculations performed for a shielding code comparison between various Department of Energy (DOE) sites. These shielding calculations were performed using an experiment drawn from the International Handbook of Evaluated Criticality Safety Benchmark Experiments, published by the Organisation for Economic Cooperation and Development/Nuclear Energy Agency (OECD/NEA). The benchmark selected for comparison is ALARM-CF-AIR-LAB-001 (“Neutron Fields in the Three-Section Concrete Labyrinth from Cf-252 Source, Benchmark ALARM CF AIR LAB-001”). The Y-12 submission for this code comparison was performed with Monte Carlo N Particle (MCNP) Transport Code System, Version 6.2 and Automated Variance Reduction Generator (ADVANTG).

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

wa-hls4ml and lui-gnn: A benchmark and GNN-based surrogate model for hls4ml resource and latency estimation

As machine learning (ML) increasingly serves as a tool for addressing real-time challenges in scientific applications, the development of advanced tooling has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as model synthesis, are now becoming limiting factors in the rapid iteration of designs. To reduce these emerging constraints, multiple efforts are being launched toward designing an ML-based surrogate model that estimates resource usage of synthesized accelerator architectures. This model would reduce the design iteration time, especially when designing within a set of given hardware constraints. This approach shows considerable potential, but as it stands, the effort is early and would benefit from coordination and standardization to assist future work as it emerges. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of more than 100,000 fully connected neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. In addition to the resource utilization and latency data provided, the dataset includes generated artifacts and log files for many of the synthesized neural networks, in order to support future research in ML-based code generation. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, as well as the average performance across a subset of the dataset. We measure the performance of a given predictor model through multiple metrics, including $R^2$ score and SMAPE on regression tasks, as well as inference time to further characterize the estimator under test. Additionally, we introduce the latency/utilization inference graph neural network (lui-gnn), a surrogate model that uses a graph neural network to represent input architectures in the form of a directed graph. This graph representation allows for a diverse set of model architectures to all be effectively handled by a surrogate model. We present the architecture and performance of the model, as evaluated by the new proposed benchmark, including SMAPE, $R^2$ score, and inference times, and find that lui-gnn generally predicts latency and utilization for the 75\% quantile within several percent of the synthesized resources on the synthetic test dataset, indicating that this approach of estimating resource and latency via a surrogate models has promise and warrants further research.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Using FIPD and OPTD to Benchmark Metallic Fuel Performance

This report serves as an introduction, tutorial, and benchmark specification for out-of-pile tests on metallic fuel. It introduces a new user to the EBR-II legacy fuel performance test program and the fast reactor fuel performance databases built to preserve the records. It then details the information stored in each database and how to find it. A benchmark specification is included for a small set of out-of-pile tests on U-10Zr fuel to function as a tutorial demonstrating how the legacy fuel performance data sets stored in the FIPD and OPTD databases can be used together to benchmark fuel performance models for steady-state and transient performance.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Continuous-Energy Verification of MCNP Calculations Using One-Group Spherical and Slab Criticality Benchmarks

This work presents a continuous-energy Monte Carlo verification study of one group spherical and slab criticality benchmarks using the MCNP ®1 code. Classical tabulations and newly generated benchmark solutions obtained by direct numerical evaluation by the author are considered. The benchmarks span weakly to strongly multiplying regimes and provide analytically defined critical radii as functions of a single parameter, c .

22 GENERAL STUDIES OF NUCLEAR REACTORS

The Aerosol Model Benchmarking Repository: A toolkit for model intercomparison

The Aerosol Model Benchmarking Repository and Standards (AMBRS) project was initiated to provide tools and to establish community standards for benchmarking aerosol models. This report describes a set of open-source tools for building, running, and analyzing aerosol box model simulations in a standardized framework. The framework consists of three core components: AMBuilder, a CMake-based build system that compiles supported models consistently; AMBRS, a Python module that defines unified numerical experiments and executes them with aligned inputs; and PyParticle, an aerosol analysis package that standardizes output, computes diagnostics, and visualizes simulation results. Together, these tools enable reproducible intercomparison of aerosol schemes and support process-level evaluation of how model simplifications affect predictions of size distributions, cloud condensation nuclei activity, and other relevant properties relevant for the Earth-Energy system. Beyond its role in benchmarking, AMBRS provides a platform for studying aerosol processes across scales and can be used to generate training data for AI/ML applications in support of a broader hierarchical aerosol modeling strategy.

54 ENVIRONMENTAL SCIENCES

ILAMBv2.7 benchmarking results comparing E3SMv2.1 land-atmosphere coupled (BGCv2LNDATM) and stand alone land (ELM) simulations with CMIP6 emission driven historical simulations

This dataset contains land model benchmarking results for the Energy Exascale Earth System Model version 2.1 (E3SMv2.1), including outputs from both coupled biogeochemistry simulations and stand-alone land model simulations. These results are compared against several emission-driven historical simulations from the Coupled Model Intercomparison Project Phase 6 (CMIP6). Benchmarking was conducted using the International Land Model Benchmarking (ILAMB) package, version 2.7 (ILAMBv2.7). CMIP6 model outputs were sourced from the Earth System Grid Federation (ESGF), while the E3SMv2.1 results were derived from raw model outputs. These outputs underwent processing steps such as time serialization, conservative regridding, and data standardization to ensure comparability. For spatial interpolation, the Earth System Modeling Framework (ESMF) tool, ESMF_RegridWeightGen, was employed to generate regridding weights, enabling the transformation of E3SM’s native cubed-sphere grid to a regular latitude-longitude grid.

Feng, Sha [PNNL]

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark -- A Bayesian Inverse UQ-based Approach for Data Assimilation

The Organization for Economic Cooperation and Development (OECD) Working Party on Nuclear Criticality Safety (WPNCS) proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian Inverse Uncertainty Quantification (IUQ) as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of Generalized Linear Least Squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. When comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that GLLS predictions fail to replicate computed response distributions for nonlinear applications, while MOCABA shows near agreement, and IUQ uses computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

FOS: Computer and information sciences

A Benchmarking Framework for Evaluating Large Language Model Capabilities in Nuclear Reactor Safety Applications

Large language models (LLMs) are increasingly capable of answering technical questions, synthesizing domain knowledge, and supporting engineering workflows. For nuclear science and engineering, these capabilities require careful, domain-specific evaluation before they can be credibly incorporated into safety-related activities, regulatory review, or technical decision support. This paper presents preliminary results from benchmarking framework for evaluating LLM capabilities in nuclear contexts. The framework is organized into three evaluation categories: nuclear fundamentals, general dual-use knowledge, and plant specific knowledge. These categories are intended to distinguish general nuclear engineering competence from broader technical reasoning and more context-dependent nuclear knowledge. Initial evaluations focus on nuclear fundamentals using questions representative of the knowledge expected of a nuclear professional engineer. Results indicate that contemporary frontier models perform at a high level and substantially exceed the performance of older model generations, with some models approaching saturation of the current benchmark. These findings suggest both the rapid improvement of LLM capabilities in specialized technical domains and the need for more discriminating evaluation methods. The paper presents the benchmark structure, preliminary model-comparison results, and ongoing work. This work supports development of verifiable, responsible, and safety-conscious methods for assessing AI systems in nuclear engineering applications.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Reproducible benchmark for the SNAP 8 experimental reactor at operating conditions

This work presents fully reproducible multiphysics benchmark models of the Systems for Nuclear Auxiliary Power (SNAP) 8 Experimental Reactor at operating conditions with coolant flow. Wet experiment (with coolant, at power) validation benchmarks are presented using both deterministic (Serpent-Griffin) and Monte-Carlo (OpenMC-Cardinal) multiphysics frameworks coupled with thermal-hydraulic solvers in MOOSE. Reactivity coefficient measurements including fuel temperature, isothermal temperature, and power coefficients show good agreement with experiments, with discrepancies within experimental uncertainty. Reactivity worth experiments for coolant, samarium, and xenon poisoning are reproduced with differences under 200 pcm. Comparison between Serpent-Griffin and OpenMC-Cardinal frameworks reveal multiphysics coupling introduces positive reactivity effects (100-200 pcm) compared to uniform temperature and density fields at nominal operating conditions. Comparison between Serpent-Griffin and reference Serpent solution shows that power distributions maintain consistent radial and axial peaking behavior. All models, assumptions, thermophysical and thermomechanical properties, and material definitions are thoroughly documented with cited references; model inputs and model generating scripts are stored in the snapReactors GitHub repository.

SNAP

Electrocatalytic benchmarking of ruthenium-based bimetallic anodes for the electrocatalytic oxidation of biomass-derived wastewater

In this paper, we report on the synthesis, characterization, and use of ruthenium oxide (RuO 2 ) doped with a secondary metal (M2) to enhance electrochemical activity and stability for the electrocatalytic oxidation (ECO) of biomass-derived wastewaters. We used different electrochemical methods such as cyclic voltammetry (CV), electrochemical surface area (ECSA), and Tafel analysis as well as physical characterization such as grazing incidence X-ray diffraction, X-ray photoelectron spectroscopy, and scanning electron microscopy to understand how the introduction of M2 affects electrochemical performance. Our results show that including an M2 improves the ECO performance regardless of the composition of the electrolyte. Specifically, we saw increase in ECSA, which could be due to enhanced charge transfer for the pH ranges evaluated. Furthermore, the introduction of organic compounds in wastewater generated during the hydrothermal liquefaction of food waste affected the ECO performance differently, depending on M2, the electrolyte composition, and anodic half-cell potential, highlighting the need to properly control the reaction conditions when testing and characterizing the electrocatalysts under different reaction regimes. We developed an in situ electrocatalytic benchmarking protocol, Boruah – Lopez-Ruiz – Strange (BLoRS), to quickly assess if the presence of M2 improves the ECO performance; thus, saving time and resources in ex situ characterization, testing, and product analysis. This foundational work provides the basis for characterization and benchmarking of electrodes of the ECO of organic compounds.

Electrocatalysis

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Practical Introduction to Benchmarking and Characterization of Quantum Computers

Rapid progress in quantum technology has transformed quantum computing and quantum information science from theoretical possibilities into tangible engineering challenges. Breakthroughs in quantum algorithms, quantum simulations, and quantum error correction are bringing useful quantum computation closer to fruition. These remarkable achievements have been facilitated by advances in quantum characterization, verification, and validation (QCVV). QCVV methods and protocols enable scientists and engineers to scrutinize, understand, and enhance the performance of quantum information-processing devices. In this tutorial, we review the fundamental principles underpinning QCVV, and introduce a diverse array of QCVV tools used by quantum researchers. We define and explain QCVV’s core models and concepts—quantum states, measurements, and processes—and illustrate how these building blocks are leveraged to examine a target system or operation. We survey and introduce protocols ranging from simple qubit characterization to advanced benchmarking methods. Along the way, we provide illustrated examples and detailed descriptions of the protocols, highlight the advantages and disadvantages of each, and discuss their potential scalability to future large-scale quantum computers. This tutorial serves as a guidebook for researchers unfamiliar with the benchmarking and characterization of quantum computers, and also as a detailed reference for experienced practitioners.

open quantum systems & decoherence

Code Benchmark of Depressurized Conduction Cooldown Transient in the High Temperature Test Facility

This paper presents results from modeling of a depressurized conduction cooldown (DCC) transient at the High Temperature Test Facility (HTTF) as part of the OECD-NEA Thermal Hydraulics Code Validation Benchmark for High-Temperature Gas-Cooled Reactors using HTTF Data . This paper briefly describes the benchmark and the models being used. It then presents a comparison of steady state and transient results based on the Problem 2 Exercise 1A and 1B definitions. We compare block and helium temperature distributions, mass flow distribution, and energy balance in steady state. All models show comparable mass flow distributions and energy balances. The temperatures within the core and outer regions are comparable in all models too, but inner reflector temperatures can vary significantly. Despite that, we find that the models are in good agreement for the full-power steady state. In the DCC, we look at block temperature at the core midplane and RCCS water exit temperature. The INL and ANL models are found to be in excellent agreement with one another on block temperature over time, while the agreement when the KAERI and NRG models are added into consideration is good. Differences in the transient heat removal from the RCCS cause the differences in block temperature over time in these models. The CNL models show similar trends to the INL, ANL, KAERI, and NRG models, but the temperatures are high because the volumes used in calculating the average temperature include the heater rods in the CNL models only. The HUN-REN model shows results that suggest significantly lower heat removal in the RCCS which merit further investigation.

22 GENERAL STUDIES OF NUCLEAR REACTORS