Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Performance benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

Random circuit block-encoded matrix and a proposal of quantum LINPACK benchmark

The LINPACK benchmark reports the performance of a computer for solving a system of linear equations with dense random matrices. Although this task was not designed with a real application directly in mind, the LINPACK benchmark has been used to define the list of TOP500 supercomputers since the debut of the list in 1993. We propose that a similar benchmark, called the quantum LINPACK benchmark, could be used to measure the whole machine performance of quantum computers. The success of the quantum LINPACK benchmark should be viewed as the minimal requirement for a quantum computer to perform a useful task of solving linear algebra problems, such as linear systems of equations. We propose an input model called the Random Circuit Block-Encoded Matrix (RACBEM), which is a proper generalization of a dense random matrix in the quantum setting. The RACBEM model is efficient to be implemented on a quantum computer and can be designed to optimally adapt to any given quantum architecture, with relying on a black-box quantum compiler. Besides solving linear systems, the RACBEM model can be used to perform a variety of linear algebra tasks relevant to many physical applications, such as computing spectral measures, time series generated by a Hamiltonian simulation, and thermal averages of the energy. We implement these linear algebra operations on IBM Q quantum devices as well as quantum virtual machines, and demonstrate their performance in solving scientific computing problems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Leveraging observed soil heterotrophic respiration fluxes as a novel constraint on global-scale models

Microbially-explicit models may improve understanding and projections of carbon dynamics in response to future climate change, but their fidelity in simulating global-scale soil heterotrophic respiration (RH), a stringent test for soil biogeochemical models, has never been evaluated. We used statistical global RH products, as well as 7,821 daily site-scale RH measurements, to evaluate the spatio-temporal performance of one first-order decay model (CASA-CNP) and two microbially-explicit biogeochemical models (CORPSE and MIMICS) that were forced by two different climate datasets. CORPSE and MIMICS did not provide any measurable performance improvement; instead, the models were highly sensitive to the meteorological input data used to drive them. Spatial RH variability was generally well simulated except in the northern middle latitudes (~50°N) and arid regions; models captured the seasonal variability of RH well, but showed more divergence in tropic and arctic regions. Our results demonstrate that the next generation of biogeochemical models shows promise, but also needs to be improved for realistic spatio-temporal variability of RH. Finally, we emphasize the importance of net primary production, soil moisture, and soil temperature inputs, and that jointly evaluating soil models for their spatial (global scale) and temporal (site scale) performance provides crucial benchmarks for improving biogeochemical models.

Jian, Jinshi↗

Deep-Learning-Derived Evaluation Metrics Enable Effective Benchmarking of Computational Tools for Phosphopeptide Identification

Tandem mass spectrometry (MS/MS)-based phosphoproteomics is a powerful technology for global phosphorylation analysis. However, applying four computational pipelines to a typical mass spectrometry (MS)-based phosphoproteomic dataset from a human cancer study, we observed a large discrepancy among the reported phosphopeptide identification and phosphosite localization results, underscoring a critical need for benchmarking. While efforts have been made to compare performance of computational pipelines using data from synthetic phosphopeptides, evaluations involving real application data have been largely limited to comparing the numbers of phosphopeptide identifications due to the lack of appropriate evaluation metrics. We investigated three deep learning-derived features as potential evaluation metrics: phosphosite probability, Delta RT and spectral similarity. Predicted phosphosite probability is computed by MusiteDeep, which provides high accuracy as previously reported; Delta RT is defined as the absolute retention time (RT) difference between RTs observed and predicted by AutoRT; and spectral similarity is defined as the Pearson’s correlation coefficient between spectra observed and predicted by pDeep2. Using a synthetic peptide dataset, we found that both Delta RT and spectral similarity provided excellent discrimination between correct and incorrect peptide-spectrum matches (PSMs) both when incorrect PSMs involved wrong peptide sequences and even when incorrect PSMs were caused by only incorrect phosphosite localization. Based on these results, we used all the three deep learning-derived features as evaluation metrics to compare different computational pipelines on diverse set of phosphoproteomic datasets and showed their utility in benchmarking performance of the pipelines. The benchmark metrics demonstrated in this study will enable users to select computational pipelines and parameters for routine analysis of phosphoproteomics data and will offer guidance for developers to improve computational methods.

59 BASIC BIOLOGICAL SCIENCES↗

Expanding Benchmarks for Nuclear Data Validation [Slides]

The presentation discusses benchmarking in comparison to experimental truth, the integral experiments that were performed, and that benchmarks are evaluated experiments. It also discusses validation testing, and that the ultimate goal of the validation end product is to improve evaluated nuclear data for applications It does state that validation highlights errors in the Nuclear Data Pipeline that affect applications, for example: Missing Cd Capture Gammas in ENDF/B-VII.1. It acknowledges that additional types of experiments are needed to test data used in applications, and that validation experiments do not have to be complicated and expensive. The presentation also discusses different types of measurements including activation foil and fission chamber measurements, reactor kinetics measurements, and subcritical measurements. The presentation concludes with an experiment summary and states that all applications need validation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmarking a Proof-of-Concept Performance Portable SYCL-based Fast Fourier Transformation Library

ABSTRACT In this paper, we present an early version of a SYCL-based FFT library, capable of running on all major vendor hardware, including CPUs and GPUs from AMD, ARM, Intel and NVIDIA. The current limitations of our library is it supports single-dimension FFTs up to 211 in length and base-2 input sequences. Although preliminary, the aim of this work is to seed further developments for a rich set of features for calculating FFTs. The library has the advantage over existing portable FFT libraries in that it is single-source, and there- fore removes the complexities that arise due to abundant use of pre-processor macros and auto-generated kernels to target different architectures. We exercise two SYCL-enabled compilers, Codeplay ComputeCpp and Intel's open-source LLVM project, to evaluate performance portability of our SYCL-based FFT on various hetero- geneous architectures.We provide studies comparing our portable library with highly optimized vendor-specific FFT libraries, and discuss potential sources hindering performance.

97 MATHEMATICS AND COMPUTING↗

LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the LLaMA 2-70B model in automating these tasks for scientific applications written in commonly used programming languages. Using representative test problems, we assess the model's capacity to generate code, documentation, and unit tests, as well as its ability to translate existing code between commonly used programming languages. Our comprehensive analysis evaluates the compilation, runtime behavior, and correctness of the generated and translated code. Additionally, we assess the quality of automatically generated code, documentation, and unit tests. Here, our results indicate that while LLaMA 2-70B frequently generates syntactically correct and functional code for simpler numerical tasks, it encounters substantial difficulties with more complex, parallelized, or distributed computations, requiring considerable manual corrections. We identify key limitations and suggest areas for future improvements to better leverage AI-driven automation in scientific computing workflows.

97 MATHEMATICS AND COMPUTING↗

The ENDF/B Nuclear Data Library and Its Impact on Reactor Simulations

The ENDF/B library, which is developed, maintained, and distributed by the Cross Section Evaluation Working Group, is the main source of nuclear data for analyses and computational simulations in nuclear applications, like different nuclear reactors concepts, radiation shielding, medical applications, astrophysics, etc. The library is constantly being improved and updated, with each release bringing an optimal representation of nuclear interactions as they are understood in their time. The most recent release, ENDF/B-VIII.1, represents a significant improvement in terms of the performance and consistency of the measured differential data relative to previous versions, as it combines the most recent experimental differential data and advanced theoretical nuclear models. As one of the many highlights, ENDF/B-VIII.1 restores a high-burnup depletion performance, comparable to ENDF/B-VII.1, that had been degraded in ENDF/B-VIII.0, while further improving the performance in criticality benchmarks, such as those in the Mosteller’s suite. Additionally noteworthy is the improved performance of ENDF/B-VIII.1 in radiation shielding and thick-target leakage spectrum integral experiments, which are also important for fusion and reactor applications. In this work we present in a very summarized way the main updates implemented in the ENDF/B-VIII.1 release and its main impacts, and also begin to delineate the path forward as to what to expect in the future for the next ENDF/B release, which, based on the timeline of the past few releases, is estimated to happen around 5 years from now. We emphasize that the ENDF/B-VIII.1 release was the product of an enormous collaborative effort among many authors and that for a complete detailed picture, the reader is strongly encouraged to refer to the article accompanying the release, which is currently in the publication process, but available as preprint [G. P. A. Nobre et al. arXiv Preprint arXiv: 2511.03564 (2025)].

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmark Calculation for the Hatch Unit 1 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

This study was the performance of the benchmark calculation for the Hatch Unit 1 cycles 1–3, to validate the SCALE 6.3/Polaris–PARCS v3.4.2 with the ENDF/B-VII.1 AMPX 56-group library by comparing the simulated results with the measured data. The benchmark results will be used in evaluating uncertainties of the SCALE/Polaris–PARCS code package for boiling water reactor (BWR) physics analysis for key nuclear parameters such as reactivity and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS and PARCS, and additionally, detailed information is provided for all the input and output files produced for the benchmark calculations. The benchmark results were summarized such that they can be used in evaluating uncertainties for key nuclear parameters with other BWR benchmark results.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Benchmark of Neutron Thermalization in Graphite Using a Pulsed Slowing-Down-Time Experiment

A benchmark has been developed using a pulsed slowing-down-time experiment to isolate the thermalization process in graphite. The experiment was conducted at the Oak Ridge Electron Linear Accelerator facility at Oak Ridge National Laboratory, and it measured the time spectrum of neutrons leaking from a graphite pile during slowing down and thermalization within graphite. Simulations of the benchmark experiment were performed using the MCNP6.1 Monte Carlo code and the ENDF/B-VII.1 and ENDF/B-VIII.0 cross-section databases. The benchmark provides a time spectrum (i.e., time-dependent counts in a detector) that allows for validation of the graphite thermal scattering libraries (TSLs). The impact on the simulations using a suite of graphite TSLs was compared with the experimental results. Given the density of nuclear graphite, the TSL corresponding to graphite with 30% porosity, as implemented in ENDF/B-VIII.0, was found to most accurately represent the measured time spectrum corresponding to the thermal energy range with an average deviation of ±1.7%.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Carbon Emissions in a Typical New Production Home: A Case Study

This report is intended to serve as a foundational document for residential construction executives and leadership teams to understand how a typical home currently performs relative to building decarbonization goals set for 2030 and 2050. The intent is to help homebuilders benchmark their current performance and better understand the largest sources of carbon emissions in their homes. There is a need to understand and assess current performance against net zero carbon goals to define achievable paths forward. This case study highlights opportunities for innovation and carbon reduction through advanced technologies and material choices.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Carbon Emissions in a Typical New Production Home: A Case Study

This report, prepared by IBACOS, is intended to serve as a foundational document for the residential construction industry to understand how a typical home currently performs relative to building decarbonization goals set for 2030 and 2050. The intent is to help homebuilders benchmark their current performance and better understand the largest sources of carbon emissions in their homes.

Carbon Emissions↗

Bramblett and Czirr Self-Shielded Fission Rates for 235 U Physical and Analytic Benchmark

These experiments were performed by Bramblett and Czirr at UCRL (i.e., LLNL) in the late 1960s through the 1970s and evaluated by Mark Lee (LFO, retired) for the ICSBEP Handbook. Dimensions, geometry, and benchmark RSFR values are taken from the Handbook. All sample calculational results used COG 11.3 with ENDF/B-VIII.0 cross-sections. A sample input listing is provided in Appendix A.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Uniformly Ordered Binary Decision Algorithm for Benchmark Experiment Correlations in Whisper Validation

When performing a validation exercise for determining the upper subcritical limit of a nuclear criticality safety application, an analyst should select and perform a statistical analysis on a population of benchmark experiments that are neutronically similar to the application. The size of this population should be sufficiently large such that the statistical analysis has a high degree of confidence that the bias plus bias uncertainty (calculational margin) has been accurately quantified. A complication arises because many benchmark experiments share common components, leading to correlations in their measured effective multiplication factors. Correlations between benchmark experiments within the population reduces its predictive power. This motivates the need for methods that consider benchmark experiment correlations and ensure adequate statistical significance of results. The Whisper code is a statistical analysis pack- age that incorporates nuclear data sensitivity coefficients from MCNP to assess benchmark experiment similarity and then performs an extreme-value analysis to estimate the bias plus bias uncertainty. The original methodology in Whisper does not consider the effect of benchmark experiment correlations when making this estimation, and this summary proposes the uniformly ordered binary decision algorithm to address this shortcoming. The original methodology in Whisper computes similarity coefficients ck for an application compared to all benchmark experiments in its library and develops weighting factors for a selected population proportional to the ck values. The methodology can be interpreted as statistically emulating a validation exercise for a particular application where the weighting factors may be viewed as the likelihood that an analyst would include a particular benchmark experiment within the population. The effective sample size of the population is the expected or mean number of benchmark experiments in the population. The uniformly ordered binary decision algorithm identifies clusters of correlated benchmark experiments within the population and then computes adjusted weighting factors based on the magnitude of the correlation coefficients within the cluster to compute a reduced effective sample size accounting for the lower information content because of correlations. Benchmark experiments within the cluster are ordered randomly with equal probability and probabilistic decisions are made as to whether a benchmark. experiment within the cluster should treated as redundant with a previous one; if two redundant benchmark experiments are included, then the conservative worst case bias plus bias uncertainty is used and the pair is counted as a single benchmark experiment in the population. Results are provided for HEU solutions in a research version of the Whisper software using benchmark experiment correlations provided by DICE, the Database for the International Criticality Safety Benchmark Evaluation Project (ICSBEP). These show that there can be a significant increase in the bias plus bias uncertainty because the effective sample size is reduced, and therefore the algorithm, needing to meet sample size requirements, expands the benchmark experiment population by accepting less similar benchmark experiments that would have otherwise not been included.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A volumetric framework for quantum computer benchmarks

We propose a very large family of benchmarks for probing the performance of quantum computers. We call them volumetric benchmarks (VBs) because they generalize IBM's benchmark for measuring quantum volume \cite{Cross18}. The quantum volume benchmark defines a family of square circuits whose depth d and width w are the same. A volumetric benchmark defines a family of rectangular quantum circuits, for which d and w are uncoupled to allow the study of time/space performance trade-offs. Each VB defines a mapping from circuit shapes — ( w , d ) pairs — to test suites C ( w , d ) . A test suite is an ensemble of test circuits that share a common structure. The test suite C for a given circuit shape may be a single circuit C , a specific list of circuits { C 1 … C N } that must all be run, or a large set of possible circuits equipped with a distribution P r ( C ) . The circuits in a given VB share a structure, which is limited only by designers' creativity. We list some known benchmarks, and other circuit families, that fit into the VB framework: several families of random circuits, periodic circuits, and algorithm-inspired circuits. The last ingredient defining a benchmark is a success criterion that defines when a processor is judged to have ``passed'' a given test circuit. We discuss several options. Benchmark data can be analyzed in many ways to extract many properties, but we propose a simple, universal graphical summary of results that illustrates the Pareto frontier of the d vs w trade-off for the processor being benchmarked.

97 MATHEMATICS AND COMPUTING↗

ChemGraph as an agentic framework for computational chemistry workflows

Atomistic simulations are essential in chemistry and materials science but remain challenging to run due to the expert knowledge required for the setup, execution, and validation stages of these calculations. We present ChemGraph, an agentic framework powered by artificial intelligence and state-of-the-art simulation tools to streamline and automate computational chemistry and materials science workflows. ChemGraph leverages graph neural network-based foundation models for accurate yet computationally efficient calculations and large language models (LLMs) for natural language understanding, task planning, and scientific reasoning to provide an intuitive and interactive interface. We evaluate ChemGraph across 13 benchmark tasks and demonstrate that smaller LLMs (GPT-4o-mini, Claude-3.5-haiku, Qwen-2.5-14B) perform well on simple workflows, while more complex tasks benefit from using larger models. Importantly, we show that decomposing complex tasks into smaller subtasks through a multi-agent framework enables GPT-4o to reach perfect accuracy and smaller LLMs to match or exceed single-agent GPT-4o's performance in these benchmarks.

Computational chemistry↗

Health Physics Research Reactor Criticality Accident Alarm System Benchmark Overview

From the countless critical experiments performed in the world during the past century, high-quality integral benchmarks experiments have been collected and gathered into the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook), managed by the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Working Group. This information preservation and dissemination effort is crucial for reactor licensing as well as criticality and radiation transport modeling validation. This summary reports on the status of a tentative benchmark addition to the ICSBEP Handbook. The proposed benchmark arises from legacy operation data of the Oak Ridge National Laboratory (ORNL) Health Physics Research Reactor (HPRR). The HPRR was a small, unmoderated, unshielded fast burst reactor that was used for research in health physics and radiobiology as well as teaching and training. As part of a comprehensive investigation of the available HPRR operation data and characteristics, different possibilities for use of the valuable results were studied. A critical experiment benchmark evaluation was performed, analyzing data coming from sub-critical and critical operation of the HPRR during operator training, steady-state irradiation of samples and before critical bursts. The results of the evaluation do not satisfy for the ICSBEP standards as the benchmark relative standard uncertainty is of about 4% for k eff , and the relative difference between sample calculations and expected k eff results is of about 1.5%. Due to those unsatisfactory results, it was decided not to pursue critical experiments evaluation of the HPRR presently and to focus instead on shielding type data for the creation of a criticality accident alarm system (CAAS) and shielding category benchmark, which is currently very scarce in the ICSBEP handbook—especially concerning critical, pulsed assembly, or reactor operation data. Several dosimetry and shielding experiments from HPRR burst operation were evaluated, with different benchmark metrics as sulfur fluence, Element 57 dose, or neutron fluence at different distances and under different shield materials conditions. An evaluation focusing on the Element 57 neutron dose as a benchmark metric was submitted to the ICSBEP Technical Review Group (TRG) meeting in October 2021, and the inclusion of the evaluation in the ICSBEP Handbook was deferred. The main change proposed by the international experiment evaluation experts is to use the neutron fluence measured by Bonner spheres as a benchmark metric. This represents a quantity closer to that actually measured by the experimentalists of the HPRR compared to the Element 57 dose, which adds another step of data transformation, thus potentially adding uncertainty to the benchmark. The evaluation has been updated and will be submitted to the 2022 ICSBEP TRG meeting for inclusion in the 2023 edition of the ICSBEP Handbook. The evaluation is performed using the KENO and MAVRIC combination from the SCALE 6.2.4 code suite which was previously used in similar CAAS benchmarks to allow for the use of variance reduction techniques.

61 RADIATION PROTECTION AND DOSIMETRY↗

Case Study of Using Kokkos and SYCLs Performance-Portable Frameworks for Milc-Dslash Benchmark on NVIDIA, AMD and Intel GPUs

Six of the top ten supercomputers in the TOP500 list from June 2021 rely on NVIDIA GPUs to achieve their peak compute bandwidth. With the announcement of Aurora, Frontier, and El Capitan, Intel and AMD have also entered the domain of providing GPUs for scientific computing. A consequence of the increased diversity in the GPU landscape is the emergence of portable programming models such as Kokkos, SYCL, OpenCL, and OpenMP, which allow application developers to maintain a single-source code across a diverse range of hardware architectures. While the portable frameworks try to optimize the compute resource usage on a given architecture, it is the programmers responsibility to expose parallelism in an application that can take advantage of thousands of processing elements available on GPUs. In this paper, we introduce a GPU-friendly parallel implementation of Milc-Dslash that exposes multiple hierarchies of parallelism in the algorithm. Milc-Dslash was designed to serve as a benchmark with highly optimized matrix-vector multiplications to measure the resource utilization on the GPU systems. The parallel hierarchies in the Milc-Dslash algorithm are mapped onto a target hardware using Kokkos and SYCL programming models. We present the performance achieved by Kokkos and SYCL implementations of Milc-Dslash on NVIDIA A100 GPU, AMD MI100 GPU, and Intel Gen9 GPU. Additionally, we compare the Kokkos and SYCL performances with those obtained from the versions written in CUDA and HIP programming models on NVIDIA A100 GPU and AMD MI100 GPU, respectively.

Dufek, Amanda S↗