Engineering PapersSearch

SEARCH · Engineering Papers

Results for “benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Ecological Benchmark for Radionuclides

The Ecological Benchmark Tool for radionuclides dataset serves as a comprehensive repository of benchmarks designed to assess ecological risks at contaminated sites. This tool facilitates the evaluation of various environmental media and contaminants, supporting regulatory compliance and ecological protection. Benchmarks are available for sediment, soil, and surface water. The dataset also provides species-specific benchmarks for fish, plants, birds, mammals, and invertebrates. Users can select benchmark sources, media, individual radionuclides, and retrieve results in tabular or spreadsheet formats for analysis. Benchmarks are derived from authoritative sources, including government agencies, scientific councils, and academic publications. The dataset supports ecological risk assessments, regulatory decision-making, and environmental planning, with tools for benchmarking against radiological thresholds, sensitive species protection, and habitat impact evaluations. This structured approach ensures a robust evaluation of ecological risks tailored to site-specific and regulatory needs.

Stewart, Debra [Oak Ridge National Laboratory (ORN

Discrete fracture network model benchmarks developed and applied in a DECOVALEX-2023 repository performance assessment study

This study presents newly developed benchmarks for modeling flow and transport within discrete fracture networks (DFNs) and useful methods for analyzing the results. The new benchmarks are designed to test modeling approaches for use in probabilistic performance assessment models of deep geologic repositories in fractured rock. The benchmarks simulate flow and transport through a 1 km 3 block of fractured rock. The first simulates migration of a short pulse of tracer through a simple network of four intersecting fractures. The second adds 1089 stochastically generated fractures. The third changes the pulse to a continuous point source. Evaluation of model performance relies on moment analysis and comparison of the results of different models. The expected nondimensional first moment of the conservative tracer for each benchmark is 1. The benchmarks were simulated by teams from Canada, Czechia, Germany, Korea, Sweden, Taiwan, and the United States as part of a DECOVALEX-2023 study (decovalex.org). The teams used various approaches, including explicit DFN modeling, DFN upscaling to an equivalent continuous porous medium (ECPM), and a combination of both methods. Transport mechanisms are modeled using either the advection-dispersion equation or particle tracking. Results demonstrate strong agreement among the models in breakthrough behavior up to the 75th percentile. Significant deviations in first moments and well-clustered outputs led to the identification of inaccuracies in several models. Such findings exemplify the benefit of exercising these benchmarks and using the presented methods to test DFN flow and transport models.

Benchmark

Depletion Benchmark of the AFIP-7 Experiment in the Advanced Test Reactor

Reactor physics depletion benchmarks for low-enriched uranium fuel are limited in number. In particular, there is very limited data for LEU benchmarks for U-10Mo (Uranium-10% Molybdenum) plate fuel developed for use in U.S. high-performance research reactors (USHPRR). USHPRR includes the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), University of Missouri Research Reactor (MURR), Massachusetts Institute of Technology Reactor (MITR), and National Bureau of Standards Reactor (NBSR) at the National Institute of Science and Technology. These reactors are fueled with high-enriched uranium dispersed fuel in a silicon/aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiated fueled specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This report documents the development of a benchmark model of the irradiation of the ATR Full -size plate In center flux trap Position 7 (AFIP-7) experiment.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS

Depletion benchmark for a high-assay low-enriched uranium fuel experiment in the advanced test reactor

Reactor physics depletion benchmarks for high-assay low-enriched uranium (HALEU) fuel are limited in number. In particular, there is limited data for HALEU benchmarks for U-10Mo (uranium-10% molybdenum) plate fuel that is being developed for use in the United States’ high performance research reactors including the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), Massachusetts Institute of Technology Reactor (MITR), University of Missouri Research Reactor (MURR), National Bureau of Standards Reactor (NBSR). These six reactors currently operate with highly enriched uranium dispersed fuel in an aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiating fuel specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This paper documents the development of a benchmark model of the irradiation of the ATR Full-size plate In center flux trap Position 7 (AFIP-7) experiment using the depletion codes MC21 and Advanced Dimensional Depletion for Engineering of Reactors (ADDER).

Nielsen, Joseph W. [Idaho National Laboratory (INL

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris

Ranking and Classifying AI Benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece C. [Cornell U.]

A High-Fidelity Model of the Peach Bottom 2 Turbine-Trip Benchmark Using VERA

This work presents a high-fidelity simulation of the Peach Bottom turbine trip (PBTT) benchmark using the Virtual Environment for Reactor Applications (VERA), a multiphysics reactor modeling tool developed by the U.S. Department of Energy’s Consortium for Advanced Simulation of Light Water Reactors energy innovation hub. The PBTT benchmark, based on a 1977 transient event at the end of cycle 2 in a General Electric Type-4 boiling water reactor (BWR), is a critical test case for validating core physics models with thermal feedback during rapid reactivity events. VERA was employed to perform end-to-end, pin-resolved simulations from conditions at the beginning of cycle 1 through the turbine-trip transient, incorporating detailed neutron transport, fuel depletion, and subchannel thermal hydraulics. The simulation reproduced key benchmark observables with high accuracy: the peak power excursion occurred at 0.75 s, matching the scram time and closely aligning with the benchmark average of 0.742 s; the simulated maximum power spike was approximately 7600 MW, which is within 3% of the benchmark average of 7400 MW; and void-collapse dynamics were consistent with benchmark expectations. Reactivity predictions during cycles 1 and 2 remained within 1500 pcm and 400 pcm of criticality, respectively. These results confirm VERA’s ability to model complex coupled neutronic and thermal hydraulic behavior in a BWR turbine-trip transient, which will support its use in future studies of modeling dryout, fuel performance, and uncertainty quantification for transients of this type.

BWR

Status of the International Criticality Safety Benchmark Evaluation Project

The International Criticality Safety Benchmark Evaluation Project (ICSBEP) has continued its work generating evaluations of new and historical benchmark experiments since the last update to the nuclear criticality safety (NCS) community at the 12th International Conference on Nuclear Criticality Conference held in 2023. One additional version of the ICSBEP Handbook has been published since that update, and the Technical Review Group (TRG) held two in-person meetings to review and approve additional benchmarks. The 2022 and 2023 editions of the handbook were combined into one release (published in November 2024) and contained 13 new evaluations with 46 different configurations and two major revisions to existing evaluations. The 2024 version of the handbook, currently under publication review, will contain two new evaluations with 15 new configurations and one major revision to HEU-MET-FAST-028, the evaluation of Flattop with a uranium core. The ICSBEP TRG met again in person in April 2025 to review benchmarks for the 2025 ICSBEP Handbook and final comment resolution is currently ongoing. Many of the new benchmarks represent contemporaneous experiments that have been specifically optimized to provide validation cases relevant to the NCS community. One major area of focus for new critical experiments is to target the sparsely populated intermediate energy (or resonance) region. Another focus of many of the new benchmarks is to provide experiments sensitive to different materials, such as chlorine, hafnium, tantalum, titanium, molybdenum, chromium, and polymethyl methacrylate (PMMA, or Lucite). The ICSBEP continues to deliver high-quality, peer reviewed evaluations of integral experiments relevant to the nuclear data community.

HEU-MET-FAST-028

Intern-Artificial Intelligence Benchmarking

Benchmarks provide a standardized method for evaluating different AI models, enabling reproducibility and comparison between models, and facilitating scientific progress. As AI models continue to develop rapidly, incorporating new datasets, capabilities, and architectures becomes more complicated. Therefore, the current static benchmarks become increasingly irrelevant. The MLCommons team argues that to make AI benchmarks more relevant, it involves making the benchmarks themselves more dynamic, as well as technical innovations that make it easier for scientists and researchers at all levels to use and contribute to the benchmarks. The current progress in technical innovation is a software that allows for a detailed view of a collection of AI benchmarks to be output in various formats that are easily readable and accessible.

Krishnan, Anjay [Fermilab]

Classifying and rating AI benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark’s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece [Fermilab]

An Analytic Benchmark for Neutron Boltzmann Transport with Downscattering—Part IV: PFNS and $\bar{ν}$ Uncertainty Propagation

An analytic benchmark with continuous-energy cross sections was previously derived to validate criticality calculations. Here, to extend the utility of the analytic benchmark to verify the implementation of $\bar{ν}$ and prompt fission neutron spectrum (PFNS) uncertainty propagation methods, new simplified forms that are dependent on the incident (fission-causing) neutron energy, as well as the outgoing neutron energy for the PFNS, are introduced in this work. The analytical forms for the flux and adjoint flux are derived for the extended benchmark and used to determine the 𝑘-eigenvalue sensitivity to $\bar{ν}$ and PFNS. The 𝑘-eigenvalue uncertainty due to $\bar{ν}$ and PFNS is calculated for the analytic benchmark using simplified$\bar{ν}$ and PFNS representations based on the ENDF-B/VIII.0 239 Pu evaluation. Because of the low sensitivity of the analytic benchmark to the physical PFNS, a nonphysical high-sensitivity PFNS is also presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Extending quantum-mechanical benchmark accuracy to biological ligand-pocket interactions

Predicting the binding affinity of ligands to protein pockets is key in the drug design pipeline. The flexibility of ligand-pocket motifs arises from a range of attractive and repulsive electronic interactions during binding. Accurately accounting for all interactions requires robust quantum-mechanical (QM) benchmarks, which are scarce for ligand-pocket systems. Additionally, disagreement between “gold standard” Coupled Cluster (CC) and Quantum Monte Carlo (QMC) methods casts doubt on many benchmarks for larger non-covalent systems. We introduce the “QUantum Interacting Dimer” (QUID) benchmark framework containing 170 non-covalent (non-)equilibrium systems modeling chemically and structurally diverse ligand-pocket motifs. Symmetry-adapted perturbation theory shows that QUID broadly covers non-covalent binding motifs and energetic contributions. Robust binding energies are obtained using complementary CC and QMC methods, achieving agreement of 0.5 kcal/mol. The benchmark data analysis reveals that several dispersion-inclusive density functional approximations provide accurate energy predictions, though their atomic van der Waals forces differ in magnitude and orientation. Contrarily, semiempirical methods and empirical force fields require improvements in capturing non-covalent interactions (NCIs) for out-of-equilibrium geometries. The wide span of NCIs, highly accurate interaction energies, and analysis of molecular properties take QUID beyond the “gold standard” for QM benchmarks of ligand-protein systems.

Puleva, Mirela [University of Luxembourg, Luxembou

Benchmark Calculation for Turkey Point Unit 3 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

Benchmark calculations were performed for Turkey Point Unit 3 cycles 1–3 to validate the SCALE 6.3/Polaris–PARCS v3.4.2 code with the ENDF/B–VII.1 56–group library by comparing the simulated results with the measured data. The benchmark results will be used in evaluating the SCALE/Polaris–PARCS code package’s uncertainties for pressurized water reactor physics analysis. That future analysis will include key nuclear parameters such as reactivity, control bank worth, temperature coefficients, and pin and assembly power peaking factors. The present document details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS, and PARCS. Additional details are provided with respect to the input and output files produced for the benchmark calculations. The benchmark results are summarized such that they can be used in evaluating uncertainties with other benchmark results for key nuclear parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Benchmark Calculation for Surry Unit 1 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

The benchmark calculations were performed for Surry Unit 1 cycles 1–3 to validate the SCALE 6.3/Polaris–Purdue Advanced Reactor Core Simulator (PARCS) v3.4.2 with the ENDF/B–VII.1 56–group library by comparing the simulated results with the measured data. The benchmark results will be used to evaluate uncertainties of the SCALE/Polaris–PARCS code package for pressurized water reactor physics analysis for key nuclear parameters such as reactivity, control bank worth, temperature coefficients, and pin and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS, and PARCS. Additional details are provided for the input and output files produced for the benchmark calculations. The benchmark results were summarized such that they can be used in evaluating uncertainties with other benchmark results for key nuclear parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Benchmark Calculation for the Quad Cities Unit 1 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

In this study, benchmark calculations were performed for the Quad Cities Unit 1 cycles 1–3 to validate the SCALE 6.3/Polaris–PARCS v3.4.2 code package with the ENDF/B-VII.1 AMPX 56-group library by comparing the simulated results with the measured data. The benchmark results will be used in evaluating uncertainties of the SCALE/Polaris–PARCS code package for boiling water reactor physics analysis for key nuclear parameters such as reactivity and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS, and PARCS; additionally, detailed information is provided for all the input and output files produced for the benchmark calculations. The benchmark results are summarized herein so that they can be used to evaluate uncertainties with other benchmark results for key nuclear parameters.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Benchmark Calculation for the Hatch Unit 1 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

This study was the performance of the benchmark calculation for the Hatch Unit 1 cycles 1–3, to validate the SCALE 6.3/Polaris–PARCS v3.4.2 with the ENDF/B-VII.1 AMPX 56-group library by comparing the simulated results with the measured data. The benchmark results will be used in evaluating uncertainties of the SCALE/Polaris–PARCS code package for boiling water reactor (BWR) physics analysis for key nuclear parameters such as reactivity and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS and PARCS, and additionally, detailed information is provided for all the input and output files produced for the benchmark calculations. The benchmark results were summarized such that they can be used in evaluating uncertainties for key nuclear parameters with other BWR benchmark results.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Benchmark Calculation for the Peach Bottom Unit 2 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

In this study, benchmark calculations were performed for Peach Bottom Unit 2 cycles 1–3 to validate the SCALE 6.3/Polaris–PARCS v3.4.2 with the ENDF/B-VII.1 AMPX 56-group library by comparing the simulated results with the measured data. The benchmark results will be used to evaluate uncertainties of the SCALE/Polaris–PARCS code package for boiling water reactor physics analysis for key nuclear parameters such as reactivity and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS and PARCS. Additionally, detailed information is provided for all the input and output files produced for the benchmark calculations. The benchmark results were summarized such that they can be used to evaluate uncertainties with other benchmark results for key nuclear parameters.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS

Benchmark Calculation for BEAVRS Cycles 1 and 2 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

This study performed the benchmark calculation for BEAVRS to validate the SCALE 6.3/Polaris–PARCS v3.4.2 code package with the ENDF/B-VII.1 56-group library, comparing the simulated results with the measured data. The benchmark results will be used in evaluating uncertainties of the SCALE/Polaris–PARCS code package for pressurized water reactor physics analysis for key nuclear parameters such as reactivity, control bank worth, temperature coefficients, and pin and assembly power peaking factors. This report details plant and fuel design specifications and input data for SCALE/Polaris as well as GenPMAXS and PARCS. Additionally, detailed information is provided for all the input and output files produced for the benchmark calculations. The benchmark results are summarized herein so that they can be used in evaluating uncertainties for key nuclear parameters with other benchmark results.

22 GENERAL STUDIES OF NUCLEAR REACTORS