Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “BENCHMARKS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Results Oriented Benchmarking: The Evolution of Benchmarking at NASA from Competitive Comparisons to World Class Space Partnerships

Informal benchmarking using personal or professional networks has taken place for many years at the Kennedy Space Center (KSC). The National Aeronautics and Space Administration (NASA) recognized early on, the need to formalize the benchmarking process for better utilization of resources and improved benchmarking performance. The need to compete in a faster, better, cheaper environment has been the catalyst for formalizing these efforts. A pioneering benchmarking consortium was chartered at KSC in January 1994. The consortium known as the Kennedy Benchmarking Clearinghouse (KBC), is a collaborative effort of NASA and all major KSC contractors. The charter of this consortium is to facilitate effective benchmarking, and leverage the resulting quality improvements across KSC. The KBC acts as a resource with experienced facilitators and a proven process. One of the initial actions of the KBC was to develop a holistic methodology for Center-wide benchmarking. This approach to Benchmarking integrates the best features of proven benchmarking models (i.e., Camp, Spendolini, Watson, and Balm). This cost-effective alternative to conventional Benchmarking approaches has provided a foundation for consistent benchmarking at KSC through the development of common terminology, tools, and techniques. Through these efforts a foundation and infrastructure has been built which allows short duration benchmarking studies yielding results gleaned from world class partners that can be readily implemented. The KBC has been recognized with the Silver Medal Award (in the applied research category) from the International Benchmarking Clearinghouse.

Bell, Michael A.↗

DOE FAIR Surrogate Benchmarks Supporting AI and Simulation Research (SBI Surrogate Benchmark Initiative) (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia(UVA). SBI repositories include data, code, and all relevant collateral artifacts, that the science and engineering community needs to use and reuse these data sets and surrogates. SBI repositories generate active research from both participants in SBI and the broader AI and domain science communities. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and capture them as surrogate benchmarks with a rich set of metadata, covering. Data; Model; Metrics specification; Machine specification; Science, Speed, Power Results, We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non Surrogate benchmarks that have many common features and similar issues regarding FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, benchmarks have datasets, models, and metadata, and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates, including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

An international benchmark for wind plant wakes from the American WAKE ExperimeNt (AWAKEN)

This article introduces the first benchmark study within the International Energy Agency Wind Task 57 framework, focusing on wind plant wakes. Leveraging data from the American WAKE ExperimeNt (AWAKEN), the benchmark aims to assess the accuracy of simulation tools in modeling wind plant wakes and their impact on the downstream flow under diverse inflow conditions. The AWAKEN field campaign, conducted in Oklahoma from 2022 to 2024, provides unprecedented observations of wind plant-atmosphere interactions, thus offering a large dataset to validate numerical models of different complexity. The benchmark will include three phases—code calibration, blind comparison, and iteration—allowing participants to refine their numerical models based on the feedback from the benchmark team. This article describes the benchmark case study selected from observations providing details on atmospheric conditions, wake evidence, and wind turbine operation. The benchmark’s structure and timeline, along with the expected publication of results, are discussed as well. This collaborative effort aims to enhance the accuracy of wind plant wake simulations, thus contributing to the improvement of wind energy production estimates.

17 WIND ENERGY↗

Design, Manufacture, and Testing of an Open-Source Benchmark Composite Hydrokinetic Turbine Blade: Preprint

In a trend toward clean energy alternatives, recent years have seen great strides in the marine energy space. Consequently, there is a pressing need for the design, development, and validation of novel energy harvesting technologies such as hydrokinetic devices, which capture kinetic energy from waves, tides, and currents. However, these devices span numerous concepts and designs that often lack solid benchmark research that can be freely referenced throughout their development. This work focuses on the design process of an open-source composite hydrokinetic turbine blade for a three-bladed marine turbine rotor assembly with a diameter of 2.5 m. The proposed blade consists of two structural composite skins that are bonded with an adhesive and filled with a foam core. This study also explores and contrasts the efficiency and resolution of low-fidelity rapid design methodologies and comprehensive high-fidelity approaches in the context of blade design, modeling, and analysis efforts, a key objective in this research. Blade hydrodynamic loads were modeled and applied to finite-element blade models to study deformations and potential failure. Ongoing and upcoming efforts will result in blade manufacture and structural testing at the National Renewable Energy Laboratory. In future work, multiple blades will be deployed at the Living Bridge site at the University of New Hampshire and will be compared to rigid aluminum blades of the same geometry, developed by Sandia National Laboratories. Ultimately, this research will lay foundational groundwork for researchers and manufacturers, establishing a baseline composite blade design that will serve as a benchmark in the development of future hydrokinetic turbine blades.

blade design↗

Application-level benchmarking of quantum computers using nonlocal game strategies

In a nonlocal game, two noncommunicating players cooperate to convince a referee that they possess a strategy that does not violate the rules of the game. Quantum strategies allow players to optimally win some games by performing joint measurements on a shared entangled state, but computing these strategies can be challenging. We present a variational quantum algorithm to compute quantum strategies for nonlocal games by encoding the rules of a nonlocal game into a Hamiltonian. We show how this algorithm can generate a short-depth optimal quantum strategy for a graph coloring game with a quantum advantage. This quantum strategy is then evaluated on fourteen different quantum hardware platforms to demonstrate its utility as a benchmark. Finally, we discuss potential sources of errors that can explain the observed decreased performance of the executed task and derive an expression for the number of samples required to accurately estimate the win rate in the presence of noise.

nonlocal games↗

SCALE HTR-PROTEUS Benchmark Model

This dataset contains input and result files of computational simulations of HTR-PROTEUS benchmark with the latest version of SCALE code system. The simulations cover criticality control rod worth calculations as well as sensitivity analysis and uncertainty quantification. Users wanting to reproduce results from this dataset are required to obtain a license to the SCALE code system for which details on the distribution can be found here: https://www.ornl.gov/scale/releases

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Glassy Carbon Substrate Oxidation Effects on Electrode Stability for Oxygen Evolution Reaction Catalysis Stability Benchmarking

Employing benchmarking metrics to capture the activity and stability of electrocatalysts for the oxygen evolution reaction (OER) in acid is a critical practice that enables meaningful comparison of catalyst material candidates reported throughout the literature. In this work, we find that ubiquitously used glassy carbon electrode substrates oxidize under typical OER operating conditions, forming a pacified, electrically insulating, and oxygen-rich surface layer that causes drastic loss of current density over the course of extended chronoamperometric stability tests at an anodic potential of 1.7 VRHE. We show that the experimentally observed stability of glassy carbon-based electrodes is approximately two orders of magnitude lower than that expected solely from dissolution-based catalyst intrinsic stability of Ir-based catalysts. We additionally find that glassy carbon-based electrode stability measured by chronoamperometric holds is greatly impacted by catalyst loading, with high catalyst loadings improving the stability of the overall electrode via a protective effect on the glassy carbon substrate. Altogether, our investigation highlights that glassy carbon is not electrochemically inert under OER conditions on the timescale of common stability tests, which can cause electrodes to exhibit performance losses that do not reflect the intrinsic stability of the actual catalyst material being investigated. In light of our findings, we underscore the usefulness of metrics, such as the S-number, to reflect intrinsic catalyst material stability.

25 ENERGY STORAGE↗

NAS Parallel Benchmark. Results 11-96: Performance Comparison of HPF and MPI Based NAS Parallel Benchmarks

High Performance Fortran (HPF), the high-level language for parallel Fortran programming, is based on Fortran 90. HALF was defined by an informal standards committee known as the High Performance Fortran Forum (HPFF) in 1993, and modeled on TMC's CM Fortran language. Several HPF features have since been incorporated into the draft ANSI/ISO Fortran 95, the next formal revision of the Fortran standard. HPF allows users to write a single parallel program that can execute on a serial machine, a shared-memory parallel machine, or a distributed-memory parallel machine. HPF eliminates the complex, error-prone task of explicitly specifying how, where, and when to pass messages between processors on distributed-memory machines, or when to synchronize processors on shared-memory machines. HPF is designed in a way that allows the programmer to code an application at a high level, and then selectively optimize portions of the code by dropping into message-passing or calling tuned library routines as 'extrinsics'. Compilers supporting High Performance Fortran features first appeared in late 1994 and early 1995 from Applied Parallel Research (APR) Digital Equipment Corporation, and The Portland Group (PGI). IBM introduced an HPF compiler for the IBM RS/6000 SP/2 in April of 1996. Over the past two years, these implementations have shown steady improvement in terms of both features and performance. The performance of various hardware/ programming model (HPF and MPI (message passing interface)) combinations will be compared, based on latest NAS (NASA Advanced Supercomputing) Parallel Benchmark (NPB) results, thus providing a cross-machine and cross-model comparison. Specifically, HPF based NPB results will be compared with MPI based NPB results to provide perspective on performance currently obtainable using HPF versus MPI or versus hand-tuned implementations such as those supplied by the hardware vendors. In addition we would also present NPB (Version 1.0) performance results for the following systems: DEC Alpha Server 8400 5/440, Fujitsu VPP Series (VX, VPP300, and VPP700), HP/Convex Exemplar SPP2000, IBM RS/6000 SP P2SC node (120 MHz) NEC SX-4/32, SGI/CRAY T3E, SGI Origin2000.

Saini, Subash↗

U.S. Solar Photovoltaic System and Energy Storage Cost Benchmarks, With Minimum Sustainable Price Analysis: Q1 2023

The U.S. Department of Energy's (DOE's) Solar Energy Technologies Office (SETO) aims to accelerate the advancement and deployment of solar technology in support of an equitable transition to a decarbonized economy no later than 2050, starting with a decarbonized power sector by 2035. Its approach to achieving this goal includes driving innovations in technology, hardware, and soft cost reductions to make solar affordable and accessible for all. As part of this effort, SETO must track solar cost trends so it can focus its research and development (R&D) on the highest-impact activities. The benchmarks in this report are bottom-up cost estimates of all major inputs to PV and energy storage system installations. Bottom-up costs are based on national averages and do not necessarily represent typical costs in all local markets. Like last year's report, this year's report includes two distinct sets of benchmarks: minimum sustainable price (MSP) benchmarks and modeled market price (MMP) benchmarks. MSP benchmarks can be interpreted as the minimum price a company needs to charge to remain financially solvent in the long term based on the minimum sustainable prices of all inputs including minimum sustainable profit margins. MMP benchmarks can be interpreted as the actual cash sales price a company charges in the given benchmark period. These simplified estimates are useful for tracking technological progress, but they do not reflect all experiences. In fact, no individual estimate under any approach can reflect the diversity of the PV and storage manufacturing and installation industries. Our residential MMP benchmark ($2.90 per watt direct current [Wdc]) is 24% higher than the MSP benchmark ($2.34/Wdc) and 9% lower than our MMP benchmark ($3.18/Wdc) from Q1 2022 in 2022 U.S. dollars (USD). For community solar, our MMP benchmark ($1.75/Wdc) is 18% higher than our MSP benchmark ($1.49/Wdc). Our Q1 2022 benchmark report has no community solar system for comparison. For utility-scale systems with one-axis tracking, our MMP benchmark ($1.17/Wdc) is 22% higher than our MSP benchmark ($0.96/Wdc) and 10% higher than its counterpart ($1.07/Wdc) in Q1 2022 in 2022 USD.

14 SOLAR ENERGY↗

AI Benchmark Democratization and Carpentry

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model architectures, scale, datasets, and deployment contexts makes evaluation a moving target. Large language models often memorize static benchmarks, causing a gap between benchmark results and real-world performance. Beyond traditional static benchmarks, continuous adaptive benchmarking frameworks are needed to align scientific assessment with deployment risks. This calls for skills and education in AI Benchmark Carpentry. From our experience with MLCommons, educational initiatives, and programs like the DOE's Trillion Parameter Consortium, key barriers include high resource demands, limited access to specialized hardware, lack of benchmark design expertise, and uncertainty in relating results to application domains. Current benchmarks often emphasize peak performance on top-tier hardware, offering limited guidance for diverse, real-world scenarios. Benchmarking must become dynamic, incorporating evolving models, updated data, and heterogeneous platforms while maintaining transparency, reproducibility, and interpretability. Democratization requires both technical innovation and systematic education across levels, building sustained expertise in benchmark design and use. Benchmarks should support application-relevant comparisons, enabling informed, context-sensitive decisions. Dynamic, inclusive benchmarking will ensure evaluation keeps pace with AI evolution and supports responsible, reproducible, and accessible AI deployment. Community efforts can provide a foundation for AI Benchmark Carpentry.

von Laszewski, Gregor [Virginia U.]↗

Health Physics Research Reactor Criticality Accident Alarm System Benchmark Overview

From the countless critical experiments performed in the world during the past century, high-quality integral benchmarks experiments have been collected and gathered into the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook), managed by the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Working Group. This information preservation and dissemination effort is crucial for reactor licensing as well as criticality and radiation transport modeling validation. This summary reports on the status of a tentative benchmark addition to the ICSBEP Handbook. The proposed benchmark arises from legacy operation data of the Oak Ridge National Laboratory (ORNL) Health Physics Research Reactor (HPRR). The HPRR was a small, unmoderated, unshielded fast burst reactor that was used for research in health physics and radiobiology as well as teaching and training. As part of a comprehensive investigation of the available HPRR operation data and characteristics, different possibilities for use of the valuable results were studied. A critical experiment benchmark evaluation was performed, analyzing data coming from sub-critical and critical operation of the HPRR during operator training, steady-state irradiation of samples and before critical bursts. The results of the evaluation do not satisfy for the ICSBEP standards as the benchmark relative standard uncertainty is of about 4% for k eff , and the relative difference between sample calculations and expected k eff results is of about 1.5%. Due to those unsatisfactory results, it was decided not to pursue critical experiments evaluation of the HPRR presently and to focus instead on shielding type data for the creation of a criticality accident alarm system (CAAS) and shielding category benchmark, which is currently very scarce in the ICSBEP handbook—especially concerning critical, pulsed assembly, or reactor operation data. Several dosimetry and shielding experiments from HPRR burst operation were evaluated, with different benchmark metrics as sulfur fluence, Element 57 dose, or neutron fluence at different distances and under different shield materials conditions. An evaluation focusing on the Element 57 neutron dose as a benchmark metric was submitted to the ICSBEP Technical Review Group (TRG) meeting in October 2021, and the inclusion of the evaluation in the ICSBEP Handbook was deferred. The main change proposed by the international experiment evaluation experts is to use the neutron fluence measured by Bonner spheres as a benchmark metric. This represents a quantity closer to that actually measured by the experimentalists of the HPRR compared to the Element 57 dose, which adds another step of data transformation, thus potentially adding uncertainty to the benchmark. The evaluation has been updated and will be submitted to the 2022 ICSBEP TRG meeting for inclusion in the 2023 edition of the ICSBEP Handbook. The evaluation is performed using the KENO and MAVRIC combination from the SCALE 6.2.4 code suite which was previously used in similar CAAS benchmarks to allow for the use of variance reduction techniques.

61 RADIATION PROTECTION AND DOSIMETRY↗

Benchmark Tracking System for Performance Monitoring

Benchmarking is essential for high-performance software development, particularly for monitoring performance across code iterations. This project focused on enhancing the benchmarking process for Lamellar, an asynchronous runtime for High-Performance Computing (HPC) systems developed at Pacific Northwest National Laboratory. Prior to this work, benchmark results were difficult to track and compare across code versions, presenting significant challenges in identifying performance regressions and long-term trends. The primary objective was to establish a systematic, reproducible approach for measuring performance and detecting regressions following code commits. Our methodology involved three key components: standardizing benchmark outputs, implementing data versioning, and developing analysis tools. We standardized the benchmark output format to JSON Line records containing specific fields (execution time, hardware specifications, and environmental variables). To address data management challenges, we evaluated several options and eventually chose a git repository dedicated to benchmark data. We developed a suite of Python tools that processed benchmark results, enriched them with metadata, and facilitated search in the repository. The resulting system enables more efficient filtering and comparison of performance metrics across commit histories, hardware configurations, and benchmark variants through a unified query interface. Our implementation reduces computational overhead by first checking for existing results through configuration matching before initiating new benchmark runs, thereby conserving resources. The system has been validated by Lamellar developers. It organizes results by benchmark type and build configurations for efficient retrieval. Future developments include a planned Large Language Model interface for predicting benchmark performance, incorporating the criterion package for statistical analysis, which will enable automated detection of statistically significant performance changes, and integration with continuous integration pipelines. Despite these enhancements being reserved for future work, this project has successfully provided the Lamellar development team with a framework for maintaining consistent performance standards and identifying optimization opportunities across workloads and hardware environments.

97 MATHEMATICS AND COMPUTING↗

MLCommons Science Benchmarks

Benchmarks are a cornerstone of modern machine learning practice, providing standardized eval- uations that enable reproducibility, comparison, and scientific progress. Yet, as AI systems particularly deep learning models become increasingly dynamic, traditional static benchmarking approaches are losing their relevance. Models rapidly evolve in architecture, scale, and capability; datasets shift; and deployment contexts continuously change, creating a moving target for evaluation. Without adaptive benchmarking frame- works, both scientific assessment and real-world de- ployment risk becoming misaligned with actual system behavior. Drawing on our experience from MLCommons, educa- tional initiatives, and government programs such as the DOE s Million Parameter Consortium, we identify key barriers that hinder the broader adoption and utility of benchmarking in AI. These include substantial resource demands, limited access to specialized hardware, lack of expertise in benchmark design, and uncertainty among practitioners about how to relate benchmark results to their own application domains. Moreover, current benchmarks often emphasize peak performance on leadership-class hardware, offering limited guidance for more diverse, real-world deployment scenarios. We argue that benchmarking itself must become dy- namic in order to incorporate evolving models, updated data, and heterogeneous computational platforms while maintaining transparency, reproducibility, and inter- pretability. Democratizing this process requires not only technical innovation, but also systematic educational efforts spanning undergraduate to professional levels to develop sustained expertise in benchmark design and use. Finally, benchmarks should be framed and com- municated to support application-relevant comparisons, enabling both developers and users to make informed, context-sensitive decisions. Advancing dynamic and inclusive benchmarking practices will be essential to ensure that evaluation keeps pace with the evolving AI landscape and supports responsible, reproducible, and accessible AI deployment.

Hawks, Benjamin G. [Fermilab]↗

Updated Godiva-IV Benchmark Preview

A note of errata prepended to the Godiva-IV delayed-critical benchmark (HEU-MET-FAST- 086) identifies two corrections to be made to the model: The glory hole in the spindle should be made larger, and the height of the safety block should be made smaller (and therefore its density made larger). In addition, the safety block at its full-in position is closer to the inner subassembly plate than was modeled in the benchmark. These changes have been made to HEU-MET-FAST-086 Case 4 in order to estimate the effect on $\kappa$ eff and on the neutron flux spectrum. Using smaller separation, a smaller safety block, and a larger glory hole caused keff to increase by 453 ± 1 pcm from the benchmark. The latest nuclear data, ENDF/B-VIII.0, have also been used, causing $\kappa$ eff to increase another 37 ± 1 pcm. Flux spectra were compared in a modeled fission foil (near the center of Godiva-IV in the glory hole) and at three external (point) detectors. Within the fission foil, using ENDF/B-VIII.0 induced changes in the flux spectrum similar in size to the changes due to using smaller separation, a smaller safety block, and a larger glory hole. At the detectors, using ENDF/B- VIII.0 induced changes in the flux spectrum much larger than those due to changing the model. In other words, the corrections to the benchmark model cause a large increase in $\kappa$ eff , but the changes to the neutron flux spectrum are small compared to those caused by using the latest nuclear data. This study presents a preview of results expected during the reevaluation of the Godiva IV benchmark, but it is not a substitute for the full reevaluation.A note of errata prepended to the Godiva-IV delayed-critical benchmark (HEU-MET-FAST- 086) identifies two corrections to be made to the model: The glory hole in the spindle should be made larger, and the height of the safety block should be made smaller (and therefore its density made larger). In addition, the safety block at its full-in position is closer to the inner subassembly plate than was modeled in the benchmark. These changes have been made to HEU-MET-FAST-086 Case 4 in order to estimate the effect on $\kappa$ eff and on the neutron flux spectrum. Using smaller separation, a smaller safety block, and a larger glory hole caused $\kappa$ eff to increase by 453 ± 1 pcm from the benchmark. The latest nuclear data, ENDF/B-VIII.0, have also been used, causing $\kappa$ eff to increase another 37 ± 1 pcm. Flux spectra were compared in a modeled fission foil (near the center of Godiva-IV in the glory hole) and at three external (point) detectors. Within the fission foil, using ENDF/B-VIII.0 induced changes in the flux spectrum similar in size to the changes due to using smaller separation, a smaller safety block, and a larger glory hole. At the detectors, using ENDF/B- VIII.0 induced changes in the flux spectrum much larger than those due to changing the model. In other words, the corrections to the benchmark model cause a large increase in $\kappa$ eff , but the changes to the neutron flux spectrum are small compared to those caused by using the latest nuclear data. This study presents a preview of results expected during the reevaluation of the Godiva IV benchmark, but it is not a substitute for the full reevaluation.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Summary of LANL Critical Benchmark Comparison Study and Revisions for Cases Involving HEU, LEU, MIX, and Pu

This report documents results obtained for revisions made to cases involving Highly Enriched Uranium (HEU), Intermediate Enriched Uranium (IEU), a mixture of Pu and Uranium (MIX), as well as Pu cases. A previous summary of revisions for HEU an Pu cases was reported and additional investigations into four cases originally presented therein uncovered further revisions which led to better agreement with other transport codes, those cases are updated in this report. The summary of all cases reported in Reference 2 is updated in this report. In addition, a previous summary of revisions for LEU and MIX was reported, a summary of those revisions in reproduced in this report for a comprehensive summary of changes to benchmarks beginning in fiscal year 2020 to current date. The report focuses on the changes made to LANL benchmarks modeled with MCNP6 using ENDF/B-VII.1 nuclear data that appeared to have discrepant results when compared with results of other codes. Feedback was used to pinpoint review of benchmark input files and to revise them when necessary. This report documents the results of review and revision of specific benchmarks highlighted as possibly discrepant in the comparison study. In addition, there is an effort tied to this work involving collaboration between LANL XCP and NCS Divisions in the development of a shared review/revision procedure and use of a new benchmark repository. LANL has a benchmark library of critical experiments from the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook modeled for use with MCNP. This collection is now over 1100 benchmarks, referred to as the Whisper-1.1 library because it is used with the sensitivity/uncertainty package, Whisper, which supports nuclear criticality safety validation and is released with MCNP6.2. The collection, originally created several decades ago, is a combination of smaller collections, which has been revised and expanded, by various groups at LANL over the years. The original authors are no longer at the laboratory and little formal documentation of review and revision of these benchmarks exists today. A branch of the benchmark collection was already the subject of a formal review undertaken by the LANL NCS Division and expanded to include XCP Division.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)↗

Sensitivity and uncertainty of the IFR-1 BISON benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered for adoption by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This paper summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment.For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. Here, the resulting variations in benchmark predictions were normalized to calculate sensitivities. These sensitivities were then used to inform input selections for the UQ.The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of the benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because the confidence intervals associated with benchmark predictions were wide. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗