Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarks”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multiple efforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680,000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin [Fermilab] (ORCID:0000000157000288↗

Benchmarking the performance of uncertainty quantification methods for neural network-based interatomic potentials

Machine-learned interatomic potentials (ML-IAPs) continue to gain popularity as accurate, computationally efficient replacements for traditional, physics-based interatomic potentials and expensive ab initio methods. Uncertainty quantification (UQ) of ML-IAPs is a growing area of research as UQ is critical in many applications of IAPs, such as developing curated datasets, active learning-based data augmentation, self-improving models, and estimating the uncertainty of molecular dynamics simulations. In this paper, we construct and benchmark a series of different neural network potentials (NNPs) with varying network architectures to determine the performance of these models with respect to both the mean and uncertainty calibration error. Each NNP method is specifically designed to predict either epistemic or aleatoric uncertainty with particular focus on the differences in behavior between the epistemic and aleatoric uncertainty estimates. We benchmark these methods using multiple datasets common in the ML-IAP literature. The results show that the aleatoric uncertainty from single-shot model architectures is a competitive alternative to ensemble-based epistemic uncertainty predictions in regions of sufficient data-density. However, in regions where the representative data is sparse, aleatoric uncertainty models tend to overpredict and epistemic methods tend to underpredict the actual model error. We conclude that the type of UQ is crucial when discussing performance of probabilistic model results as different methods have different performance characteristics depending on the regime in which they are evaluated. Therefore, the type of UQ method should be carefully evaluated against both the data characteristics and requirements for the intended application.

97 MATHEMATICS AND COMPUTING↗

Reference solutions for linear radiation transport: the Hohlraum and Lattice Benchmarks

Radiation transport describes the propagation of energetic particles through space as they interact with a surrounding material medium. In a kinetic description, radiation transport is modeled by a radiation transport equation (RTE) that prescribes the density of the radiation in position-momentum phase space. The purpose of this dataset is to provide highly resolved solutions to two benchmark problems. These two benchmarks do not possess exact solutions; moreover, the construction of a manufactured solution may require a non-physical source that is not desirable, especially if it spoils the physical nature of the solution. Thus the goal of this computational study is to provide a highly resolved reference solution for testing newer, more cost efficient methods that are currently being developed in the research community.

97 MATHEMATICS AND COMPUTING↗

Li1−xNiO2 Many-body DMC Benchmark Dataset

The dataset contains all numerical data generated in support of the manuscript “Many‑body Benchmark of Electronic Charge and Spin Densities for Li1–xNiO2​” (Journal of Chemical Theory and Computation, DOI: 10.1021/acs.jctc.5c02097, URL: https://pubs.acs.org/doi/10.1021/acs.jctc.5c02097). The materials included in this repository are: 1. Data files used to produce all figures and tables in the main manuscript and supporting information. 2. Benchmark density‑functional theory (DFT) datasets used for the charge‑ and spin‑density analyses. 3. Reference many‑body diffusion Monte Carlo (DMC) calculations and associated input/output files.

36 MATERIALS SCIENCE↗

Wire-arc Additive Manufacturing Benchmark

This is the dataset associated with the 2022 SRP Additive Manufacturing Prediction Challenge, originally hosted on Github at https://github.com/SRP-AM/SRP_AM_Prediction_Challenge. The benchmark was designed for validating prediction for the temperature history, residual stress, and distortion of an additively manufactured metal part with relatively simple geometry. A calibration problem with the same as-built geometry is provided with measured quantities of interest; including temperature histories at selective locations, post-build residual stress at selective locations, and overall distortion measurements. The challenge problem is presented with a different build sequence (i.e. thermal history). In this dataset, we include the actual recorded calibration and challenge measurements, as well as benchmark template files for testing predictions without incorporating the challenge data. Supplementary files around the materials and setup are available for transparency and reproducibility.

Bachus, Nicholas [UC Davis, Davis, CA]↗

Studying the Random Number Generators in MCNP6 using an Analytic Benchmark

An analytic solution to a previously studied toy problem is derived and used as a code verification benchmark. Using various Random Number Generators (RNGs) in MCNP6, including the newest SFC64 RNG available in MCNP6.3.1, and their various properties (e.g., RNG stride), we show how these RNGs perform and how to correct or workaround potential issues with respect to the analytic benchmark problem.

97 MATHEMATICS AND COMPUTING↗

Modeling Enhancements, Cross-Section Generation Updates, and Benchmarking with Shift

This technical report documents the modeling enhancements, cross-section generation updates, and bench marking with the Shift Monte Carlo code performed under the US Department of Energy Nuclear Energy Advanced Modeling and Simulation Program in FY 2024. The work performed included several modeling enhancements, such as integration of cross-section generation in Titan and the ability to produce microscopic multigroup cross sections with Shift. Benchmarking of the cross sections produced by Shift and the two-step workflow with Griffin was performed for three problems: the Advanced Breeder Test Reactor, a generic pebble bed reactor, and a TRISO heat pipe microreactor. Comparisons of results from these benchmark problems were done with Serpent, OpenMC, and Griffin. These enhancements provide a robust foundation for applying Shift for both reference and two-step neutronics analysis for advanced reactor simulation.

97 MATHEMATICS AND COMPUTING↗

Benchmark Study Matrix for Microreactor Geometries Relevant to Multiple Developers

A benchmark study was developed that include design, development, manufacturing, and performance measurement of agnostic reactor relevant geometries to support industry’s adoption of advanced manufacturing in a variety of structures. A matrix of five microreactor component geometries, specifically based on the recent feasibility study on a Marvel microreactor liner, was developed and the Pacific Northwest National Laboratory team initiated one material/process combination, namely 316H using laser powder directed energy deposition (DED). Although the benchmark starts initially with simplistic cubical and cylindrical forms, it builds up to a mock-up of a non-proprietary design that can demonstrate a variety of features potentially useful for presenting knowledge to specific designers of microreactors and for the matter also for other reactor type designers. The initial cubical and cylindrical forms are initial steps to obtain surface features and dimensional responses to the identified process parameters to be used in the non-proprietary design mockup.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Final Design for Additional Thermal/Epithermal eXperiments (TEX) with Sodium Chloride Absorbers to Provide Validation Benchmarks for TerraPower

The first set of Thermal/Epithermal eXperiments (TEX) with chlorine absorbers (TEX-Cl) were executed in Q4FY24 and are in the process of being benchmarked for the ICSBEP. TEX-Cl builds upon the TEX-HEU baseline cases that were published in the 2022 International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. TEX-HEU, like TEX-Pu, was designed to be modular to allow for the incorporation of various absorbers and reflectors to test nuclear data and application case needs. For example, TEX-HEU with hafnium (TEX-Hf) utilizes hafnium plates as both absorbers and reflectors, depending on the tested configuration. A second set of chlorine experiments, dubbed More TEX-Cl, are laid out in this report to meet the needs of TerraPower for chlorine validation for their Molten Chloride Fast Reactor (MCFR) systems. TerraPower’s Molten Chloride Reactor Experiment (MCRE) and MCFR are fast molten salt reactors that utilize sodium chloride (NaCl) salt eutectics as the fuel and coolant. The MCRE eutectic is a mixture of NaCl and uranium trichloride (UCl 3 ). An abundant need for chlorine absorption validation has been expressed by multiple members of the community, including Y-12 (whose needs were addressed with the first set of experiments), LANL (whose needs were addressed with the Chlorine Worth Study (CWS)), TerraPower, institute de radioprotection et de sûreté nucléaire (IRSN), Savannah River Nuclear Solutions (SNRS), and others. Of the members who have expressed interest in this validation, most are interested in the fast neutron energy region, where the 35 Cl(n,p) reaction is most prominent. New 35 Cl(n,p) differential cross section measurements performed by LANL at LANCSE show substantial changes to the cross sections (Figure 1) and may be validated through these experiments as some configurations are optimally sensitive to this cross section.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multipleefforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of synthesized ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680 000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin G. [Fermilab]↗

TEX-Chlorine: Highly Enriched Uranium with Chloride Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility

This report documents the final benchmark of the TEX-Chlorine (IER-499) Thermal Epithermal eXperiments (TEX) with highly enriched uranium with chlorine absorbers and high-density polyethylene reflectors and moderators. TEX-Chlorine is a variation of the TEX-HEU baseline assembly with the addition of sodium chloride absorber plates. This evaluation contains three experimental configurations that were performed on Comet at NCERC between July and August 2024. The three configurations, which were acceptable as benchmark cases, spanned from thermal (first two configurations) to fast (third configuration). All three cases were reviewed and accepted by the ICSBEP TRG in April 2025 and was submitted to the ICSBEP in August 2025 after receiving subgroup approval.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Final Design for Thermal/Epithermal eXperiments (TEX) with Lithium Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility

One of the main goals of the Thermal/Epithermal eXperiments (TEX) project is to use existing Nuclear Criticality Safety Program (NCSP) assets to create critical experiment plutonium and uranium test beds for materials important to criticality safety that have insufficient benchmark evaluations. The plutonium test bed experiments were completed in 2018 and are published in the 2020 edition of the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The uranium test bed assemblies were completed in 2023 and accepted in the 2024 edition of the ICSBEP Handbook. The Nuclear Criticality Safety (NCS) group at Y-12 National Security Complex has identified programmatic need for validation cases for uranium electrorefining operations at Y-12. The electrorefining operation credits lithium enriched in 6 Li in addition to 35 Cl as absorbers in the design criticality safety evaluation for precluding criticality under upset conditions in the large and geometrically unfavorable electro-refiner. There is, however, inadequate experimental validation for the 6 Li absorbers. As an extension of the TEX uranium test bed, TEX-Cl critical experiments were performed with sodium chloride salt to address the 35 Cl thermal absorption as well as other validation needs at Los Alamos National Laboratory (LANL). These experiments were completed in 2024 and accepted into the 2025 ICSBEP Handbook. To continue the methodology used in TEX-Cl, TEX-Li aims to accomplish the same. The overall design of both experiments was to use commercially available, high purity, salts with polyethylene moderator and HEU plates to configure a critical assembly. Three experiments are planned for TEX-Li using encapsulated lithium carbonate (Li 2 CO 3 ), with natural 6 Li abundance. For all experiments, the highly enriched uranium (HEU) Jemima plates will be used as fissile material. Multiple layers will be stacked together with encapsulated Li 2 CO 3 alternated with polyethylene in standard configurations. Standard stacking was found to be optimal in matching the different sensitivity profiles provided by the Y-12 models. Three configurations are proposed with varying polyethylene moderation and a constant 1/4” absorber thickness. The first uses 11 layers of 5/4” polyethylene, the second uses 9 layers of 3/4” polyethylene, and the third uses 10 layers of 1/2” polyethylene. Calculations showed that some alternative forms of lithium-based materials provided slightly less-optimal sensitivity profiles when compared to lithium carbonate but come with other drawbacks. These alternatives included lithium aluminate (LiAlO 2 ), Aluminum-2050 alloy, Aluminum-8090, Aluminum-2095, lithium hydride (LiH), and lithium fluoride (LiF). Lithium aluminate and aluminum-2050 provided comparable sensitivity profiles when compared to lithium carbonate and can be used instead if lithium carbonate cannot be readily procured. After a broad material study, lithium carbonate outperformed any alternative material with a balance in affordability and workability. The assessment of experimental uncertainties of the non-absorber and absorber components was predicted to be 0.00089 and 0.00093 Δk eff , respectively. The largest uncertainties may be reduced with precision dimensional inspection of the components. Many of the parts and equipment for IER 575 have already been fabricated or procured for previous projects and therefore do not contribute significantly to the overall cost of this experiment. This includes the Jemima plates and Comet critical assembly machine, which are existing NCSP assets, as well as the aluminum platen and polyethylene reflector rings, which were fabricated and authorized for the TEX experiment involving HEU with polyethylene. Lithium carbonate containers will be procured by LANL and will be filled by LLNL. The total material costs for TEX-Li experiments are estimated to be on the order of $\$$47,400. Precision inspection, including dimensional, mass, density, and impurity, is recommended for all components for an estimated cost of $\$$12,000.

35Cl↗

The Risk Assessment Information System Compendium of Ecological Screening Benchmarks for Chemicals and Radionuclides (2025) (Volume VIII – Chemicals for Sediment: General)

Volume VIII presents the chemical benchmarks for sediment: general. The benchmarks here are applicable to both freshwater sediments and marine sediments or where no differentiation was made by the source. These values are not repeated in the freshwater sediment and marine sediment categories. Likewise, values listed in the freshwater sediment and marine sediment categories are not repeated in the generic categories. Consult Volume I for information on the sources.

54 ENVIRONMENTAL SCIENCES↗

The Risk Assessment Information System Compendium of Ecological Screening Benchmarks for Chemicals and Radionuclides (2025) (Volume V – Chemicals for Surface Water: General)

Volume V presents the chemical benchmarks for surface water: general. The benchmarks here are applicable to both fresh water and marine or where no differentiation was made by the source. These values are not repeated in the fresh water and marine categories. Likewise, values listed in the fresh water and marine categories are not repeated in the generic categories. Consult Volume I for information on the sources.

54 ENVIRONMENTAL SCIENCES↗

ORNL Testing of Multiple Graphite Benchmarks [Slides]

Integral criticality safety and reactor physics benchmark experiments from the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook) and the International Handbook of Reactor Physics Experiments (IRPhE Handbook) are essential for nuclear data validation and testing

ENDF↗

Benchmark Exercise Report for Experimental Study of Bubble Scrubbing in Water Coolant Pool

Mechanistic assessments of radionuclide release during postulated accidents are expected to be included in advanced reactor license applications. The mechanistic source term (MST) provides an opportunity for vendors to realistically evaluate the radiological consequences of an incident, and may aid in justifying reduced emergency planning zones and plant sites. However, the development of MSTs for advanced nuclear reactors is challenging because there are numerous phenomena that can affect the transport and retention of radionuclides. As part of a trial MST assessment for a metal-fueled, pool-type sodium cooled fast reactor (SFR), led by Argonne National Laboratory, a simplified radionuclide transport code (SRT code) was developed, which includes models to estimate the quantity of fission product aerosols scrubbed in the sodium pool during postulated accident scenarios. In a pool-type SFR, when fission products are released into the coolant pool due to failure of fuel pins, most of the radionuclides are scrubbed by the coolant pool, but some have the potential to migrate to the cover gas region through entrainment within gas bubbles. The SRT code contains a model that evaluates this scrubbing behavior and calculates the fraction of fission product aerosols that reach the cover gas. Due to a lack of available validation data for sodium pool scrubbing, the U.S. Department of Energy funded an experiment at the University of Wisconsin-Madison to measure aerosol scrubbing by injecting air bubbles containing aerosol into a coolant pool. Prior to performing an experiment with liquid sodium, a water loop experiment was performed. Their experiment evaluated the effect of changing the aerosol size, aerosol density, aerosol concentration, bubble size, and pool depth on the aerosol scrubbing efficiency of the pool. In this benchmark experiment, the base tests were conducted by repeated tests of isolated bubbles. Afterwards, more prototypic tests with bubble swarms were performed to evaluate the interactions between the bubbles. The bubble swarm test was able to confirm that a larger amount of aerosol scrubbing occurred than the single bubble test. It was also confirmed that as the bubble size, aerosol density, and pool height increase, the extent of pool scrubbing also increases and does not change with the aerosol concentration. In addition, since the degree of scrubbing is the lowest at aerosol sizes between 0.01 and 1 μm, that is, the largest amount of aerosol is emitted, it was confirmed that the analysis of this size in MST is the most important. This benchmark experiment informs the direction of future sodium experiments.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Artificial Intelligence Benchmarking

AI benchmarking is a method for evaluating the effectiveness of an AI model using a set of standardized metrics, for example, high school-level math exams. These benchmarks and their results will enable ranking various AI models based on their effectiveness in performing a specific task.

Krishnan, Anjay [Fermilab]↗