Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

wa-hls4ml: A Benchmark and Surrogate Models for hls4ml Resource and Latency Estimation

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs. These advancements have solved major obstacles, but also exposed new challenges. For example, processes that were not previously considered bottlenecks, such as hardware synthesis, are becoming limiting factors in the rapid iteration of designs. To mitigate these emerging constraints, multipleefforts have been undertaken to develop an ML-based surrogate model that estimates resource usage of synthesized ML accelerator architectures. We introduce wa-hls4ml, a benchmark for ML accelerator resource and latency estimation, and its corresponding initial dataset of over 680 000 fully connected and convolutional neural networks, all synthesized using hls4ml and targeting Xilinx FPGAs. The benchmark evaluates the performance of resource and latency predictors against several common ML model architectures, primarily originating from scientific domains, as exemplar models, and the average performance across a subset of the dataset. Additionally, we introduce GNN- and transformer-based surrogate models that predict latency and resources for ML accelerators. We present the architecture and performance of the models and find that the models generally predict latency and resources for the 75% percentile within several percent of the synthesized resources on the synthetic test dataset.

Hawks, Benjamin G. [Fermilab]↗

TEX-Chlorine: Highly Enriched Uranium with Chloride Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility

This report documents the final benchmark of the TEX-Chlorine (IER-499) Thermal Epithermal eXperiments (TEX) with highly enriched uranium with chlorine absorbers and high-density polyethylene reflectors and moderators. TEX-Chlorine is a variation of the TEX-HEU baseline assembly with the addition of sodium chloride absorber plates. This evaluation contains three experimental configurations that were performed on Comet at NCERC between July and August 2024. The three configurations, which were acceptable as benchmark cases, spanned from thermal (first two configurations) to fast (third configuration). All three cases were reviewed and accepted by the ICSBEP TRG in April 2025 and was submitted to the ICSBEP in August 2025 after receiving subgroup approval.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Final Design for Thermal/Epithermal eXperiments (TEX) with Lithium Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility

One of the main goals of the Thermal/Epithermal eXperiments (TEX) project is to use existing Nuclear Criticality Safety Program (NCSP) assets to create critical experiment plutonium and uranium test beds for materials important to criticality safety that have insufficient benchmark evaluations. The plutonium test bed experiments were completed in 2018 and are published in the 2020 edition of the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The uranium test bed assemblies were completed in 2023 and accepted in the 2024 edition of the ICSBEP Handbook. The Nuclear Criticality Safety (NCS) group at Y-12 National Security Complex has identified programmatic need for validation cases for uranium electrorefining operations at Y-12. The electrorefining operation credits lithium enriched in 6 Li in addition to 35 Cl as absorbers in the design criticality safety evaluation for precluding criticality under upset conditions in the large and geometrically unfavorable electro-refiner. There is, however, inadequate experimental validation for the 6 Li absorbers. As an extension of the TEX uranium test bed, TEX-Cl critical experiments were performed with sodium chloride salt to address the 35 Cl thermal absorption as well as other validation needs at Los Alamos National Laboratory (LANL). These experiments were completed in 2024 and accepted into the 2025 ICSBEP Handbook. To continue the methodology used in TEX-Cl, TEX-Li aims to accomplish the same. The overall design of both experiments was to use commercially available, high purity, salts with polyethylene moderator and HEU plates to configure a critical assembly. Three experiments are planned for TEX-Li using encapsulated lithium carbonate (Li 2 CO 3 ), with natural 6 Li abundance. For all experiments, the highly enriched uranium (HEU) Jemima plates will be used as fissile material. Multiple layers will be stacked together with encapsulated Li 2 CO 3 alternated with polyethylene in standard configurations. Standard stacking was found to be optimal in matching the different sensitivity profiles provided by the Y-12 models. Three configurations are proposed with varying polyethylene moderation and a constant 1/4” absorber thickness. The first uses 11 layers of 5/4” polyethylene, the second uses 9 layers of 3/4” polyethylene, and the third uses 10 layers of 1/2” polyethylene. Calculations showed that some alternative forms of lithium-based materials provided slightly less-optimal sensitivity profiles when compared to lithium carbonate but come with other drawbacks. These alternatives included lithium aluminate (LiAlO 2 ), Aluminum-2050 alloy, Aluminum-8090, Aluminum-2095, lithium hydride (LiH), and lithium fluoride (LiF). Lithium aluminate and aluminum-2050 provided comparable sensitivity profiles when compared to lithium carbonate and can be used instead if lithium carbonate cannot be readily procured. After a broad material study, lithium carbonate outperformed any alternative material with a balance in affordability and workability. The assessment of experimental uncertainties of the non-absorber and absorber components was predicted to be 0.00089 and 0.00093 Δk eff , respectively. The largest uncertainties may be reduced with precision dimensional inspection of the components. Many of the parts and equipment for IER 575 have already been fabricated or procured for previous projects and therefore do not contribute significantly to the overall cost of this experiment. This includes the Jemima plates and Comet critical assembly machine, which are existing NCSP assets, as well as the aluminum platen and polyethylene reflector rings, which were fabricated and authorized for the TEX experiment involving HEU with polyethylene. Lithium carbonate containers will be procured by LANL and will be filled by LLNL. The total material costs for TEX-Li experiments are estimated to be on the order of $\$$47,400. Precision inspection, including dimensional, mass, density, and impurity, is recommended for all components for an estimated cost of $\$$12,000.

35Cl↗

The Risk Assessment Information System Compendium of Ecological Screening Benchmarks for Chemicals and Radionuclides (2025) (Volume VIII – Chemicals for Sediment: General)

Volume VIII presents the chemical benchmarks for sediment: general. The benchmarks here are applicable to both freshwater sediments and marine sediments or where no differentiation was made by the source. These values are not repeated in the freshwater sediment and marine sediment categories. Likewise, values listed in the freshwater sediment and marine sediment categories are not repeated in the generic categories. Consult Volume I for information on the sources.

54 ENVIRONMENTAL SCIENCES↗

The Risk Assessment Information System Compendium of Ecological Screening Benchmarks for Chemicals and Radionuclides (2025) (Volume V – Chemicals for Surface Water: General)

Volume V presents the chemical benchmarks for surface water: general. The benchmarks here are applicable to both fresh water and marine or where no differentiation was made by the source. These values are not repeated in the fresh water and marine categories. Likewise, values listed in the fresh water and marine categories are not repeated in the generic categories. Consult Volume I for information on the sources.

54 ENVIRONMENTAL SCIENCES↗

ORNL Testing of Multiple Graphite Benchmarks [Slides]

Integral criticality safety and reactor physics benchmark experiments from the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook) and the International Handbook of Reactor Physics Experiments (IRPhE Handbook) are essential for nuclear data validation and testing

ENDF↗

Benchmark Exercise Report for Experimental Study of Bubble Scrubbing in Water Coolant Pool

Mechanistic assessments of radionuclide release during postulated accidents are expected to be included in advanced reactor license applications. The mechanistic source term (MST) provides an opportunity for vendors to realistically evaluate the radiological consequences of an incident, and may aid in justifying reduced emergency planning zones and plant sites. However, the development of MSTs for advanced nuclear reactors is challenging because there are numerous phenomena that can affect the transport and retention of radionuclides. As part of a trial MST assessment for a metal-fueled, pool-type sodium cooled fast reactor (SFR), led by Argonne National Laboratory, a simplified radionuclide transport code (SRT code) was developed, which includes models to estimate the quantity of fission product aerosols scrubbed in the sodium pool during postulated accident scenarios. In a pool-type SFR, when fission products are released into the coolant pool due to failure of fuel pins, most of the radionuclides are scrubbed by the coolant pool, but some have the potential to migrate to the cover gas region through entrainment within gas bubbles. The SRT code contains a model that evaluates this scrubbing behavior and calculates the fraction of fission product aerosols that reach the cover gas. Due to a lack of available validation data for sodium pool scrubbing, the U.S. Department of Energy funded an experiment at the University of Wisconsin-Madison to measure aerosol scrubbing by injecting air bubbles containing aerosol into a coolant pool. Prior to performing an experiment with liquid sodium, a water loop experiment was performed. Their experiment evaluated the effect of changing the aerosol size, aerosol density, aerosol concentration, bubble size, and pool depth on the aerosol scrubbing efficiency of the pool. In this benchmark experiment, the base tests were conducted by repeated tests of isolated bubbles. Afterwards, more prototypic tests with bubble swarms were performed to evaluate the interactions between the bubbles. The bubble swarm test was able to confirm that a larger amount of aerosol scrubbing occurred than the single bubble test. It was also confirmed that as the bubble size, aerosol density, and pool height increase, the extent of pool scrubbing also increases and does not change with the aerosol concentration. In addition, since the degree of scrubbing is the lowest at aerosol sizes between 0.01 and 1 μm, that is, the largest amount of aerosol is emitted, it was confirmed that the analysis of this size in MST is the most important. This benchmark experiment informs the direction of future sodium experiments.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Artificial Intelligence Benchmarking

AI benchmarking is a method for evaluating the effectiveness of an AI model using a set of standardized metrics, for example, high school-level math exams. These benchmarks and their results will enable ranking various AI models based on their effectiveness in performing a specific task.

Krishnan, Anjay [Fermilab]↗

Final Design for Thermal/Epithermal eXperiments (TEX) with Lithium Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility (IER 575, CED-2 Report)

One of the main goals of the Thermal/Epithermal eXperiments (TEX) project is to use existing Nuclear Criticality Safety Program (NCSP) assets to create critical experiment plutonium and uranium test beds for materials important to criticality safety that have insufficient benchmark evaluations. The plutonium test bed experiments were completed in 2018 and are published in the 2020 edition of the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The uranium test bed assemblies were completed in 2023 and accepted in the 2024 edition of the ICSBEP Handbook.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmarking of ENDF/B-VIII.1 Thermal Scattering Library for Hydrogen

In correctly characterizing the energy and momentum transfer at thermal and cold energies between neutrons and its interacting medium, thermal scattering libraries, which details energy states due to the intra- and inter-molecular bond effects for the medium materials, are applied in place of free-gas cross section libraries in a particle transport simulation code. They are essential to the neutron performance of a neutron facility like SNS, where thermalized neutrons from 20 K liquid hydrogen and ambient (~300 K) water are transported to the beamlines for neutron scattering experiments in studying materials. Recently, a new version of thermal scattering library for parahydrogen and orthohydrogen at 14-20 K was developed and to be released in ENDF/B-VIII.1. It is, therefore, important to benchmark its impacts on the prediction of moderator performance due to the updates in the thermal scattering library. In this study, the recent ENDF/B-VIII.1 thermal scattering library was compared to the current ENDF/B-VII.1 one in the neutron performance calculations for the decoupled and coupled hydrogen moderators at SNS under theorized and real working conditions. In addition, the predictions using both thermal scattering libraries were benchmarked to the measurements of moderator performance. The consistency between the libraries was observed mostly for parahydrogen and the difference in orthohydrogen at cold neutron energies was noted.

Lu, Wei [Oak Ridge National Laboratory (ORNL), Oak↗

HamLib: A library of Hamiltonians for benchmarking quantum algorithms and hardware

In order to characterize and benchmark computational hardware, software, and algorithms, it is essential to have many problem instances on-hand. This is no less true for quantum computation, where a large collection of real-world problem instances would allow for benchmarking studies that in turn help to improve both algorithms and hardware designs. To this end, here we present a large dataset of qubit-based quantum Hamiltonians. The dataset, called HamLib (for Hamiltonian Library), is freely available online and contains problem sizes ranging from 2 to 1000 qubits. HamLib includes problem instances of the Heisenberg model, Fermi-Hubbard model, Bose-Hubbard model, molecular electronic structure, molecular vibrational structure, MaxCut, Max- k -SAT, Max- k -Cut, QMaxCut, and the traveling salesperson problem. The goals of this effort are (a) to save researchers time by eliminating the need to prepare problem instances and map them to qubit representations, (b) to allow for more thorough tests of new algorithms and hardware, and (c) to allow for reproducibility and standardization across research studies.

97 MATHEMATICS AND COMPUTING↗

Benchmarking a Tunable Quantum Neural Network on Trapped-Ion and Superconducting Hardware

We implement a quantum generalization of a neural network on trapped-ion and IBM superconducting quantum computers to classify MNIST images, a common benchmark in computer vision. The network feedforward involves qubit rotations whose angles depend on the results of measurements in the previous layer. The network is trained via simulation, but inference is performed experimentally on quantum hardware. The classical-to-quantum correspondence is controlled by an interpolation parameter, $a$, which is zero in the classical limit. Increasing $a$ introduces quantum uncertainty into the measurements, which is shown to improve network performance at moderate values of the interpolation parameter. We then focus on particular images that fail to be classified by a classical neural network but are detected correctly in the quantum network. For such borderline cases, we observe strong deviations from the simulated behavior. We attribute this to physical noise, which causes the output to fluctuate between nearby minima of the classification energy landscape. Such strong sensitivity to physical noise is absent for clear images. We further benchmark physical noise by inserting additional single-qubit and two-qubit gate pairs into the neural network circuits. Our work provides a springboard toward more complex quantum neural networks on current devices: while the approach is rooted in standard classical machine learning, scaling up such networks may prove classically non-simulable and could offer a route to near-term quantum advantage.

FOS: Physical sciences↗

The AWAKEN wind farm benchmark, Part 2: Modeling results

Accurately modeling wind farm performance in complex atmospheric flows remains a challenge. This paper presents the modeling results of the American WAKE experimeNt (AWAKEN) wind farm benchmark, a collaborative effort involving 16 research groups from academia and industry within the International Energy Agency Wind Technology Collaboration Programme Task 57. The study evaluates a diverse suite of simulation tools, ranging from fast-running engineering wake models to high-fidelity large-eddy simulations, against a diurnal case study observed during the AWAKEN campaign. The benchmark utilized a three-phase structure to progressively assess model performance as observational data availability increased. Initial blind predictions showed that higher-fidelity models did not uniformly outperform simpler simulation tools. A distinct spatial bias was observed where models struggled to resolve the interplay between a low-level jet, wakes, and terrain-induced flow acceleration. In subsequent phases, leveraging additional measurements for model improvement led to a reduction in mean absolute error across the model ensemble; however, this effect was most pronounced in engineering wake models, where targeted calibration reduced error by up to 40~\%. Overall, the study demonstrates that inflow characterization remains a primary prerequisite for accuracy, particularly for models relying on coarse forcing datasets. While the limited ability to resolve local terrain-flow interactions under single-day conditions represent a recognized constraint, the overall findings on wake modeling and real-world validation still provide valuable guidance for model application and for mitigating this limitation.

Bodini, Nicola↗

Q1-2024 Solar Cost Benchmarks

Each year, the U.S. Department of Energy’s (DOE) Solar Energy Technologies Office (SETO) and its national laboratory partners develop cost benchmarks for U.S. solar photovoltaic (PV) systems. These benchmarks track progress toward reducing solar costs and guide R&D priorities. Unlike typical studies that report only $/W, SETO uses intrinsic units (e.g., $/m² for mounting structures) to better capture how technology improvements such as module efficiency would impact system costs. This allows flexible modeling where inputs can vary significantly to assess cost sensitivity. Costs are reported in two ways: Minimum Sustainable Price (MSP): Long term, financially viable price under stable market conditions. Modeled Market Price (MMP): Actual market price, influenced by short term distortions such as tariffs or subsidies. Three national labs collect cost data from industry stakeholders, ensuring no duplication in outreach to stakeholders. Data reflects real transactions (primarily from Q1) and is weighted based on the number of sources per cost element. The PV System Cost Model (PVSCM) divides total installed system cost into eight categories: 1. Module (PV) 2. Inverter 3. Energy Storage System (ESS) 4. Structural BOS (SBOS) 5. Electrical BOS (EBOS) 6. Fieldwork 7. Office work 8. Other (developer/EPC costs) The first five are hardware costs, while the last three are soft costs. Each category includes fixed and variable cost components, where “size” depends on context (e.g., manufacturing capacity for modules vs. system capacity for installation costs). Variable costs are expressed using appropriate intrinsic units. The model reflects the owner’s upfront overnight capital cost, excluding tax credits. Tariffs and subsidies are treated as temporary market distortions affecting MMP but not MSP. PVSCM is implemented in Excel, where cost elements are aggregated into total system cost. Additional sheets handle unit conversions and operation & maintenance (O&M), with O&M costs levelized over the system’s lifetime.

14 SOLAR ENERGY↗

The NAS kernel benchmark program

A collection of benchmark test kernels that measure supercomputer performance has been developed for the use of the NAS (Numerical Aerodynamic Simulation) program at the NASA Ames Research Center. This benchmark program is described in detail and the specific ground rules are given for running the program as a performance test.

Bailey, D. H.↗

Benchmarks of programming languages for special purposes in the space station

Although Ada is likely to be chosen as the principal programming language for the Space Station, certain needs, such as expert systems and robotics, may be better developed in special languages. The languages, LISP and Prolog, are studied and some benchmarks derived. The mathematical foundations for these languages are reviewed. Likely areas of the space station are sought out where automation and robotics might be applicable. Benchmarks are designed which are functional, mathematical, relational, and expert in nature. The coding will depend on the particular versions of the languages which become available for testing.

Knoebel, Arthur↗

Benchmarking hypercube hardware and software

It was long a truism in computer systems design that balanced systems achieve the best performance. Message passing parallel processors are no different. To quantify the balance of a hypercube design, an experimental methodology was developed and the associated suite of benchmarks was applied to several existing hypercubes. The benchmark suite includes tests of both processor speed in the absence of internode communication and message transmission speed as a function of communication patterns.

Grunwald, Dirk C.↗