Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Final Design for Thermal/Epithermal eXperiments (TEX) with Lithium Absorbers to Provide Validation Benchmarks for Y-12 Electrorefining Facility (IER 575, CED-2 Report)

One of the main goals of the Thermal/Epithermal eXperiments (TEX) project is to use existing Nuclear Criticality Safety Program (NCSP) assets to create critical experiment plutonium and uranium test beds for materials important to criticality safety that have insufficient benchmark evaluations. The plutonium test bed experiments were completed in 2018 and are published in the 2020 edition of the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook. The uranium test bed assemblies were completed in 2023 and accepted in the 2024 edition of the ICSBEP Handbook.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Benchmarking of ENDF/B-VIII.1 Thermal Scattering Library for Hydrogen

In correctly characterizing the energy and momentum transfer at thermal and cold energies between neutrons and its interacting medium, thermal scattering libraries, which details energy states due to the intra- and inter-molecular bond effects for the medium materials, are applied in place of free-gas cross section libraries in a particle transport simulation code. They are essential to the neutron performance of a neutron facility like SNS, where thermalized neutrons from 20 K liquid hydrogen and ambient (~300 K) water are transported to the beamlines for neutron scattering experiments in studying materials. Recently, a new version of thermal scattering library for parahydrogen and orthohydrogen at 14-20 K was developed and to be released in ENDF/B-VIII.1. It is, therefore, important to benchmark its impacts on the prediction of moderator performance due to the updates in the thermal scattering library. In this study, the recent ENDF/B-VIII.1 thermal scattering library was compared to the current ENDF/B-VII.1 one in the neutron performance calculations for the decoupled and coupled hydrogen moderators at SNS under theorized and real working conditions. In addition, the predictions using both thermal scattering libraries were benchmarked to the measurements of moderator performance. The consistency between the libraries was observed mostly for parahydrogen and the difference in orthohydrogen at cold neutron energies was noted.

Lu, Wei [Oak Ridge National Laboratory (ORNL), Oak↗

HamLib: A library of Hamiltonians for benchmarking quantum algorithms and hardware

In order to characterize and benchmark computational hardware, software, and algorithms, it is essential to have many problem instances on-hand. This is no less true for quantum computation, where a large collection of real-world problem instances would allow for benchmarking studies that in turn help to improve both algorithms and hardware designs. To this end, here we present a large dataset of qubit-based quantum Hamiltonians. The dataset, called HamLib (for Hamiltonian Library), is freely available online and contains problem sizes ranging from 2 to 1000 qubits. HamLib includes problem instances of the Heisenberg model, Fermi-Hubbard model, Bose-Hubbard model, molecular electronic structure, molecular vibrational structure, MaxCut, Max- k -SAT, Max- k -Cut, QMaxCut, and the traveling salesperson problem. The goals of this effort are (a) to save researchers time by eliminating the need to prepare problem instances and map them to qubit representations, (b) to allow for more thorough tests of new algorithms and hardware, and (c) to allow for reproducibility and standardization across research studies.

97 MATHEMATICS AND COMPUTING↗

Benchmarking a Tunable Quantum Neural Network on Trapped-Ion and Superconducting Hardware

We implement a quantum generalization of a neural network on trapped-ion and IBM superconducting quantum computers to classify MNIST images, a common benchmark in computer vision. The network feedforward involves qubit rotations whose angles depend on the results of measurements in the previous layer. The network is trained via simulation, but inference is performed experimentally on quantum hardware. The classical-to-quantum correspondence is controlled by an interpolation parameter, $a$, which is zero in the classical limit. Increasing $a$ introduces quantum uncertainty into the measurements, which is shown to improve network performance at moderate values of the interpolation parameter. We then focus on particular images that fail to be classified by a classical neural network but are detected correctly in the quantum network. For such borderline cases, we observe strong deviations from the simulated behavior. We attribute this to physical noise, which causes the output to fluctuate between nearby minima of the classification energy landscape. Such strong sensitivity to physical noise is absent for clear images. We further benchmark physical noise by inserting additional single-qubit and two-qubit gate pairs into the neural network circuits. Our work provides a springboard toward more complex quantum neural networks on current devices: while the approach is rooted in standard classical machine learning, scaling up such networks may prove classically non-simulable and could offer a route to near-term quantum advantage.

FOS: Physical sciences↗

The AWAKEN wind farm benchmark, Part 2: Modeling results

Accurately modeling wind farm performance in complex atmospheric flows remains a challenge. This paper presents the modeling results of the American WAKE experimeNt (AWAKEN) wind farm benchmark, a collaborative effort involving 16 research groups from academia and industry within the International Energy Agency Wind Technology Collaboration Programme Task 57. The study evaluates a diverse suite of simulation tools, ranging from fast-running engineering wake models to high-fidelity large-eddy simulations, against a diurnal case study observed during the AWAKEN campaign. The benchmark utilized a three-phase structure to progressively assess model performance as observational data availability increased. Initial blind predictions showed that higher-fidelity models did not uniformly outperform simpler simulation tools. A distinct spatial bias was observed where models struggled to resolve the interplay between a low-level jet, wakes, and terrain-induced flow acceleration. In subsequent phases, leveraging additional measurements for model improvement led to a reduction in mean absolute error across the model ensemble; however, this effect was most pronounced in engineering wake models, where targeted calibration reduced error by up to 40~\%. Overall, the study demonstrates that inflow characterization remains a primary prerequisite for accuracy, particularly for models relying on coarse forcing datasets. While the limited ability to resolve local terrain-flow interactions under single-day conditions represent a recognized constraint, the overall findings on wake modeling and real-world validation still provide valuable guidance for model application and for mitigating this limitation.

Bodini, Nicola↗

Q1-2024 Solar Cost Benchmarks

Each year, the U.S. Department of Energy’s (DOE) Solar Energy Technologies Office (SETO) and its national laboratory partners develop cost benchmarks for U.S. solar photovoltaic (PV) systems. These benchmarks track progress toward reducing solar costs and guide R&D priorities. Unlike typical studies that report only $/W, SETO uses intrinsic units (e.g., $/m² for mounting structures) to better capture how technology improvements such as module efficiency would impact system costs. This allows flexible modeling where inputs can vary significantly to assess cost sensitivity. Costs are reported in two ways: Minimum Sustainable Price (MSP): Long term, financially viable price under stable market conditions. Modeled Market Price (MMP): Actual market price, influenced by short term distortions such as tariffs or subsidies. Three national labs collect cost data from industry stakeholders, ensuring no duplication in outreach to stakeholders. Data reflects real transactions (primarily from Q1) and is weighted based on the number of sources per cost element. The PV System Cost Model (PVSCM) divides total installed system cost into eight categories: 1. Module (PV) 2. Inverter 3. Energy Storage System (ESS) 4. Structural BOS (SBOS) 5. Electrical BOS (EBOS) 6. Fieldwork 7. Office work 8. Other (developer/EPC costs) The first five are hardware costs, while the last three are soft costs. Each category includes fixed and variable cost components, where “size” depends on context (e.g., manufacturing capacity for modules vs. system capacity for installation costs). Variable costs are expressed using appropriate intrinsic units. The model reflects the owner’s upfront overnight capital cost, excluding tax credits. Tariffs and subsidies are treated as temporary market distortions affecting MMP but not MSP. PVSCM is implemented in Excel, where cost elements are aggregated into total system cost. Additional sheets handle unit conversions and operation & maintenance (O&M), with O&M costs levelized over the system’s lifetime.

14 SOLAR ENERGY↗

Status of the CERBERUS Evaluation for the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook

Modeling & Simulation (M&S) tools are used to analyze advanced reactor designs and the safety of current nuclear operations. As computers continue to improve, we are able to enhance resolution in our calculations. Therefore, the limitations of simulation capability are in the quality of data that is being used, including our ability to quantify the uncertainty and sensitivity of that data. In order to model systems of interest with increasing accuracy, the industry must improve key nuclear data measurements. The International Criticality Safety Benchmark Evaluation Project (ICSBEP) compiles and evaluates experiment data in a handbook that can be used by criticality safety engineers and others to validate computer codes and cross section libraries at nuclear facilities. Both critical and subcritical experiments are included in the handbook. These experiments, along with differential measurements, can help improve the quality of nuclear data. Concerns regarding the accuracy of Cu nuclear data have been published. The large values and trend of C-E for the Zeus intermediate energy benchmark, being one of the primary examples. Furthermore, very few experiments have been designed to be sensitive to Cu (as shown in Figure 1), so an integral, critical experiment is needed to help resolve these differences.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

NEA HTTR LOFC Project Test#3 Benchmark Results

In the second half of FY23, the High Temperature Engineering Test Reactor (HTTR) Loss Of Forced Cooling (LOFC)#3 data for the 9 MW test case with Vessel Cooling System (VCS) off were made available through the Nuclear Energy Agency (NEA) LOFC project framework; the neutronic model developed for the initial test (LOFC#1) achieved a satisfactory level of maturity, demonstrating its accuracy in predicting power evolution and core re-criticality, but LOFC#3 should be used primarily to investigate thermal hydraulic phenomena, as the reactor was shut down prematurely due to overheating in the upper reactor components, which prevented re-criticality; this report focuses on advancing the HTTR thermal hydraulic model to accurately simulate the LOFC#3 scenario, including simulating the LOFC#3 benchmark and generating the corresponding benchmark specifications, aiming to ensure consistency across participant models and provide essential data for future participants, including private industry stakeholders seeking to validate their computational tools.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers

Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.

Naug, Avisek [Hewlett Packard Enterprise]↗

Benchmarking quantum trial wavefunctions for phaseless auxiliary-field quantum Monte Carlo

The phaseless auxiliary-field quantum Monte Carlo (ph-AFQMC) method is a stochastic imaginary-time projection technique for computing ground-state properties of strongly correlated quantum systems, with accuracy that depends critically on the choice of trial wavefunction. Here, we investigate ph-AFQMC with trial states prepared using parameterized quantum circuits. In this work, we present a comprehensive benchmarking study of quantum trial wavefunctions spanning unitary coupled-cluster, Hamiltonian-informed, Jastrow-inspired, and adaptively constructed ansatze. The benchmarking evaluates accuracy, expressibility, and scalability of these ansatze within the QC-AFQMC framework. We test these ansatze on linear hydrogen chains under bond stretching and find that several ansatz families produce chemically accurate ph-AFQMC energies across the dissociation curve. We have performed simulations using the CUDA-Q quantum development platform on the GPU partition of the Perlmutter supercomputer. When comparing ansatze at similar numbers of variational parameters, we find that different ansatz families yield comparable ph-AFQMC results despite exhibiting substantially different variational energies, optimization costs, and circuit depths. Our results indicate that the variational energy of an ansatz is not always a reliable indicator of its quality for ph-AFQMC and reveal instances of over-parameterization. In the strongly correlated regime, trial wavefunctions obtained from adaptive ansatze, exemplified here by ADAPT-VQE with the UCCSD operator pool, can outperform their fixed-ansatz counterparts (UCCSD) in terms of projected energies while using substantially more compact circuits, providing a flexible route to optimize quantum resources within the ph-AFQMC framework.

Rofougaran, Rod [LBNL, Berkeley; Columbia U.; PNL,↗

Hydrogen Bond Benchmark: Focal‐Point Analysis and Assessment of DFT Functionals

We performed a hierarchical, convergent ab initio benchmark study and systematically analyzed the performance of density functional approximations for describing hydrogen bonds in small neutral, cationic, and anionic complexes, as well as in larger systems involving amide, urea, deltamide, and squaramide moieties. Focal point analyses (FPA), extrapolating to the ab initio limit, were carried out using correlated wave function methods up to CCSDT(Q) for the small complexes and CCSD(T) for the larger systems, together with correlation-consistent Gaussian basis sets up to the complete basis set limit. Optimized geometries and vibrational frequencies were obtained at the CCSD(T) level. The resulting FPA hydrogen-bond energies converge within a few tenths of a kcal mol −1 . These reference data were used to evaluate 60 density functionals (including 12 dispersion-corrected), spanning the local-density approximation (LDA), generalized gradient approximations (GGAs), meta-GGAs, hybrids, meta-hybrids, double-hybrids, and range-separated hybrids. Overall, the meta-hybrid M06-2X provides the best performance for both hydrogen bond energies and geometries, while the dispersion-corrected GGAs BLYP-D3(BJ) and BLYP-D4 also yield accurate hydrogen-bond data and can serve as cost-effective options for studying large and complex systems.

coupled cluster theory↗

Benchmark study of the DTU OWC chamber with both two-way and one-way absorption

This paper reports on a benchmark study based on small-scale (1:50) measurements of a single, oscillating water column chamber mounted sideways in a long flume. The geometry of the OWC chamber is extracted from a barge-like, attenuator-type floating concept “KNSwing” with 40 chambers targeted for deployment in the Danish part of the North Sea. In addition to traditional two-way energy extraction we also consider one-way energy extraction with passive venting and compare chamber response, pressures and total absorbed energy between the two methods. A blind study was established for the numerical modeling, with participants applying several implementations of weakly nonlinear potential flow theory and commercial Navier–Stokes solvers (CFD). Both compressible and incompressible models were used for the air phase. Potential flow calculations predict more energy absorption near the chamber resonance for one-way absorption than for two-way absorption, but the opposite is found from the experimental measurements. This outcome is mainly attributed to energy losses in the experimental passive valve system, but this conclusion must be confirmed by better experimental measurements. Modeling the one-way valve in CFD proved to be very challenging and only one team was able to provide results which were generally closer to the experiments. The study illustrates the challenges associated with both numerical and experimental analysis of OWC chambers. Air compressibility effects were not found to be important at this scale, even with the large volume of additional air used for the one-way case.

16 TIDAL AND WAVE POWER↗

Using Electrochemistry to Benchmark, Understand, and Develop Noble Metal Nanoparticle Syntheses

The complex chemical nature of metal nanoparticle synthesis presents obstacles for the mechanistic understanding of nanoparticle growth and predictive synthesis design, despite significant progress in this area. Real-time characterization of the chemical processes that take place throughout nanoparticle growth will enable progress toward addressing outstanding challenges in metal nanoparticle synthesis, such as mitigating synthetic reproducibility issues, defining chemical mechanisms that direct nanoparticle growth, and designing synthetic conditions for previously unachievable combinations of nanoparticle shape and composition. In this Perspective, we present open-circuit potential (OCP) measurements as an in situ, real-time method for characterizing chemical changes during nanoparticle growth and discuss the method’s strengths in comparison to and in combination with other characterization techniques. We propose the use of OCP measurements as benchmarks for troubleshooting irreproducibility and streamlining synthetic optimization. Finally, we explore possibilities for using the increased parameter space accessible by electrodeposition to accelerate the development of shape-selective nanoparticle syntheses.

benchmarking↗

Integral Nuclear Data and Benchmarking Needs for Fusion Energy Systems

Fusion energy systems are currently being designed and optimized using radiation transport codes. To deal with the unique environment inside a fusion-based system, many of these designs incorporate novel materials able to withstand the high radiation fields, ensure adequate cooling and thermal protection, and produce tritium. Validation plays a vital role in building trust in the predictive power of these models and computational methods. Validation of a code consists of modeling documented real-world experiments and comparing the code-predicted response to the measured response. Adequate validation requires measured responses from real-world experiments, also known as integral data, that mimic the system being designed, including materials, impinging radiation, and temperature, among other variables. The most trusted integral data are experimental responses that have been through a rigorous benchmarking process that develops a recommended computational model and evaluates all experimental uncertainties. Finally, there are a few research groups around the world that have been producing integral data for fusion applications, but a substantial investment is needed to address the unique validation needs of the fusion community.

Fusion↗

A code-to-code benchmark for magneto-convection in a horizontal duct

Liquid metals and magnetic fields are used in many technical applications such as metallurgy, crystal growth and nuclear fusion reactors. When an electrically conducting fluid moves in a magnetic environment, electric currents and electromagnetic forces are generated that affect velocity and pressure losses in the flow. These magnetohydrodynamic (MHD) interactions have to be investigated to optimize the engineering processes. The characteristics of MHD flows depend on the geometrical configuration, the strength of the applied magnetic field, the electrical properties of fluid and structural materials and the thermal conditions. In the so-called blankets for fusion reactors, where liquid metals are used to breed the plasma fuel component tritium and to extract the generated heat, magneto-convective flows play a crucial role in determining heat and mass transfer. Therefore, the availability of numerical codes to simulate this type of flow is mandatory and their validation is a necessary step to guarantee the reliability of the results. For that reason, a benchmark problem has been defined to simulate liquid metal flows in a horizontal rectangular duct heated from below and exposed to a non-uniform magnetic field. Results obtained by five research groups using different codes are compared.

benchmark↗

Benchmarking of three DWM-based wake models at below-rated wind speeds

Wind turbine wake models are essential tools for predicting power losses and structural loads in wind farms. Among these, the dynamic wake meandering (DWM) model, included as a recommended approach in the International Electrotechnical Commission design standard, is a widely used engineering-fidelity method that balances accuracy and computational cost. This study compares the performance of three DWM-based wake model implementations (from the Technical University of Denmark, the National Renewable Energy Laboratory, and the Institute for Energy Technology) under below-rated wind speed conditions. Model predictions of wake flow, power output, and structural loads for a four-turbine row are evaluated across different ambient turbulence levels and wind-direction misalignments and compared against high-fidelity large-eddy simulation results. All three models captured the overall wake evolution and mean turbine performance with reasonable accuracy; their predicted time-averaged thrust and power were typically within 5 %–10 % of the large-eddy simulation benchmark. However, notable differences emerged in wake structure and unsteady load predictions, with discrepancies increasing for turbines further downstream. These differences highlight the importance of modelling choices such as wake summation and turbulence treatment, which strongly influence power-deficit and fatigue-load predictions. Comparison with large-eddy simulations reveals each approach's strengths and weaknesses, indicating where improvements are needed. Overall, the findings point to specific refinements for DWM models to improve their fidelity, ultimately enabling more robust wake predictions for wind farm design and operation.

17 WIND ENERGY↗