Engineering PapersSearch

SEARCH · Engineering Papers

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

A Benchmarking Framework for Evaluating Large Language Model Capabilities in Nuclear Reactor Safety Applications

Large language models (LLMs) are increasingly capable of answering technical questions, synthesizing domain knowledge, and supporting engineering workflows. For nuclear science and engineering, these capabilities require careful, domain-specific evaluation before they can be credibly incorporated into safety-related activities, regulatory review, or technical decision support. This paper presents preliminary results from benchmarking framework for evaluating LLM capabilities in nuclear contexts. The framework is organized into three evaluation categories: nuclear fundamentals, general dual-use knowledge, and plant specific knowledge. These categories are intended to distinguish general nuclear engineering competence from broader technical reasoning and more context-dependent nuclear knowledge. Initial evaluations focus on nuclear fundamentals using questions representative of the knowledge expected of a nuclear professional engineer. Results indicate that contemporary frontier models perform at a high level and substantially exceed the performance of older model generations, with some models approaching saturation of the current benchmark. These findings suggest both the rapid improvement of LLM capabilities in specialized technical domains and the need for more discriminating evaluation methods. The paper presents the benchmark structure, preliminary model-comparison results, and ongoing work. This work supports development of verifiable, responsible, and safety-conscious methods for assessing AI systems in nuclear engineering applications.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Reproducible benchmark for the SNAP 8 experimental reactor at operating conditions

This work presents fully reproducible multiphysics benchmark models of the Systems for Nuclear Auxiliary Power (SNAP) 8 Experimental Reactor at operating conditions with coolant flow. Wet experiment (with coolant, at power) validation benchmarks are presented using both deterministic (Serpent-Griffin) and Monte-Carlo (OpenMC-Cardinal) multiphysics frameworks coupled with thermal-hydraulic solvers in MOOSE. Reactivity coefficient measurements including fuel temperature, isothermal temperature, and power coefficients show good agreement with experiments, with discrepancies within experimental uncertainty. Reactivity worth experiments for coolant, samarium, and xenon poisoning are reproduced with differences under 200 pcm. Comparison between Serpent-Griffin and OpenMC-Cardinal frameworks reveal multiphysics coupling introduces positive reactivity effects (100-200 pcm) compared to uniform temperature and density fields at nominal operating conditions. Comparison between Serpent-Griffin and reference Serpent solution shows that power distributions maintain consistent radial and axial peaking behavior. All models, assumptions, thermophysical and thermomechanical properties, and material definitions are thoroughly documented with cited references; model inputs and model generating scripts are stored in the snapReactors GitHub repository.

SNAP

Electrocatalytic benchmarking of ruthenium-based bimetallic anodes for the electrocatalytic oxidation of biomass-derived wastewater

In this paper, we report on the synthesis, characterization, and use of ruthenium oxide (RuO 2 ) doped with a secondary metal (M2) to enhance electrochemical activity and stability for the electrocatalytic oxidation (ECO) of biomass-derived wastewaters. We used different electrochemical methods such as cyclic voltammetry (CV), electrochemical surface area (ECSA), and Tafel analysis as well as physical characterization such as grazing incidence X-ray diffraction, X-ray photoelectron spectroscopy, and scanning electron microscopy to understand how the introduction of M2 affects electrochemical performance. Our results show that including an M2 improves the ECO performance regardless of the composition of the electrolyte. Specifically, we saw increase in ECSA, which could be due to enhanced charge transfer for the pH ranges evaluated. Furthermore, the introduction of organic compounds in wastewater generated during the hydrothermal liquefaction of food waste affected the ECO performance differently, depending on M2, the electrolyte composition, and anodic half-cell potential, highlighting the need to properly control the reaction conditions when testing and characterizing the electrocatalysts under different reaction regimes. We developed an in situ electrocatalytic benchmarking protocol, Boruah – Lopez-Ruiz – Strange (BLoRS), to quickly assess if the presence of M2 improves the ECO performance; thus, saving time and resources in ex situ characterization, testing, and product analysis. This foundational work provides the basis for characterization and benchmarking of electrodes of the ECO of organic compounds.

Electrocatalysis

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE

Surrogate model evaluation and building energy benchmarking for commercial buildings

Building energy consumption benchmarking involves challenges associated with various energy patterns for different building types; heating, ventilating, and air-conditioning (HVAC) system types; and climates. Given significant variation in energy use patterns, accurate prediction of long-term energy use using surrogate models remains challenging. Multiple linear regression (MLR) is commonly used for building energy benchmarking because of its simple structure; however, it lacks accuracy compared to other black-box models. Although many studies have compared surrogate models and offer guidance on model selection based on metrics, they do not provide detailed analysis on improving the surrogate model accuracy. In this paper, we implement a surrogate model using polynomial ridge regression (i.e., MLR with interaction terms combined with ridge regularization) for small office and retail strip mall buildings across six HVAC system types and all climate zones, for electricity and natural gas in baseline and proposed scenarios. A simulation workflow is developed using OpenStudio TM /EnergyPlus TM to generate simulation data using measures over a wide range of efficiency inputs. Enhancements based on statistical insights are used for improving the model accuracy using filters, input transformations, and change points. Surrogate models achieved average coefficient of variation of the root mean squared error (CVRMSE) values of 2.17, 1.06, 2.05, and 3.26 for proposed electricity, proposed natural gas, baseline electricity, and baseline natural gas, respectively, with enhancements reducing CVRMSE by an average of 14.9% across all combinations. We provide model interpretation via Shapley additive explanations to determine which input variables most influence energy consumption and provide supportive arguments for enhancements.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Practical Introduction to Benchmarking and Characterization of Quantum Computers

Rapid progress in quantum technology has transformed quantum computing and quantum information science from theoretical possibilities into tangible engineering challenges. Breakthroughs in quantum algorithms, quantum simulations, and quantum error correction are bringing useful quantum computation closer to fruition. These remarkable achievements have been facilitated by advances in quantum characterization, verification, and validation (QCVV). QCVV methods and protocols enable scientists and engineers to scrutinize, understand, and enhance the performance of quantum information-processing devices. In this tutorial, we review the fundamental principles underpinning QCVV, and introduce a diverse array of QCVV tools used by quantum researchers. We define and explain QCVV’s core models and concepts—quantum states, measurements, and processes—and illustrate how these building blocks are leveraged to examine a target system or operation. We survey and introduce protocols ranging from simple qubit characterization to advanced benchmarking methods. Along the way, we provide illustrated examples and detailed descriptions of the protocols, highlight the advantages and disadvantages of each, and discuss their potential scalability to future large-scale quantum computers. This tutorial serves as a guidebook for researchers unfamiliar with the benchmarking and characterization of quantum computers, and also as a detailed reference for experienced practitioners.

open quantum systems & decoherence

Final Report of the NASA Office of Safety and Mission Assurance Agile Benchmarking Team

To ensure that the NASA Safety and Mission Assurance (SMA) community remains in a position to perform reliable Software Assurance (SA) on NASAs critical software (SW) systems with the software industry rapidly transitioning from waterfall to Agile processes, Terry Wilcutt, Chief, Safety and Mission Assurance, Office of Safety and Mission Assurance (OSMA) established the Agile Benchmarking Team (ABT). The Team's tasks were: 1. Research background literature on current Agile processes, 2. Perform benchmark activities with other organizations that are involved in software Agile processes to determine best practices, 3. Collect information on Agile-developed systems to enable improvements to the current NASA standards and processes to enhance their ability to perform reliable software assurance on NASA Agile-developed systems, 4. Suggest additional guidance and recommendations for updates to those standards and processes, as needed. The ABT's findings and recommendations for software management, engineering and software assurance are addressed herein.

Software Assurance

IFAR Liner Benchmark Challenge #1 - DLR Impedance Eduction of Uniform and Axially Segmented Liners and Comparison with NASA Results

This paper presents the contribution from the German Aerospace Center (DLR) to the first liner benchmark challenge under the framework of the International Forum for Aviation Research (IFAR).Therefore, two sets of acoustically damping wall treatment, called ’liner samples’, have been produced by additive manufacturing based on the design data provided by NASA coordinating this benchmark. These liner samples have been integrated and acoustically characterized in the liner flow test facility DUCT-R at DLR Berlin as well as in the liner flow test facility GFIT at NASA Langley. Besides the dissipation coefficients and the axial pressure profiles, the liner wall impedance was educed by first determining the axial wave numbers and then applying a straightforward method based on the one-dimensional Convected Helmholtz Equation. Finally, the comparison of the liner impedance values to the NASA results show a fairly good agreement.

liner characterization

Code Benchmark of Depressurized Conduction Cooldown Transient in the High Temperature Test Facility

This paper presents results from modeling of a depressurized conduction cooldown (DCC) transient at the High Temperature Test Facility (HTTF) as part of the OECD-NEA Thermal Hydraulics Code Validation Benchmark for High-Temperature Gas-Cooled Reactors using HTTF Data . This paper briefly describes the benchmark and the models being used. It then presents a comparison of steady state and transient results based on the Problem 2 Exercise 1A and 1B definitions. We compare block and helium temperature distributions, mass flow distribution, and energy balance in steady state. All models show comparable mass flow distributions and energy balances. The temperatures within the core and outer regions are comparable in all models too, but inner reflector temperatures can vary significantly. Despite that, we find that the models are in good agreement for the full-power steady state. In the DCC, we look at block temperature at the core midplane and RCCS water exit temperature. The INL and ANL models are found to be in excellent agreement with one another on block temperature over time, while the agreement when the KAERI and NRG models are added into consideration is good. Differences in the transient heat removal from the RCCS cause the differences in block temperature over time in these models. The CNL models show similar trends to the INL, ANL, KAERI, and NRG models, but the temperatures are high because the volumes used in calculating the average temperature include the heater rods in the CNL models only. The HUN-REN model shows results that suggest significantly lower heat removal in the RCCS which merit further investigation.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Progress in the HTTF Benchmark and RELAP5-3D Gas-Cooled Reactor Validation

This slide set presents results of analysis done in the first year after the kickoff of the HTTF benchmark. This work includes the development of a new RELAP5-3D model of the facility and collection and presentation of results from one of the exercises in the benchmark. Highlights include good agreement between the new and legacy models for full-power steady state and similar predictions of maximum block temperature during a pressurized conduction cooldown. Additionally, code-to-code comparisons of a full-power steady state and depressurized conduction cooldown (DCC) from full-power steady state as part of Problem 2 Exercises 1A and 1B show good agreement on temperatures predicted in the core but considerable spread in central reflector temperatures. During the DCC, temperature behavior is similar across all models, but cooldown rates are dictated by the performance of the reactor cavity cooling system

22 GENERAL STUDIES OF NUCLEAR REACTORS

Benchmarking Monte Carlo codes for the modelling of low-energy neutron production target reactions

The increasing adoption of accelerator-based neutron sources (ABNS) for applications including neutron capture therapy (NCT) research has highlighted the need for accurate simulation tools. Precise modelling of the neutron production target is crucial to ensure that simulated predictions of neutron beam characteristics used for subsequent beam shaping assembly design are reliable. This work presents a comprehensive benchmarking of four widely-used Monte Carlo codes - Geant4, PHITS, FLUKA (CERN), and MCNP - for modelling low-energy neutron production target reactions. Using their recommended physics models and cross-section libraries, we evaluate each code’s performance in simulating four beam-target reactions: 7 Li(p,n) 7 Be, 9 Be(p,n) 9 B, 9 Be(d,n) 10 B, and C(d,n)N. Predictions of neutron yield, angular distributions, and energy spectra are compared against available thick target experimental data. Results show varying levels of agreement between the codes depending on the reaction type, energy range, and beam characteristics. Geant4, MCNP and PHITS are the overall best performing codes for the simulation of total neutron yield and yield in the forward direction across most reactions. Across energies where experimental benchmarks exist, inter-code discrepancies in total and forward-directed yield are typically 10 to 30%, with larger deviations at near-threshold incident ion energies. PHITS provides the best overall reproduction of experimental spectra, particularly for the 9 Be(p,n) 9 B reaction. Additionally, PHITS demonstrates superior computational performance for most reactions. These findings provide valuable guidance for ABNS design, highlighting the strengths and limitations of each code for the simulation of low-energy neutron production reactions.

43 PARTICLE ACCELERATORS

Development, validation, and verification of multi-pass thermo-mechanical welding simulations using the open-source MOOSE framework: NeT TG4 benchmark weldment

This study develops and validates a sequentially coupled thermo-mechanical welding simulation for the three-pass 316L stainless steel NeT TG4 benchmark weldment using the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) and the Nuclear Engineering Material model Library (NEML). A diffused ellipsoidal heat source was calibrated against thermocouple data and weld macrographs to accurately model the fusion zone geometry and transient thermal fields. Material hardening is represented using the Lemaitre-Chaboche mixed isotropic-kinematic hardening model, while four annealing models - no annealing, single-stage at 1050 °C and 1300 °C, and two-stage at 800 °C/1300 °C - were implemented to assess the impact of annealing models on the accuracy of the predicted welding-induced plasticity, distortions, and residual stresses. The predictions were validated against experimental measurements and benchmarked against results from commercial software, demonstrating that thermo-mechanical MOOSE welding simulations achieve comparable accuracy with enhanced computational efficiency. This work highlights the potential of using open-source finite element frameworks like MOOSE for advanced manufacturing simulations.

Ji, Wendy [Australian Nuclear Science and Technolo

Flow reversal benchmark of a one-sided heated narrow rectangular channel with CATHARE and RELAP5

Flow reversal in narrow coolant channels can be a crucial phenomenon for the safety of research reactors with a downward nominal flow direction. During a loss of forced flow accident, the downward flow stagnates briefly before transitioning into an upward natural circulation flow. The fuel may be damaged if dryout occurs and threshold fuel and/or cladding temperatures are exceeded. A comprehensive study is provided for flow reversal in narrow rectangular channels by examining experimental data and conducting software model analyses. The literature on flow reversal was reviewed, and selected experimental datasets were used to benchmark against CATHARE and RELAP5 models and also compare the code calculations with each other. The experimental data comes from flow reversal tests conducted with a narrow rectangular channel with one-sided heating. The results were compared with experimental data for successful flow reversal tests and predicted dryout power for dryout conditions. Also, the study examined the effects of the pump coastdown period, inlet liquid temperature, system pressure, and localized pressure drops. The experimental results showed that shorter coastdown periods, reduced pressure drops, and lower coolant inlet temperatures increased the dryout power. However, the system pressure did not noticeably affect the results. The simulation results showed that both CATHARE and RELAP5 agreed with experimental data, capturing the trends of the experimental results. Slight differences between each code calculation, as well as the predicted and measured dryout powers, were attributed to experimental uncertainties and the modeling of physical phenomena such as wall nucleation, interfacial heat transfer, drag coefficients, and critical heat flux. Overall, this study provides an understanding of flow reversal and the prediction capabilities of thermal-hydraulics software models. In conclusion, a future study of the flow reversal benchmark of a narrow rectangular channel with two-sided heating may provide additional valuable insights.

CATHARE

Concordant Mode Approach (CMA): Vibrational Analysis of New and Upgraded Intermolecular Benchmarks for Noncovalent Bonding

The Concordant Mode Approach (CMA) is a novel method that offers tremendous potential for increasing the system size and the level of theory attainable in quantum chemical computations of molecular vibrational frequencies. To investigate the extension of CMA to intermolecular vibrations, computations with coupled cluster singles and doubles with perturbative triples theory [CCSD(T)] using two augmented correlation-consistent polarized-valence triple-ζ basis sets (aug-cc-pVTZ or h-aug-ccpVTZ) were performed on 17 prototypical loosely bound complexes of hydrogen-bonded, dispersion, and mixed character. These Level A results provide new and upgraded benchmarks for noncovalent bonding and a severe test for CMA vibrational analyses. The Level A target frequencies were recovered remarkably well using second-order Møller−Plesset perturbation theory (MP2) with h-aug-cc-pVTZ for generating the underlying (Level B) normal modes of the CMA scheme. Employing this Level B within the lowest-rung CMA-0A method reproduces the 435 benchmark frequencies with a mean absolute error (MAE) of 0.23 cm −1 and a corresponding standard deviation (σ) of 0.84 cm −1 ; strikingly, the corresponding subset of 106 interfragment frequencies exhibits MAE = 0.34 cm −1 and σ = 0.90 cm −1 . Subsequent application of the higher-rung CMA-2A scheme eliminates all outliers and reduces the overall MAE to a minuscule 0.08 cm −1 with the inclusion of only 3.0% of the off-diagonal couplings not accounted for by CMA-0A. Accordingly, the highly efficient CMA methodology proves to be robust even for vibrations on flat potential energy surfaces.

Aromatic compounds

Diffusion Quantum Monte Carlo Benchmarking of Magnetic Moments in MnBi 2 Te 4

The intrinsically antiferromagnetic topological insulator, MnBi 2 Te 4 (MBT), has garnered significant attention recently due to its potential to host numerous exotic topological quantum states. Unfortunately, their consistent realization has been hindered by intrinsic antisite defects among the Mn and Bi sublattices. In this work, we establish Mn magnetization of pristine MBT through high level diffusion Monte Carlo calculations, which can serve as a precise starting point for various models to estimate antisite defect concentrations in actual MBT samples. The benchmark quality of DMC calculations is further identified from out model estimating antisite defect concentrations, which combines the benchmarked Mn magnetization with data from magnetic susceptibility and intermediate field magnetization measurements. This reproduces well Bi Mn and Mn Bi concentrations measured in the experiments. Here, we anticipate these theoretically based magnetic purity measures may be used as minimization targets in cycles of refinement to synthesize MBT with low antisite defect concentrations and more reproducible topological properties.

Defects

Quantum annealing for combinatorial optimization: a benchmarking study

Quantum annealing (QA) has the potential to significantly improve solution quality and reduce time complexity in solving combinatorial optimization problems compared to classical optimization methods. However, due to the limited number of qubits and their connectivity, the QA hardware did not show such an advantage over classical methods in past benchmarking studies. Recent advancements in QA with more than 5000 qubits, enhanced qubit connectivity, and the hybrid architecture promise to realize the quantum advantage. Here, we use a quantum annealer with state-of-the-art techniques and benchmark its performance against classical solvers. To compare their performance, we solve over 50 optimization problem instances represented by large and dense Hamiltonian matrices using quantum and classical solvers. The results demonstrate that a state-of-the-art quantum solver has higher accuracy (~0.013%) and a significantly faster problem-solving time (~6561×) than the best classical solver. Our results highlight the advantages of leveraging QA over classical counterparts, particularly in hybrid configurations, for achieving high accuracy and substantially reduced problem solving time in large-scale real-world optimization problems.

97 MATHEMATICS AND COMPUTING

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang

Phonon Olympics: Phonon property and lattice thermal conductivity benchmarking from open-source packages

Three widely used open-source packages for determining phonon properties and lattice thermal conductivities (ALAMODE, phono3py, and ShengBTE) are benchmarked by teams of expert users and the package developers. The phonons for Ge, RbBr, monolayer MoSe 2 , and AlN are modeled at zero temperature, and they scatter through three-phonon and phonon-isotope processes, with thermal conductivities obtained from the linearized Peierls–Boltzmann transport equation with input from density functional theory calculations. Over a wide range of temperatures, the thermal conductivities calculated by the teams fall within at most ±15% of their mean values for each of the four materials. The phonon frequencies, obtained from the harmonic force constants, do not show large differences between the calculations, indicating that the modal heat capacities and group velocities are not responsible for the thermal conductivity variations. It is the lifetimes associated with three-phonon scattering, obtained from the cubic force constants, that drive the variations. The many decisions required to calculate the cubic force constants (e.g., supercell size, atomic displacement, neighbor cutoff, and application of symmetries) make identification of the precise origin of the thermal conductivity variations challenging. The calculated thermal conductivities do not generally show agreement with experimental measurements, which is attributed to the limitations of the density functional theory calculations. Guidance for the development of best practices is provided, which will help to standardize protocols needed for building thermal conductivity databases. The results provide a baseline for future benchmarking of other packages and more advanced calculations.

McGaughey, Alan J. H. [Carnegie Mellon Univ., Pitt