Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Progress in the HTTF Benchmark and RELAP5-3D Gas-Cooled Reactor Validation

This slide set presents results of analysis done in the first year after the kickoff of the HTTF benchmark. This work includes the development of a new RELAP5-3D model of the facility and collection and presentation of results from one of the exercises in the benchmark. Highlights include good agreement between the new and legacy models for full-power steady state and similar predictions of maximum block temperature during a pressurized conduction cooldown. Additionally, code-to-code comparisons of a full-power steady state and depressurized conduction cooldown (DCC) from full-power steady state as part of Problem 2 Exercises 1A and 1B show good agreement on temperatures predicted in the core but considerable spread in central reflector temperatures. During the DCC, temperature behavior is similar across all models, but cooldown rates are dictated by the performance of the reactor cavity cooling system

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Benchmarking Monte Carlo codes for the modelling of low-energy neutron production target reactions

The increasing adoption of accelerator-based neutron sources (ABNS) for applications including neutron capture therapy (NCT) research has highlighted the need for accurate simulation tools. Precise modelling of the neutron production target is crucial to ensure that simulated predictions of neutron beam characteristics used for subsequent beam shaping assembly design are reliable. This work presents a comprehensive benchmarking of four widely-used Monte Carlo codes - Geant4, PHITS, FLUKA (CERN), and MCNP - for modelling low-energy neutron production target reactions. Using their recommended physics models and cross-section libraries, we evaluate each code’s performance in simulating four beam-target reactions: 7 Li(p,n) 7 Be, 9 Be(p,n) 9 B, 9 Be(d,n) 10 B, and C(d,n)N. Predictions of neutron yield, angular distributions, and energy spectra are compared against available thick target experimental data. Results show varying levels of agreement between the codes depending on the reaction type, energy range, and beam characteristics. Geant4, MCNP and PHITS are the overall best performing codes for the simulation of total neutron yield and yield in the forward direction across most reactions. Across energies where experimental benchmarks exist, inter-code discrepancies in total and forward-directed yield are typically 10 to 30%, with larger deviations at near-threshold incident ion energies. PHITS provides the best overall reproduction of experimental spectra, particularly for the 9 Be(p,n) 9 B reaction. Additionally, PHITS demonstrates superior computational performance for most reactions. These findings provide valuable guidance for ABNS design, highlighting the strengths and limitations of each code for the simulation of low-energy neutron production reactions.

43 PARTICLE ACCELERATORS↗

Development, validation, and verification of multi-pass thermo-mechanical welding simulations using the open-source MOOSE framework: NeT TG4 benchmark weldment

This study develops and validates a sequentially coupled thermo-mechanical welding simulation for the three-pass 316L stainless steel NeT TG4 benchmark weldment using the open-source Multiphysics Object-Oriented Simulation Environment (MOOSE) and the Nuclear Engineering Material model Library (NEML). A diffused ellipsoidal heat source was calibrated against thermocouple data and weld macrographs to accurately model the fusion zone geometry and transient thermal fields. Material hardening is represented using the Lemaitre-Chaboche mixed isotropic-kinematic hardening model, while four annealing models - no annealing, single-stage at 1050 °C and 1300 °C, and two-stage at 800 °C/1300 °C - were implemented to assess the impact of annealing models on the accuracy of the predicted welding-induced plasticity, distortions, and residual stresses. The predictions were validated against experimental measurements and benchmarked against results from commercial software, demonstrating that thermo-mechanical MOOSE welding simulations achieve comparable accuracy with enhanced computational efficiency. This work highlights the potential of using open-source finite element frameworks like MOOSE for advanced manufacturing simulations.

Ji, Wendy [Australian Nuclear Science and Technolo↗

Flow reversal benchmark of a one-sided heated narrow rectangular channel with CATHARE and RELAP5

Flow reversal in narrow coolant channels can be a crucial phenomenon for the safety of research reactors with a downward nominal flow direction. During a loss of forced flow accident, the downward flow stagnates briefly before transitioning into an upward natural circulation flow. The fuel may be damaged if dryout occurs and threshold fuel and/or cladding temperatures are exceeded. A comprehensive study is provided for flow reversal in narrow rectangular channels by examining experimental data and conducting software model analyses. The literature on flow reversal was reviewed, and selected experimental datasets were used to benchmark against CATHARE and RELAP5 models and also compare the code calculations with each other. The experimental data comes from flow reversal tests conducted with a narrow rectangular channel with one-sided heating. The results were compared with experimental data for successful flow reversal tests and predicted dryout power for dryout conditions. Also, the study examined the effects of the pump coastdown period, inlet liquid temperature, system pressure, and localized pressure drops. The experimental results showed that shorter coastdown periods, reduced pressure drops, and lower coolant inlet temperatures increased the dryout power. However, the system pressure did not noticeably affect the results. The simulation results showed that both CATHARE and RELAP5 agreed with experimental data, capturing the trends of the experimental results. Slight differences between each code calculation, as well as the predicted and measured dryout powers, were attributed to experimental uncertainties and the modeling of physical phenomena such as wall nucleation, interfacial heat transfer, drag coefficients, and critical heat flux. Overall, this study provides an understanding of flow reversal and the prediction capabilities of thermal-hydraulics software models. In conclusion, a future study of the flow reversal benchmark of a narrow rectangular channel with two-sided heating may provide additional valuable insights.

CATHARE↗

Concordant Mode Approach (CMA): Vibrational Analysis of New and Upgraded Intermolecular Benchmarks for Noncovalent Bonding

The Concordant Mode Approach (CMA) is a novel method that offers tremendous potential for increasing the system size and the level of theory attainable in quantum chemical computations of molecular vibrational frequencies. To investigate the extension of CMA to intermolecular vibrations, computations with coupled cluster singles and doubles with perturbative triples theory [CCSD(T)] using two augmented correlation-consistent polarized-valence triple-ζ basis sets (aug-cc-pVTZ or h-aug-ccpVTZ) were performed on 17 prototypical loosely bound complexes of hydrogen-bonded, dispersion, and mixed character. These Level A results provide new and upgraded benchmarks for noncovalent bonding and a severe test for CMA vibrational analyses. The Level A target frequencies were recovered remarkably well using second-order Møller−Plesset perturbation theory (MP2) with h-aug-cc-pVTZ for generating the underlying (Level B) normal modes of the CMA scheme. Employing this Level B within the lowest-rung CMA-0A method reproduces the 435 benchmark frequencies with a mean absolute error (MAE) of 0.23 cm −1 and a corresponding standard deviation (σ) of 0.84 cm −1 ; strikingly, the corresponding subset of 106 interfragment frequencies exhibits MAE = 0.34 cm −1 and σ = 0.90 cm −1 . Subsequent application of the higher-rung CMA-2A scheme eliminates all outliers and reduces the overall MAE to a minuscule 0.08 cm −1 with the inclusion of only 3.0% of the off-diagonal couplings not accounted for by CMA-0A. Accordingly, the highly efficient CMA methodology proves to be robust even for vibrations on flat potential energy surfaces.

Aromatic compounds↗

Diffusion Quantum Monte Carlo Benchmarking of Magnetic Moments in MnBi 2 Te 4

The intrinsically antiferromagnetic topological insulator, MnBi 2 Te 4 (MBT), has garnered significant attention recently due to its potential to host numerous exotic topological quantum states. Unfortunately, their consistent realization has been hindered by intrinsic antisite defects among the Mn and Bi sublattices. In this work, we establish Mn magnetization of pristine MBT through high level diffusion Monte Carlo calculations, which can serve as a precise starting point for various models to estimate antisite defect concentrations in actual MBT samples. The benchmark quality of DMC calculations is further identified from out model estimating antisite defect concentrations, which combines the benchmarked Mn magnetization with data from magnetic susceptibility and intermediate field magnetization measurements. This reproduces well Bi Mn and Mn Bi concentrations measured in the experiments. Here, we anticipate these theoretically based magnetic purity measures may be used as minimization targets in cycles of refinement to synthesize MBT with low antisite defect concentrations and more reproducible topological properties.

Defects↗

Quantum annealing for combinatorial optimization: a benchmarking study

Quantum annealing (QA) has the potential to significantly improve solution quality and reduce time complexity in solving combinatorial optimization problems compared to classical optimization methods. However, due to the limited number of qubits and their connectivity, the QA hardware did not show such an advantage over classical methods in past benchmarking studies. Recent advancements in QA with more than 5000 qubits, enhanced qubit connectivity, and the hybrid architecture promise to realize the quantum advantage. Here, we use a quantum annealer with state-of-the-art techniques and benchmark its performance against classical solvers. To compare their performance, we solve over 50 optimization problem instances represented by large and dense Hamiltonian matrices using quantum and classical solvers. The results demonstrate that a state-of-the-art quantum solver has higher accuracy (~0.013%) and a significantly faster problem-solving time (~6561×) than the best classical solver. Our results highlight the advantages of leveraging QA over classical counterparts, particularly in hybrid configurations, for achieving high accuracy and substantially reduced problem solving time in large-scale real-world optimization problems.

97 MATHEMATICS AND COMPUTING↗

A curated benchmark for cofolding models on kinase conformational states

Abstract Protein kinases are critical drug targets, requiring therapeutics that can modulate their active and inactive conformational states. While cofolding models can generate global folds directly from kinase sequences and ligand SMILES strings, these models have not yet been tested on their ability to recover ligand-induced-fit conformational states of the kinase proteins. Here, we introduce KinConfBench, a curated benchmark of 2225 high-quality human kinase chains to evaluate the ability of four state-of-the-art cofolding models—Boltz-2, Chai-1, Protenix, and RoseTTAFold-All-Atom—to recover both canonical and rare conformational states. We show that geometric success metrics of a ligand pose in the active site do not correlate strongly with the correct kinase conformational state, motivating a new set of dynamical benchmarks for assessing cofolding models. While all four cofolding models achieve ~60–80% prediction accuracy for kinase conformational classification, they exhibit severe mode collapse when performing multiple inferences, show negligible structural diversity in sampling induced-fit motions, and display a prevalent “apo-drift” in which most cofolding models predominantly predict the kinase to be in its ligand-free state. Our results highlight that capturing ligand-induced protein conformational diversity, not just geometric fit, is critical for next-generation structure-based drug discovery.

Sun, Kunyang↗

Phonon Olympics: Phonon property and lattice thermal conductivity benchmarking from open-source packages

Three widely used open-source packages for determining phonon properties and lattice thermal conductivities (ALAMODE, phono3py, and ShengBTE) are benchmarked by teams of expert users and the package developers. The phonons for Ge, RbBr, monolayer MoSe 2 , and AlN are modeled at zero temperature, and they scatter through three-phonon and phonon-isotope processes, with thermal conductivities obtained from the linearized Peierls–Boltzmann transport equation with input from density functional theory calculations. Over a wide range of temperatures, the thermal conductivities calculated by the teams fall within at most ±15% of their mean values for each of the four materials. The phonon frequencies, obtained from the harmonic force constants, do not show large differences between the calculations, indicating that the modal heat capacities and group velocities are not responsible for the thermal conductivity variations. It is the lifetimes associated with three-phonon scattering, obtained from the cubic force constants, that drive the variations. The many decisions required to calculate the cubic force constants (e.g., supercell size, atomic displacement, neighbor cutoff, and application of symmetries) make identification of the precise origin of the thermal conductivity variations challenging. The calculated thermal conductivities do not generally show agreement with experimental measurements, which is attributed to the limitations of the density functional theory calculations. Guidance for the development of best practices is provided, which will help to standardize protocols needed for building thermal conductivity databases. The results provide a baseline for future benchmarking of other packages and more advanced calculations.

McGaughey, Alan J. H. [Carnegie Mellon Univ., Pitt↗

Benchmarking core turbulence and transport predictions for an inductive compact tokamak reactor plasma

Motivated by the need for accurate, timely, and efficient calculations of plasma transport, predictions of plasma turbulence properties made using different TGLF saturation rules are benchmarked against corresponding predictions from linear and nonlinear gyrokinetic CGYRO simulations. This benchmarking is carried out using parameters taken from an inductive burning plasma scenario in a hypothetical compact high-field (R maj = 4 m, B T = 8 T) tokamak, lying in a much different regime of parameter space than either the TGLF calibration regime or current-day experiments. The core turbulent transport in this scenario is predicted to be dominated by ion temperature gradient (ITG) turbulence. In general, the ITG critical gradients predicted by various TGLF saturation rules are quite close to the CGYRO predictions. Both codes predict similar linear ITG growth rates and frequency spectra, as well as their scaling with R/L T i = −Rd ln(T i )/dr. However, TGLF systematically predicts unstable trapped-electron modes (TEMs) above k y ρ s ≃ 0.5 not seen by CGYRO for the same parameters, due to TGLF predicting a lower threshold in R/L T e than CGYRO for TEM onset. It is shown that for this scenario, nonlinear CGYRO simulations predict stiffer ITG turbulence than the TGLF SAT0 and SAT1 saturation rules, with energy fluxes close in magnitude and scaling with R/L T i to what is predicted by the SAT2 saturation rule. Self-consistent core profiles calculated using nonlinear CGYRO flux predictions and the PORTALS transport solver are shown to agree fairly well with corresponding predictions made using the TGLF SAT2 model, including a similar level of density peaking.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hofstadter butterfly and quantum transport benchmarks in PVA-exfoliated graphene heterostructures

Polymer exposure during van der Waals heterostructure fabrication is widely regarded as compromising the integrity of the electronic system required for hosting emergent quantum physics. This assumption has persisted largely because electronic benchmarking of heterostructures from polymer-exposed graphene has remained limited to only foundational transport metrics, such as mobility and charge inhomogeneity. Here, we challenge this assumption by establishing that graphene heterostructures produced by polyvinyl alcohol (PVA)-assisted exfoliation and encapsulated in hexagonal boron nitride using elevated-temperature lamination satisfy demanding quantum transport benchmarks. Beyond exhibiting ultra-high mobility and ballistic transport, these heterostructures yield quantum scattering times comparable to the best polymer-free devices. Most demanding of all, moiré superlattices from PVA-exposed graphene exhibit Hofstadter butterfly spectra, confirming spatially uniform interlayer coupling across the device area. These results establish that PVA exposure is compatible with low-disorder electronic systems, relaxing the trade-off between scalable fabrication and low-disorder quantum transport. This study motivates further development of polymer-assisted assembly with engineered residue-removal protocols.

36 MATERIALS SCIENCE↗

Role of electron correlation on the adenine dimer interaction for non-equilibrium geometries: a benchmark Quantum Monte Carlo study

The accurate description of non-covalent interactions is critical for understanding the structure, dynamics, and eventual function of biomolecules. The adenine dimer serves as a benchmark system for computational methods due to its role in nucleic acid structures and its rich conformational landscape. In this study, we employ benchmark diffusion quantum Monte Carlo (DMC) methods to investigate the relative energies and role of electron correlation on a set of adenine dimer conformations generated via a search of the potential energy landscape using the global optimizer algorithm. Relative DMC energies are compared against a wide range of density functional theory (DFT) approximation results. We find that although most of the DFT functionals perform well for low-energy structures, their accuracy varies significantly for higher-energy conformations, including stacked and T-shaped structures. A large fraction of the variation is due to the treatment of the van der Waals interaction. BLYP, B3LYP, and PBE0 significantly improve with added D4 dispersion, while the recent r2SCAN-D4 and ωB97M-V functionals show the least scatter and closest agreement with the DMC. These findings highlight the delicate nature of these interactions in biomolecular systems and provide guidance for simulations of their structure and dynamics and for the development of machine learned interatomic potentials.

Washburn, Laurel [ORNL] (ORCID:0000000324179335)↗

The Role of Nuclear Data Sensitivities in Prompt α-Eigenvalue Predictions of Delayed Critical Benchmarks

Alpha (α) eigenvalues, which describe the logarithmic time derivative of the neutron population in a multiplying system, are integral to time-dependent behavior and diagnostic applications. However, uncertainties in the evaluated nuclear data can significantly impact the accuracy of transport simulations for such quantities. This work explores the use of machine learning models to predict two key outputs, α-eigenvalues and keff bias, using input features derived from α-eigenvalue sensitivities to nuclear data. The criticality safety benchmark models used in this study come from the International Handbook of Evaluated Criticality Safety Benchmark Experiments. Three models, random forest, XGBoost, and NGBoost, are trained on both energy-resolved and energy-summed α sensitivities. For the α-eigenvalue bias prediction, NGBoost achieved the highest R 2 (0.9476) using energy-resolved features, while XGBoost performed best using summed sensitivities. In contrast, when predicting the keff bias, all the models showed moderate predictive capability (best R 2 ≈ 0.72), as the mapping from the static α-sensitivities to the static keff bias was less direct. SHAP (SHapley Additive exPlanations) analysis was used to interpret the model predictions. Across both prediction tasks, the features associated with neutron capture [H-1 (n, γ)], uranium scattering reactions (such as 235 U elastic/inelastic), and actinide capture/fission reactions (such as 239 Pu and 234 U) were consistently identified as the most impactful. This highlights the key role of specific nuclear reactions and energy ranges in shaping both time-dependent and steady-state criticality behavior. These results demonstrated that α-sensitivities, despite being computed for time-dependent metrics, can provide valuable insights for predicting both α-eigenvalues and the keff bias. Moreover, machine learning models offer a promising pathway for uncovering important nuclear data dependencies and guiding future data evaluation efforts.

Nuclear data↗

Benchmarking machine learning strategies for phase-field problems

Abstract We present a comprehensive benchmarking framework for evaluating machine-learning approaches applied to phase-field problems. This framework focuses on four key analysis areas crucial for assessing the performance of such approaches in a systematic and structured way. Firstly, interpolation tasks are examined to identify trends in prediction accuracy and accumulation of error over simulation time. Secondly, extrapolation tasks are also evaluated according to the same metrics. Thirdly, the relationship between model performance and data requirements is investigated to understand the impact on predictions and robustness of these approaches. Finally, systematic errors are analyzed to identify specific events or inadvertent rare events triggering high errors. Quantitative metrics evaluating the local and global description of the microstructure evolution, along with other scalar metrics representative of phase-field problems, are used across these four analysis areas. This benchmarking framework provides a path to evaluate the effectiveness and limitations of machine-learning strategies applied to phase-field problems, ultimately facilitating their practical application.

36 MATERIALS SCIENCE↗

Real singlet scalar benchmarks in the multi-TeV resonance regime

Scalar extensions of the Standard Model (SM) are of much interest at the Large Hadron Collider (LHC) and future colliders. In particular, these models can give rise to resonant di-Higgs production and alter the Higgs trilinear coupling. In this paper, we study di-Higgs production in the Standard Model extended by a real scalar singlet with no additional symmetries. We determine how large the resonant di-Higgs rate and variation in the Higgs trilinear coupling can be in four scenarios: current LHC results and projected results at the high luminosity LHC (HL-LHC), the HL-LHC combined with a circular 𝑒 − ⁢𝑒 + collider such as the Circular Electron Positron Collider or Future Circular Collider with electron-positron collisions, and the HL-LHC combined with a linear 𝑒 − ⁢𝑒 + collider such as the International Linear Collider. While these are updated results from a previous study by [I. M. Lewis and M. Sullivan, Benchmarks for double Higgs production in the singlet extended standard model at the LHC, Phys. Rev. D 96, 035037 (2017).] using current LHC data, we go further and find benchmark points in the multi-TeV resonance regime for future colliders beyond the HL-LHC. Considering current LHC results, the resonant di-Higgs rate can still be an order of magnitude larger than the SM predicted di-Higgs rate. In the HL-LHC scenario, the Higgs trilinear coupling can still be a factor of three larger than the SM prediction for resonance masses in the 1.5–3.5 TeV range, where resonant searches may have less reach. This enhancement is just at the projected 2⁢𝜎 sensitivity of the HL-LHC. We find there are resonance masses for which the change in the Higgs trilinear is maximized while the resonant rate is negligible. We provide an analytical understanding of these effects with a discussion on the interplay of various constraints on the parameter space and the Higgs trilinear coupling.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Benchmark Suite for Evaluating Scientific AI Workloads on GPUs

AI applications have been steadily increasing in the allocation portfolio among leadership computing facilities. These applications depend on deep learning frameworks with hardware acceleration and underlying software systems. With the rapid development of applications, software stacks, and hardware devices, it is essential to evaluate the performance of core operations in AI workloads for direction of optimizations and procurement of next-generation high-performance computing (HPC) infrastructures. Currently, most benchmarks lack scientific AI workloads. So, we present DeepKernelBench and the experimental results of evaluating the benchmark suite for early observations and performance comparisons on datacenter GPUs using representative workloads for scientific AI, including Attentions, General matrix multiplications, Geometrics and Fourier neural operations.

Jin, Zheming [Advanced Micro Devices (AMD)]↗

Benchmark High-Fidelity EMT Models for Power Grid with PV Plants

In recent times electromagnetic transient (EMT) modeling tools have been identified as one of the most important requirements in replicating, analyzing, and investigating the dynamics of the power grid with photovoltaic (PV) plants. However, there are no benchmark models for power grid with PVs to investigate emerging challenges with higher penetration of PVs (like trips and momentary cessations during faults from a region far away). To this end, in this paper, synthetic benchmark high-fidelity EMT dynamic models of power grid with large-scale PV plants are presented. The models are developed in PSCAD and PSCAD/Fortran. Simulation results for different use cases (events) and scenarios are presented.

Marthi, Phani Ratna Vanamali↗

Lifetimes of excited states in $^{16}$C as a benchmark for ab initio developments

Lifetimes of higher-lying states ($2_2^+$ and $4_1^+$) in 16 C have been measured, employing the Gammasphere and Microball detector arrays, as key observables to test and refine ab initio calculations based on interactions developed within chiral Effective Field Theory. The presented experimental constraints to these lifetimes of $\tau ({2_2^+}) = [244, 446]\,~\textrm{fs}$ and $\tau ({4_1^+}) = [1.8, 4]\,~\textrm{ps}$, combined with previous results on the lifetime of the $2_1^+$ state of 16 C, provide a rather complete set of key observables to benchmark the theoretical developments. We present No-Core Shell-Model calculations using state-of-the-art chiral 2- (NN) and 3-nucleon (3N) interactions at next-to-next-to-next-to-leading order for both the NN and the 3N contributions and a generalized natural-orbital basis (instead of the conventional harmonic-oscillator single-particle basis) which reproduce, for the first time, the experimental findings remarkably well. The level of agreement of the new calculations as compared to the CD-Bonn meson-exchange NN interaction is notable and presents a critical benchmark for theory.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗