Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Benchmarking”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Investigation of Benchmark $k$ eff Sensitivity and Uncertainty for 239 Pu fission in Specific Energy Ranges

Nuclear data at intermediate energies (from 1 to 100s of keV) are evaluated based on scarce differential data and theory unable to capture physics’ expected structure. There is also a lack of integral data. This is a known deficiency and is challenging to address. Calculated effective multiplication factor, k eff , values for intermediate energy experiments are ~25× further from experiment than for fast energies and are often well outside the experimental uncertainties. The goal of the PARADIGM (PARallel Approach of Differential and InteGral Measurements) project is to significantly re duce the uncertainties of intermediate energy nuclear data for 239 Pu. To this end, PARADIGM simultaneously optimizes experiments at both the Los Alamos Neutron Science Center (LANSCE) and National Criticality Experiments Research Center (NCERC). The combined set of data will inform new intermediate-energy nuclear data. By execution of differential and integral experiments, establishment of new theory, and undertaking nuclear data evaluation in parallel, the timeline to deliver improved nuclear data to users will be reduced significantly that is to three years. For the PARADIGM project, it was decided to optimize an integral experiment for two neutron energy ranges, within the full intermediate energy range. The low energy range goes from 1 to 30 keV, while the higher energy range goes from 30 to 600 keV. This work focuses on nuclear data sensitivities and uncertainties for 239 Pu fission for existing experiments in the International Criticality Safety Benchmark Evaluation Project (ICSBEP). When designing new experiments, it is important to understand what benchmarks currently exist. For a more traditional experiment design (in which a specific application model(s) exists), comparisons would be made between the application model(s) and existing benchmarks. For PARADIGM, there is no specific application model, but instead the specific nuclear data reaction and energy ranges of interest can be explored for existing benchmarks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Towards robust surrogate models: Benchmarking machine learning approaches to expediting phase field simulations of brittle fracture

Data-driven approaches have the potential to make modeling complex, nonlinear physical phenomena significantly more computationally tractable. For example, computational modeling of fracture is a core challenge where machine learning techniques have the potential to provide a much needed speedup that would enable progress in areas such as multi-scale modeling and uncertainty quantification. Currently, phase field modeling (PFM) of fracture is one such approach that offers a convenient variational formulation to model crack nucleation, branching and propagation. To date, machine learning techniques have shown promise in approximating PFM simulations. While standard fracture benchmarks represent realistic scenarios frequently observed in practice, they typically do not provide sufficiently challenging tests for data-driven methods. Here, to address this gap, we introduce a challenging dataset based on PFM simulations designed to benchmark and advance ML methods for fracture modeling. This dataset includes three energy decomposition methods, two boundary conditions, and 1000 random initial crack configurations for a total of 6000 simulations. Each sample contains 100 time steps capturing the temporal evolution of the crack field. Alongside this dataset, we also implement and evaluate Physics Informed Neural Networks (PINN), Fourier Neural Operators (FNO), and UNet models as baselines, and explore the impact of ensembling strategies on prediction accuracy. With this combination of our dataset and baseline models drawn from the literature we aim to provide a standardized and challenging benchmark for evaluating machine learning approaches to solid mechanics. Our results highlight both the promise and limitations of popular current models, and demonstrate the utility of this dataset as a testbed for advancing machine learning in fracture mechanics research.

Benchmark dataset

Benchmark study of a new simplified DFN model for shearing of intersecting fractures and faults

It is challenging to quantitatively predict shearing of intersecting fractures/faults because of dynamic frictional contacts accompanied by possible nonlinear rock deformation. To address such challenges, a new conceptual model—the simplified DFN model—was proposed and validated by Hu et al. 46 to use major paths (MPs) to represent complicated DFNs for calculation of shearing. In this work, we conducted a benchmark study for three examples that involve different levels of complexity of intersecting fractures, and correspondingly different numbers of MPs. The codes and software that were used in the benchmark cover a range of continuum, discontinuum and hybrid numerical methods: NMM (LBNL), FLAC3D (LBNL), GBDEM (KIGAM), FRACOD (DynaFrax), and CASRock (CAS). The general consistency between DFN and MP cases as predicted by all the codes/software demonstrates that major paths can be used to simplify the geometry of DFNs in a wide range of software. Disagreement in results made by some software and potential future improvements are discussed. We show that (1) shearing of one or multiple major fractures can be reduced if there are multiple smaller intersecting fractures in that area, which is a useful basis for understanding and controlling induced seismicity and merits further analysis, and (2) the agreement achieved in the benchmark examples provide confidence that the simplified DFN model is a promising conceptual model that can be used for different types of numerical approaches and software for simplifying the analysis of the shearing of intersecting fractures and faults.

58 GEOSCIENCES

Benchmark of the Chlorine Worth Study Experiments in Support of Chlorine Nuclear Data Validation for Nuclear Criticality Safety

The Chlorine Worth Study (CWS) was a critical experiment to address an urgent need for thermal chlorine nuclear data validation in plutonium systems. This urgent need is tied directly to plutonium recycle and recovery operations in the plutonium facility at Los Alamos National Laboratory, where exceptionally conservative criticality safety limits are used because no credit is taken for the neutron capture by chlorine. The experiment used weapons-grade plutonium metal plates clad in stainless steel, known as the PANN (plutonium aluminum no nickel) ZPPR (zero power physics reactor) plates. The plutonium was reflected and moderated by high-density polyethylene and included combinations of polyvinyl chloride (PVC) and chlorinated polyvinyl chloride (CPVC) as absorbers. The experiment and benchmark included three configurations mimicking 30 g 239 Pu/L plutonium, 300 g 239 Pu/L plutonium, and 600 g 239 Pu/L plutonium in an aqueous chloride solution. Uncertainties in the benchmark included five broad categories: (1) criticality measurement, (2) mass and density, (3) dimensions, (4) material compositions, and (5) positioning. The largest contribution to the overall uncertainties for all three cases came from the material compositions, in particular the PVC and CPVC absorber compositions. A detailed model was created to be a near match (that is within expectations of transport code users) and a simplified model was created to minimize offset dimensions and expedite modeling for code validation. Sample calculations were completed in MCNP6.3 with ENDF/B-VIII.0 and ENDF/B-VII.1 nuclear data. For the detailed and simplified models, the average difference between the computed and experimental k eff was 951 pcm. CWS will serve as the key validation experiment for nuclear criticality safety in support of aqueous chloride operations. The sensitivity to the chlorine capture cross section is orders of magnitude greater than other existing benchmarks. The current limits, as defined by nuclear criticality safety, are 520 g Pu per batch, i.e. the minimum critical mass of the Pu solution infinitely reflected by water [Criticality Handbook: Volume II, (1969)]. This extremely conservative critical mass limit does not credit any neutron capture by chlorine (in particular neutron capture by 35 Cl) and greatly impedes the throughput required for current and future operations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Benchmark for two-dimensional large scale coherent structures in partially magnetized E × B plasmas—community collaboration & lessons learned

Low-temperature plasmas (LTPs) are essential to both fundamental scientific research and critical industrial applications. As in many areas of science, numerical simulations have become a vital tool for uncovering new physical phenomena and guiding technological development. Code benchmarking remains crucial for verifying implementations and evaluating performance. This work continues the Landmark benchmark initiative, a series specifically designed to support the verification of LTP codes. In this study, seventeen simulation codes from a collaborative community of nineteen international institutions modeled a partially magnetized E × B Penning discharge. The emergence of large scale coherent structures, or rotating plasma spokes, endows this configuration with an enormous range of time scales, making it particularly challenging to simulate. The codes showed excellent agreement on the rotation frequency of the spoke as well as key plasma properties, including time-averaged ion density, plasma potential, and electron temperature profiles. Achieving this level of agreement came with challenges, and we share lessons learned on how to conduct future benchmarking campaigns. Comparing code implementations, computational hardware, and simulation runtimes also revealed interesting trends, which are summarized with the aim of guiding future plasma simulation software development.

benchmarking

ICSBEP Benchmarking Tutorial for DNCSH

As part of the DOE/NRC Collaboration for Criticality Safety Support for Commercial-Scale HALEU for Fuel Cycles and Transportation (DNCSH), a workshop was held on Teams on how to write a ICSBEP benchmark that meets modern standards. The workshop prepared people performing experiments and writing benchmarks so that they have an increased chance of submitting a benchmark that will be accepted. The workshop was led by experts from LANL, LLNL and Sandia.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

A large-scale benchmarking of deterministic and stochastic derivative-free optimization algorithms

This presentation summarizes our work in the PrOMMiS project on benchmarking of data-driven optimization algorithms and their applications in self-driving laboratories. This work supports the broader project goal of accelerating the identification of promising separation methods and operating conditions for critical minerals separation processes. We present a systematic benchmarking study of 42 data-driven optimization algorithms on a broad collection of 502 test problems. The results identify BAM, GLCCLUSTER, and MULTIMIN as the most effective optimization solvers, with BAM showing the highest overall performance and solving more than 80% of the benchmark problems. The study also shows that no single solver consistently outperforms the others across all problem types, indicating that our future laboratory applications may benefit from using a small set of strong solvers rather than relying on a single method. The presentation also illustrates an in-silico chemical reactor case study showing that data-driven optimization methods can guide autonomous experimentation in a self-driving laboratory and identify optimal operating conditions within a small number of experiments. Overall, the results provide a basis for selecting efficient optimization methods and demonstrate the practical use of data-driven optimization in self-driving laboratory workflows.

36 MATERIALS SCIENCE

The HTR-Proteus Benchmark: Analysis and Use as a Verification and Validation Case

This presentation outlines the evaluation and application of the HTR-Proteus benchmark as a verification and validation (V&V) case for advanced reactor modeling tools. The work supports the U.S. Department of Energy’s HALEU Availability Program (HAP) and the joint DOE/NRC DNCSH project, which aims to reduce criticality safety uncertainties in commercial-scale HALEU fuel cycle and transportation systems. The HTR-Proteus experiments, conducted at the Paul Scherrer Institute, provide high-fidelity data for TRISO-fueled, graphite-moderated pebble bed reactors with high neutron leakage—conditions relevant to HALEU transport scenarios. This study focuses on Cores 4.2 and 4.3 of the HTR-Proteus benchmark, analyzing key sources of uncertainty including pebble packing, TRISO particle positioning, and core height. Using Project Chrono for realistic pebble geometries and Serpent, SHIFT, and MCNP for neutronics simulations, the study quantifies the impact of these uncertainties on the effective multiplication factor (keff). Results show that a sample size of 110 pebble configurations is sufficient to converge keff, with ±30 pcm uncertainty due to packing randomness. TRISO positioning and core height variations also significantly influence keff, highlighting the importance of detailed modeling in V&V efforts. The benchmark serves as a valuable test case for validating the Griffin reactor physics code and improving confidence in HALEU system simulations.

73 - NUCLEAR PHYSICS AND RADIATION PHYSICS

Benchmarking Operators in Deep Neural Networks for Improving Performance Portability of SYCL

SYCL is a portable programming model for heterogeneous computing, so it is important to obtain reasonable performance portability of SYCL. Towards the goal of better understanding and improving performance portability of SYCL for machine learning workloads, we have been developing benchmarks for basic operators in deep neural networks (DNNs). These operators could be offloaded to heterogeneous computing devices such as graphics processing units (GPUs) to speed up computation. In this paper, we introduce the benchmarks, evaluate the performance of the operators on GPU-based systems, and describe the causes of the performance gap between the SYCL and Compute Unified Device Architecture (CUDA) kernels. We find that the causes are related to the utilization of the texture cache for read-only data, optimization of the memory accesses with strength reduction, use of local memory, and register usage per thread. We hope that the efforts of developing benchmarks for studying performance portability will stimulate discussion and interactions within the community.

Jin, Zheming [ORNL] (ORCID:000000027197780X)

SMR safety through HTTF modeling and benchmark efforts for code validation for gas-cooled reactor applications

Accurate modeling and simulation tools for thermal-hydraulics calculations are a key element needed to design and license new advanced reactors including Small Modular Reactors (SMR) and Microreactors. Uncertainties in modeling and simulation can have significant safety and economic implications. The High Temperature Test Facility (HTTF) at Oregon State University (OSU) is a scaled integral effects experiment designed to investigate transient behavior in high-temperature gas-cooled prismatic-block nuclear reactors. High-quality measurement data is available from the HTTF that is suitable for a thermal-hydraulics code validation benchmark for gas-cooled reactor simulations. Here, this paper summarizes individual HTTF modeling efforts to date for tool validation at Idaho National Laboratory (INL), Argonne National Laboratory (ANL), Oregon State University (OSU) and Canadian Nuclear Laboratories (CNL) using system thermal-hydraulics codes, Computational Fluid Dynamics (CFD) codes and system-CFD code couplings. Also, the paper introduces the ongoing OECD Nuclear Energy Agency (NEA) High Temperature Gas Reactor Thermal-Hydraulics (HTGR T/H) benchmark that allows for better comparisons of results between different international modeling teams. The benchmark provides well defined computational problems that include code-to-code comparisons and comparisons to measured data. These problems provide an avenue for quantifying accuracy and identifying sources of uncertainty in thermal-hydraulics calculations, including in measured thermophysical properties, as part of validation for gas-cooled reactor simulation tools.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN

Comparison of URANS and LES predictions for the open phase of the OECD NEA CSNI fluid structure interaction CFD benchmark

The OECD NEA CSNI WGAMA CFD Task Group ran a benchmark in 2020 and 2021 to assess the predictive capabilities of coupled fluid structure interaction (FSI) CFD analysis methods. This paper presents the predictions made for the open phase of the benchmark using URANS and LES turbulence modelling approaches, and a comparison of the results to the experimental data. The benchmark comprised a channel containing two inline cylinders in cross-flow. The cylinders were fixed at one end, free at the other, and had measured resonant frequencies and damping properties. The URANS modelling used ANSYS Fluent 2-way coupled to ANSYS Mechanical. The LES modelling used Nek5000, 1-way coupled to Diablo. Comparisons with cross-channel velocity profiles are presented, both for the mean flow and its RMS. Comparisons are also made to the frequency spectra for point measurements of fluid velocity and pressure, and for the accelerations of the free end of each cylinder. URANS predicts the average velocity profiles relatively well, and is able to predict the velocity and acceleration spectra at the shedding frequency. However, the frequency content at the 4th harmonic of the shedding frequency is low in the URANS flow fields, and so does not excite accelerations at the resonant frequency of the cylinders. LES makes better predictions of the average profiles, and the velocity spectra agree well at both the shedding frequency and at higher frequencies. In conclusion, the 1-way coupled LES results show good agreement for acceleration spectra.

22 GENERAL STUDIES OF NUCLEAR REACTORS

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

Many-Body Benchmark of Electronic Charge and Spin Densities for Li 1–x NiO 2

Accurate benchmarks are particularly important for highly correlated oxides as mean-field approximations often fail to describe the subtle balance of charge transfer and magnetism in these materials with an accuracy comparable to experimental needs. Here we present accurate diffusion Monte Carlo (DMC) results of the electronic charge and spin densities for the tunable highly correlated oxide Li 1–x NiO 2 for x = 0, 1/2, and 1. To enable quantitative comparisons, we introduce a robust density-partitioning scheme, extending Voronoi analysis to assign atomic charges from spatially noisy DMC densities. We then benchmark common approximations used in density functional theory (DFT). Comparison against DMC shows that r 2 SCAN delivers the most balanced performance across charge, spin, and radial density descriptors, nearly reproducing DMC results for LiNiO 2 and apical Ni sites in Li 0.5 NiO 2 . Hybrid functionals (PBE0, SCAN0) perform unexpectedly poorly, and PBE + U + V yields inconsistent trends between charge and spin densities. Therefore, the r 2 SCAN functional minimizes errors relative to DMC while capturing the variable valence of the Ni ion and also retaining the computational efficiency of DFT for large-scale simulations of the tunable structural and electronic phases of Li1−xNiO2. Our study highlights the importance of accurate benchmarking of the fundamental quantities involved in DFT to select appropriate DFT approximations in order to advance the predictive modeling of charge-transfer-driven phenomena in correlated electron systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

The neurobench framework for benchmarking neuromorphic computing algorithms and systems

Neuromorphic computing shows promise for advancing computing efficiency and capabilities of AI applications using brain-inspired principles. However, the neuromorphic research field currently lacks standardized benchmarks, making it difficult to accurately measure technological advancements, compare performance with conventional methods, and identify promising future research directions. This article presents NeuroBench, a benchmark framework for neuromorphic algorithms and systems, which is collaboratively designed from an open community of researchers across industry and academia. NeuroBench introduces a common set of tools and systematic methodology for inclusive benchmark measurement, delivering an objective reference framework for quantifying neuromorphic approaches in both hardware-independent and hardware-dependent settings. For latest project updates, visit the project website (neurobench.ai).

Yik, Jason [Harvard Univ., Cambridge, MA (United S

Citation network datasets for benchmarking spiking graph neural networks on experimental neuromorphic hardware

Spiking neural networks (SNNs) running on neuromorphic computers offer an energy-efficient alternative for AI tasks. Recently, spiking graph neural networks (S-GNNs) have been shown to produce encouraging results on benchmark citation network datasets such as Cora, CiteSeer, and PubMed for node classification tasks. These S-GNNs were run on SNN simulators only because they contain up to tens of thousands of neurons and up to millions of synapses, translating poorly to neuromorphic hardware. Therefore, in this paper, we create a suite of benchmark datasets from the CiteSeer dataset that can be accommodated on current neuromorphic hardware platforms. Our contribution consists of a collection of three datasets. First, we have an induced subgraph of CiteSeer, which we call MiniSeer, containing 2110 papers, 3604 binary features, and 6 topics. Second, MicroSeer is a very small dataset consisting of 84 papers, 1227 features, and 6 topics. Lastly, BiteSeer is a collection of 15 binary classification datasets. We present creation of these datasets along with accuracies, running times, and spike counts when simulated. We believe that our results in this paper will be used by the neuromorphic community to benchmark, test, and develop neuromorphic hardware and simulators.

Zhu, Kevin [George Mason University, Virginia]

Nuclear Data Adjustment for Nonlinear Applications in the OECD/NEA WPNCS SG14 Benchmark—A Bayesian Inverse UQ-Based Approach for Data Assimilation

The Organisation for Economic Co-operation and Development Working Party on Nuclear Criticality Safety has proposed a benchmark exercise to assess the performance of current nuclear data adjustment techniques applied to nonlinear applications and experiments with low correlation to applications. This work introduces Bayesian inverse uncertainty quantification (IUQ) employing scientific machine learning surrogate models as a method for nuclear data adjustments in this benchmark, and compares IUQ to the more traditional methods of generalized linear least squares (GLLS) and Monte Carlo Bayes (MOCABA). Posterior predictions from IUQ showed agreement with GLLS and MOCABA for linear applications. Here, when comparing GLLS, MOCABA, and IUQ posterior predictions to computed model responses using adjusted parameters, we observe that the GLLS predictions failed to replicate the computed response distributions for nonlinear applications, while MOCABA showed near agreement, and IUQ used the computed model responses directly. We also discuss observations on why experiments with low correlation to applications can be informative to nuclear data adjustments and identify some properties useful in selecting experiments for inclusion in nuclear data adjustment. Performance in this benchmark indicates potential for Bayesian IUQ in nuclear data adjustments.

Bayesian calibration

Benchmarking the performance of a high-Q cavity qudit using random unitaries

High-coherence cavity resonators are excellent resources for encoding quantum information in higher-dimensional Hilbert spaces, moving beyond traditional qubit-based platforms. A natural strategy is to use the Fock basis to encode information in qudits. One can perform quantum operations on the cavity mode qudit by coupling the system to a non-linear ancillary transmon qubit. However, the performance of the cavity-transmon device is limited by the noisy transmons. It is, therefore, important to develop practical benchmarking tools for these qudit systems in an algorithm-agnostic manner. We gauge the performance of these qudit platforms using sampling tests such as the heavy output generation test as well as the linear cross-entropy benchmark, by way of simulations of such a system subject to realistic dominant noise channels. We use selective number-dependent arbitrary phase and unconditional displacement gates as our universal gateset. Our results show that contemporary transmons comfortably enable controlling a few tens of Fock levels of a cavity mode. This framework allows benchmarking even higher dimensional qudits as those become accessible with improved transmons.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND

Benchmarking Variables for Checkpointing in HPC Applications

Checkpoint/Restart (C/R) is a widely used fault tolerance mechanism in converged systems of cloud, edge, and HPC. However, users often rely on their experience to determine which variables to checkpoint, as there is currently no benchmark that can provide a reference. This can result in checkpointing redundant or even incorrect variables. To address this issue, we propose a benchmark suite that includes critical variables for checkpointing, which have been manually identified, and a method for identifying those critical variables, with 20 representative HPC applications. Our method involves analyzing data dependency between variables to identify critical variables analytically. We verify the identified variables' correctness with a widely used C/R library FTI by an ablation study. With our benchmark suite and data dependency analysis, HPC practitioners now have a reference for identifying checkpointing variables and better knowledge of what kind of variables to checkpoint.

Fu, Xiang