Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmark testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

matsim-agents v1.0

matsim-agents is a multi-agent AI framework for atomistic materials simulation and discovery. It orchestrates large language models (LLMs), machine-learned interatomic potentials (MLIPs), and DFT codes into a single agentic loop running on laptops and DOE leadership-class supercomputers. MULTI-AGENT ORCHESTRATION A LangGraph state machine with three nodes: a Planner that converts a natural-language research objective into structured tasks; an Executor that dispatches atomistic tools and loops until the queue is empty; and an Analyst that summarizes results into a human-readable report. State is checkpointed after every step and human-in-the-loop gates can be inserted at any edge. HYPOTHESIS-DRIVEN DISCOVERY CHAT An interactive REPL (matsim-agents chat) that couples LLM dialogue with atomistic simulation. Chemical formulas are automatically detected in conversation turns and trigger a full crystal-phase exploration: structure generation → relaxation → stability scoring → result injection back into the conversation, creating a closed hypothesis-refinement loop. CRYSTAL PHASE ENUMERATION Given a composition, the phase explorer enumerates prototypes by stoichiometry: elemental (fcc/bcc/hcp/sc/diamond), binary 1:1 (rocksalt/CsCl/zincblende/ wurtzite/fluorite/rutile), ternary 1:1:3 (cubic perovskite), ternary 1:2:4 (perovskite + spinel), quaternary 1:1:2:6 (Fm-3m double perovskite). 2-D prototypes (graphene, h-BN, MoS2 2H/1T) and multilayer stacking are also supported via --include-2d and --num-layers. SUPERCELL GENERATION AND SITE DECORATION Auto-tiling to a minimum atom count (--min-atoms), explicit NxNxN tiling (--supercell), symmetry-distinct site decorations (--n-orderings), and isotropic lattice-scale sweeps (--lattice-scales) for volume bracketing. MLFF RELAXATION AND STABILITY SCORING HydraGNN (multi-headed GNN) drives structure relaxation via ASE with FIRE, BFGS, or BFGSLineSearch. Stability output: delta-E/atom ranking across phases and a max-residual-force dynamical-stability proxy. Other MLIPs (MACE, NequIP, Orb) can be plugged in through the same interface. DFT BACKENDS Quantum ESPRESSO pw.x and VASP 6.6 are first-class labellers. Both have validated GPU builds and SLURM/PBS launchers for three DOE platforms: Frontier (AMD MI250X, ROCm), Aurora (Intel PVC, oneAPI), Perlmutter (NVIDIA A100, CUDA). QE produces ~100 binaries (pw.x, ph.x, epw.x, ...). VASP supports scf, relax, vc-relax, and vc-relax-shape run types. ACTIVE-LEARNING LOOP matsim-agents al run CONFIG.yaml drives an iterative HydraGNN-DFT loop: MD generates candidates → ensemble/MC-dropout uncertainty selects the most informative → DFT labels them in parallel inside one allocation → dataset grows → HydraGNN retrains → repeat. DFT backend is a single YAML toggle (dft.backend: vasp | qe). LLM-generated seed structures are supported (no curated POSCAR library needed). Config uses ${VAR}, ${VAR:-default}, ${VAR:?msg} shell-style substitution for cross-user/cross-site portability. LLM BACKENDS Ollama (local, default), vLLM (HPC multi-GPU serving), OpenAI, Anthropic, HuggingFace Transformers+Accelerate. Selected at runtime via flag or env var with no code changes. HPC PORTABILITY Same Python entry points run on Frontier (ROCm 7.2), Aurora (oneAPI), and Perlmutter (CUDA 12). DFT and ML stacks are never co-loaded in the same shell; they couple through the scheduler and filesystem. Advanced multi-node launchers (serve, discovery-chat, single-relaxation, active-learning, QE warm-start) are provided for all three platforms. CODABENCH COMPETITION BUNDLE A self-contained benchmark: 159 atomistic test structures across 11 material classes, 5 tasks (formation energy, forces, ML relaxation, AI-DFT relaxation, phase stability ranking), public/private leaderboard split (30/70), and four ready-to-run baselines: MACE-MP-0, HydraGNN, UMA, AllScAIP.

Lupo Pasini, Massimiliano [Oak Ridge National Labo↗

CephFS experiments on stria.sandia.gov

This report is an institutional record of experiments conducted to explore performance of a vendor installation of CephFS on the SNL stria cluster. Comparisons between CephFS, the Lustre parallel file system, and NFS were done using the IOR and MDTEST benchmarking tools, a test program which uses the SEACAS/Trilinos IOSS library, and the checkpointing activity performed by the LAMMPS molecular dynamics simulation.

74 ATOMIC AND MOLECULAR PHYSICS↗

Assessment of ORIGEN Reactor Library Development for Pebble-Bed Reactors Based on the PBMR-400 Benchmark

This report provides an evaluation of present SCALE capabilities for modeling depletion of pebble-bed reactor systems, using the PBMR-400 benchmark as a test case. A specific aim of this work is to understand the system characteristics required to generate production-quality reactor data libraries for rapid depletion calculations with ORIGEN. This report includes a discussion of present SCALE capabilities for modeling doubly heterogeneous fuels, prior SCALE work modeling pebble bed–type reactors, and a detailed neutronic analysis of the PBMR-400 core for both fresh and equilibrium-composition core conditions.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

CephFS experiments on stria.sandia.gov

This report is an institutional record of experiments conducted to explore performance of a vendor installation of CephFS on the SNL stria cluster. Comparisons between CephFS, the Lustre parallel file system, and NFS were done using the IOR and MDTEST benchmarking tools, a test program which uses the SEACAS/Trilinos IOSS library, and the checkpointing activity performed by the LAMMPS molecular dynamics simulation.

97 MATHEMATICS AND COMPUTING↗

Energy Characterization of the UMD Electron Accelerator

Multiple different groups at Los Alamos National Laboratory have a cited interest in understanding the effects of ionizing radiation in materials. Of especial interest to the ISR division is understanding the effects in materials that are used in the space environment, where radiation is omnipresent and impossible to fully shield from. Development is under way to simulate the effects of radiation damage in detectors and electronics, primarily with particle transport codes such as Geant4 and RAMPART, a code that provides a user friendly interface to Geant4. However, these simulations need benchmarking to experimental tests, some of which have been performed at a variety of small electron accelerator facilities. One of these facilities is the Linear Electron Accelerator at the University of Maryland Radiation Facilities. This facility can provide electron beams on the order of 100 mA delivered in tight macropulse bunches, which allows for experiments where the behavior of electronics is monitored as a function of pulse number. While the beam rates per macropulse of this accelerator is monitored during an experiment via a current pick-off connected to a large graphite beamstop, the energy distribution of this electron beam is not measured during the experiment and has not been well characterized. The energy distribution of the beam is pertinent for these types of experiments, since dose is directly related to energy deposited in a medium. Further, these experiments are performed in open air and sample placement in the room can dictate what dose rate they receive per macropulse. We have created simulated dosemaps for the UMD vault room, but these simulations assume a given energy of the beam, both for calculating the dose and to simulate how much the beam diverges in the air. There have been two attempts now to characterize this accelerator beam energy with a plate type spectrometer, including the most recent experiment earlier this year to quantify the beam energy distribution. These results will also be compared to a dosimetry experiment performed by collaborators at the UMD facility, the details of which are described in Reference. Understanding the electron beam energy is important for future experiments, because the energy directly impacts the dose and radiation damage given to a sample and is important for properly simulating the electron beam when comparing simulation to experiment.

43 PARTICLE ACCELERATORS↗

HTGR Multiphysics Application Drivers FY26 Updates

This report summarizes FY26 progress under the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program's high-temperature gas-cooled reactor (HTGR) application driver work, covering a wide range of activities such as code validation and multi-physics code assessment. 1) A detailed SAM model of the High-Temperature Engineering Test Reactor (HTTR) was developed using a unique-block grouping approach, with an extended parallel thermal network method to capture block-to-block conduction and radiation heat transfer, and applied to steady-state simulations of the HTTR 30~MW and 9~MW cases. 2) In another activity, SAM's newly implemented multi-component gas flow model was validated against the Natural convection Shutdown heat removal Test Facility (NSTF) argon ingress experiment, correctly capturing the density-driven suppression and thermal recovery of natural circulation observed when argon is introduced into the air-cooled Reactor Cavity Cooling System (RCCS) loop. 3) For the OECD/NEA High Temperature Test Facility (HTTF) benchmark, we co-led the international benchmark activities as well as the OECD/NEA final benchmark report to be released at the end of this year. 4) Finally, the coupled Griffin-SAM modeling capability for pebble-bed HTGRs was advanced by verifying the Griffin neutronics solution against Serpent Monte Carlo for a realistic non-uniform temperature distribution, resolving several deficiencies in the SAM-to-Griffin temperature transfer scheme, and enabling distinct fuel kernel, moderator, and coolant temperatures for cross section feedback. These new features were demonstrated in a PBR load-following transient.

Lee, Alvin↗

Multi-Area Distribution System State Estimation Using Decentralized Physics-Aware Neural Networks

The development of active distribution grids requires more accurate and lower computational cost state estimation. In this paper, the authors investigate a decentralized learning-based distribution system state estimation (DSSE) approach for large distribution grids. The proposed approach decomposes the feeder-level DSSE into subarea-level estimation problems that can be solved independently. The proposed method is decentralized pruned physics-aware neural network (D-P2N2). The physical grid topology is used to parsimoniously design the connections between different hidden layers of the D-P2N2. Monte Carlo simulations based on one-year of load consumption data collected from smart meters for a three-phase distribution system power flow are developed to generate the measurement and voltage state data. The IEEE 123-node system is selected as the test network to benchmark the proposed algorithm against the classic weighted least squares and state-of-the-art learning-based DSSE approaches. Numerical results show that the D-P2N2 outperforms the state-of-the-art methods in terms of estimation accuracy and computational efficiency.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancement of Distribution System State Estimation Using Pruned Physics-Aware Neural Networks: Preprint

Realizing complete observability in the three-phase distribution system remains a challenge that hinders the implementation of classical state estimation algorithms. In this paper, a new method so-called pruned physics-aware neural network (P2N2) is developed to improve the voltage estimation accuracy in the distribution system. The method relies on the physical grid topology, which is used to design the connections between different hidden layers of a neural network model. To verify the proposed method, a numerical simulation based on one-year smart meter data of load consumptions for threephase power flow is developed to generate the measurement and voltage state data. The IEEE 123 node system is selected as the test network to benchmark the proposed algorithm against the classical weighted least squares (WLS). Numerical results show that P2N2 outperforms WLS, in terms of data redundancy and estimation accuracy.

distribution systems state estimation↗

An Efficient, Scalable IO Framework for Sparse Data: larcv3

Neutrino physics is one of the fundamental areas of research into the origins and properties of the Universe. Many experimental neutrino projects use sophisticated detectors to observe properties of these particles, and have turned to deep learning and artificial intelligence techniques to analyze their data. From this, we have developed \texttt{larcv}, a \texttt{C++} and \texttt{Python} based framework for efficient IO of sparse data with particle physics applications in mind. We describe in this paper the \texttt{larcv} framework and some benchmark IO performance tests. \texttt{larcv} is designed to enable fast and efficient IO of ragged and irregular data, at scale on modern HPC systems, and is compatible with the most popular open source data analysis tools in the Python ecosystem.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Validation of OpenPronghorn for Periodic Hill Flow Separation

OpenPronghorn is an open-source, MOOSE-based thermal-hydraulics simulation tool used for advanced reactor analysis. As an open-source code, it offers a transparent framework for validating governing equations, assumptions, and numerical methods against established benchmarks. This study evaluates OpenPronghorn's RANS turbulence model against the ERCOFTAC Case 81 periodic hill benchmark, a standard test case for separated turbulent flow featuring curved-wall separation, recirculation, shear-layer development, and reattachment. Streamwise velocity profiles predicted by OpenPronghorn were compared to reference LES data at multiple x/h locations. Results show that OpenPronghorn captures the overall trend of the velocity profiles, but the largest discrepancies occur in the separated-flow region, where turbulence is highly anisotropic and strongly affected by adverse pressure gradients and wall curvature—conditions that are inherently difficult for standard RANS models to resolve. Future work will test alternative k-e model variants and correction terms to improve prediction accuracy in this region.

42 - ENGINEERING↗

HELIUM LEAK TEST MODELING OF A SPENT NUCLEAR FUEL CANISTER

The U.S. Department of Energy (DOE) is considering the development of one or more federal consolidated interim storage facilities (CISFs) to be used to store commercial spent nuclear fuel (SNF) at locations in the U.S. One of the first technical challenges of a CISF is performing an inspection of SNF canisters upon their receipt to confirm they can be placed into the CISF’s licensed storage configuration. The canister receipt inspection is critical to CISF site operations. The test is conceived as being a helium (He) leak check, intended to confirm that the confinement boundary of a SNF canister is intact. SNF canisters are filled with He when they are sealed, so detection of a He leak indicates that a through-wall flaw has occurred in the canister confinement boundary. Other measurements are planned to occur upon canister receipt in addition to the He leak check such as krypton-85 measurements, which would indicate confinement breaches of one or more fuel rods in addition to a breach of the SNF canister. However, the He leak check has been identified as one of such high importance and has such significant technical challenges that a full-scale demonstration is needed to confirm the He leak test’s viability and to assist in planning relative to its operational requirements. A modeling methodology for simulating the He detection test was developed to help inform the test plan and the design of the test vessels. To develop the modeling methodology a detailed computational fluid dynamics (CFD) benchmark model was constructed to compare against leak rate test data from a transportation package for radioactive material. This report is focused on modeling efforts to simulate the benchmark leak test.

Suffield, Sarah R.↗

Code Benchmark of Depressurized Conduction Cooldown Transient in the High Temperature Test Facility

This paper presents results from modeling of a depressurized conduction cooldown (DCC) transient at the High Temperature Test Facility (HTTF) as part of the OECD-NEA Thermal Hydraulics Code Validation Benchmark for High-Temperature Gas-Cooled Reactors using HTTF Data . This paper briefly describes the benchmark and the models being used. It then presents a comparison of steady state and transient results based on the Problem 2 Exercise 1A and 1B definitions. We compare block and helium temperature distributions, mass flow distribution, and energy balance in steady state. All models show comparable mass flow distributions and energy balances. The temperatures within the core and outer regions are comparable in all models too, but inner reflector temperatures can vary significantly. Despite that, we find that the models are in good agreement for the full-power steady state. In the DCC, we look at block temperature at the core midplane and RCCS water exit temperature. The INL and ANL models are found to be in excellent agreement with one another on block temperature over time, while the agreement when the KAERI and NRG models are added into consideration is good. Differences in the transient heat removal from the RCCS cause the differences in block temperature over time in these models. The CNL models show similar trends to the INL, ANL, KAERI, and NRG models, but the temperatures are high because the volumes used in calculating the average temperature include the heater rods in the CNL models only. The HUN-REN model shows results that suggest significantly lower heat removal in the RCCS which merit further investigation.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Benchmark Specification for Select Experiments Conducted at the University of Wisconsin-Madison Thermal Stratification Test Facility

The Department of Energy (DOE)-Nuclear Energy University Programs (NEUP) supported the creation and operation of the Thermal Stratification Test Facility (TSTF) at the University of Wisconsin Madison (UWM) as part of a larger effort to understand thermal stratification behavior in liquid-metal-cooled reactors. The TSTF was designed to simulate transients in a reactor plenum that are known to cause thermal stratification. High-reliability and high-resolution measurements of the flow and temperature were collected for use as experimental benchmarks to support validation efforts for computational models. The results of these tests contribute to the greater understanding of thermal stratification behavior of liquid sodium under various configurations and operating conditions. The six TSTF tests selected for benchmarking are a set of forced circulation tests at a fixed flow rate with different Upper Internal Structure (UIS) configurations in the test section (no UIS, solid UIS, and a UIS with flow area of 4, 8, 12, and 100%). This report provides a complete description of the benchmark problems, including all key test facility details, descriptions of each test condition, and measured data for comparison with modeled results.

42 ENGINEERING↗

Thermal Neutron Scattering Law Benchmark and Validation at NCSU [Slides]

The presentation discusses how the NCSU graphite SDT ORELA experiment is under development as a benchmark. It discusses the testing of various ENDF/B-VII.1 and ENDF/B-VIII.0 graphite libraries that have been performed using the PROTEUS benchmark. The presentation demonstrates that the NSCU Polyethylene ENDF/B-VIII.0 TSL improves agreement with total cross section measurements. The presentation concludes by stating that the NCSU single crystal sapphire TSL library is based on realistic data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Development of the TREAT M2 experiment as a transient benchmark

A transient benchmark based on the Transient Reactor Test Facility (TREAT) M2 Calibration experiment (M2-CAL) is under development. TREAT, at Idaho National Laboratory, is a graphite moderated air-cooled research reactor which has been used extensively for fuel material testing under extreme and accident conditions. Accurate benchmark models are a beneficial component in the operations, experimental planning, and development of TREAT. In this work, we present a transient benchmark model for the M2-CAL experiment core loading. The benchmark model incorporates the coupling between the Monte Carlo Code SERPENT and the computational fluid dynamics code OpenFOAM to capture the temperature feedback mechanism. The M2-CAL transient 2580 was simulated in this work. A pre-transient analysis was performed to determine the optimum core composition and conditions before the beginning of the transient. The analysis was validated against historic TREAT kinetic measurements and the worth of the transient rod T-2. The axial power distribution in the flux wire for the M2-CAL experiment was determined and contrasted with the experiment. The model was then used to simulate the M2-CAL transient 2580 experiment based on the reported pre-transient and transient conditions. Several transient observables were calculated and compared to the experiment. The model shows good agreement with the experimental power traces, period, and average increase of power at the power ramp (time > 7.8 s). Inverse point kinetics analysis was performed during the period of the power ramp for the model and the experiment based on that, the temperature feedback component was isolated and contrasted with the experimental feedback. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

DECOVALEX-2023: Task F1 Final Report

DECOVALEX-2023 Task F is a comparison of models and methods for post-closure performance assessment (PA) of a deep geologic repository for radioactive waste. The general aims of Task F are to build confidence in the models, methods, and software used for PA and to stimulate additional research and development in PA methodologies. The task objectives are to motivate development of PA modelling skills and capabilities, to examine the influence of model choices on calculated repository performance, and to compare the uncertainties introduced by model choices to other sources of uncertainty. Task F involves no actual experiment or site. It is a PA modelling exercise that requires the conceptual development of hypothetical repository designs and geologic settings. Because three of the teams were interested in salt and the rest of the teams were interested in crystalline rock, Task F was split into two branches: Task F1 for crystalline rock and Task F2 for salt. This report is for Task F1, crystalline rock. Teams from seven countries (Canada, Czech Republic, Germany, Korea, Sweden, Taiwan, and United States) participated in Task F1. The teams worked together to define the features, events, and processes of the reference case repository and established a set of performance measures. In addition, they defined a set of benchmark problems designed to test and compare modelling capabilities for fracture flow and transport at different scales. The repository design and benchmark problems are documented in a Task Specification that evolved over time as the group honed the specifications. The benchmark problems verified that each team can aptly model flow and transport in fractured media in 1-, 2-, and 3-dimensions. Two general approaches were used for the 3-dimensional benchmarks: discrete fracture network (DFN) and equivalent continuous porous medium (ECPM). DFN modelling involves explicit meshing of each fracture while ECPM modelling aims to capture the effective porosity and directional permeability of each cell in a space-filling mesh as affected by intersecting fractures. In some models, a combination of the two is used, i.e., DFN for large known fractures and ECPM for the rest of the domain. Transport is solved by using either the advection-dispersion equation or particle tracking. Although some variation is observed among model breakthrough curves in the benchmark problems, there is strong agreement in breakthrough behaviour up to at least the 75 th percentile for all benchmarks. At the 90 th percentile, breakthrough results show larger differences, suggesting several models retain substantially higher fractions of tracer in regions of slower moving water. In addition to the flow and transport benchmarks, several teams completed the source term benchmark, verifying capabilities for modelling radionuclide decay and ingrowth, waste package breach, instant release fractions, fuel matrix degradation rates, and radionuclide solubility limitations. The reference case is conceptualized as a generic spent fuel repository at a depth of 450 m in fractured crystalline rock. The repository has 50 parallel backfilled drifts, each with 50 deposition holes 6 m apart. Each deposition hole contains a 4-PWR waste package and bentonite buffer. The rock domain is 5 km in length, 2 km in width, and 1 km in depth. It has 6 deterministic fractured deformation zones and a multitude of stochastic fractures. Teams generally used the ECPM approach for the entire rock or a hybrid approach in which the deterministic fracture zones are modelled with a DFN and the rest of the rock is modelled by ECPM. Of the reference case problems specified, only the results of the initial reference case problem are compared in this report. The initial problem focuses on transport from the deposition holes to the surface, i.e., it neglects waste package performance. Tracers are released at all waste package locations at time zero and tracked for their releases to the near field and ground surface. The water fluxes calculated at the ground surface entry and exit regions of the domain are similar for all models except for two that have considerably lower fluxes. For tracer transport, large differences are observed among models in the magnitude of tracer transported. Much of the difference appears to be due to how the repository is implemented and hence the different degrees of repository simplification. Models that exclude the drifts, buffer, and backfill from the domain tend to show greater release of tracers and radionuclides from the repository. The initial study presented here indicates that major differences in modelling important processes within the repository (e.g., diffusion through buffer and backfill) can produce broadly different release and transport results, especially when those processes are excluded. Even for the models that included all specified features, events, and processes, the results show significant differences and demonstrate the importance of examining multiple modelling approaches in performance assessment. The differences in results observed in this study are expected to motivate teams to either increase complexity in future versions of the reference case models or to improve methods to account for the effects of simplified features and processes. Either way, future improvements in these models are expected to produce results that more closely agree.

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗