Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “model based definitions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Kinematic and dynamical origins of mean-p T fluctuations in heavy-ion collisions

Event-by-event fluctuations of the mean transverse momentum (mean-p T ) provide a sensitive probe of collective dynamics beyond single-particle spectra and anisotropic flow. We present a systematic study of mean-p T fluctuation observables using a Bayesian-calibrated multistage hydrodynamic framework, including quantitative comparisons to RHIC measurements and model-based investigations of beam-energy and kinematic-acceptance effects. The experimental definitions employed by the STAR and ALICE Collaborations are implemented explicitly and found to yield consistent results within controlled limits. We study the centrality and beam-energy dependence of the observable, its sensitivity to key soft-sector ingredients, and the impact of the kinematic p T acceptance. By introducing scaled-p T cuts, we demonstrate that a part of the apparent energy dependence arises from kinematic projection effects, while the remaining trends reflect genuine collective dynamics. Our results establish mean-p T fluctuations as a nontrivial and independent validation of calibrated hydrodynamic descriptions of the quark–gluon plasma.

Event-by-event correlations↗

High dimensional predictions of suicide risk in 4.2 million US Veterans using ensemble transfer learning

We present an ensemble transfer learning method to predict suicide from Veterans Affairs (VA) electronic medical records (EMR). A diverse set of base models was trained to predict a binary outcome constructed from reported suicide, suicide attempt, and overdose diagnoses with varying choices of study design and prediction methodology. Each model used twenty cross-sectional and 190 longitudinal variables observed in eight time intervals covering 7.5 years prior to the time of prediction. Ensembles of seven base models were created and fine-tuned with ten variables expected to change with study design and outcome definition in order to predict suicide and combined outcome in a prospective cohort. The ensemble models achieved c-statistics of 0.73 on 2-year suicide risk and 0.83 on the combined outcome when predicting on a prospective cohort of ~4.2 M veterans. The ensembles rely on nonlinear base models trained using a matched retrospective nested case-control (Rcc) study cohort and show good calibration across a diversity of subgroups, including risk strata, age, sex, race, and level of healthcare utilization. In addition, a linear Rcc base model provided a rich set of biological predictors, including indicators of suicide, substance use disorder, mental health diagnoses and treatments, hypoxia and vascular damage, and demographics. Similar content being viewed by others

60 APPLIED LIFE SCIENCES↗

Calculation of neutron flux spectra of the VVER-1000 mock-up shielding benchmark with Monte Carlo code MCS utilizing mesh-based weight window

The measurements of neutron spectra compiled inside the NEA-1517/82 package from the Shielding Integral Benchmark Archive and Database (SINBAD) are chosen as benchmark cases to validate the variance reduction technique based on weight window in Monte Carlo Code MCS. A full 3D model for fixed source mode calculation with hexagonal lattice source definition is developed to simulate total of 6 points of measurements at the vicinity of the reactor and the reactor pressure vessel region. A code/code comparison against MCNP6 code is first conducted as verification element for the mesh-based weight window capability in MCS. Finally, the validation results are presented against measurements. The verification against MCNP6 code gives good agreement in addition of the insight to the importance of user understanding to determine proper reference point and reference lower weight bound for scaling which is not required in MCS code due to its capability of automatic scaling. The comparison of neutron spectra between MCS and measurements shows good agreement within 3 standard deviations for all of six detector positions.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

A Unified Workflow for Sensitivity-Based Kinetic Analysis in Microkinetic Models

Degrees of rate control (DRC), apparent activation energies, and apparent reaction orders are established local sensitivity diagnostics for interpreting microkinetic models, but applying them routinely to large mechanisms often requires substantial reaction-specific bookkeeping, perturbation design, and postprocessing. Here, in this study, we present a unified derivative-based workflow that evaluates these quantities from a single compiled reaction-network model and target-rate definition. For any user-provided microkinetic model, the workflow compiles the mechanism into stoichiometrically consistent mass-action rate equations, solves the surface dynamics, and uses automatic differentiation to compute sensitivities with respect to rate constants, temperature, and gas partial pressures. By combining their calculations in the same framework, the workflow clearly demonstrates the relationships between different DRCs and the apparent activation energy. Using existing examples of propylene partial oxidation and methane oxidation on Pd(100), we verify expected transient redistribution of rate control, distinguish net Campbell DRCs from one-sided directional sensitivities, and show how apparent activation energy can be reconstructed either from one-sided DRCs or from state-based DRCs while critical mechanistic insights are obtained consistently. In the methane oxidation case, a pathway-subset test further illustrates how a simplified mechanism preserves key kinetic signatures of a full model, showing the potential of our user-friendly tool for model construction beyond kinetic analysis.

36 MATERIALS SCIENCE↗

Technical Report on Waveform Fit Metrics for Global Models

The new WAVEFORMS Initiative in the Ground-based Nuclear Detonation Detection (GNDD) program includes an increased emphasis on the development of Earth models and methods to predict entire seismic and acoustic waveforms more accurately. In general, this increased emphasis is predicated on the need to better characterize seismic events and provide improved model-based discrimination between event types including earthquakes and explosions. More specifically, while current moment tensor inversion methods tend to work well for larger events (M>~4) using tuned 1-D Earth models, the development of state-of-the-art 3-D models and methods is required for the prediction of shorter period waves over large areas for discrimination of smaller events. There is no standard metric for model-based waveform prediction accuracy used in the waveform modeling/inversion community. However, there are several popular waveform misfit definitions; and minimizing the corresponding objective functions is the goal of waveform inversion. Some example misfit definitions employed for adjoint waveform tomography include measures of simple travel time differences (e.g. Tape et al., 2010), cross-correlation travel time differences (e.g. Luo and Schuster, 1991), multi-taper frequency dependent methods (e.g. Lei et al., 2020), time-frequency phase misfit functions (e.g. Fichtner 2010; Rodgers et al., 2022), normalized cross-correlation methods (e.g. Tao et al., 2018), and others. In some cases, these misfit definitions also involve complicated weighting schemes and summations over multiple frequency bands making it difficult to duplicate the misfit measurement with alternative models and datasets. Although each of the misfit definitions mentioned above are useful for developing waveform models, the actual misfit values are not usually meaningful outside of a given project, model, and/or dataset. Therefore, it is difficult to understand and communicate model performance for predicting waveforms and comparing to other models and/or new model iterations with a different dataset. Therefore, there is a need for a generalized method for evaluating overall model performance that is independent from the specific misfit chosen to develop the waveform models that is also intuitive and meaningful. In this report, we describe a new metric we refer to as ‘Percent of Correlated Signal’. The following sections describe and demonstrate the metric with a case study event and a more rigorous test using a random selection of globally distributed events. While the focus here is on global tomography models, the metric is meant to applicable to regional ‘wiggle-for-wiggle’ waveform models/studies as well.

58 GEOSCIENCES↗

Generating An Advanced Cross-section Library For HTGR Pebble Bed Depletion Calculations Using Reduced-Order Model Generation Techniques

For code development, Advanced Reactor Technologies - Gas Cooled Reactors Program (ART-GCR) rely on a collaboration with the Nuclear Energy Advanced Modeling and Simulation (NEAMS) program, but the cross sections generation and the methodology definition is part of this program area goals. Based on previous studies in FY23, the size of microscopic cross section libraries increases rapidly with the number of tabulations, requiring significant amount of memory and drastically slowing down the Griffin calculations when evaluating cross sections via the multivariate linear interpolation approach. Rising to these challenges, this work investigates constructing Reduced-order Models (ROMs) for the multi-group microscopic cross sections to accelerate the cross section evaluation in Griffin. A database of multigroup cross sections is first collected considering all possible parameters that a designer could change for optimization. Down-selection of the ROM techniques afterward shows Deep Neural Network (DNN) as the best candidate when jointly consider memory efficiency, predictive accuracy, computational cost, scalability, flexibility and ease of implementation of the algorithms in comparison to the multidimensional interpolation. This work develops a specific interface that enables the cross section predictions using pre-trained DNN models into Griffin leveraging the existing ROM capabilities. DNNs have been trained for all isotopes for use in Griffin. Preliminary Griffin testing shows that DNNs exhibit exceptional predictive accuracy and the use of DNNs provides orders of magnitude improvement in memory efficiency compared to conventional interpolation techniques. With such ROM techniques, it holds great promise to further increase the fidelity of the Pebble Bed Reactor (PBR) simulation by increasing the number of tabulations/state variables during cross section evaluation, while maintaining the computational cost affordable in Griffin.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Dynamics-based halo model for large scale structure

Accurate modeling of the one-to-two halo transition has long been difficult to achieve. Here, we demonstrate that physically motivated halo definitions that respect the bimodal phase-space distribution of dark matter particles near halos resolves this difficulty. Specifically, the two phase-space components are overlapping and correspond to (1) particles orbiting the halo and (2) particles infalling into the halo for the first time. Motivated by this decomposition, García et al. [Mon. Not. R. Astron. Soc. 521, 2464 (2023)] advocated for defining halos as the collection of particles orbiting their self-generated potential. This definition identifies the traditional one-halo term of the halo-mass correlation function with the distribution of orbiting particles around a halo, while the two-halo term governs the distribution of infalling particles. We use dark matter simulations to demonstrate that the distribution of orbiting particles is finite and can be characterized by a single physical scale 𝑟 h , which we refer to as the halo radius. The two-halo term is described using a simple yet accurate empirical model based on the Zel’dovich correlation function. We further demonstrate that the halo radius imprints itself on the distribution of infalling particles at small scales. Our final model for the halo-mass correlation function is accurate at the ≈ 2% level for 𝑟∈ [0.1, 50] ℎ −1 Mpc. The Fourier transform of our best-fit model describes the halo-mass power spectrum with comparable accuracy for 𝑘 ∈ [0.06, 6.0] ℎ Mpc −1 .

79 ASTRONOMY AND ASTROPHYSICS↗

Sharp decline of dust events induces regional wetting over arid and semi-arid Northwest China in the NCAR Community atmosphere model

Abstract Multiple lines of observational evidence have indicated a significant wetting over the arid and semi-arid Northwest China (NWC) during recent decades, coinciding with a simultaneous sharp decline of dust events. Although recent studies have attributed NWC wetting to different anthropogenic and natural forcings, the mechanisms are not definitive and the regional wetting has been greatly underestimated in the Coupled Model Intercomparison Project historical simulations. Based on sensitivity experiments with different dust emission amounts using the NCAR Community Atmospheric Model version 5 (CAM5), here we find that decreasing dusts exert significant impacts on mixed-phase clouds through reducing the concentration of ice nucleating particles, increase the NWC precipitation and thus induce regional wetting through enhancing convection precipitation. A possible convection invigoration mechanism whereby the atmospheric vertical temperature gradient and convective instability are strengthened by reduced dusts, leading to convection invigoration and increased precipitation. These results are reinforced by simulations over the dust region in North Africa where mixed-phase and ice clouds are rare and reduced dusts do not increase precipitation. This study highlights the possible mechanism of dust-ice cloud interactions in recent NWC wetting and future regional climate change.

54 ENVIRONMENTAL SCIENCES↗

Towards 20 T Hybrid Accelerator Dipole Magnets

We report the most effective way to achieve very high collision energies in a circular particle accelerator is to maximize the field strength of the main bending dipoles. In dipole magnets using Nb-Ti superconductor the practical field limit is considered to be 8-9 T. When Nb 3 Sn superconductor material is utilized, a field level of 15-16 T can be achieved. To further push the magnetic field beyond the Nb 3 Sn limits, High Temperature Superconductors (HTS) need to be considered in the magnet design. The most promising HTS materials for particle accelerator magnets are Bi2212 and REBCO. However, their outstanding performance comes with a significantly higher cost. Therefore, an economically viable option towards 20 T dipole magnets could consist in an “hybrid” solution, where both HTS and Nb3Sn materials are used. We discuss in this paper preliminary conceptual designs of various 20 T hybrid magnet concepts. After the definition of the overall design criteria, the coil dimensions and parameters are investigated with finite element models based on simple sector coils. Preliminary 2D cross-section computation results are then presented and three main layouts compared: cos-theta, block, and common-coil. Both traditional designs and more advanced stress-management options are considered.

43 PARTICLE ACCELERATORS↗

Harmonising the land-use flux estimates of global models and national inventories for 2000–2020

As the focus of climate policy shifts from pledges to implementation, there is a growing need to track progress on climate change mitigation at the country level, particularly for the land-use sector. Despite new tools and models providing unprecedented monitoring opportunities, striking differences remain in estimations of anthropogenic land-use CO 2 fluxes between, on the one hand, the national greenhouse gas inventories (NGHGIs) used to assess compliance with national climate targets under the Paris Agreement and, on the other hand, the Global Carbon Budget and Intergovernmental Panel on Climate Change (IPCC) assessment reports, both based on global bookkeeping models (BMs). Recent studies have shown that these differences are mainly due to inconsistent definitions of anthropogenic CO 2 fluxes in managed forests. Countries assume larger areas of forest to be managed than BMs do, due to a broader definition of managed land in NGHGIs. Additionally, the fraction of the land sink caused by indirect effects of human-induced environmental change (e.g. fertilisation effect on vegetation growth due to increased atmospheric CO 2 concentration) on managed lands is treated as non-anthropogenic by BMs but as anthropogenic in most NGHGIs. We implement an approach that adds the CO 2 sink caused by environmental change in countries' managed forests (estimated by 16 dynamic global vegetation models, DGVMs) to the land-use fluxes from three BMs. This sum is conceptually more comparable to NGHGIs and is thus expected to be quantitatively more similar. Our analysis uses updated and more comprehensive data from NGHGIs than previous studies and provides model results at a greater level of disaggregation in terms of regions, countries and land categories (i.e. forest land, deforestation, organic soils, other land uses). Our results confirm a large difference (6.7 GtCO 2 yr —1 ) in global land-use CO 2 fluxes between the ensemble mean of the BMs, which estimate a source of 4.8 GtCO 2 yr —1 for the period 2000–2020, and NGHGIs, which estimate a sink of —1.9 GtCO 2 yr —1 in the same period. Most of the gap is found on forest land (3.5 GtCO 2 yr —1 ), with differences also for deforestation (2.4 GtCO 2 yr —1 ), for fluxes from other land uses (1.0 GtCO 2 yr —1 ) and to a lesser extent for fluxes from organic soils (0.2 GtCO 2 yr —1 ). By adding the DGVM ensemble mean sink arising from environmental change in managed forests (—6.4 GtCO 2 yr —1 ) to BM estimates, the gap between BMs and NGHGIs becomes substantially smaller both globally (residual gap: 0.3 GtCO 2 yr —1 ) and in most regions and countries. However, some discrepancies remain and deserve further investigation. For example, the BMs generally provide higher emissions from deforestation than NGHGIs and, when adjusted with the sink in managed forests estimated by DGVMs, yield a sink that is often greater than NGHGIs. In summary, this study provides a blueprint for harmonising the estimations of anthropogenic land-use fluxes, allowing for detailed comparisons between global models and national inventories at global, regional and country levels. This is crucial to increase confidence in land-use emissions estimates, support investments in land-based mitigation strategies and assess the countries' collective progress under the Global Stocktake of the Paris Agreement.

54 ENVIRONMENTAL SCIENCES↗

Woven ceramic matrix composite surrogate model based on physics-informed recurrent neural network

A recurrent neural network (RNN) based surrogate model is developed to emulate the nonlinear constitutive behavior of woven ceramic matrix composites (CMCs) driven by matrix damage at multiple length scales. Physics-informed constraints are introduced into the surrogate model through regularization to ground the prediction in physics and improve its predictive capabilities. Training data is generated using the multiscale generalized method of cells (MSGMC) approach coupled with a matrix damage model. This coupling permits simulating the nonlinear behavior of woven CMCs based on constituent response at the micro-, meso-, and macroscales. The multiscale repeating unit cell is loaded under non-monotonic conditions including multiple load / unload cycles and tension / compression. The fiber volume fraction as well as the intra- and intertow void volume fractions are also varied in the generation of training data. Therefore, the RNN-based surrogate model is tasked with predicting, as a function of variable input strain sequence and fiber and void volume fractions, the resulting stress versus strain response while satisfying physical constraints such as positive semi-definiteness of the tangent stiffness matrix and linear elastic unloading. Further, the trained surrogate model effectively matches the stress versus strain response and successfully predicts the tangent modulus throughout the loading regime. Neural network based surrogate models can offer efficient alternatives to running computationally intensive multiscale material models to simulate the nonlinear response of large structural models. Therefore the presented work provides evidence towards the feasibility of developing, training, and running such models for CMCs with complex architectures, nonlinear multiaxial material response, and under non-monotonic loading conditions.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Wind and solar energy droughts: Potential impacts on energy system dynamics and research needs

This Perspective article provides a brief overview of the topic of wind and solar energy droughts (henceforth WSDs). It does not attempt to provide a complete literature review of the subject but rather highlights some of the main concepts associated with WSDs. These include wind and solar energy drought definitions and metrics; meteorological conditions producing WSDs; a comparison of their characteristics with hydrologic droughts and hydropower droughts; model-based and observational datasets useful for WSD analyses; the linkage of WSDs to transmission, storage, and demand response; the potential impacts of WSDs vs energy demand variations; wind and solar flood events; WSD predictability; WSD dependency on climate modes of variability; climate change impacts on WSDs; and the special challenge of evaluating the characteristics of WSDs in developing countries that have limited historical data available. Finally, the manuscript identifies research areas that the authors believe would provide immediate benefit to energy system planners.

14 SOLAR ENERGY↗

Mechanistic understanding of pH effects on the oxygen evolution reaction

The oxygen-evolution reaction (OER) is pivotal in many energy-conversion technologies as it is an important counter reaction to others that convert stable chemicals to higher-value products using electrochemistry. The local microenvironment and pH for the anode OER can vary from acidic to neutral to alkaline depending on the system being explored, making definitive mechanistic insights difficult. In this paper, we couple experiments, first-principles calculations based on density functional theory, microkinetics, and transport modeling to explore the entire pH range of the OER. At low current densities, neutral pH values unexpectedly perform better than the acidic and alkaline conditions, and this trend is reversed at higher current densities (> 20 mA cm -2 ). Using multiscale modeling, this switch is rationalized by a change from a dual-reaction mechanism to a single rate-determining step. The model also shows how the alkaline reaction rates dominate in the middle to high pH range. Furthermore, we explore that the local pH for near-neutral conditions is much different (e.g., 2.4 at the reaction surface vs. 9 in the bulk) than the pH extremes, demonstrating the criticality that transport phenomena plays in kinetic activity.

36 MATERIALS SCIENCE↗

Robustness of Vacancy-Bound Non-Abelian Anyons in the Kitaev Model in a Magnetic Field

Non-Abelian anyons in quantum spin liquids (QSLs) provide a promising route to fault-tolerant topological quantum computation. In the exactly solvable Kitaev honeycomb model, such anyons of the QSL state can be bound to nonmagnetic spin vacancies and endowed with non-Abelian statistics by an infinitesimal magnetic field. Here, we investigate how this approach for stabilizing non-Abelian anyons extends to a finite magnetic field represented by a proper Zeeman term. Through large-scale density-matrix renormalization group simulations, we compute the vacancy-anyon binding energy as a function of magnetic field for both the ferromagnetic and antiferromagnetic Kitaev models. Here, we find that anyon binding remains robust within the entire QSL phase for the ferromagnetic Kitaev model but breaks down already inside this phase for the antiferromagnetic Kitaev model. To compute a binding energy several orders of magnitude below the magnetic energy scale, we introduce both a refined definition and an extrapolation scheme based on carefully tailored perturbations.

Xiao, Bo [Oak Ridge National Laboratory (ORNL), Oa↗

Ab initio calculations of third-order elastic coefficients

Third-order elasticity (TOE) theory predicts strain-induced changes in second-order elastic coefficients (SOECs) and can model elastic wave propagation in stressed media. Although third-order elastic tensors have been determined based on first principles in previous studies, their current definition is based on an expansion of thermodynamic energy in terms of the Lagrangian strain near the natural, or zero pressure, reference state. This definition is inconvenient for predictions of SOECs under significant initial stresses. Therefore, when TOE theory is necessary to study the strain dependence of elasticity, the seismological community has resorted to an empirical version of the theory. This study reviews the thermodynamic definition of the third-order elastic tensor and proposes using an “effective” third-order elastic tensor. An explicit expression for the effective third-order elastic tensor is given and verified. Additionally, we extend the ab initio approach to calculate third-order elastic tensors under finite pressure and apply it to two cubic systems, namely, NaCl and MgO. As applications and validations, we evaluate (a) strain-induced changes in SOECs and (b) pressure derivatives of SOECs based on ab initio calculations. Good agreement between third-order elasticity-based predictions and numerically calculated values confirms the validity of our theory.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Accelerating full-waveform inversion using source stacking: synthetic experiments at the global scale in a realistic 3-D earth model

SUMMARY The spectral element method is currently the method of choice for computing accurate synthetic seismic wavefields in realistic 3-D earth models at the global scale. However, it requires significantly more computational time, compared to normal mode-based approximate methods. Source stacking, whereby multiple earthquake sources are aligned on their origin time and simultaneously triggered, can reduce the computational costs by several orders of magnitude. We present the results of synthetic tests performed on a realistic radially anisotropic 3-D model, slightly modified from model SEMUCB-WM1 with three component synthetic waveform ‘data’ for a duration of 10 000 s, and filtered at periods longer than 60 s, for a set of 273 events and 515 stations. We consider two definitions of the misfit function, one based on the stacked records at individual stations and another based on station-pair cross-correlations of the stacked records. The inverse step is performed using a Gauss–Newton approach where the gradient and Hessian are computed using normal mode perturbation theory. We investigate the retrieval of radially anisotropic long wavelength structure in the upper mantle in the depth range 100–800 km, after fixing the crust and uppermost mantle structure constrained by fundamental mode Love and Rayleigh wave dispersion data. The results show good performance using both definitions of the misfit function, even in the presence of realistic noise, with degraded amplitudes of lateral variations in the anisotropic parameter ξ. Interestingly, we show that we can retrieve the long wavelength structure in the upper mantle, when considering one or the other of three portions of the cross-correlation time series, corresponding to where we expect the energy from surface wave overtone, fundamental mode or a mixture of the two to be dominant, respectively. We also considered the issue of missing data, by randomly removing a successively larger proportion of the available synthetic data. We replace the missing data by synthetics computed in the current 3-D model using normal mode perturbation theory. The inversion results degrade with the proportion of missing data, especially for ξ, and we find that a data availability of 45 per cent or more leads to acceptable results. We also present a strategy for grouping events and stations to minimize the number of missing data in each group. This leads to an increased number of computations but can be significantly more efficient than conventional single-event-at-a-time inversion. We apply the grouping strategy to a real picking scenario, and show promising resolution capability despite the use of fewer waveforms and uneven ray path distribution. Source stacking approach can be used to rapidly obtain a starting 3-D model for more conventional full-waveform inversion at higher resolution, and to investigate assumptions made in the inversion, such as trade-offs between isotropic, anisotropic or anelastic structure, different model parametrizations or how crustal structure is accounted for.

Geochemistry & Geophysics↗

STARTR: An Open-Source MARVEL model for the NRIC Virtual Test Bed [Poster]

The National Reactor Innovation Center (NRIC) seeks to improve the understanding of microreactor physics in industry and academia through the development of a Microreactor Applications Research Validation and Evaluation (MARVEL) reactor-based model, published on the Virtual Test Bed (VTB). To achieve this goal, the Sodium-cooled Thermal-spectrum Advanced Research Test Reactor (STARTR) model was built using publicly available MARVEL specifications where possible and approximations where applicable, and was optimized for fast runtimes for researchers to receive rapid simulation feedback. STARTR will fill a gap between stakeholder interest and available models, as the first Sodium-cooled Thermal Reactor (STR) hosted on the VTB with baseline performance sanctioned by INL. This project involved the definition of all materials used in the reactor, geometry and all reactor subcomponents, and assertion of tallies and simulation settings within OpenMC 0.13.3. This poster details a small subset of the overall reactor physics testing: the two-dimensional power peaking factors and the flux energy spectrum, as well as plots of the created geometry. Future work includes code-to-code verification between the OpenMC-based model and a separately designed MCNP 6.2-based model.

21 - SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLAN↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗