Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Benchmark data”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

MS25: Materials Science-Focused Benchmark Data Set for Machine Learning Interatomic Potentials

Here, we present MS25, a benchmark data set for evaluating machine learning interatomic potentials (MLIPs) across diverse materials-relevant systems including MgO surfaces, liquid water, zeolites, a catalytic Pt surface reaction, high-entropy alloys (HEAs), and disordered Zr-oxides. Five MLIP architectures (MACE, NequIP, Allegro, MTP, and Torch-ANI) are trained and tested, focusing not only on traditional metrics (energies, forces, and stresses) but also explicitly validating derived physical observables such as lattice constants, volumes, and reaction barriers. We find that most models reach comparable accuracy on standard error metrics across the simple systems, although equivariant MLIPs offer 1.5–2× improvements over nonequivariant MLIPs in energy and force error for structurally complex or compositionally disordered environments such as HEAs and Zr–O systems. Our analysis highlights that low errors in energy and force predictions do not guarantee reliable observables, emphasizing the necessity of explicit validation. We demonstrate limitations in cross-framework transferability, as models trained on one zeolite framework (CHA) fail to reliably generalize to predictions of structurally distinct frameworks (e.g., MFI). Size-extensive tests show some dependence on system size for MgO, resulting from forced periodicity. The HEA and Zr–O data sets are identified as challenging tests for future benchmarks and MLIP model architecture developments as they show significant differentiation in error between MLIP architectures and are still relatively difficult at 1000 training images. Moving forward, we recommend that benchmarking efforts shift their focus from marginal accuracy improvements in energy and force errors toward identifying and understanding model failure modes, rigorously assessing transferability, and evaluating how their errors affect observable predictions. For researchers looking to choose an MLIP architecture, we suggest selecting equivariant MLIP architectures if the complexity of the system is a challenge. For simple materials problems, auxiliary features such as integration with molecular dynamics engines, trade-offs between computational data set generation cost vs MLIP inference speed, and framework integration may play a more important decision factor than small differences in error metrics that are unlikely to matter for production-level research.

chemical structure

DeepBench: A simulation package for physical benchmarking data

We introduce **DeepBench**, a python library that generates simple simulated image data from first principles, such as basic geometric shapes and astronomical objects. These data are highly valuable for developing (calibration, testing, and benchmarking) statistical and machine learning models because they make it possible to connect the final data product to physically interpretable inputs. This software includes tools to curate and store the datasets to maximize reproducibility.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Validation Data for Benchmarking Wire Arc Additive Manufacturing Process Simulations

Residual stresses cause geometric distortion and affect mechanical performance of additively manufactured structures, yet they are notoriously difficult to assess and predict. Distortion (warpage) can drive parts outside dimensional tolerance limits, leading to part rejection or rework. For parts that meet tolerance, locked-in residual stress fields can affect structural integrity during operation, particularly subcritical cracking by fatigue, creep, or corrosion. This work develops benchmark data for a common additive manufacturing process (Wire Arc Additive Manufacturing) that can be applied for calibration and validation of physical process models that predict residual stress fields. The work includes design of two different samples of differing geometry, detailed manufacturing records for a set of physical samples, and an extensive set of residual stress measurement data developed using two diverse techniques (the contour method and neutron diffraction). An initial application of the work is also reported, where a modeling challenge was issued to secure residual stress model predictions from two independent laboratories that were blind to residual stress measurement data. These initial blind residual stress predictions show significant discrepancies relative to the measurement data, illustrating the potential value of the underlying validation data. An open repository for this work, including the sample designs, manufacturing process records, and the residual stress data, is also provided for future application in non-blind validation efforts.

36 MATERIALS SCIENCE

DAISY Benchmark Performance Data

This repository contains the underlying data from benchmark experiments for Drifting Acoustic Instrumentation SYstems (DAISYs) in waves and currents described in "Performance of a Drifting Acoustic Instrumentation SYstem (DAISY) for Characterizing Radiated Noise from Marine Energy Converters" (https://link.springer.com/article/10.1007/s40722-024-00358-6). DAISYs consist of a surface expression connected to a hydrophone recording package by a tether. Both elements are instrumented to provide metadata (e.g., position, orientation, and depth). Information about how to build DAISYs is available at https://www.pmec.us/research-projects/daisy. The repository's primary content is three compressed archives (.zip format), each containing multiple MATLAB binary data files (.mat format). A table relating individual data files to figures in the paper, as well as the structure of each file, is included in the repository as a Word document (Data Description MHK-DR.docx). Most of the files contain time series information for a single DAISY deployment (file naming convention: [site]_DAISY_[Drift #].mat) consisting of processed hydrophone data and associated metadata. For a limited number of DAISY deployments, the hydrophone package was replaced with an acoustic Doppler velocimeter (file naming convention: [site]_DAISY_[Drift #]_ADV.mat). Data were collected over several years at three locations: (1) Sequim Bay at Pacific Northwest National Laboratory's Marine & Coastal Research Laboratory (MCRL) in Sequim, WA, the energetic tidal channel in Admiralty Inlet, WA (Admiralty Inlet), and the U.S. Navy's Wave Energy Test Site (WETS) in Kaneohe, HI. Brief descriptions of data files at each location follow. - MCRL - (1) Drift #4 and #16 contrast the performance of a DAISY and a reference hydrophone (icListen HF Reson), respectively, in the quiescent interior of Sequim Bay (September 2020). (2) Drift #152 and #153 are velocity measurements for a drifting acoustic Doppler velocimeter in in the tidally-energetic entrance channel inside a flow shield and exposed to the flow, respectively (January 2018). (3) Two non-standard files are also included: DAISY_data.mat corresponds to a subset of a DAISY drift over an Adaptable Monitoring Package (AMP) and AMP_data.mat corresponds to approximately co-temporal data for a stationary hydrophone on the AMP (February 2019). - Admiralty Inlet - (1) Drift #1-12 correspond to tests with flow shielded DAISYs, unshielded DAISYs, a reference hydrophone, and drifting acoustic Doppler velocimeter with 5, 10, and 15 m tether lengths between surface expression and hydrophone recording package (July 2022). (2) Drift #13-20 correspond to tests of flow shielded DAISYs with three different tether materials (rubber cord, nylon line, and faired nylon line) in lengths of 5, 10, and 15 m (July 2022). - WETS - (1) Drift #30-32 correspond to tests with a heave plate incorporated into the tether (standard configuration for wave sites), rubber cord only, and rubber cord, but with a flow shielded hydrophone (November 2022). (2) Drift #49-58 and Drift #65-68 correspond to measurements around mooring infrastructure at the 60 m berth where time-delay-of-arrival localization was demonstrated for different DAISY arrangements and hydrophone depths (November 2022).

16 TIDAL AND WAVE POWER

The MUSIC Critical Benchmark and Nuclear Data

The Measurement of Uranium Subcritical and Critical (MUSIC) experiment was a series of measurements of critical and subcritical configurations of bare highly enriched uranium. The goal was to compare measurement methods, analysis techniques, and simulation methods across regimes of criticality and to provide high-quality validation of 235 U nuclear data. A benchmark evaluation of the two critical configurations of the MUSIC experiments will soon be published in the release of the International Criticality Safety Benchmark Evaluation Project Handbook. The recent execution of the experiment aids in proper quantification of model simplifications and all uncertainties associated with the experiment. Historical benchmark evaluations are heavily relied on for uranium nuclear data validation despite the fact that the same level of documentation and comparable uncertainty analysis may not be present. The MUSIC evaluation is less likely to include “unknown unknowns” that could impede accurately modeling the system. Presented are both highly detailed and very simplified models, which represent the experimental configurations accurately, aiding the users of the benchmark for nuclear data or transport code validation. The sensitivities of k eff to nuclear data and nuclear data–related uncertainties are very similar between this experiment and previous bare uranium sphere experiments. In addition, the nuclear data uncertainties to any nuclides other than 235 U are small. For all these reasons, the recently evaluated MUSIC benchmark critical configurations could prove very useful for 235 U nuclear data validation. Currently, major libraries have good agreement with the experimental results, within 200 pcm for all nuclear data libraries, and within one standard deviation of the experimental result for most. Suggested nuclear data adjustments based on MUSIC and Lady Godiva are also presented, with posterior improvements to both the agreement in k eff and the uncertainty associated with the nuclear data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Validation of Pronghorn for Natural-Circulation Molten Salt Loops

This paper presents the development and validation of a high-fidelity thermal-hydraulic model of a molten salt natural circulation flow loop, designed for integration within a digital twin framework. The study evaluates the performance of Idaho National Laboratory’s Pronghorn against experimental data from Texas A&M University Molten Salt Flow Loop (MSFL) four Hitec-salt test benchmark data. Natural circulation of high-Prandtl-number fluids exhibits complex, counter-intuitive flow patterns that make pointwise thermocouple readings unreliable. Experimental work at TAMU’s MSFL provides benchmark data, including flow visualization at a test-section and centerline steady-state temperature measurements along the loop. Validation includes four single-phase natural circulation test cases with Hitec salt. Key metrics include flow profile agreement and steady-state temperature accuracy. Pronghorn results for two-dimensional single-phase agree qualitatively with the experimental flow profile. This paper illustrates the importance of Computational Fluid Dynamics (CFD) in elucidating the behavior of high-Prandtl-number thermal-hydraulics, along with how misleading centerline temperature measurements can be. Pronghorn reproduces the axial and radial stratification that makes single thermocouple readings unreliable. Future research will focus on reduced-order modeling techniques to enable rapid simulation suitable for real-time digital twin applications. The validated cases provide a basis for developing reduced-order surrogates aimed at real-time digital-twin applications.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Integral Nuclear Data and Benchmarking Needs for Fusion Energy Systems

Fusion energy systems are currently being designed and optimized using radiation transport codes. To deal with the unique environment inside a fusion-based system, many of these designs incorporate novel materials able to withstand the high radiation fields, ensure adequate cooling and thermal protection, and produce tritium. Validation plays a vital role in building trust in the predictive power of these models and computational methods. Validation of a code consists of modeling documented real-world experiments and comparing the code-predicted response to the measured response. Adequate validation requires measured responses from real-world experiments, also known as integral data, that mimic the system being designed, including materials, impinging radiation, and temperature, among other variables. The most trusted integral data are experimental responses that have been through a rigorous benchmarking process that develops a recommended computational model and evaluates all experimental uncertainties. Finally, there are a few research groups around the world that have been producing integral data for fusion applications, but a substantial investment is needed to address the unique validation needs of the fusion community.

Fusion

Multi-resolution Arctic Shrub Cover Dataset Derived from UAS and Airborne SfM and LiDAR (2013-2025)

We synthesized 177 unoccupied aerial system flights and 77 airborne flights across the Arctic and created a multi-resolution benchmark data of low-to-tall shrub fractional cover leveraging Structure-from-Motion and Light Detection and Ranging. The resulting dataset covered a total of 1899 km2 across Alaska, Western Canada, Sweden, and Siberian Arctic, including key sites from the Oro Arctic to the High Arctic. The dataset is organized into 6 primary data collection directories (“Abisko,” “AWI,” “ERE,” “Fairbanks,” “NGEE,” “Toolik”), each containing site and flight subdirectories. Flight directories include shrub cover rasters (*.tifs) at 1 m, 5 m, and 30 m resolution, the canopy height model at 1 m resolution (*.tifs), and a bounding box *.kml file. For the AWI, Abisko, NGEE, and Fairbanks collections, we also include the GCC raster at 1 m resolution (*.tif). Files are organized by Collection > Site > Flight Name > Data Files. Flight rasters are in the local UTM zone and the .kml files are in the geographic coordinate system EPSG 4326. We also include a .csv file that details the source datasets for every flight. The Next-Generation Ecosystem Experiments in the Arctic (NGEE Arctic) project is a research effort to reduce uncertainty in the Department of Energy’s Energy Exascale Earth System Model (E3SM) by developing a predictive understanding of Arctic tundra ecosystems underlain by permafrost and to quantify feedbacks from the Arctic tundra to the Earth system. NGEE Arctic is supported by the Department of Energy's Office of Biological and Environmental Research. Over Phases 1–3, observations made by the NGEE Arctic team across a gradient of permafrost landscapes in Arctic Alaska improved the representation of tundra processes in the land surface component of E3SM (the E3SM Land Model, ELM). Model improvements emphasized unique aspects of permafrost environments and explored reductions in model complexity while retaining predictive power. The Arctic-informed ELM developed by NGEE Arctic has been used to make novel predictions on processes ranging from permafrost thaw to soil biogeochemical cycling to Earth system feedbacks associated with the unique characteristics of tundra plants. In Phase 4, the NGEE Arctic team is evaluating our new predictive understanding under novel conditions across the Arctic domain. In collaboration with partners at long-term pan-Arctic research sites we are examining whether an Arctic-informed ELM can faithfully simulate interactions among surface and subsurface processes at site, regional, and pan-Arctic scales. In turn, we are using variety of tools to dynamically extend and evaluate ELM inference, with an emphasis on data synthesis and pan-Arctic model evaluation, reintegration of code with an evolving E3SM, scaling across heterogeneous Arctic landscapes, and the appropriate representation of the impacts of increasingly frequent Arctic disturbances.

canopy height model

Sum-of-Fractions Method

Sum-of-fractions is a method intended to make sure a subcritical margin for aqueous solutions and slurries of fissionable isotopes exists. The method indicates that a system is subcritical if the sum of the ratios of the mass of each isotope (in a mixture) to its individual minimum subcritical mass limit is less than or equal to one. Historically, the basis of the sum-of-fractions has been derived from allowances given in the American National Standards Institute (ANSI)/ American Nuclear Society (ANS)-8.15-1981. However, the allowance was removed in ANSI/ANS-8.15-2014 due to a lack of technical basis. A methodology was developed to assess the validity of using the sum-of-fractions for water- or polyethylene-moderated systems for the following nuclides: 232U, 233U, 234U, 235U, 237Np, 236Pu, 238Pu, 239Pu, 240Pu, 241Pu, 242Pu, 241Am, 242mAm, 243Am, 242Cm, 243Cm, 244Cm, 245Cm, 246Cm, 247Cm, 249Cf, and 251Cf. The methodology uses available benchmark data for mixtures of 233U, 235U, and 239Pu to establish the calculational margin, and a mass limit reduction to establish the margin of subcriticality. Water- or polyethylene-moderated and -reflected mixtures containing the nuclides are evaluated with the code system, SCALE 6.2.4. Including the calculational margin, subcritical mass limits for each nuclide were computed for optimally water- or polyethylene-moderated and fully reflected systems. These masses were used to create nuclide mixtures in which the sum of the mass to subcritical mass limit ratios is one. The various nuclide mixtures were modeled over a range of moderation and demonstrate the keff does not exceed the calculational margin. For additional assurance of subcriticality, a significant mass reduction is applied to each computed minimum critical mass of the nuclides without adequate benchmark data consistent with the method in ANSI/ANS-8.15-2014.

criticality safety, Actinide

High-Fidelity, Large-Scale, Realistic Dataset Development

The final report summarizes the work performed for supporting the ARPA-E Grid Optimization Competition (Challenge 2 and Challenge 3) within the stated period. Challenge 2 For the challenge period, the main responsibility of the team is to investigate, gen- erate, and deliver parts of the data sets for the competition, based on the competition model for Challenge 2, existing data sets from Challenge 1, and data source supplied by other data set teams. Challenge 3 For the challenge period, the main responsibility of the team is to propose, create, deliver, and maintain the data format during the competition period. The data format will specify how the benchmark data will be represented and communicated to competitors. It will also specify how competitors should report back the solutions. The data format will be closely aligned with the problem formulation (maintained by the formulation team) and the solution validation process (maintained by the validation team). Our team is also responsible in investigating, generating, and delivering parts of the data sets for the competition. The data sets will be created based on the competition model for Challenge 3, existing data sets from Challenge 1 and Challenge 2, and data source supplied by other data set teams.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI

Toward Verification of RANS Simulations of the T-Tube Modular Divertor Using Large Eddy Simulations of Impinging Turbulent Plane Jets

Turbulent impinging jets have been proposed to cool high heat flux plasma-facing components such as the solid tungsten target plates of the divertor in long-pulse magnetic fusion energy reactors. In particular, the T-tube modular divertor, originally developed by the ARIES Team, consists of two concentric cylindrical tubes where helium flows through a slot in the inner tube, forming an approximately planar jet that impinges upon and cools the inner surface of the pressure boundary (namely, the outer tube) and the ~15-cm 2 plasma-facing W target. The objective of this work is to demonstrate that large eddy simulations (LESs) accurately simulate the thermal transport in canonical flows that comprise the cooling flow in the T-tube, as well as validate temperatures from LES with experimental measurements in a simplified T-tube geometry. Wall‑resolved LESs, validated by experimental data and verified by direct numerical simulations (DNSs), provide benchmark data for two canonical flows in the T‑tube, namely, planar impinging and wall jets, for Reynolds numbers Re B = 4 × 10 3 to 2 × 10 4 . Our LES results are within 4% to 12% root-mean-square error (RMSE) of surface Nusselt number distributions (Nu) from experiments and DNSs. The validated LES results are then used as the ground truth to evaluate four Reynolds‑averaged Navier-Stokes (RANS) turbulence closures, namely, the k‑ω SST, realizable k‑ε, GEKO, and γ‑SST models. The k‑ω SST model has the best overall performance in terms of heat transfer, giving surface Nu within 12% RMSE of the LES results for high‑ReB impinging jets and reduced overprediction in the wall‑jet region. The GEKO model with default constants has the next best performance, providing slightly better Nu predictions for low ReB impinging jets (versus k-ω SST) but worse overall performance over the full range of ReB studied here. The realizable k‑ε turbulence model significantly overestimates turbulence near the stagnation point, while the γ‑SST model suppresses near‑wall production, biasing the simulations toward simulating laminar surface heat transfer. Simulations of the simplified T‑tube show that LES and RANS simulations with the k‑ω SST model give nearly identical average heat transfer coefficients (HTCs) over the impingement surface. The realizable k‑ε model predicts significantly lower wall temperatures due to overestimation of HTC in the outlet flow.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Acoustic sensing and autoencoder approach for abnormal gas detection in a spent nuclear fuel canister mock-up

Currently, spent nuclear fuel (SNF) from commercial nuclear power plants is stored in stainless-steel canisters for interim dry storage. To provide an inert environment, these canisters are backfilled with helium after vacuum drying. However, the helium environment may be contaminated during extended storage because of the material degradation. For example, the heavier fission gas xenon may be released from the fuel rods into the canister cavity should the fuel cladding be breached. Other gases such as air and water vapor may also be present as a result of leakage caused by chloride-induced stress corrosion cracking on the canister walls or by insufficient vacuum drying. Therefore, monitoring the gas composition can provide critical information about the health of SNF canisters. In this study, noninvasive testing was conducted on a 2/3-scaled SNF canister mock-up using acoustic sensing. Ultrasonic transducers were placed on the exterior surface of the canister to probe the gas composition. A dataset was collected by sealing the canister mock-up and introducing up to 1.53% argon or 1.29% air into the helium background gas. Three methods were used to detect changes in the gas composition: the time-of-flight (TOF) method, the differential method, and the autoencoder method. Results showed that the TOF method had sufficient resolution to detect abnormal gas concentrations of less than 1.0%. The differential method demonstrated a periodic in-phase and out-of-phase behavior between the benchmark (i.e., pure helium) and abnormal (i.e., with argon or air) state signals. The variational autoencoder (VAE) and the Wasserstein autoencoder (WAE) were trained on the benchmark data and were applied directly to the abnormal state data. It was found that both the unsupervised VAE and the WAE were able to distinguish the benchmark and abnormal states of the canister mock-up based on the reconstruction error.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS

ENDF/B-VIII.1: Updated Nuclear Reaction Data Library for Science and Applications

The ENDF/B-VIII.1 library is the newest recommended evaluated nuclear data file by the Cross Section Evaluation Working Group (CSEWG) for use in nuclear science and technology applications, and incorporates advances made in the six years since the release of ENDF/B-VIII.0. Among key advances made are that the 239 Pu file was reevaluated by a joint international effort and that updated 16,18 O, 19 F, 28–30 Si, 50–54 Cr, 55 Mn, 54,56,57 Fe, 63,65 Cu, 139 La, 233,235,238 U, and 240,241 Pu neutron nuclear data from the IAEA coordinated INDEN collaboration were adopted. Over 60 neutron dosimetry cross sections were adopted from the IAEA's IRDFF-II library. In addition, the new library includes significant changes for 3 He, 6 Li, 9 Be, 51 V, 88 Sr, 103 Rh, 140,142 Ce, Dy, 181 Ta, Pt, 206–208 Pb, and 234,236 U neutron data, and new nuclear data for the photonuclear, charged-particle and atomic sublibraries. Numerous thermal neutron scattering kernels were reevaluated or provided for the very first time. On the covariance side, work was undertaken to introduce better uncertainty quantification standards and testing for nuclear data covariances. The significant effort to reevaluate important nuclides has reduced bias in the simulations of many integral experiments with particular progress noted for fluorine, copper, and stainless steel containing benchmarks. Data issues hindered the successful deployment of the previous ENDF/B-VIII.0 for commercial nuclear power applications in high burnup situations. These issues were addressed by improving the 238 U and 239,240,241 Pu evaluated data in the resonance region. The new library performance as a function of burnup is similar to the reference ENDF/B-VII.1 library.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Uncertainty Quantification in GADRAS Inverse Modeling

The Gamma Detector Response and Analysis Software (GADRAS) package includes an inverse modeling tool that is helpful in identifying characteristics of unknown radioactive materials. Traditionally, uncertainties in this analysis were derived solely from measurement data quality and the fit of synthetic spectra. This paper aims to rigorously quantify additional sources of uncertainty, focusing on uncertainties arising from measurements being analyzed, Detector Response Function (DRF) characterization, and DRF extrapolation. Applying these findings to the BeRPBall benchmark data set, we demonstrated the impact of these uncertainties on plutonium and polyethylene estimates. The results underscore the importance of incorporating diverse uncertainty sources to enhance the accuracy and reliability of GADRAS’s inverse modeling capabilities.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Airborne imaging spectroscopy surveys of Arctic and boreal Alaska and northwestern Canada 2017–2023

Since 2015, NASA’s Arctic Boreal Vulnerability Experiment (ABoVE) has investigated how climate change impacts the vulnerability and/or resilience of the permafrost-affected ecosystems of Alaska and northwestern Canada. ABoVE conducted extensive surveys with the Next Generation Airborne Visible/Infrared Imaging Spectrometer (AVIRIS-NG) during 2017, 2018, 2019, and 2022 and with AVIRIS-3 in 2023 to characterize tundra, taiga, peatlands, and wetlands in unprecedented detail. The ABoVE AVIRIS dataset comprises ~1700 individual flight lines covering ~120,000 km 2 with nominal 5 m × 5 m spatial resolution. Data include individual transects to capture important gradients like the tundra-taiga ecotone and maps of up to 10,000 km 2 for key study areas like the Mackenzie Delta. The ABoVE AVIRIS surveys enable diverse ecosystem science, provide crucial benchmark data for validating retrievals from the PACE, PRISMA, and EnMAP satellite sensors and help prepare for the SBG and CHIME missions. This paper guides interested researchers to fully explore the ABoVE AVIRIS spectral imagery and complements our guide to the ABoVE airborne synthetic aperture radar surveys.

Miller, Charles E. [California Institute of Techno

Empirically-calibrated H100 node power models for accurate AI training energy estimation

Accurately quantifying the energy use of artificial intelligence (AI) training is critical for infrastructure planning, carbon accounting, and sustainable data center operation, but few studies have directly measured the power consumption of production workloads on contemporary hardware. By combining empirical measurements from Brookhaven National Laboratory during AI training on 8-graphics-processing-unit H100 systems with open-source benchmarking data, we develop statistical models relating computational intensity to node-level power consumption. We measure the gap between manufacturer-rated thermal design power (TDP) and actual power demand during AI training. Our analysis reveals that even computationally intensive workloads operate at only 76% of the 10.2 kW TDP rating. Our architecture-specific model, calibrated to floating-point operations, predicts energy consumption with 11.4% mean absolute percentage error, significantly outperforming TDP-based approaches (27%–37% error). We identified distinct power signatures between transformer and convolutional neural network architectures, with transformers showing characteristic fluctuations that may impact grid stability. These results provide a measurement-grounded basis for improving AI training energy estimates, enabling more reliable infrastructure sizing, cost projections, and environmental impact assessments.

Newkirk, Alex C

Multi‐Model Ensembles in Ecosystem Modeling: Challenges and Best Practices for Decision‐Making

Ecosystem models are increasingly central to the decision-making for environmental policy, conservation planning, and climate-related investments. Yet, the growing reliance on Multi-Model Ensembles (MMEs) of ecosystem models by practitioners and policymakers, sometimes under tight timelines and imperfect information, has frequently outpaced the scientific rigor required to ensure ensemble reliability. Here, MMEs refer to approaches that combine targeted predictions from multiple models with the expectation of improving robustness and quantifying predictive uncertainty. Poorly designed MMEs may create a false sense of confidence and lead to suboptimal policy and market decisions. This perspective argues that robust decision-making-relevant MMEs must be grounded on two pillars: (1) rigorous Model Intercomparison Projects (MIPs), which identify inter-model agreement and disagreement, characterize model uncertainties, and evaluate robustness with observationally based benchmarks—MIPs' diagnostic evaluation is so critical that it must be needed to drive MME's decision in model selection and weighting, especially when only a limited number of models available; and (2) co-design by both stakeholders and scientists to ensure that scenarios, metrics and uncertainty requirements provide decision-relevant information. Building upon the past success and lessons from the existing MIPs-MMEs efforts (e.g., climate/Earth system/crop), we derived the theoretical basis for MMEs, addressed their specific challenges in ecosystem modeling, and highlighted proper consideration of model numbers and diversity, risk of model inter-dependence, effective calibration of model parameters, possible overdue of some ecosystem model development, critical roles of open benchmark data across a wide range of conditions, and suggested use of Artificial Intelligence to support MIPs-MMEs. We highlighted the under-recognized opportunity for MIPs and MMEs to drive scientific progress and innovation through identifying better performing models, systematic benchmarking, feedback loops, and targeted model improvement. By following actionable best practice guidelines, MMEs can evolve from ad hoc aggregation of models into a trusted backbone of environmental policy and decision-making.

ecosystem modeling

ENDF/B-VIII.1

The ENDF/B-VIII.1 release is the newest evaluated nuclear data library produced, distributed, and recommended by CSEWG for use in nuclear science and technology applications. Among the many key advances, relative to the previous version ENDF/B-VIII.0, are: re-evaluation of 239Pu file by a joint international effort; updated 16,18O, 19F, 28-30Si, 50-54Cr, 55Mn, 54,56,57Fe, 63,65Cu, 139La, 233,235,238U, and 240,241Pu neutron nuclear data by the IAEA-coordinated INDEN collaboration; significant changes for 3He, 6Li, 9Be, 51V, 88Sr, 103Rh, 140,142Ce, Dy, 181Ta, Pt, 206-208Pb, and 234,236U neutron data; new nuclear data for the photo-nuclear, being 196 adopted from the IAEA2019 Photonuclear Data Library and one new file from JENDL-5; and new evaluations for the charged-particle and atomic sublibraries. Numerous thermal neutron scattering kernels were re-evaluated or provided for the very first time. Additionally, new covariance testing was implemented. ENDF/B-VIII.1 reduced bias in the simulations of many integral experiments with particular progress noted for fluorine, copper and stainless steel containing benchmarks. Data issues which had hindered the deployment of ENDF/B-VIII.0 for commercial nuclear power applications in high burn-up situations, were addressed. ENDF/B-VIII.1 data are distributed in both ENDF-6 and GNDS formats.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS