Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmark testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

A High-Fidelity Cyber-Physical Testbed-Based Benchmarking Dataset For Testing Operational Technology Specific Intrusion Detection Systems

Quality datasets serve a critical purpose in cyber security research. Data is needed to understand system behavior and develop security controls to protect critical systems. However, for critical infrastructure operational environments there is a lack of available datasets to study because of the high cost and specialized capabilities necessary to generate them. This paper documents the development of a dataset of high fidelity hardware in the loop laboratory simulated models of electric and natural gas distribution systems with real cyber attack test cases. A deep dive discussion for the experimental setup and controls for generating the data is provided along with observations from using the data in evaluating intrusion detection approaches.

Ashok, Aditya↗

Comparative analysis of energy deposition modes available in Serpent 2 within the framework of the supercritical water reactor - Fuel qualification test reactor physics benchmark

A joint European Canadian Chinese development of a supercritical water-cooled small modular reactor (SCW-SMR) technology is in progress since September 2020 in the framework of a Horizon 2020 project called ECC-SMART. As a main purpose of the project, proper estimates of energy deposition and its spatial distribution are prerequisites for the accurate analysis of safety related parameters of the SCW-SMR concept under development. A supercritical water reactor fuel computational benchmark model, provided by Canadian Nuclear Laboratories, was applied for detailed comparison of different energy deposition calculation options available in the Serpent 2 Monte Carlo code. The effect of energy deposition options on the normalization of the results as well as on the spatial distribution of the energy deposition are discussed. Consistent energy deposition calculation methods are presented between three Monte Carlo codes, viz., Serpent 2, MCNP6 and OpenMC. Although resource-intensive, the use of the coupled neutron-photon transport mode of Serpent 2 is recommended for accurate spatial and quantitative characterization of energy deposition in the SCW-SMR fuel assemblies, accounting for both neutron and photon heating of all the materials. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Design, Manufacture, and Testing of an Open-Source Benchmark Composite Hydrokinetic Turbine Blade: Preprint

In a trend toward clean energy alternatives, recent years have seen great strides in the marine energy space. Consequently, there is a pressing need for the design, development, and validation of novel energy harvesting technologies such as hydrokinetic devices, which capture kinetic energy from waves, tides, and currents. However, these devices span numerous concepts and designs that often lack solid benchmark research that can be freely referenced throughout their development. This work focuses on the design process of an open-source composite hydrokinetic turbine blade for a three-bladed marine turbine rotor assembly with a diameter of 2.5 m. The proposed blade consists of two structural composite skins that are bonded with an adhesive and filled with a foam core. This study also explores and contrasts the efficiency and resolution of low-fidelity rapid design methodologies and comprehensive high-fidelity approaches in the context of blade design, modeling, and analysis efforts, a key objective in this research. Blade hydrodynamic loads were modeled and applied to finite-element blade models to study deformations and potential failure. Ongoing and upcoming efforts will result in blade manufacture and structural testing at the National Renewable Energy Laboratory. In future work, multiple blades will be deployed at the Living Bridge site at the University of New Hampshire and will be compared to rigid aluminum blades of the same geometry, developed by Sandia National Laboratories. Ultimately, this research will lay foundational groundwork for researchers and manufacturers, establishing a baseline composite blade design that will serve as a benchmark in the development of future hydrokinetic turbine blades.

blade design↗

Initial Benchmarks of UV LEDs and Comparisons with White LEDs

The primary goal of this report is to benchmark the initial level of performance of a selection of commercial UV LEDs across all three bands (i.e., UV-A, UV-B, UV-C). To provide the initial performance benchmarks, a test matrix containing 13 different UV LED products was created in association with the LED Systems Reliability Consortium (LSRC). The products in this test matrix were all commercially available as of June 2021, and at least 22 samples of each product were tested. In addition, two common, commercial white LEDs were tested to provide a benchmark against blue-pumped white LEDs. Testing of the samples included electrical performance testing (e.g., current-voltage measurements) and photometric testing in a calibrated integrating sphere capable of measuring devices in the UV-A, UV-B, and UV-C bands. The electrical testing provided insights into the performance of the semiconductor layers in the LEDs, allowing parameters such as the threshold voltage (V th ) and serial resistance (R serial ) of each sample to be determined. The photometric testing provided insights into emission wavelengths, peak shapes, and radiant efficiencies for each sample. Combined, the information from these tests permits the overall device efficiencies to be compared and provides insights into the electrical and optical performance of the technology.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm

Abstract We present a benchmark test suite and an automated machine learning procedure for evaluating supervised machine learning (ML) models for predicting properties of inorganic bulk materials. The test suite, Matbench, is a set of 13 ML tasks that range in size from 312 to 132k samples and contain data from 10 density functional theory-derived and experimental sources. Tasks include predicting optical, thermal, electronic, thermodynamic, tensile, and elastic properties given a material’s composition and/or crystal structure. The reference algorithm, Automatminer, is a highly-extensible, fully automated ML pipeline for predicting materials properties from materials primitives (such as composition and crystal structure) without user intervention or hyperparameter tuning. We test Automatminer on the Matbench test suite and compare its predictive power with state-of-the-art crystal graph neural networks and a traditional descriptor-based Random Forest model. We find Automatminer achieves the best performance on 8 of 13 tasks in the benchmark. We also show our test suite is capable of exposing predictive advantages of each algorithm—namely, that crystal graph methods appear to outperform traditional machine learning methods given ~10 4 or greater data points. We encourage evaluating materials ML algorithms on the Matbench benchmark and comparing them against the latest version of Automatminer.

36 MATERIALS SCIENCE↗

OptiBench: An Optimization Benchmark Tool for Renewable Energy Problems

We propose a benchmark framework and visualization tool, OptiBench, for analyzing the performance of state-of-the-art optimization solvers across a variety of optimization problems in renewable energy research. Our framework is designed from the ground up in the Julia programming language and enables analysis at scale on high performance computing (HPC) systems. Our visualization tool allows effortless evaluation of optimization solver performance, robustness, and accuracy through intuitive plots, e.g., performance profiles, heat maps, and distribution plots. We have tested three benchmark suites relevant to the modeling of renewable energy systems, viz., CUTEst, PGLib-OPF, and WaterTAP water treatment optimization problems. We illustrate benchmarking of CUTEst using OptiBench on the National Laboratory of the Rockies's (NLR) HPC Kestrel. Our findings indicate that MA57 HSL linear solver demonstrated the best overall performance for an experimental IPOPT implementation. Our work is ongoing and we intend to add support for more optimization solvers and benchmark test suites in the future.

97 MATHEMATICS AND COMPUTING↗

OptiBench: An Optimization Benchmark Tool for Renewable Energy Problems

We propose a benchmark framework and visualization tool, OptiBench, for analyzing the performance of state-of-the-art optimization solvers across a variety of optimization problems in renewable energy research. Our framework is designed from the ground up in the Julia programming language and enables analysis at scale on high performance computing (HPC) systems. Our visualization tool allows effortless evaluation of optimization solver performance, robustness, and accuracy through intuitive plots, e.g., performance profiles, heat maps, and distribution plots. We have tested three benchmark suites relevant to the modeling of renewable energy systems, viz., CUTEst, PGLib-OPF, and WaterTAP water treatment optimization problems. We illustrate benchmarking of CUTEst using OptiBench on the National Renewable Energy Laboratory's (NREL) HPC Kestrel. Our findings indicate that MA57 HSL linear solver demonstrated the best overall performance for an experimental IPOPT implementation. Our work is ongoing and we intend to add support for more optimization solvers and benchmark test suites in the future.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Benchmark Specifications for TREAT Tests M5, M6, and M7

Detailed information describing three TREAT tests performed on metallic fuels has been collected and organized for use as a benchmark in evaluating the performance of metallic fuel models and codes. The tests, designated M5, M6, and M7, subjected EBR-II-irradiated fuel pins to a single type of overpower transient (at prototypical conditions with full coolant flow and an exponential power rise on an 8 second period) in flowing sodium loops. Six fuel pins were tested; five were ternary (U-19Pu-10Zr) alloy fuel clad in D9 with burnup ranging from 0.8 to 9.8 at. %, and one was binary alloy fuel (U-10Zr) clad in HT9 with 2.9 at. % burnup. The information gathered from the test records is expected to be useful for pre-transient characterization of the irradiated fuel pins as well as the transient analysis of the metallic fuel when subjected to severe accident conditions. This report presents benchmark specifications for the M5, M6, and M7 TREAT tests, identifies where primary sources of benchmark-related information can be found, and includes background information to help a user of the data understand their applicability and limitations.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Advances in benchmarking and round robin testing for PEM water electrolysis: Reference protocol and hardware

While the number of publications in the PEM water electrolysis community increases each year, no common ground concerning reference hardware (test cells and test bench) and testing protocols has been yet established. This would, however, be necessary for the comparability of experimental results. First attempts for such reference hardware and procedures have been made in the framework of the Task 30 Electrolysis within the Technology Collaboration Programme on Advanced Fuel Cells (AFC TCP) of the International Energy Agency (IEA). Since then, improvements of both the test hardware (test cell and components) as well as the measurement protocol were identified, and a revised methodology and key results based on a comprehensive measurement series have been obtained. A detailed protocol for testing commercial reference components with a reference laboratory test cell developed in-house by Fraunhofer ISE is presented. For evaluation of the protocol and the hardware, it was tested at three different institutions at the same time. Impedance spectroscopic and polarization data was acquired and analyzed. The obtained differences in performance were calculated to give the community an expectation window to compare own data to. Finally, the importance of a thorough temperature control and the conditioning phase are demonstrated.

08 HYDROGEN↗

ORNL Testing of Multiple Graphite Benchmarks [Slides]

Integral criticality safety and reactor physics benchmark experiments from the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook) and the International Handbook of Reactor Physics Experiments (IRPhE Handbook) are essential for nuclear data validation and testing

ENDF↗

OECD-NEA HTTF Benchmark Progress and Updates

General slides discussing High Temperature Test Facility benchmark progress, Prismatic HTGR deployment, modeling and simulation tools for validating systems, and verification and validation issues. Also discuss collaboration and conclusions of RELAP5-3D validation activities based on HTTF with trends in data and international impact to accelerate deployment of prismatic HTGR microreactors by providing an opportunity for designers to assess their codes against experimental data and solutions from other codes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Depletion Benchmark of the AFIP-7 Experiment in the Advanced Test Reactor

Reactor physics depletion benchmarks for low-enriched uranium fuel are limited in number. In particular, there is very limited data for LEU benchmarks for U-10Mo (Uranium-10% Molybdenum) plate fuel developed for use in U.S. high-performance research reactors (USHPRR). USHPRR includes the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), University of Missouri Research Reactor (MURR), Massachusetts Institute of Technology Reactor (MITR), and National Bureau of Standards Reactor (NBSR) at the National Institute of Science and Technology. These reactors are fueled with high-enriched uranium dispersed fuel in a silicon/aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiated fueled specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This report documents the development of a benchmark model of the irradiation of the ATR Full -size plate In center flux trap Position 7 (AFIP-7) experiment.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Benchmark Specification of Advanced Burner Test Reactor

As an effort to assess the Argonne Reactor Computation (ARC) suite of fast reactor analysis codes, a numerical benchmark problem was developed using the reference 250 MWt Advanced Burner Test Reactor (ABTR) metallic core fueled with beginning of equilibrium cycle compositions (Chang et al., 2006). Material thermal expansion at operating condition was modeled by adjusting the hexagonal pitch, axial meshes, and the fuel and structure material densities appropriately. Irradiation swelling of metal fuel was considered, and the bond sodium was displaced into the lower part of fission gas plenum.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Corrosion Control in Carbon Fiber Reinforced Plastic (CFRP) Composite-Aluminum Closure Panel Hem Joints

This project targeted the implementation of CFRP/Al closure panels as a drop-in replacement for all-aluminum closures used in automobile production today. The project focused on the development of corrosion fundamentals involving CFRP materials and the use of those materials in mixed-material joints. Numerous studies were completed to benchmark current materials and test mixed metal joints. Good correlation was found between outdoor exposures and laboratory cyclic corrosion testing. The insights from the corrosion testing fundamentals and benchmarking phases informed the development of low-cure adhesive and electrocoats and conductive primers for the CFRP. Adhesive and electrocoat materials were successfully developed meeting low-cure targets for bake (150C/10min). These materials were scaled-up and used to produce 5 prototype liftgates to further evaluate CFRP/Al joints in a full-scale part. The lower cure materials enabled mitigation of CTE mismatch and reduced the required bake temperature, resulting in reduced part distortion. Corrosion testing of the liftgates revealed unexpected increases in corrosion relative to the standard bake materials. While rigorous evaluation of the parts was not possible due to numerous complexities associated with prototype testing, the results provided significant insights into the material behavior of CFRP/Al joints. Key learnings from the project included a better understanding of the system complexity (interplay of substrate/primer conductivity, production of full-scale prototypes), advances in coating/adhesive design for low-cure applications involving CFRP, and advancement of corrosion test protocols for CFRP.

36 MATERIALS SCIENCE↗

IER 480: TEX-Pu Benchmark (PU-MET-THERM-004) to Test Polyethylene and Lucite Thermal Scattering Laws [Slides]

This presentation finds that PMT-004 has four new benchmark cases highly sensitive to PE TSL (2 cases) and PMMA TSL (2 Cases). PE cases were well predicted using MCNP6.2 and ENDF/B-VIII.0. PMMA cases overpredicted by approximately 0.6-0.7% in k eff at 20°C. Accepted into 2023 version of the ICSBEP Handbook. Temperature had a large impact on reactivity of the critical configuration. Implications for validation work for thermal cases- need to adjust TSL data to correct temperature as it can have hundreds of pcm effects for a few °C. Future thermal experiments should try and measure reactivity at multiple temperatures to aid in data testing and benchmark adjustment.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Open Source Evaluation Framework for Solar Forecasting

The Solar Forecast Arbiter is an open-source evaluation framework for solar forecasting. The framework enables evaluations of solar irradiance, solar power, and net-load forecasts that are impartial, repeatable and auditable. The Solar Forecast Arbiter addresses stakeholder-informed use cases including evaluation of forecast skill, comparisons to reference data sets, private forecast trials, and evaluation of probabilistic forecast skill. The framework includes a data validation toolkit, reference data sources, data privacy protocols, and benchmark forecast capabilities for intra-hour and day ahead forecast horizons. Reports and metrics communicate the relative merits of the test and benchmark forecasts. The reports are created from standardized templates and include graphics for qualitatively evaluating deterministic and probabilistic forecasts and standard metrics for quantitatively evaluating forecasts. The Solar Forecast Arbiter is designed to support all solar forecasting stakeholders, including Solar Forecasting 2 Topic Area 2 and Topic Area 3 teams.

14 SOLAR ENERGY↗