Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “BENCHMARKS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Summary of LANL Critical Benchmark Comparison Study and Revisions for Cases Involving HEU, LEU, MIX, and Pu

This report documents results obtained for revisions made to cases involving Highly Enriched Uranium (HEU), Intermediate Enriched Uranium (IEU), a mixture of Pu and Uranium (MIX), as well as Pu cases. A previous summary of revisions for HEU an Pu cases was reported and additional investigations into four cases originally presented therein uncovered further revisions which led to better agreement with other transport codes, those cases are updated in this report. The summary of all cases reported in Reference 2 is updated in this report. In addition, a previous summary of revisions for LEU and MIX was reported, a summary of those revisions in reproduced in this report for a comprehensive summary of changes to benchmarks beginning in fiscal year 2020 to current date. The report focuses on the changes made to LANL benchmarks modeled with MCNP6 using ENDF/B-VII.1 nuclear data that appeared to have discrepant results when compared with results of other codes. Feedback was used to pinpoint review of benchmark input files and to revise them when necessary. This report documents the results of review and revision of specific benchmarks highlighted as possibly discrepant in the comparison study. In addition, there is an effort tied to this work involving collaboration between LANL XCP and NCS Divisions in the development of a shared review/revision procedure and use of a new benchmark repository. LANL has a benchmark library of critical experiments from the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook modeled for use with MCNP. This collection is now over 1100 benchmarks, referred to as the Whisper-1.1 library because it is used with the sensitivity/uncertainty package, Whisper, which supports nuclear criticality safety validation and is released with MCNP6.2. The collection, originally created several decades ago, is a combination of smaller collections, which has been revised and expanded, by various groups at LANL over the years. The original authors are no longer at the laboratory and little formal documentation of review and revision of these benchmarks exists today. A branch of the benchmark collection was already the subject of a formal review undertaken by the LANL NCS Division and expanded to include XCP Division.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

An MLCommons Scientific Benchmarks Ontology

Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transformative applications of machine learning to critical scientific use-cases more fragmented and less clear in pathways to impact. This paper introduces an ontology for scientific benchmarking developed through a unified, community-driven effort that extends the MLCommons ecosystem to cover physics, chemistry, materials science, biology, climate science, and more. Building on prior initiatives such as XAI-BENCH, FastML Science Benchmarks, PDEBench, and the SciMLBench framework, our effort consolidates a large set of disparate benchmarks and frameworks into a single taxonomy of scientific, application, and system-level benchmarks. New benchmarks can be added through an open submission workflow coordinated by the MLCommons Science Working Group and evaluated against a six-category rating rubric that promotes and identifies high-quality benchmarks, enabling stakeholders to select benchmarks that meet their specific needs. The architecture is extensible, supporting future scientific and AI/ML motifs, and we discuss methods for identifying emerging computing patterns for unique scientific workloads. The MLCommons Science Benchmarks Ontology provides a standardized, scalable foundation for reproducible, cross-domain benchmarking in scientific machine learning. A companion webpage for this work has also been developed as the effort evolves: https://mlcommons-science.github.io/benchmark/

Hawks, Ben [Fermilab] (ORCID:0000000157000288)↗

Sensitivity and uncertainty of the IFR-1 BISON benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered for adoption by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This paper summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment.For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. Here, the resulting variations in benchmark predictions were normalized to calculate sensitivities. These sensitivities were then used to inform input selections for the UQ.The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of the benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because the confidence intervals associated with benchmark predictions were wide. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Informing Robust Functional Relationship Benchmarks: An Evaluation of the Temperature Sensitivity of Ecosystem Respiration Across the Arctic-Boreal Region

During land model development, simulated carbon dynamics are often benchmarked against observational data sets to evaluate model performance. Functional relationship benchmarks are the relationship between a driving variable (e.g., temperature) and a response variable (e.g., ecosystem respiration) and are a promising tool for assessing model performance by evaluating modeled sensitivities to changing environmental conditions. However, observed functional relationships can be influenced by choices made during data collection and throughout the benchmarking process, impacting the inferred skill of land models. To avoid misrepresenting a model's true performance, it is necessary to systematically evaluate best practices when constructing functional relationship benchmarks. We developed a set of guidelines for constructing functional relationship benchmarks, considering the choice of data set, number of daily observations, temporal extent, and temporal resolution across Alaska and Canada over a 20-year period from 2001 to 2020. The temperature sensitivity of ecosystem respiration from observations, evaluated through an apparent Q 10 , is highly variable both spatially and as a result of the data processing approach applied in the benchmark formation. When benchmarking 13 models from the Warming Permafrost Model Intercomparison Project (WrPMIP), the range in inferred model skill is substantially impacted by the choices applied in constructing functional relationship benchmarks. The inferred performance of a given model is most sensitive to the number of daily observations and temporal extent, followed by choice of benchmark data set and temporal averaging. Results from this analysis can guide the development of consistent and robust functional relationships for future model evaluation studies.

Poe, Jeralyn [Northern Arizona University, Flagsta↗

Benchmarking quantum computers

The rapid pace of development in quantum computing technology has sparked a proliferation of benchmarks to assess the performance of quantum computing hardware and software. However, not all benchmarks are of equal merit. Good ones empower scientists, engineers, programmers and users to understand the power of a computing system, whereas bad ones can misdirect research and inhibit progress. In this Perspective, we survey the science of quantum computer benchmarking. Here, we discuss the role of benchmarks and benchmarking and how good benchmarks can drive and measure progress towards the long-term goal of useful quantum computations, known as quantum utility. We explain how different kinds of benchmark quantify the performance of different parts of a quantum computer, discuss existing benchmarks, examine recent trends in benchmarking, and highlight important open research questions in this field.

Proctor, Timothy James [Sandia National Laboratori↗

Benchmarking for AI for Science

AI has been instrumental for recent developments in a number of domains of the sciences. With several hundred machine learning (ML) algorithms and models, and numerous AI-specific hardware platforms, a common quest for all scientists working on AI for Science is around the selection of machine learning algorithm(s) to solve their domain-specific scientific problems. A number of different initiatives around AI Benchmarking have been set up and have been useful in understanding the benefits of different ML algorithms for different tasks.However, with the majority of these AI Benchmarking initiatives focusing on the conventional notions of benchmarking, where the focus is purely runtime performance (such as training time or inference time), their suitability for benchmarking different ML algorithms for solving scientific problems has been viewed as a performance problem even though both are hardly the same. To make reasonable, explainable, and justifiable advancements in science using AI, it is critical to focus on the merits of these algorithms in handling different domain science problems. In other words, more emphasis must be given on Benchmarking for AI for Science than AI Benchmarking. The vision of the former is not only to assess the performance of ML algorithms, but also to assess, and understand the benefits and merits of different ML algorithms in handling scientific problems. Benchmarking for AI for Science, instead of pure performance focused AI Benchmarking, has several benefits: (i) it has the potential to offer advances in the sciences, much more rapidly than through pure performance-based AI methods, (ii) it will encourage the community to focus on developing better domain-specific AI techniques, particularly given the provision for being able to benchmark different techniques, and (iii) it will encourage hardware manufacturers to focus on developing science-specific hardware subsystems.

Thiyagalingam, Jeyan↗

Sensitivity and Uncertainty of the IFR-1 BISON Benchmark

The fuel performance code BISON is being used to evaluate metallic fuel for a new fast-spectrum test reactor called the Versatile Test Reactor, which is being considered by the US Department of Energy. To quantify the accuracy of BISON predictions, researchers at Oak Ridge National Laboratory have been developing a series of benchmarks based on legacy metallic fuel experiments. As part of this effort, the sensitivity of BISON predictions to variations in model inputs and the uncertainties associated with BISON predictions must be established. This report summarizes efforts to perform a comprehensive sensitivity analysis (SA) and uncertainty quantification (UQ) on a benchmark based on the IFR-1 experiment. For the SA, at least one input was chosen from every BISON model and physics module used in the benchmark. The inputs were varied individually in a series of BISON simulations. The resulting variations in benchmark predictions were normalized to calculate sensitivities. The strongest sensitivities were identified and used to inform input selections for the UQ. The UQ was performed using the Monte Carlo UQ method. A literature review was conducted to estimate uncertainty distributions for the selected inputs, and values were sampled randomly from each distribution in a series of BISON simulations. Variations in the benchmark predictions were used to estimate uncertainty distributions and confidence intervals. It was found that nearly 100% of benchmark predictions matched the corresponding legacy values within the confidence intervals. However, this is at least partially because of the wide confidence intervals associated with the benchmark predictions. The uncertainty contributions of assumptions in the benchmark, experimental uncertainties, and BISON models were quantified. Some analysis was performed to identify inputs that contributed to the uncertainties. Finally, recommendations are made for future benchmark development and future BISON development.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

FAIR Surrogate Benchmarks Supporting AI and Simulation Research (Final Report)

Computational Science is being revolutionized by integrating AI and simulation and, in particular, by deep learning surrogate models that can replace all or part of traditional large‐scale HPC computations. Such surrogates can achieve remarkable performance improvements, as much as several orders of magnitude, and save both compute time and energy. The Surrogate Benchmark Initiative (SBI) project creates a community repository and FAIR (Findable, Accessible, Interoperable, and Reusable) data ecosystem for HPC application surrogate benchmarks. The SBI team comes from Argonne National Laboratory (ANL), Indiana University (IU), Rutgers University, the University of Tennessee, Knoxville (UTK), and the University of Virginia (UVA). SBI repositories include data, code, and all relevant collateral artifacts that the science and engineering community need to use and reuse these data sets and surrogates. SBI repositories generate active research from both the participants in SBI and the broad community of AI and domain scientists. This project develops surrogates that use several different neural nets to learn and quickly infer the results of simulations and data systems and captures them as surrogate benchmarks with a rich set of metadata covering: Data; Model; Metrics specification; Machine specification; and Science, Speed, and Power Results. We research FAIR metadata for these benchmarks. We develop application surrogate examples as benchmarks across many fields (ANL, UTK, IU, UVA). We also study non-Surrogate benchmarks that have many common features and similar issues as regards FAIRness. We work with MLCommons (UVA, UTK), which is a major machine learning benchmarking activity where we get metadata ontologies, software, and benchmarks, Benchmarks have datasets, models, and metadata and they need a technical framework developed by UTK and Rutgers and deployed by UVA. We study features of Surrogates including performance, training set size, and uncertainty quantification (Rutgers, UVA and IU).

97 MATHEMATICS AND COMPUTING↗

Creating Benchmark Data for Artificial Intelligence and Machine Learning Space Biology Research

To identify an appropriate AI/ML approach for a specific problem, the best practice is to measure algorithm performance through the benchmarking process. A scientific benchmark consists of an AI-ready dataset and a reference implementation on a specific scientific question. The NASA Science Mission Directorate (SMD) has started the “Benchmark Initiative for AI/ML to create scientific benchmark datasets in three applications: 1) scientific benchmarking, which finds the best algorithm for a specific problem; 2) application benchmarking, which measures algorithm performance against a set of parameters; and 3) system benchmarking, which evaluates performance of hardware and software architecture. Currently, there are no standardized datasets available to benchmark AI/ML algorithms in the domain of space biology. In this work, we constructed two AI/ML-ready biological datasets from experiments in space-flown mice: cellular imaging and RNA-seq. First, radiation-exposed immune cells harbor DNA damage foci that can be fluorescently marked to visualize the amount of damage following exposure to ionizing radiation. However, such large datasets are difficult to analyze visually, due to imaging inconsistencies and human bias, and classical image processing approaches can fail on imaging artifacts. AI/ML are therefore exciting alternative, providing the speed of machines and the accuracy of humans. We have made this dataset available at https://registry.opendata.aws/bps_microscopy/. Second, high-throughput nucleic acid sequencing (DNA-seq, RNA-seq) has become widespread in biomedical research due to the growing availability and affordability of these assays. However, most sequencing datasets suffer from high dimensionality and low sample count. In this work, we used a generative adversarial network to synthesize a standardized, AI-ready, publicly available benchmark dataset for space biology RNA-seq data with sufficient space-flown and ground control mouse liver samples from NASA GeneLab. This dataset is available at https://registry.opendata.aws/bps_rnaseq/. These datasets are now fully open the Space Biology community to test their favorite AI/ML approaches.

James Casaletto↗

EUI Benchmarks for Net-Zero Energy Buildings in India

In our study we present EUI benchmarks for NZEBs for six building types across residential and non-residential typologies and for India's five climate zones. This approach is similar to the simulation-based benchmarks used by the B3 program in Minnesota, the Cal-Arch methodology in California, and the US Solar Decathlon approach, which combine simulations with actual building data. Of the six building types we explore one in detail with a range of operation scenarios, specifically narrowing down the EUI benchmarks for mixed-mode building operation and for variable temperature set-points as prescribed in the National Building Code of India.The contribution of this work is to provide rigorous end-use level EUI benchmarks for six building types, and to describe a method for simulation-based EUI benchmarks for mixed-mode operation with variable setpoints to highlight the difference between the standard approach used for the five building types and the On-Site Construction Worker Housing which additionally has the mixed-mode variable setpoint approach.On-Site Construction Worker Housings are typically poorly constructed temporary structures, without adequate thermal comfort, It is critical to provide adequate thermal comfort to protect people from the warming effects of climate change, and to discover super efficient and cost-effective ways to do so.The methodology used for the On-Site Construction Worker Housing (CWH) results in an 80% acceptability according to the India Model for Adaptive Comfort in the National Building Code of India. EUIs of all six buildings are 60% lower than the minimum compliance with India's Energy Conservation Building Codes, providing benchmarks for efficiency levels. The end-use level EUI benchmarks are now provided to over 1800 Solar Decathlon India participants so that they can compare the performance of their NZEB designs.In particular, the CWH results provide an insight into the importance of the mixed-mode operation with variable temperature set-points. The results from the simulation study show that for an NZEB target, the EUI with standard thermal comfort model and without mixed operation is 58.26 kWh/m2*year, while that with the variable set-points of the adaptive model with mixed mode operation is 24.23 kWh/m2*year. This is a 58% reduction in EUI. Given that many building types including residential, and non-residential operate in mixed mode, it is important to take this work further to develop mixed-mode operation NZEB benchmarks so that the carbon intensity of these buildings could be lower than the standard thermal comfort model approach.

benchmarking↗

Opportunities for enhancing MLCommons efforts while leveraging insights from educational MLCommons earthquake benchmarks efforts

MLCommons is an effort to develop and improve the artificial intelligence (AI) ecosystem through benchmarks, public data sets, and research. It consists of members from start-ups, leading companies, academics, and non-profits from around the world. The goal is to make machine learning better for everyone. In order to increase participation by others, educational institutions provide valuable opportunities for engagement. In this article, we identify numerous insights obtained from different viewpoints as part of efforts to utilize high-performance computing (HPC) big data systems in existing education while developing and conducting science benchmarks for earthquake prediction. As this activity was conducted across multiple educational efforts, we project if and how it is possible to make such efforts available on a wider scale. This includes the integration of sophisticated benchmarks into courses and research activities at universities, exposing the students and researchers to topics that are otherwise typically not sufficiently covered in current course curricula as we witnessed from our practical experience across multiple organizations. As such, we have outlined the many lessons we learned throughout these efforts, culminating in the need for benchmark carpentry for scientists using advanced computational resources. The article also presents the analysis of an earthquake prediction code benchmark while focusing on the accuracy of the results and not only on the runtime; notedly, this benchmark was created as a result of our lessons learned. Energy traces were produced throughout these benchmarks, which are vital to analyzing the power expenditure within HPC environments. Additionally, one of the insights is that in the short time of the project with limited student availability, the activity was only possible by utilizing a benchmark runtime pipeline while developing and using software to generate jobs from the permutation of hyperparameters automatically. It integrates a templated job management framework for executing tasks and experiments based on hyperparameters while leveraging hybrid compute resources available at different institutions. The software is part of a collection called cloudmesh with its newly developed components, cloudmesh-ee (experiment executor) and cloudmesh-cc (compute coordinator).

58 GEOSCIENCES↗

A Benchmark Example for Delamination Propagation Predictions Based on the Calibrated End-Loaded Split Specimen

A benchmark example based on the Calibrated End-Loaded Split (C-ELS) specimen is developed and used to assess the performance of recently developed delamination propagation capabilities in the Abaqus/Standard finite element code. The C-ELS specimen has the advantage of a longer region of stable delamination propagation compared to the existing mode II benchmark case. The new benchmark example may therefore provide a better assessment tool by enabling more stable crack growth in regions further away from the boundary conditions or load application. First, a benchmark result is created manually using two-dimensional finite element models of the C-ELS specimen with different delamination lengths. Second, the performance of the delamination propagation capabilities in the Abaqus/Standard finite element code are assessed by comparing the results to the benchmark case. Two examples with different starter delamination lengths are studied. A shorter starter length is chosen to create a scenario with unstable delamination propagation. A longer delamination causes stable delamination propagation. Detailed results from three-dimensional analyses with aligned and mis-aligned meshes and two levels of mesh refinement are provided for several permutations of numerical input parameters. In general, good agreement can be achieved between the results obtained from the quasi-static propagation analysis and the benchmark analysis. However, particular non-default settings are found to be most reliable, accurate, and numerically efficient. Numerical artifacts including anomalous unreleased nodes in the crack wake and zig-zag crack fronts occur, and further development of the Abaqus/Standard VCCT propagation may be required. Use of the benchmark case to assess the continuous improvements in one finite element code illustrates the value of establishing benchmark solutions.

delamination↗

Impacts of benchmarking choices on inferred model skill of the Arctic–Boreal terrestrial carbon cycle

Abstract Land surface models require continuous validation against observations to improve and reduce simulation uncertainty. However, inferred model performance can be heavily influenced by subjective choices made in the selection and application of observational data products. A key area often misrepresented by models is the Arctic–Boreal region, which is a potential tipping point region in Earth’s climate system due to large permafrost carbon stocks that are vulnerable to release with climate warming. We use the International Land Model Benchmarking (ILAMB) framework to evaluate how the model skill of TRENDY-v9 models varies based on the choice of observational-based benchmark and how benchmarks are applied in model evaluation. This analysis uses global datasets integrated into ILAMB and new, regionally-specific observational products from the Arctic–Boreal Vulnerability Experiment. Our results cover the overall time period of 1979–2019 and show that model scores can vary substantially depending on the data product applied, with higher model scores indicating better model performance against observations. The lowest model scores occur when benchmarked against regional, compared to global, datasets. We also evaluate observed and modeled functional relationships between ecosystem respiration and air temperature and between gross primary production and precipitation. Here, we find that the magnitude and shape of the responses are strongly impacted by the choice of observational dataset and the approach used to construct the functional relationship benchmark. These results suggest that model evaluation studies could conclude a false sense of model skill if only using a single benchmark data product or if not applying regional data products when performing a regional model analysis. Collectively, our findings highlight the influence of benchmarking choices on model evaluation and point to the need for benchmarking guidelines when assessing model skill.

Poe, Jeralyn (ORCID:0000000318495278)↗

IER 555: Godiva Benchmark Update CED-2 (Final Design Report)

The International Criticality Safety Benchmark Evaluation Project (ICSBEP) evaluation of the Godiva IV critical assembly, HEU-MET-FAST-086: GODIVA-IV DELAYED-CRITICAL EXPERIMENTS (HMF-086), was completed by Russ Mosteller. Five critical experiment configurations performed at the Los Alamos National Laboratory (LANL) Technical Area (TA)-18 were evaluated as acceptable benchmark cases. The five cases consist of four delayed critical configurations which differ in control rod positions and one prompt critical configuration. All cases calculated a lower $k_{eff}$ than measured by experiment. This data is referred to as the TA-18 Godiva IV benchmark in this report. In 2005, Godiva IV was disassembled for relocation to the Nevada Test Site (NTS), now Nevada National Security Site (NNSS), at the National Criticality Experiments Research Center (NCERC). Following the disassembly and subsequent reassembly and startup of Godiva IV at NCERC, additional information about the Godiva IV components was obtained. An errata note was added to the HMF-086 evaluation in the ICSBEP handbook to provide this new information until a revision to the benchmark evaluation could be performed. In addition to those corrections, there are differences between the Godiva IV assembly at TA-18 and the Godiva IV assembly at NCERC. These differences include both assembly-specific differences (differences in the safety block gap, differences in the control rod positions, a new NCERC Top Hat and contamination shield) as well as environmental differences, such as the size and shape of the experimental building where the assembly is located. An additional model with similar cases, referred to as the NCERC Godiva IV benchmark in this report, will be added to the revised HMF-086 to capture these additional differences. This will provide the best benchmark model of Godiva for use by those performing experiments at NCERC. The IER 555 CED-2 report documents the information that will be updated in the HMF-086 revision, both the corrections to the TA-18 Godiva IV benchmark and the subsequent changes to create a NCERC Godiva IV benchmark. It describes the measurements that will be performed for cases similar to those performed at TA-18. It also describes measurements that will be included in the evaluation as additional data.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Comparison of 250 MHz R10K Origin 2000 and 400 MHz Origin 2000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on Steger, a 250 MHz Origin 2000 system with R10K processors, currently installed at the NASA Ames National Advanced Supercomputing (NAS) facility. For comparison purposes, the tests were also run on Lomax, a 400 MHz Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to measure system performance. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versions used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.↗

Comparison of Origin 2000 and Origin 3000 Using NAS Parallel Benchmarks

This report describes results of benchmark tests on the Origin 3000 system currently being installed at the NASA Ames National Advanced Supercomputing facility. This machine will ultimately contain 1024 R14K processors. The first part of the system, installed in November, 2000 and named mendel, is an Origin 3000 with 128 R12K processors. For comparison purposes, the tests were also run on lomax, an Origin 2000 with R12K processors. The BT, LU, and SP application benchmarks in the NAS Parallel Benchmark Suite and the kernel benchmark FT were chosen to determine system performance and measure the impact of changes on the machine as it evolves. Having been written to measure performance on Computational Fluid Dynamics applications, these benchmarks are assumed appropriate to represent the NAS workload. Since the NAS runs both message passing (MPI) and shared-memory, compiler directive type codes, both MPI and OpenMP versions of the benchmarks were used. The MPI versions used were the latest official release of the NAS Parallel Benchmarks, version 2.3. The OpenMP versiqns used were PBN3b2, a beta version that is in the process of being released. NPB 2.3 and PBN 3b2 are technically different benchmarks, and NPB results are not directly comparable to PBN results.

Turney, Raymond D.↗

Development of a C-ELS Specimen-Based Numerical Benchmark for Mode II Delamination and Assessment of Two VCCT-Based Propagation Strategies

A finite element (FE) benchmark example inspired by the calibrated end-loaded split (C-ELS) specimen is developed and used to assess the performance of delamination propagation capabilities based on linear elastic fracture mechanics (LEFM). The C-ELS specimen has the advantage of a longer region of stable delamination propagation compared to the existing mode II benchmark case. The new benchmark example may therefore provide a better assessment tool by enabling more stable crack growth in regions further away from the boundary conditions or load application. First, a benchmark result is created manually using two-dimensional finite element models of the C-ELS specimen with different delamination lengths. Second, the performance of the virtual crack closure technique (VCCT) delamination propagation capabilities in the Abaqus/Standard®1 FE code and the recently developed Progressive Release eXplicit-VCCT (PRX-VCCT) method are assessed by comparing the results to the benchmark case. Two examples with different starter delamination lengths are studied. A shorter starter length is chosen to create a scenario with unstable delamination propagation. A longer delamination encourages stable delamination propagation. Detailed results from three-dimensional analyses with aligned and misaligned meshes and two levels of mesh refinement are provided. In general, good agreement can be achieved between the results obtained from the quasi-static propagation analysis and the benchmark analysis. Numerical artifacts including anomalous unreleased nodes in the crack wake and zig-zag crack fronts occur for propagation analyses using Abaqus/Standard VCCT. In comparison, continuous, smooth, delamination fronts are observed for PRX-VCCT. The use of the benchmark case to assess different VCCT-based propagation strategies illustrates the value of establishing benchmark cases.

Composite Materials↗

Results of a Geant4 benchmarking study for bio‐medical applications, performed with the G4‐Med system

Geant4, a Monte Carlo Simulation Toolkit extensively used in bio-medical physics, is in continuous evolution to include newest research findings to improve its accuracy and to respond to the evolving needs of a very diverse user community. In 2014, the G4-Med benchmarking system was born from the effort of the Geant4 Medical Simulation Benchmarking Group, to benchmark and monitor the evolution of Geant4 for medical physics applications. The G4-Med system was first described in our Medical Physics Special Report published in 2021. Results of the tests were reported for Geant4 10.5. Purpose In this work, we describe the evolution of the G4-Med benchmarking system. Methods The G4-Med benchmarking suite currently includes 23 tests, which benchmark Geant4 from the calculation of basic physical quantities to the simulation of more clinically relevant set-ups. New tests concern the benchmarking of Geant4-DNA physics and chemistry components for regression testing purposes, dosimetry for brachytherapy with a 125 I source, dosimetry for external x-ray and electron FLASH radiotherapy, experimental microdosimetry for proton therapy, and in vivo PET for carbon and oxygen beams. Regression testing has been performed between Geant4 10.5 and 11.1. Finally, a simple Geant4 simulation has been developed and used to compare Geant4 EM physics constructors and physics lists in terms of execution times. Results In summary, our EM tests show that the parameters of the multiple scattering in the Geant4 EM constructor G4EmStandardPhysics_option3 in Geant4 11.1, while improving the modeling of the electron backscattering in high atomic number targets, are not adequate for dosimetry for clinical x-ray and electron beams. Therefore, these parameters have been reverted back to those of Geant4 10.5 in Geant4 11.2.1. The x-ray radiotherapy test shows significant differences in the modeling of the bremsstrahlung process, especially between G4EmPenelopePhysics and the other constructors under study (G4EmLivermorePhysics, G4EmStandardPhysics_option3, and G4EmStandardPhysics_option4). These differences will be studied in an in-depth investigation within our Group. Improvement in Geant4 11.1 has been observed for the modeling of the proton and carbon ion Bragg peak with energies of clinical interest, thanks to the adoption of ICRU90 to calculate the low energy proton stopping powers in water and of the Linhard–Sorensen ion model, available in Geant4 since version 11.0. Nuclear fragmentation tests of interest for carbon ion therapy show differences between Geant4 10.5 and 11.1 in terms of fragment yields. In particular, a higher production of boron fragments is observed with Geant4 11.1, leading to a better agreement with reference data for this fragment. Conclusions Based on the overall results of our tests, we recommend to use G4EmStandardPhysics_option4 as EM constructor and QGSP_BIC_HP with G4EmStandardPhysics_option4, for hadrontherapy applications. The Geant4-DNA physics lists report differences in modeling electron interactions in water, however, the tests have a pure regression testing purpose so no recommendation can be formulated.

62 RADIOLOGY AND NUCLEAR MEDICINE↗