Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “BENCHMARKS”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Structural Benchmark Creep Testing for Microcast MarM-247 Advanced Stirling Convertor E2 Heater Head Test Article SN18

This report provides test methodology details and qualitative results for the first structural benchmark creep test of an Advanced Stirling Convertor (ASC) heater head of ASC-E2 design heritage. The test article was recovered from a flight-like Microcast MarM-247 heater head specimen previously used in helium permeability testing. The test article was utilized for benchmark creep test rig preparation, wall thickness and diametral laser scan hardware metrological developments, and induction heater custom coil experiments. In addition, a benchmark creep test was performed, terminated after one week when through-thickness cracks propagated at thermocouple weld locations. Following this, it was used to develop a unique temperature measurement methodology using contact thermocouples, thereby enabling future benchmark testing to be performed without the use of conventional welded thermocouples, proven problematic for the alloy. This report includes an overview of heater head structural benchmark creep testing, the origin of this particular test article, test configuration developments accomplished using the test article, creep predictions for its benchmark creep test, qualitative structural benchmark creep test results, and a short summary.

life (durability)↗

A Benchmark Example for Delamination Propagation Predictions Based on the Single Leg Bending Specimen Under Quasi-Static and Fatigue Loading

Benchmark examples based on Single Leg Bending (SLB) specimens with equal and unequal bending arm thicknesses were used to assess the performance of delamination prediction capabilities in finite element codes. First, the development of the quasi-static benchmark cases using the Virtual Crack Closure Technique (VCCT) is discussed in detail. Second, based on the quasi-static benchmark results, additional benchmark cases to assess delamination propagation under fatigue loading are created. Third, the application is demonstrated for the commercial finite element code Abaqus Standard 2018. The benchmark cases are compared to results obtained from VCCT-based, automated quasi-static propagation analysis. A comparison with results from automated fatigue propagation analysis was not performed at this point since the current version of Abaqus does not include this capability under variable mixed-mode conditions. In general, good agreement between the results obtained from the quasi-static propagation analysis and the benchmark results were achieved. Overall, the benchmarking procedure proved valuable for analysis verification.

Krueger, Ronald↗

A Benchmark Example for Delamination Propagation Predictions Based on the Single Leg Bending Specimen under Quasi-static and Fatigue Loading

Benchmark examples based on Single Leg Bending (SLB) specimens with equal and unequal bending arm thicknesses were used to assess the performance of delamination prediction capabilities in finite element codes. First, the development of the quasi-static benchmark cases using the Virtual Crack Closure Technique (VCCT) is discussed in detail. Second, based on the quasi-static benchmark results, additional benchmark cases to assess delamination propagation under fatigue loading are created. Third, the application is demonstrated for the commercial finite element code Abaqus Standard 2018. The benchmark cases are compared to results obtained from VCCT-based, automated quasi-static propagation analysis. A comparison with results from automated fatigue propagation analysis was not performed at this point since the current version of Abaqus does not include this capability under variable mixed-mode conditions. In general, good agreement between the results obtained from the quasi-static propagation analysis and the benchmark results were achieved. Overall, the benchmarking procedure proved valuable for analysis verification.

Ronald Krueger↗

Beyond Energy Efficiency: A clustering approach to embed demand flexibility into building energy benchmarking

The intermittency of carbon-free renewables and the demand changes associated with the widespread push for electrifying the transportation and building sectors provides an opportunity for buildings to go beyond energy efficiency and push towards providing demand flexibility to the electricity grid. The duality of energy efficiency and demand flexibility is necessary for success in a sustainable and reliable energy transition. Current building energy benchmarking models are limited in their ability to integrate concepts of demand flexibility and/or utilize granular smart meter data. Thus, current benchmarking methods are focused annual energy usage and fail to incorporate how the time of use of energy consumption impacts emissions in a quickly changing energy grid. Without a more comprehensive view of energy usage and associated real-time emissions, current benchmarking methods are unlikely to realize the full decarbonization potential of buildings. New emerging data streams provide an opportunity to develop a new generation of benchmarking energy models that embed dimensions of energy efficiency, grid interactivity, and demand flexibility into their analysis. In this paper, we propose a four-step method for embedding grid interactivity and demand flexibility into building benchmarking models that utilizes emerging building and time-series electricity data streams. We first engineer features to produce a mix-type dataset that encompasses many attributes of grid-interactive and efficient buildings, and then we apply K-medoids using Gower's Distance to produce peer-group clusters. We apply the method to a case study of 306 primary and secondary schools in southern California, USA. The results show that the method effectively clusters buildings by attributes of demand flexibility and energy efficiency. The clustering results reveal patterns in inefficient building operations and demand inflexibility at the building peer group level. In conclusion, the interpretation of clusters can serve as an integrated energy efficiency and demand flexibility benchmarking model and inform performance-specific policy targeting for buildings that go beyond traditional efficiency measures.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Discrete fracture network model benchmarks developed and applied in a DECOVALEX-2023 repository performance assessment study

This study presents newly developed benchmarks for modeling flow and transport within discrete fracture networks (DFNs) and useful methods for analyzing the results. The new benchmarks are designed to test modeling approaches for use in probabilistic performance assessment models of deep geologic repositories in fractured rock. The benchmarks simulate flow and transport through a 1 km 3 block of fractured rock. The first simulates migration of a short pulse of tracer through a simple network of four intersecting fractures. The second adds 1089 stochastically generated fractures. The third changes the pulse to a continuous point source. Evaluation of model performance relies on moment analysis and comparison of the results of different models. The expected nondimensional first moment of the conservative tracer for each benchmark is 1. The benchmarks were simulated by teams from Canada, Czechia, Germany, Korea, Sweden, Taiwan, and the United States as part of a DECOVALEX-2023 study (decovalex.org). The teams used various approaches, including explicit DFN modeling, DFN upscaling to an equivalent continuous porous medium (ECPM), and a combination of both methods. Transport mechanisms are modeled using either the advection-dispersion equation or particle tracking. Results demonstrate strong agreement among the models in breakthrough behavior up to the 75th percentile. Significant deviations in first moments and well-clustered outputs led to the identification of inaccuracies in several models. Such findings exemplify the benefit of exercising these benchmarks and using the presented methods to test DFN flow and transport models.

Benchmark↗

Depletion Benchmark of the AFIP-7 Experiment in the Advanced Test Reactor

Reactor physics depletion benchmarks for low-enriched uranium fuel are limited in number. In particular, there is very limited data for LEU benchmarks for U-10Mo (Uranium-10% Molybdenum) plate fuel developed for use in U.S. high-performance research reactors (USHPRR). USHPRR includes the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), University of Missouri Research Reactor (MURR), Massachusetts Institute of Technology Reactor (MITR), and National Bureau of Standards Reactor (NBSR) at the National Institute of Science and Technology. These reactors are fueled with high-enriched uranium dispersed fuel in a silicon/aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiated fueled specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This report documents the development of a benchmark model of the irradiation of the ATR Full -size plate In center flux trap Position 7 (AFIP-7) experiment.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Verifying MCNP Models of the TEX High 240 Plutonium Benchmark

Computational modeling programs are invaluable tools that allow us to understand systems, safely develop new processes, and make reliable predictions about future designs. However, the effectiveness of these codes is limited by the degree to which their parameters match the real world. In the field of nuclear engineering, cross section data is one of these vital parameters. Accurate cross section data on important fissile and fissionable isotopes promotes the design of safer and more efficient fabrication, transportation, storage, and stockpiling of nuclear fuel. Unfortunately, there are knowledge gaps in data on key isotopes. In 2011, a multinational meeting hosted by the US Department of Energy Nuclear Criticality Safety Program ranked the priority of certain cross section data needs. In response, Lawrence Livermore National Lab (LLNL) designed the Thermal and Epithermal eXperiment (TEX) series of benchmark experiments. Benchmark experiments are used to validate current cross section data. They validate data by comparing the results of an actual experiment to the predicted results from a computational model. The data a benchmark applies to depends on the isotope and energy range the experiment’s neutron multiplication factor ( k eff ) is most sensitive to. The development and testing of the TEX High 240 Plutonium Benchmark will help validate 240 Pu cross section data. The configuration and materials of this benchmark are designed to be most sensitive to 240 Pu's intermediate energy range (from 0.625 ev to 100 keV ). MCNP® models of the assembly have been developed by LLNL and the results have been written in the final design report. In order for the discrepancies between benchmark models and experiments to be attributed to cross section inaccuracies, the accuracy of the models needs to be verified. The goal of this project is to verify of the results of LLNL's modeling by creating a new set of MCNP models and comparing the results.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Depletion benchmark for a high-assay low-enriched uranium fuel experiment in the advanced test reactor

Reactor physics depletion benchmarks for high-assay low-enriched uranium (HALEU) fuel are limited in number. In particular, there is limited data for HALEU benchmarks for U-10Mo (uranium-10% molybdenum) plate fuel that is being developed for use in the United States’ high performance research reactors including the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), Massachusetts Institute of Technology Reactor (MITR), University of Missouri Research Reactor (MURR), National Bureau of Standards Reactor (NBSR). These six reactors currently operate with highly enriched uranium dispersed fuel in an aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiating fuel specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This paper documents the development of a benchmark model of the irradiation of the ATR Full-size plate In center flux trap Position 7 (AFIP-7) experiment using the depletion codes MC21 and Advanced Dimensional Depletion for Engineering of Reactors (ADDER).

Nielsen, Joseph W. [Idaho National Laboratory (INL↗

BuildingQA: A Benchmark for Natural Language Question Answering over Building Knowledge Graphs

Graph-based representations of building metadata using ontologies like Brick are vital for smart building applications, but querying them remains a challenge for practitioners. Knowledge Graph Question Answering (KGQA) systems, meant to retrieve answers from natural language questions, traditionally require large-scale training data, making them ill-suited for the specialized and data-scarce building domain. The advent of Large Language Models (LLMs) offers a paradigm shift, enabling zero-shot natural language querying without building/domain-specific training. Yet, there is no standardized benchmark for building-specific KGQA which can guide and validate research in this area. To address this gap, our work makes three primary contributions. First, we introduce the BuildingQA Benchmark Dataset, constructed through a multi-stage process of collecting practitioner data, augmenting it with LLMs for linguistic diversity, and curating a final set of 188 questions across 4 buildings. Second, we characterize the benchmark's complexity and ambiguity, introducing a novel method to quantify its "lexical gap" and providing a four-stage diagnostic framework for analyzing how systems fail. Third, we benchmark zero-shot LLM-powered KGQA systems to establish baseline performance and analyze their failure modes. Our evaluation reveals that top-performing systems achieve a maximum F1 score of only 0.38. This result does not indicate a failure of these powerful systems, but rather underscores the unique challenges posed by our benchmark. It demonstrates a critical performance gap, showing that current methods successful on general KGs struggle with the specific lexical and structural nuances of the building domain. BuildingQA1 thus provides the benchmark dataset and foundational analysis needed to drive the development of novel, domain-aware methods required to unlock the use of semantic data in buildings.

Mulayim, Ozan Baris↗

IER-517: Molybdenum Optimized Benchmark System Demonstrating Integral Correlations (MOBY DICK)

Nuclear criticality experiments are essential to the validation of nuclear data used in simulation software. The quality of nuclear data becomes paramount as simulation software becomes more relied upon for criticality safety studies and designs of nuclear systems. To improve the quality of nuclear data, experimenters can design critical experiments that are sensitive to isotope reaction pairs in materials of interest. The efforts conducted by the Organisation for Economic Co-operation and Development - Nuclear Energy Agency (OECD-NEA) Working Party on Nuclear Criticality Safety (WPNCS) Subgroup 8: Preservation of Expert Knowledge and Judgement Applied to Criticality Benchmarks (SG8) to categorize benchmarks according to their usefulness for nuclear data validation have been of great importance. Based on the OECD studies benchmark experiments included in the International Criticality Safety Benchmark Evaluation Project (ICSBEP) Handbook are concisely used by nuclear data evaluators, criticality safety engineers and others to validate nuclear data and simulation results. A lack of benchmarks sensitive to molybdenum in the (ICSBEP), particularly in the intermediate range, was noted by Los Alamos National Laboratory (LANL), the French Institut de Radioprotection et de Sûreté Nucléaire (IRSN), and Y-12 National Security Site prompting them to submit a joint integral experiment request to the Nuclear Criticality Safety Program (NCSP) in 2019. The request included both HEU and Plutonium systems in order to validate differential nuclear data focusing on the intermediate energy range but also includes thermal and fast configurations. This document represents the preliminary design work for a series of molybdenum integral experiments known as Molybdenum Optimized Benchmark System Demonstrating Integral Correlations (MOBY DICK).

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Machine Learning meets Algebraic Combinatorics: A Suite of Benchmark Datasets to Accelerate AI for Mathematics Research

The use of benchmark datasets has become an important engine of progress in machine learning (ML) over the past 15 years. Recently there has been growing interest in utilizing machine learning to drive advances in research-level mathematics. However, off-the-shelf solutions often fail to deliver the types of insights required by mathematicians. This suggests the need for new ML methods specifically designed with mathematics in mind. The question then is: what benchmarks should the community use to evaluate these? On the one hand, toy problems such as learning the multiplicative structure of small finite groups have become popular in the mechanistic interpretability community whose perspective on explainability aligns well with the needs of mathematicians. While toy datasets are a useful benchmark for initial work, they lack the scale, complexity, and sophistication of many of the principal objects of study in modern mathematics. To address this, we introduce a new collection of benchmark datasets, Algebraic Combinatorics Benchmarks (ACBench), representing either classic or open problems in algebraic combinatorics, a subfield of mathematics that studies discrete structures arising from abstract algebra. After describing the datasets, we discuss the challenges involved in constructing “good” mathematics benchmarks, describe baseline model performance, and discuss some of the insights these datasets can provide that may be of interest even to those who are not interested in mathematics research itself.

97 MATHEMATICS AND COMPUTING↗

Ranking and Classifying AI Benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece C. [Cornell U.]↗

A High-Fidelity Model of the Peach Bottom 2 Turbine-Trip Benchmark Using VERA

This work presents a high-fidelity simulation of the Peach Bottom turbine trip (PBTT) benchmark using the Virtual Environment for Reactor Applications (VERA), a multiphysics reactor modeling tool developed by the U.S. Department of Energy’s Consortium for Advanced Simulation of Light Water Reactors energy innovation hub. The PBTT benchmark, based on a 1977 transient event at the end of cycle 2 in a General Electric Type-4 boiling water reactor (BWR), is a critical test case for validating core physics models with thermal feedback during rapid reactivity events. VERA was employed to perform end-to-end, pin-resolved simulations from conditions at the beginning of cycle 1 through the turbine-trip transient, incorporating detailed neutron transport, fuel depletion, and subchannel thermal hydraulics. The simulation reproduced key benchmark observables with high accuracy: the peak power excursion occurred at 0.75 s, matching the scram time and closely aligning with the benchmark average of 0.742 s; the simulated maximum power spike was approximately 7600 MW, which is within 3% of the benchmark average of 7400 MW; and void-collapse dynamics were consistent with benchmark expectations. Reactivity predictions during cycles 1 and 2 remained within 1500 pcm and 400 pcm of criticality, respectively. These results confirm VERA’s ability to model complex coupled neutronic and thermal hydraulic behavior in a BWR turbine-trip transient, which will support its use in future studies of modeling dryout, fuel performance, and uncertainty quantification for transients of this type.

BWR↗

Machine characterization and benchmark performance prediction

From runs of standard benchmarks or benchmark suites, it is not possible to characterize the machine nor to predict the run time of other benchmarks which have not been run. A new approach to benchmarking and machine characterization is reported. The creation and use of a machine analyzer is described, which measures the performance of a given machine on FORTRAN source language constructs. The machine analyzer yields a set of parameters which characterize the machine and spotlight its strong and weak points. Also described is a program analyzer, which analyzes FORTRAN programs and determines the frequency of execution of each of the same set of source language operations. It is then shown that by combining a machine characterization and a program characterization, we are able to predict with good accuracy the run time of a given benchmark on a given machine. Characterizations are provided for the Cray-X-MP/48, Cyber 205, IBM 3090/200, Amdahl 5840, Convex C-1, VAX 8600, VAX 11/785, VAX 11/780, SUN 3/50, and IBM RT-PC/125, and for the following benchmark programs or suites: Los Alamos (BMK8A1), Baskett, Linpack, Livermore Loops, Madelbrot Set, NAS Kernels, Shell Sort, Smith, Whetstone and Sieve of Erathostenes.

Saavedra-Barrera, Rafael H.↗

Benchmark characterization

An abstract system of benchmark characteristics that makes it possible, in the beginning of the design stage, to design with benchmark performance in mind is presented. The benchmark characteristics for a set of commonly used benchmarks are then shown. The benchmark set used includes some benchmarks from the Systems Performance Evaluation Cooperative (SPEC). The SPEC programs are industry-standard applications that use specific inputs. Processor, memory-system, and operating-system characteristics are addressed.

Conte, Thomas M.↗

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.↗

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David↗

Research on computer systems benchmarking

This grant addresses the topic of research on computer systems benchmarking and is more generally concerned with performance issues in computer systems. This report reviews work in those areas during the period of NASA support under this grant. The bulk of the work performed concerned benchmarking and analysis of CPUs, compilers, caches, and benchmark programs. The first part of this work concerned the issue of benchmark performance prediction. A new approach to benchmarking and machine characterization was reported, using a machine characterizer that measures the performance of a given system in terms of a Fortran abstract machine. Another report focused on analyzing compiler performance. The performance impact of optimization in the context of our methodology for CPU performance characterization was based on the abstract machine model. Benchmark programs are analyzed in another paper. A machine-independent model of program execution was developed to characterize both machine performance and program execution. By merging these machine and program characterizations, execution time can be estimated for arbitrary machine/program combinations. The work was continued into the domain of parallel and vector machines, including the issue of caches in vector processors and multiprocessors. All of the afore-mentioned accomplishments are more specifically summarized in this report, as well as those smaller in magnitude supported by this grant.

Smith, Alan Jay↗