Engineering PapersSearch

SEARCH · Engineering Papers

Results for “benchmark”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Ranking and Classifying AI Benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece C. [Cornell U.]

A High-Fidelity Model of the Peach Bottom 2 Turbine-Trip Benchmark Using VERA

This work presents a high-fidelity simulation of the Peach Bottom turbine trip (PBTT) benchmark using the Virtual Environment for Reactor Applications (VERA), a multiphysics reactor modeling tool developed by the U.S. Department of Energy’s Consortium for Advanced Simulation of Light Water Reactors energy innovation hub. The PBTT benchmark, based on a 1977 transient event at the end of cycle 2 in a General Electric Type-4 boiling water reactor (BWR), is a critical test case for validating core physics models with thermal feedback during rapid reactivity events. VERA was employed to perform end-to-end, pin-resolved simulations from conditions at the beginning of cycle 1 through the turbine-trip transient, incorporating detailed neutron transport, fuel depletion, and subchannel thermal hydraulics. The simulation reproduced key benchmark observables with high accuracy: the peak power excursion occurred at 0.75 s, matching the scram time and closely aligning with the benchmark average of 0.742 s; the simulated maximum power spike was approximately 7600 MW, which is within 3% of the benchmark average of 7400 MW; and void-collapse dynamics were consistent with benchmark expectations. Reactivity predictions during cycles 1 and 2 remained within 1500 pcm and 400 pcm of criticality, respectively. These results confirm VERA’s ability to model complex coupled neutronic and thermal hydraulic behavior in a BWR turbine-trip transient, which will support its use in future studies of modeling dryout, fuel performance, and uncertainty quantification for transients of this type.

BWR

Machine characterization and benchmark performance prediction

From runs of standard benchmarks or benchmark suites, it is not possible to characterize the machine nor to predict the run time of other benchmarks which have not been run. A new approach to benchmarking and machine characterization is reported. The creation and use of a machine analyzer is described, which measures the performance of a given machine on FORTRAN source language constructs. The machine analyzer yields a set of parameters which characterize the machine and spotlight its strong and weak points. Also described is a program analyzer, which analyzes FORTRAN programs and determines the frequency of execution of each of the same set of source language operations. It is then shown that by combining a machine characterization and a program characterization, we are able to predict with good accuracy the run time of a given benchmark on a given machine. Characterizations are provided for the Cray-X-MP/48, Cyber 205, IBM 3090/200, Amdahl 5840, Convex C-1, VAX 8600, VAX 11/785, VAX 11/780, SUN 3/50, and IBM RT-PC/125, and for the following benchmark programs or suites: Los Alamos (BMK8A1), Baskett, Linpack, Livermore Loops, Madelbrot Set, NAS Kernels, Shell Sort, Smith, Whetstone and Sieve of Erathostenes.

Saavedra-Barrera, Rafael H.

Benchmark characterization

An abstract system of benchmark characteristics that makes it possible, in the beginning of the design stage, to design with benchmark performance in mind is presented. The benchmark characteristics for a set of commonly used benchmarks are then shown. The benchmark set used includes some benchmarks from the Systems Performance Evaluation Cooperative (SPEC). The SPEC programs are industry-standard applications that use specific inputs. Processor, memory-system, and operating-system characteristics are addressed.

Conte, Thomas M.

The NAS parallel benchmarks

A new set of benchmarks has been developed for the performance evaluation of highly parallel supercomputers in the framework of the NASA Ames Numerical Aerodynamic Simulation (NAS) Program. These consist of five 'parallel kernel' benchmarks and three 'simulated application' benchmarks. Together they mimic the computation and data movement characteristics of large-scale computational fluid dynamics applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification-all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, D. H.

The NAS parallel benchmarks

A new set of benchmarks was developed for the performance evaluation of highly parallel supercomputers. These benchmarks consist of a set of kernels, the 'Parallel Kernels,' and a simulated application benchmark. Together they mimic the computation and data movement characteristics of large scale computational fluid dynamics (CFD) applications. The principal distinguishing feature of these benchmarks is their 'pencil and paper' specification - all details of these benchmarks are specified only algorithmically. In this way many of the difficulties associated with conventional benchmarking approaches on highly parallel systems are avoided.

Bailey, David

Research on computer systems benchmarking

This grant addresses the topic of research on computer systems benchmarking and is more generally concerned with performance issues in computer systems. This report reviews work in those areas during the period of NASA support under this grant. The bulk of the work performed concerned benchmarking and analysis of CPUs, compilers, caches, and benchmark programs. The first part of this work concerned the issue of benchmark performance prediction. A new approach to benchmarking and machine characterization was reported, using a machine characterizer that measures the performance of a given system in terms of a Fortran abstract machine. Another report focused on analyzing compiler performance. The performance impact of optimization in the context of our methodology for CPU performance characterization was based on the abstract machine model. Benchmark programs are analyzed in another paper. A machine-independent model of program execution was developed to characterize both machine performance and program execution. By merging these machine and program characterizations, execution time can be estimated for arbitrary machine/program combinations. The work was continued into the domain of parallel and vector machines, including the issue of caches in vector processors and multiprocessors. All of the afore-mentioned accomplishments are more specifically summarized in this report, as well as those smaller in magnitude supported by this grant.

Smith, Alan Jay

Assessment of Static Delamination Propagation Capabilities in Commercial Finite Element Codes Using Benchmark Analysis

With capabilities for simulating delamination growth in composite materials becoming available, the need for benchmarking and assessing these capabilities is critical. In this study, benchmark analyses were performed to assess the delamination propagation simulation capabilities of the VCCT implementations in Marc TM and MD NastranTM. Benchmark delamination growth results for Double Cantilever Beam, Single Leg Bending and End Notched Flexure specimens were generated using a numerical approach. This numerical approach was developed previously, and involves comparing results from a series of analyses at different delamination lengths to a single analysis with automatic crack propagation. Specimens were analyzed with three-dimensional and two-dimensional models, and compared with previous analyses using Abaqus . The results demonstrated that the VCCT implementation in Marc TM and MD Nastran(TradeMark) was capable of accurately replicating the benchmark delamination growth results and that the use of the numerical benchmarks offers advantages over benchmarking using experimental and analytical results.

Orifici, Adrian C.

NASA Software Engineering Benchmarking Study

To identify best practices for the improvement of software engineering on projects, NASA's Offices of Chief Engineer (OCE) and Safety and Mission Assurance (OSMA) formed a team led by Heather Rarick and Sally Godfrey to conduct this benchmarking study. The primary goals of the study are to identify best practices that: Improve the management and technical development of software intensive systems; Have a track record of successful deployment by aerospace industries, universities [including research and development (R&D) laboratories], and defense services, as well as NASA's own component Centers; and Identify candidate solutions for NASA's software issues. Beginning in the late fall of 2010, focus topics were chosen and interview questions were developed, based on the NASA top software challenges. Between February 2011 and November 2011, the Benchmark Team interviewed a total of 18 organizations, consisting of five NASA Centers, five industry organizations, four defense services organizations, and four university or university R and D laboratory organizations. A software assurance representative also participated in each of the interviews to focus on assurance and software safety best practices. Interviewees provided a wealth of information on each topic area that included: software policy, software acquisition, software assurance, testing, training, maintaining rigor in small projects, metrics, and use of the Capability Maturity Model Integration (CMMI) framework, as well as a number of special topics that came up in the discussions. NASA's software engineering practices compared favorably with the external organizations in most benchmark areas, but in every topic, there were ways in which NASA could improve its practices. Compared to defense services organizations and some of the industry organizations, one of NASA's notable weaknesses involved communication with contractors regarding its policies and requirements for acquired software. One of NASA's strengths was its software assurance practices, which seemed to rate well in comparison to the other organizational groups and also seemed to include a larger scope of activities. An unexpected benefit of the software benchmarking study was the identification of many opportunities for collaboration in areas including metrics, training, sharing of CMMI experiences and resources such as instructors and CMMI Lead Appraisers, and even sharing of assets such as documented processes. A further unexpected benefit of the study was the feedback on NASA practices that was received from some of the organizations interviewed. From that feedback, other potential areas where NASA could improve were highlighted, such as accuracy of software cost estimation and budgetary practices. The detailed report contains discussion of the practices noted in each of the topic areas, as well as a summary of observations and recommendations from each of the topic areas. The resulting 24 recommendations from the topic areas were then consolidated to eliminate duplication and culled into a set of 14 suggested actionable recommendations. This final set of actionable recommendations, listed below, are items that can be implemented to improve NASA's software engineering practices and to help address many of the items that were listed in the NASA top software engineering issues. 1. Develop and implement standard contract language for software procurements. 2. Advance accurate and trusted software cost estimates for both procured and in-house software and improve the capture of actual cost data to facilitate further improvements. 3. Establish a consistent set of objectives and expectations, specifically types of metrics at the Agency level, so key trends and models can be identified and used to continuously improve software processes and each software development effort. 4. Maintain the CMMI Maturity Level requirement for critical NASA projects and use CMMI to measure organizations developing software for NASA. 5.onsolidate, collect and, if needed, develop common processes principles and other assets across the Agency in order to provide more consistency in software development and acquisition practices and to reduce the overall cost of maintaining or increasing current NASA CMMI maturity levels. 6. Provide additional support for small projects that includes: (a) guidance for appropriate tailoring of requirements for small projects, (b) availability of suitable tools, including support tool set-up and training, and (c) training for small project personnel, assurance personnel and technical authorities on the acceptable options for tailoring requirements and performing assurance on small projects. 7. Develop software training classes for the more experienced software engineers using on-line training, videos, or small separate modules of training that can be accommodated as needed throughout a project. 8. Create guidelines to structure non-classroom training opportunities such as mentoring, peer reviews, lessons learned sessions, and on-the-job training. 9. Develop a set of predictive software defect data and a process for assessing software testing metric data against it. 10. Assess Agency-wide licenses for commonly used software tools. 11. Fill the knowledge gap in common software engineering practices for new hires and co-ops.12. Work through the Science, Technology, Engineering and Mathematics (STEM) program with universities in strengthening education in the use of common software engineering practices and standards. 13. Follow up this benchmark study with a deeper look into what both internal and external organizations perceive as the scope of software assurance, the value they expect to obtain from it, and the shortcomings they experience in the current practice. 14. Continue interactions with external software engineering environment through collaborations, knowledge sharing, and benchmarking.

Rarick, Heather L.

A Benchmark Example for Delamination Growth Predictions Based on the Single Leg Bending Specimen Under Fatigue Loading

Analysis benchmarking is used to evaluate new algorithms for automated VCCT-based delamination growth analysis. First, existing benchmark cas s based on the Single Leg Bending (SLB) specimen for crack propagation prediction under quasi-static loading are summarized. Second, the development of new SLB-based benchmark cases to assess the static and fatigue growth prediction capabilities under mixed-mode I/II conditions is discussed in detail. Additionally, a scheme is proposed to interpolate between known fatigue delamination growth rates to obtain values for mixed-mode ratios for which data has not been defined in the input. Further, a comparison is presented, in which the benchmark cases are used to assess new analysis tools in ABAQUS/Standard FD03. These recently implemented tools yield results that are in good agreement with the benchmark examples. The ability to assess the implementation of new methods in one finite element code illustrates the value of establishing benchmark solutions.

Krueger, Ronald

Benchmark Problem Development for Testing Maturity of Intelligent Contingency Management Tools

This paper presents a process used to develop appropriate scenarios and metrics for evaluating the maturity of intelligent contingency management algorithms. A benchmark scenario is a reference point against which something can be measured, compared, or assessed. Creating an accurate benchmark requires considerable research and expertise. The scenario itself is an artificial representation of a real-world event, designed to achieve a set of learning objectives through experiential learning. Designing an effective benchmark simulation scenario requires careful planning, including identification of clear objectives; capability assessment of the algorithm/tool being evaluated; assessment of necessary levels of fidelity; development of a process flow map of events and event interactions; and identification of metrics that map back to objectives. Thus, a benchmark scenarios for contingency management might consist of one or several commonly used functions taken from real world applications, used for evaluation, characterization and performance measurement of a contingency management algorithm. Behavior of the contingency management algorithm under different environmental conditions should then be able to be predicted using a set of benchmark functions. The paper describes the resulting benchmark problem as an illustration of the application of this process.

Jon Holbrook

Status of the International Criticality Safety Benchmark Evaluation Project

The International Criticality Safety Benchmark Evaluation Project (ICSBEP) has continued its work generating evaluations of new and historical benchmark experiments since the last update to the nuclear criticality safety (NCS) community at the 12th International Conference on Nuclear Criticality Conference held in 2023. One additional version of the ICSBEP Handbook has been published since that update, and the Technical Review Group (TRG) held two in-person meetings to review and approve additional benchmarks. The 2022 and 2023 editions of the handbook were combined into one release (published in November 2024) and contained 13 new evaluations with 46 different configurations and two major revisions to existing evaluations. The 2024 version of the handbook, currently under publication review, will contain two new evaluations with 15 new configurations and one major revision to HEU-MET-FAST-028, the evaluation of Flattop with a uranium core. The ICSBEP TRG met again in person in April 2025 to review benchmarks for the 2025 ICSBEP Handbook and final comment resolution is currently ongoing. Many of the new benchmarks represent contemporaneous experiments that have been specifically optimized to provide validation cases relevant to the NCS community. One major area of focus for new critical experiments is to target the sparsely populated intermediate energy (or resonance) region. Another focus of many of the new benchmarks is to provide experiments sensitive to different materials, such as chlorine, hafnium, tantalum, titanium, molybdenum, chromium, and polymethyl methacrylate (PMMA, or Lucite). The ICSBEP continues to deliver high-quality, peer reviewed evaluations of integral experiments relevant to the nuclear data community.

HEU-MET-FAST-028

Intern-Artificial Intelligence Benchmarking

Benchmarks provide a standardized method for evaluating different AI models, enabling reproducibility and comparison between models, and facilitating scientific progress. As AI models continue to develop rapidly, incorporating new datasets, capabilities, and architectures becomes more complicated. Therefore, the current static benchmarks become increasingly irrelevant. The MLCommons team argues that to make AI benchmarks more relevant, it involves making the benchmarks themselves more dynamic, as well as technical innovations that make it easier for scientists and researchers at all levels to use and contribute to the benchmarks. The current progress in technical innovation is a software that allows for a detailed view of a collection of AI benchmarks to be output in various formats that are easily readable and accessible.

Krishnan, Anjay [Fermilab]

Classifying and rating AI benchmarks

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark’s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Shiraishi, Reece [Fermilab]

An Analytic Benchmark for Neutron Boltzmann Transport with Downscattering—Part IV: PFNS and $\bar{ν}$ Uncertainty Propagation

An analytic benchmark with continuous-energy cross sections was previously derived to validate criticality calculations. Here, to extend the utility of the analytic benchmark to verify the implementation of $\bar{ν}$ and prompt fission neutron spectrum (PFNS) uncertainty propagation methods, new simplified forms that are dependent on the incident (fission-causing) neutron energy, as well as the outgoing neutron energy for the PFNS, are introduced in this work. The analytical forms for the flux and adjoint flux are derived for the extended benchmark and used to determine the 𝑘-eigenvalue sensitivity to $\bar{ν}$ and PFNS. The 𝑘-eigenvalue uncertainty due to $\bar{ν}$ and PFNS is calculated for the analytic benchmark using simplified$\bar{ν}$ and PFNS representations based on the ENDF-B/VIII.0 239 Pu evaluation. Because of the low sensitivity of the analytic benchmark to the physical PFNS, a nonphysical high-sensitivity PFNS is also presented.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Spacecraft Observatory Benchmark Problem for Optical Disturbance Rejection

Since mid-1980s NASA has launched several first-generation interplanetary observatory missions. A salient feature of these missions has been the mounting of the optical payload on top of a large flexible space structure. This imposes challenging control-structure interaction problems (CSI) that have been studied over several decades. The Dec. 2021 launch of the James Webb Telescope marks the latest example of such a mission. One way to encourage development of novel control methods is to provide a benchmark problem for researchers. By the mid-2000s, there were a few well-known benchmark problems developed by the control community with a focus on CSI applications. However, in the past 15 years, the control challenges in optical observation missions have shifted from CSI to line-of-sight (LOS) precision pointing. This is because popular designs of the second-generation (Gen II) space earth missions have eliminated the large structure interface by mounting the optical payload directly on top of a rigid bus platform. The focus now is to achieve accurate pointing with maximum rejection of the surrounding disturbances. The NASA Safety and Engineering Center (NESC) has tasked T he Aerospace Corporation to develop a Benchmark Problem to study the challenging control issues that Gen II observatory missions encounter. This software tool is to be released as a public domain platform aimed at government, academic, and industry researchers. The focus of this benchmark problem is for researchers to design a set of innovative payload control laws with the maximum Optical Disturbance Rejection (ODR) capability to meet a set of pre-defined μrad-level of LOS jitter requirements. This paper describes the development of the benchmark problem. Aerospace/NASA will provide an integrated LOS plant model and disturbance/command profiles for users to run their design and simulation with as well as a user’s guide describing the required interfaces. The Aerospace Corp is also working to develop a hardware fast steering mirror testbed for users to demonstrate their innovative design in real hardware, if so desired.

Spacecraft Observatory Benchmark Problem

Extending quantum-mechanical benchmark accuracy to biological ligand-pocket interactions

Predicting the binding affinity of ligands to protein pockets is key in the drug design pipeline. The flexibility of ligand-pocket motifs arises from a range of attractive and repulsive electronic interactions during binding. Accurately accounting for all interactions requires robust quantum-mechanical (QM) benchmarks, which are scarce for ligand-pocket systems. Additionally, disagreement between “gold standard” Coupled Cluster (CC) and Quantum Monte Carlo (QMC) methods casts doubt on many benchmarks for larger non-covalent systems. We introduce the “QUantum Interacting Dimer” (QUID) benchmark framework containing 170 non-covalent (non-)equilibrium systems modeling chemically and structurally diverse ligand-pocket motifs. Symmetry-adapted perturbation theory shows that QUID broadly covers non-covalent binding motifs and energetic contributions. Robust binding energies are obtained using complementary CC and QMC methods, achieving agreement of 0.5 kcal/mol. The benchmark data analysis reveals that several dispersion-inclusive density functional approximations provide accurate energy predictions, though their atomic van der Waals forces differ in magnitude and orientation. Contrarily, semiempirical methods and empirical force fields require improvements in capturing non-covalent interactions (NCIs) for out-of-equilibrium geometries. The wide span of NCIs, highly accurate interaction energies, and analysis of molecular properties take QUID beyond the “gold standard” for QM benchmarks of ligand-protein systems.

Puleva, Mirela [University of Luxembourg, Luxembou

Benchmark Calculation for Turkey Point Unit 3 Cycles 1-3 Using the SCALE 6.3/Polaris–PARCS v3.4.2 Code Package

Benchmark calculations were performed for Turkey Point Unit 3 cycles 1–3 to validate the SCALE 6.3/Polaris–PARCS v3.4.2 code with the ENDF/B–VII.1 56–group library by comparing the simulated results with the measured data. The benchmark results will be used in evaluating the SCALE/Polaris–PARCS code package’s uncertainties for pressurized water reactor physics analysis. That future analysis will include key nuclear parameters such as reactivity, control bank worth, temperature coefficients, and pin and assembly power peaking factors. The present document details plant and fuel design specifications and input data for SCALE/Polaris, GenPMAXS, and PARCS. Additional details are provided with respect to the input and output files produced for the benchmark calculations. The benchmark results are summarized such that they can be used in evaluating uncertainties with other benchmark results for key nuclear parameters.

22 GENERAL STUDIES OF NUCLEAR REACTORS