Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Benchmark testing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Benchmark Problem Development for Testing Maturity of Intelligent Contingency Management Tools

This paper presents a process used to develop appropriate scenarios and metrics for evaluating the maturity of intelligent contingency management algorithms. A benchmark scenario is a reference point against which something can be measured, compared, or assessed. Creating an accurate benchmark requires considerable research and expertise. The scenario itself is an artificial representation of a real-world event, designed to achieve a set of learning objectives through experiential learning. Designing an effective benchmark simulation scenario requires careful planning, including identification of clear objectives; capability assessment of the algorithm/tool being evaluated; assessment of necessary levels of fidelity; development of a process flow map of events and event interactions; and identification of metrics that map back to objectives. Thus, a benchmark scenarios for contingency management might consist of one or several commonly used functions taken from real world applications, used for evaluation, characterization and performance measurement of a contingency management algorithm. Behavior of the contingency management algorithm under different environmental conditions should then be able to be predicted using a set of benchmark functions. The paper describes the resulting benchmark problem as an illustration of the application of this process.

Jon Holbrook↗

ORNL Testing of Multiple Graphite Benchmarks [Slides]

Integral criticality safety and reactor physics benchmark experiments from the International Handbook of Evaluated Criticality Safety Benchmark Experiments (ICSBEP Handbook) and the International Handbook of Reactor Physics Experiments (IRPhE Handbook) are essential for nuclear data validation and testing

ENDF↗

Aerothermal modeling program, phase 1

The physical modeling embodied in the computational fluid dynamics codes is discussed. The objectives were to identify shortcomings in the models and to provide a program plan to improve the quantitative accuracy. The physical models studied were for: turbulent mass and momentum transport, heat release, liquid fuel spray, and gaseous radiation. The approach adopted was to test the models against appropriate benchmark-quality test cases from experiments in the literature for the constituent flows that together make up the combustor real flow.

Sturgess, G. J.↗

OECD-NEA HTTF Benchmark Progress and Updates

General slides discussing High Temperature Test Facility benchmark progress, Prismatic HTGR deployment, modeling and simulation tools for validating systems, and verification and validation issues. Also discuss collaboration and conclusions of RELAP5-3D validation activities based on HTTF with trends in data and international impact to accelerate deployment of prismatic HTGR microreactors by providing an opportunity for designers to assess their codes against experimental data and solutions from other codes.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Architecture and evolution of Goddard Space Flight Center Distributed Active Archive Center

The Goddard Space Flight Center (GSFC) Distributed Active Archive Center (DAAC) has been developed to enhance Earth Science research by improved access to remote sensor earth science data. Building and operating an archive, even one of a moderate size (a few Terabytes), is a challenging task. One of the critical components of this system is Unitree, the Hierarchical File Storage Management System. Unitree, selected two years ago as the best available solution, requires constant system administrative support. It is not always suitable as an archive and distribution data center, and has moderate performance. The Data Archive and Distribution System (DADS) software developed to monitor, manage, and automate the ingestion, archive, and distribution functions turned out to be more challenging than anticipated. Having the software and tools is not sufficient to succeed. Human interaction within the system must be fully understood to improve efficiency to improve efficiency and ensure that the right tools are developed. One of the lessons learned is that the operability, reliability, and performance aspects should be thoroughly addressed in the initial design. However, the GSFC DAAC has demonstrated that it is capable of distributing over 40 GB per day. A backup system to archive a second copy of all data ingested is under development. This backup system will be used not only for disaster recovery but will also replace the main archive when it is unavailable during maintenance or hardware replacement. The GSFC DAAC has put a strong emphasis on quality at all level of its organization. A Quality team has also been formed to identify quality issues and to propose improvements. The DAAC has conducted numerous tests to benchmark the performance of the system. These tests proved to be extremely useful in identifying bottlenecks and deficiencies in operational procedures.

Bedet, Jean-Jacques↗

Depletion Benchmark of the AFIP-7 Experiment in the Advanced Test Reactor

Reactor physics depletion benchmarks for low-enriched uranium fuel are limited in number. In particular, there is very limited data for LEU benchmarks for U-10Mo (Uranium-10% Molybdenum) plate fuel developed for use in U.S. high-performance research reactors (USHPRR). USHPRR includes the Advanced Test Reactor (ATR), Advanced Test Reactor Critical Facility (ATR-C), High Flux Isotope Reactor (HFIR), University of Missouri Research Reactor (MURR), Massachusetts Institute of Technology Reactor (MITR), and National Bureau of Standards Reactor (NBSR) at the National Institute of Science and Technology. These reactors are fueled with high-enriched uranium dispersed fuel in a silicon/aluminum matrix. In support of conversion to a HALEU fuel, qualification of U-10Mo formed into a monolithic foil is being performed. Fuel qualification involves irradiated fueled specimens in the ATR. The irradiation tests provide an opportunity to benchmark depletion capabilities of reactor physics codes in support of the ATR operation, as well as develop benchmarks that can be used by other institutions to benchmark other reactor physics codes. This report documents the development of a benchmark model of the irradiation of the ATR Full -size plate In center flux trap Position 7 (AFIP-7) experiment.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Verification and Validation Studies for the LAVA CFD Solver

The verification and validation of the Launch Ascent and Vehicle Aerodynamics (LAVA) computational fluid dynamics (CFD) solver is presented. A modern strategy for verification and validation is described incorporating verification tests, validation benchmarks, continuous integration and version control methods for automated testing in a collaborative development environment. The purpose of the approach is to integrate the verification and validation process into the development of the solver and improve productivity. This paper uses the Method of Manufactured Solutions (MMS) for the verification of 2D Euler equations, 3D Navier-Stokes equations as well as turbulence models. A method for systematic refinement of unstructured grids is also presented. Verification using inviscid vortex propagation and flow over a flat plate is highlighted. Simulation results using laminar and turbulent flow past a NACA 0012 airfoil and ONERA M6 wing are validated against experimental and numerical data.

Validation↗

Benchmark Specification of Advanced Burner Test Reactor

As an effort to assess the Argonne Reactor Computation (ARC) suite of fast reactor analysis codes, a numerical benchmark problem was developed using the reference 250 MWt Advanced Burner Test Reactor (ABTR) metallic core fueled with beginning of equilibrium cycle compositions (Chang et al., 2006). Material thermal expansion at operating condition was modeled by adjusting the hexagonal pitch, axial meshes, and the fuel and structure material densities appropriately. Irradiation swelling of metal fuel was considered, and the bond sodium was displaced into the lower part of fission gas plenum.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

ART/Ada design project, phase 1. Task 3 report: Test plan

The plan is described for the integrated testing and benchmark of Phase Ada based ESBT Design Research Project. The integration testing is divided into two phases: (1) the modules that do not rely on the Ada code generated by the Ada Generator are tested before the Ada Generator is implemented; and (2) all modules are integrated and tested with the Ada code generated by the Ada Generator. Its performance and size as well as its functionality is verified in this phase. The target platform is a DEC Ada compiler on VAX mini-computers and VAX stations running the VMS operating system.

Allen, Bradley P.↗

Corrosion Control in Carbon Fiber Reinforced Plastic (CFRP) Composite-Aluminum Closure Panel Hem Joints

This project targeted the implementation of CFRP/Al closure panels as a drop-in replacement for all-aluminum closures used in automobile production today. The project focused on the development of corrosion fundamentals involving CFRP materials and the use of those materials in mixed-material joints. Numerous studies were completed to benchmark current materials and test mixed metal joints. Good correlation was found between outdoor exposures and laboratory cyclic corrosion testing. The insights from the corrosion testing fundamentals and benchmarking phases informed the development of low-cure adhesive and electrocoats and conductive primers for the CFRP. Adhesive and electrocoat materials were successfully developed meeting low-cure targets for bake (150C/10min). These materials were scaled-up and used to produce 5 prototype liftgates to further evaluate CFRP/Al joints in a full-scale part. The lower cure materials enabled mitigation of CTE mismatch and reduced the required bake temperature, resulting in reduced part distortion. Corrosion testing of the liftgates revealed unexpected increases in corrosion relative to the standard bake materials. While rigorous evaluation of the parts was not possible due to numerous complexities associated with prototype testing, the results provided significant insights into the material behavior of CFRP/Al joints. Key learnings from the project included a better understanding of the system complexity (interplay of substrate/primer conductivity, production of full-scale prototypes), advances in coating/adhesive design for low-cure applications involving CFRP, and advancement of corrosion test protocols for CFRP.

36 MATERIALS SCIENCE↗

Test Cases for Flutter of the Benchmark Models Rectangular Wings on the Pitch and Plunge Apparatus

The supercritical airfoil was chosen as a relatively modem airfoil for comparison. The BOO12 model was tested first. Three different types of flutter instability boundaries were encountered, a classical flutter boundary, a transonic stall flutter boundary at angle of attack, and a plunge instability near M = 0.9 and for zero angle of attack. This test was made in air and was Transonic Dynamics Tunnel (TDT) Test 468. The BSCW model (for Benchmark SuperCritical Wing) was tested next as TDT Test 470. It was tested using both with air and a heavy gas, R-12, as a test medium. The effect of a transition strip on flutter was evaluated in air. The B64AOlO model was subsequently tested as TDT Test 493. Some further analysis of the experimental data for the BOO12 wing is presented. Transonic calculations using the parameters for the BOO12 wing in a two-dimensional typical section flutter analysis are given. These data are supplemented with data from the Benchmark Active Controls Technology model (BACT) given and in the next chapter of this document. The BACT model was of the same planform and airfoil as the BOO12 model, but with spoilers and a trailing edge control. It was tested in the heavy gas R-12, and was instrumented mostly at the 60 per cent span. The flutter data obtained on PAPA and the static aerodynamic test cases from BACT serve as additional data for the BOO12 model. All three types of flutter are included in the BACT Test Cases. In this report several test cases are selected to illustrate trends for a variety of different conditions with emphasis on transonic flutter. Cases are selected for classical and stall flutter for the BSCW model, for classical and plunge for the B64AOlO model, and for classical flutter for the BOO12 model. Test Cases are also presented for BSCW for static angles of attack. Only the mean pressures and the real and imaginary parts of the first harmonic of the pressures are included in the data for the test cases, but digitized time histories have been archived. The data for the test cases are available as separate electronic files. An overview of the model and tests is given, the standard formulary for these data is listed, and some sample results are presented.

Robert M Bennett↗

Implementation of a First-Order Quadratic Program Solver in C

This paper details a translation of a first order quadratic program (QP) solver from MATLAB to C. NASA could use this QP solver to generate online flight path trajectories for powered descent vehicles during landing. Over 12 weeks, the team designed, implemented, and tested two iterations of the QP solver for accuracy and runtime on 104 benchmark QP tests. The final iteration was 541.07% faster than the first, handling most tests in under one second. Additionally, it solved four more QP tests for N≥1383, and all outputs for cost and D_x matched the MATLAB reference values.

Optimization↗

IER 480: TEX-Pu Benchmark (PU-MET-THERM-004) to Test Polyethylene and Lucite Thermal Scattering Laws [Slides]

This presentation finds that PMT-004 has four new benchmark cases highly sensitive to PE TSL (2 cases) and PMMA TSL (2 Cases). PE cases were well predicted using MCNP6.2 and ENDF/B-VIII.0. PMMA cases overpredicted by approximately 0.6-0.7% in k eff at 20°C. Accepted into 2023 version of the ICSBEP Handbook. Temperature had a large impact on reactivity of the critical configuration. Implications for validation work for thermal cases- need to adjust TSL data to correct temperature as it can have hundreds of pcm effects for a few °C. Future thermal experiments should try and measure reactivity at multiple temperatures to aid in data testing and benchmark adjustment.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

AI Model Benchmarking for Nonproliferation Applications: Steel Thread Benchmarking Task Force Technical Report (Rev. 2)

Steel Thread is a NA-22 venture that seeks to build trustworthy, reliable AI models that can be used in a wide variety of nonproliferation tasks. A key aspect of building these models is developing appropriate benchmarks and evaluation methods, which will enable the venture to identify and adapt models to provide the most value in the nonproliferation domain. Benchmarks must be relevant to key tasks in this domain, such as question answering, information retrieval, document summarization and classification, consensus analysis, and image and data analysis. This report 1) provides an overview of benchmark design, evaluation, and challenges; 2) reviews a variety of open benchmarks, with a focus on language models and tasks; and 3) identifies benchmarks that are most relevant to Steel Thread. This report is intended to serve as a basis for further efforts to classify and evaluate benchmarks and their correlation with success on nonproliferation-specific tasks. The Steel Thread venture has defined benchmarks to be a particular combination of a dataset (or datasets) and a metric (or metrics) conceptualized as representing one or more specific tasks or sets of abilities for a specific modality. It is adopted by a research community as a shared framework for comparing methods.1 It includes 1) Data: Labeled (a designated subset not used for training, which could be all the data), 2) Metric: A way to quantify performance, 3) Task/Ability: What the benchmark is testing, 4) Protocol: A structured and repeatable evaluation process, 5) Baseline/Reference Model: For comparison; could be statistical, rule-based, SME-derived, or another model, and 6) Maintenance Plan: to update with new information over time; important for long-term utility. For further clarity, the definition includes what a benchmark, in this context, is not. It is not a corpus of training data, specific to a model (it is intended to apply to a range of models), a universal evaluation of performance, a guarantee that the ‘top’ model on the leaderboard will be the best fit for every specific use case, an all-encompassing proof of a model’s universal quality, nor is it a one-size-fits-all measure of success. It does not cover every real-world constraint (like operational, ethical, or cost considerations), a systems integration test, or a unit test. This definition was inspired by and resulted from discussions within the Steel Thread Benchmarking Task Force. This group was formed to define what we would mean as a benchmark within Steel Thread but persisted as the need to develop a thorough understanding of the large and expanding existing benchmarking space. This technical report is a result of the group’s divide and conquer approach to exploring this space. The release of benchmarks might not be progressing as quickly as model development, but it is moving very fast, as many benchmarks quickly become saturated, when state-of-the-art models score so close to the benchmark’s ceiling that their results are virtually indistinguishable. At that point, the test no longer differentiates between new systems, so researchers usually stop reporting scores as the benchmark no longer informs about improvements from the next generation of models. In the OpenAI announcement of GPT-5, they reported results on six flagship public benchmarks (AIME 2025, SWE-bench Verified, Aider Polyglot, MMMU, HealthBench Hard, GPQA) but the full system-card covers roughly thirty-five separate evaluations, comprising hundreds of test task items in total. There have been some efforts to summarize benchmarks in specific fields, like for text-to-image generation, but these surveys have had a narrow methodology scope. Therefore, a comprehensive survey of all benchmarks or even all benchmarks that could be relevant to Steel Thread is outside of the scope of this report. We chose some specific benchmarks to investigate in detail.

97 MATHEMATICS AND COMPUTING↗

Toward a Benchmark for Multi-Threaded Testing Tools

Looking for intermittent bugs is a problem that has been getting prominence in testing. Multi-threaded code is becoming very common, mostly on the server side. As there is no silver bullet solution, research focuses on a variety of partial solutions. We outline a road map for combining the research on the different disciplines of testing multi-threaded programs and on evaluating its quality. The project goals are to create a benchmark that can be used to evaluate different solutions, to create a framework with open API's that enables combining techniques in the multithreading domain, and to create a focus for the research in this area around which a community of people who try to solve similar problems with different techniques, could congregate. The benchmark, apart from containing programs with documented bugs, includes other artifacts, such as traces, that are used for evaluating some of the technologies. We have started creating such a bench mrk and detail the lesson learned in the process. The framework will enable technology developers, for example, race detectors, to concentrate on their components and use other ready made components, (e.g., instrumentor) to create a testing solution.

Eytani, Yaniv↗

Open Source Evaluation Framework for Solar Forecasting

The Solar Forecast Arbiter is an open-source evaluation framework for solar forecasting. The framework enables evaluations of solar irradiance, solar power, and net-load forecasts that are impartial, repeatable and auditable. The Solar Forecast Arbiter addresses stakeholder-informed use cases including evaluation of forecast skill, comparisons to reference data sets, private forecast trials, and evaluation of probabilistic forecast skill. The framework includes a data validation toolkit, reference data sources, data privacy protocols, and benchmark forecast capabilities for intra-hour and day ahead forecast horizons. Reports and metrics communicate the relative merits of the test and benchmark forecasts. The reports are created from standardized templates and include graphics for qualitatively evaluating deterministic and probabilistic forecasts and standard metrics for quantitatively evaluating forecasts. The Solar Forecast Arbiter is designed to support all solar forecasting stakeholders, including Solar Forecasting 2 Topic Area 2 and Topic Area 3 teams.

14 SOLAR ENERGY↗

Performance modeling & simulation of complex systems (A systems engineering design & analysis approach)

Modeling of the Multi-mission Image Processing System (MIPS) will be described as an example of the use of a modeling tool to design a distributed system that supports multiple application scenarios. This paper examines: (a) modeling tool selection, capabilities, and operation (namely NETWORK 2.5 by CACl), (b) pointers for building or constructing a model and how the MIPS model was developed, (c) the importance of benchmarking or testing the performance of equipment/subsystems being considered for incorporation the design/architecture, (d) the essential step of model validation and/or calibration using the benchmark results, (e) sample simulation results from the MIPS model, and (f) how modeling and simulation analysis affected the MIPS design process by having a supportive and informative impact.

Hall, Laverne↗