Engineering PapersSearch

DOE OSTI · 3019412

Ranking and Classifying AI Benchmarks

Abstract

We created a set of standards to efficiently evaluate AI benchmarks through objective means. Although prevalent, especially in recent times, AI benchmarks have no single way to measure their effectiveness. The MLCommons team provided a set of criteria for evaluating benchmarks, although the criteria lacks a clearly defined set of evaluation rules. We created a rubric with preset factors to efficiently and objectively evaluate a benchmark s quality. We created a software framework for processing lists of benchmarks for visualization. The framework and rating system allows researchers to quickly check if their benchmarks are effective.

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Shiraishi, Reece C. [Cornell U.], Hawks, Benjamin G. [Fermilab]. 2025-08-06. Ranking and Classifying AI Benchmarks. https://doi.org/10.2172/3019412

Cite the original work for its findings. Save a collection to share your selection of sources.