BenchBoard
Leaderboard Overview
Public benchmarks · Updated weekly

Leaderboard Overview

Community-curated rankings of public ML benchmarks, scored across vision, language, speech and more. Tap any row to inspect a benchmark in depth.

Top-ranked public datasets
RankDatasetScoreCategory
Top 5 of 10 tracked datasets Composite 0–100 scale
Benchmark summaries
Data Explorer

Browse every benchmark

Search, filter and sort the full registry of public benchmarks. Select a row to open its detail page.

RankDatasetCategoryScoreDate
Showing 10 datasetsTap a row for full detail
Benchmark Detail

Benchmark: ImageNet-22K

Benchmark cover
Rank
1
Score
98.7 / 100
Category
Vision
Date
2024-05-12

Score progression Last 12 evaluations

How this benchmark is evaluated

Every submission runs on identical held-out splits with fixed seeds and frozen preprocessing, so scores are directly comparable across research groups and time. Models are evaluated offline; no test labels are ever released.

Scoring pipeline

  1. Submissions are containerized and executed on the same hardware pool (8×A100) with a 24-hour wall clock.
  2. Primary and auxiliary metrics are computed on the private evaluation server from raw predictions.
  3. A composite score rescales each metric to a 0–100 band and weights it per category conventions.
  4. Two independent reviewers audit anomalous runs before a score is published to the leaderboard.

Versioning & fairness

The evaluation harness is version-pinned; when the harness or split changes, historical scores are re-run or clearly flagged as legacy so ranking movement is always attributable to real model improvements.

This benchmark sits at rank 1 in its category. Nearby competitors:

RankDatasetCategoryScore
Filters applied