Public benchmarks · Updated weekly
Leaderboard Overview
Community-curated rankings of public ML benchmarks, scored across vision, language, speech and more. Tap any row to inspect a benchmark in depth.
Top-ranked public datasets
| Rank | Dataset | Score | Category |
|---|
Top 5 of 10 tracked datasets
Composite 0–100 scale
Benchmark summaries
Data Explorer
Browse every benchmark
Search, filter and sort the full registry of public benchmarks. Select a row to open its detail page.
| Rank | Dataset | Category | Score | Date |
|---|
Showing 10 datasetsTap a row for full detail
Benchmark Detail
Benchmark: ImageNet-22K
Rank
1
Score
98.7 / 100
Category
Vision
Date
2024-05-12
Score progression Last 12 evaluations
How this benchmark is evaluated
Every submission runs on identical held-out splits with fixed seeds and frozen preprocessing, so scores are directly comparable across research groups and time. Models are evaluated offline; no test labels are ever released.
Scoring pipeline
- Submissions are containerized and executed on the same hardware pool (8×A100) with a 24-hour wall clock.
- Primary and auxiliary metrics are computed on the private evaluation server from raw predictions.
- A composite score rescales each metric to a 0–100 band and weights it per category conventions.
- Two independent reviewers audit anomalous runs before a score is published to the leaderboard.
Versioning & fairness
The evaluation harness is version-pinned; when the harness or split changes, historical scores are re-run or clearly flagged as legacy so ranking movement is always attributable to real model improvements.
This benchmark sits at rank 1 in its category. Nearby competitors:
| Rank | Dataset | Category | Score |
|---|