Spring 2025 evaluation cycle
Benchmark Leaderboard
Frontier research models ranked across benchmark suites. Select models to compare, or open any model or benchmark for the full breakdown.
Tip — click a benchmark column header to read the methodology and task-level rankings, or a model name for per-benchmark scores.

About the leaderboard
One consistent harness, every suite
ModelBench evaluates every model with the same deterministic harness: identical prompt templates, fixed decoding settings, and three seeded runs averaged into a single reported score. No model-specific prompting tricks, no cherry-picked checkpoints.
All benchmarks are scored higher-is-better on a 0–100 scale. Percentiles are computed against the eight models currently on the board, and rankings refresh automatically when a new checkpoint is released.