Updated weekly
Spring 2025 evaluation cycle

Benchmark Leaderboard

Frontier research models ranked across benchmark suites. Select models to compare, or open any model or benchmark for the full breakdown.

Tip — click a benchmark column header to read the methodology and task-level rankings, or a model name for per-benchmark scores.

Source code on a monitor
About the leaderboard

One consistent harness, every suite

ModelBench evaluates every model with the same deterministic harness: identical prompt templates, fixed decoding settings, and three seeded runs averaged into a single reported score. No model-specific prompting tricks, no cherry-picked checkpoints.

All benchmarks are scored higher-is-better on a 0–100 scale. Percentiles are computed against the eight models currently on the board, and rankings refresh automatically when a new checkpoint is released.