AI Benchmark Leaderboard

Live rankings across frontier language models

Filters

Compare · 0 selected Tick two or more models to compare side-by-side.
Rank ModelOverall Score

About

Engineer working on model evaluation code
How rankings work

One score, five rigorous benchmarks

Every model is evaluated on identical prompts under identical settings. The Overall Score is a weighted aggregate of MMLU, GPQA, HumanEval, GSM8K and BBH results, refreshed after each evaluation run. Scores are normalized to a 0–100 scale, and every result links to the raw run data for full transparency.

Benchmark Data Explorer

Every evaluation run, filterable and exportable

Filters

Results

ModelBenchmarkScoreRankDate

Methodology

Circuit board macro photography
Under the hood

Identical prompts, identical settings

Each run uses a frozen prompt set with temperature 0 and majority voting where the benchmark permits it. Results are re-scored when a provider ships a silent model update, and the eval date always reflects the exact run timestamp shown in the table above. Explore freely — every cell is traceable to its source run.

Model Performance Detail

Per-benchmark breakdown and history

Model Comparison

Side-by-side across every tracked metric

Reading the results

Macro shot of a circuit board
Tips

The best value in each row is highlighted

Highlighted cells mark the strongest result per metric. Note that price and context length trade off against raw score — a slightly lower score at a fraction of the price often wins for production workloads. Click any model name in the header to open its full performance detail.