
Every model is evaluated on identical prompts under identical settings. The Overall Score is a weighted aggregate of MMLU, GPQA, HumanEval, GSM8K and BBH results, refreshed after each evaluation run. Scores are normalized to a 0–100 scale, and every result links to the raw run data for full transparency.

Each run uses a frozen prompt set with temperature 0 and majority voting where the benchmark permits it. Results are re-scored when a provider ships a silent model update, and the eval date always reflects the exact run timestamp shown in the table above. Explore freely — every cell is traceable to its source run.

Highlighted cells mark the strongest result per metric. Note that price and context length trade off against raw score — a slightly lower score at a fraction of the price often wins for production workloads. Click any model name in the header to open its full performance detail.