LLM benchmark scores (monthly)
Monthly cross-leaderboard scoreboard of large language models: every (model, benchmark) score from LMArena's official public leaderboard dataset (Bradley-Terry ratings, text/overall arena, CC-BY-4.0) and Epoch AI's Capabilities & Benchmarking data hub (56 current capability benchmarks, CC-BY), rescaled per benchmark to a comparable 0-100 (min-max within the snapshot; all benchmarks are higher-is-better) with per-benchmark ranks. Models are entity-resolved across the two sources (exact name match, then one arena-variant suffix stripped, unambiguous only; tier recorded per row). Superseded Epoch benchmarks are pruned. Columns: snapshot date (day-granular UTC), source, benchmark, benchmark release date, resolved model key, raw upstream model id, display name, resolution tier, organization, ISO alpha-3 country, raw score, score unit, normalized 0-100 score, rank in benchmark, models in benchmark, vote count (LMArena), source publish date, source URL. Primary key: (snapshot_date, source, benchmark, source_model_id). Cadence: monthly; same-day re-runs are content-hash no-ops. Caveats: the 0-100 scale is within-benchmark relative, not absolute; LMArena measures human preference, Epoch benchmarks measure task accuracy — the composite dataset blends them explicitly. Sample use: filter benchmark = 'GPQA diamond' order by rank_in_benchmark for the current reasoning-benchmark ranking.
- Rows
- 3,121
- Columns
- 18
- Source cadence
- Monthly
- Last refreshed
- Sep 25, 2026
- Theme
- technology
| Column | Type | Description |
|---|---|---|
| snapshot_date | string | Fetch date at UTC day granularity. Same-day re-runs produce identical snapshots; any upstream change yields a new content-hashed snapshot. (unit: ISO date) |
| source | string | 'lmarena' (LMArena official public leaderboard dataset) or 'epoch_ai' (Epoch AI Capabilities & Benchmarking data hub). |
| benchmark | string | Benchmark name: an Epoch AI benchmark (e.g. 'GPQA diamond') or 'LMArena Text Arena (overall)'. Superseded Epoch benchmarks are pruned. |
| benchmark_release_date | string | Benchmark release date from Epoch AI's benchmark_metadata.csv; empty for LMArena. (unit: ISO date) |
| model_key | string | Resolved canonical model key (separator-stripped lowercase name): the cross-source join key. See resolution_tier for how the join was made. |
| source_model_id | string | Raw upstream model id (LMArena model_name or Epoch AI Model display name); the per-source row identity. |
| display_name | string | Human-readable model name; Epoch AI's display name preferred when the model is resolved, else the LMArena id. |
| resolution_tier | string | Entity-resolution tier: 'exact' (identical normalized names), 'variant' (one arena-variant suffix stripped, unambiguous only), 'epoch_only', 'lmarena_only'. |
| organization | string | Model organization as reported by the source (Epoch ECI metadata preferred). |
| country_code | string | ISO alpha-3 country of the organization (Epoch AI metadata only; empty for LMArena-only models). (unit: ISO alpha-3) |
| raw_score | float | Raw upstream score: Bradley-Terry rating for LMArena, 0-1 performance fraction for Epoch AI. (unit: see score_unit) |
| score_unit | string | 'bradley-terry rating' (LMArena) or 'fraction 0-1' (Epoch AI). |
| normalized_score | float | Per-benchmark min-max rescale to 0-100 within the snapshot: (x - min) / (max - min) * 100. All benchmarks are higher-is-better. Null when undefined (single-observation benchmark). (unit: 0-100) |
| rank_in_benchmark | integer | Rank by raw score descending within the benchmark (1 = best); ties share the rank. (unit: rank) |
| n_models_in_benchmark | integer | Models scored on this benchmark in this snapshot. (unit: count) |
| vote_count | float | LMArena pairwise votes behind the rating; null for Epoch AI rows. (unit: count) |
| source_publish_date | string | Upstream publish date: LMArena leaderboard_publish_date, or the Epoch AI row date. (unit: ISO date) |
| source_url | string | Exact upstream URL behind the row: the LMArena HF dataset page or the Epoch AI benchmark_data.zip download. (unit: url) |
First 10 sample rows — a preview, not the complete dataset.
| snapshot_date | source | benchmark | benchmark_release_date | model_key | source_model_id | display_name | resolution_tier | organization | country_code | raw_score | score_unit | normalized_score | rank_in_benchmark | n_models_in_benchmark | vote_count | source_publish_date | source_url |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | gpt35turbonov2023 | GPT-3.5 Turbo (Nov 2023) | GPT-3.5 Turbo (Nov 2023) | epoch_only | — | — | 0.372 | fraction 0-1 | 100 | 1 | 9 | — | 2023-11-06 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3small74b | phi-3-small 7.4B | phi-3-small 7.4B | epoch_only | — | — | 0.372 | fraction 0-1 | 100 | 1 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | llama38b | Llama 3-8B | Llama 3-8B | variant | — | — | 0.36 | fraction 0-1 | 94.87 | 3 | 9 | — | 2024-04-18 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3medium14b | phi-3-medium 14B | phi-3-medium 14B | epoch_only | — | — | 0.337 | fraction 0-1 | 85.26 | 4 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | mixtral8x7b | Mixtral 8x7B | Mixtral 8x7B | epoch_only | — | — | 0.328 | fraction 0-1 | 81.41 | 5 | 9 | — | 2023-12-11 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3mini38b | phi-3-mini 3.8B | phi-3-mini 3.8B | epoch_only | — | — | 0.292 | fraction 0-1 | 66.03 | 6 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | gemma7b | Gemma 7B | Gemma 7B | epoch_only | — | — | 0.231 | fraction 0-1 | 39.74 | 7 | 9 | — | 2024-02-21 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | mistral7bv01 | Mistral 7B v0.1 | Mistral 7B v0.1 | epoch_only | — | — | 0.207 | fraction 0-1 | 29.49 | 8 | 9 | — | 2023-09-27 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi2 | Phi-2 | Phi-2 | epoch_only | — | — | 0.138 | fraction 0-1 | 0 | 9 | 9 | — | 2023-12-12 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | APEX-Agents | 2026-01-21 | claudefable51 | Claude Fable 5.1 | Claude Fable 5.1 | variant | — | — | 0.686 | fraction 0-1 | 100 | 1 | 32 | — | 2026-09-01 | https://epoch.ai/data/benchmark_data.zip |
Profiled Sep 25, 2026 from snapshot 20260925T034920Z-f70e069d920c
Measured- Completeness
- 95.2%
- Rows
- 3,121
- Columns
- 18
- Columns with gaps
- 1
| Column | Missing | Distinct | Range | Distribution |
|---|---|---|---|---|
| snapshot_datevarchar | 0% | 1 | — |
|
| sourcevarchar | 0% | 2 | — |
|
| benchmarkvarchar | 0% | 63 | — |
|
| benchmark_release_datevarchar | 0% | 52 | — |
|
| model_keyvarchar | 0% | 582 | — |
|
| source_model_idvarchar | 0% | 793 | — |
|
| display_namevarchar | 0% | 578 | — |
|
| resolution_tiervarchar | 0% | 4 | — |
|
| organizationvarchar | 0% | 32 | — |
|
| country_codevarchar | 0% | 1 | — |
|
| raw_scoredouble | 0% | 2,722 | 0 → 1,508median 0.525 | 32 outside 1st–99th percentile |
| score_unitvarchar | 0% | 2 | — |
|
| normalized_scoredouble | 0% | 2,070 | 0 → 100median 55.85 | |
| rank_in_benchmarkbigint | 0% | 440 | 1 → 402median 35 | 32 outside 1st–99th percentile |
| n_models_in_benchmarkbigint | 0% | 39 | 6 → 402median 76 | 30 outside 1st–99th percentile |
| vote_countdouble | 87.1% | 409 | 791 → 194,909median 21,609 | 10 outside 1st–99th percentile |
| source_publish_datevarchar | 0% | 206 | — |
|
| source_urlvarchar | 0% | 2 | — |
|
- Current
20260925T034920Z-f70e069d920c · sha256 f70e069d920c…
3,121 rows · first snapshot
Point any LLM at the metadata endpoint — the documentation above is machine-readable too (JSON-LD + Croissant).
curl "https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly" | jq '{title, rows, columns_count, license}'import requests
ds = requests.get("https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly").json()
print(ds["title"], ds["rows"], "rows")
# Sample rows for an LLM context window
for row in ds.get("sample_rows", [])[:5]:
print(row)API endpoint: https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly
Tip: fetch /llms.txt for the full machine-readable catalog.
Where this data comes from and what was made from it. Other people's work shows as counts; only shared projects are named.
Cite this snapshot
Pinned to snapshot 20260925T034920Z-f70e069d920c and its content hash, so readers get exactly the data you used.
LLM benchmark scoreboard (agent-curated). (2026). LLM benchmark scores (monthly) [Data set, snapshot 20260925T034920Z-f70e069d920c, sha256 f70e069d920c]. Datazimuts. Retrieved 2026-09-25, from https://datazimuts.com/en/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly?snapshot=20260925T034920Z-f70e069d920c
@misc{dz_llm_benchmark_signals_llm_benchmark_scor_f70e069d,
title = {{LLM benchmark scores (monthly)}},
author = {{LLM benchmark scoreboard (agent-curated)}},
year = {2026},
publisher = {Datazimuts},
howpublished = {\url{https://datazimuts.com/en/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly?snapshot=20260925T034920Z-f70e069d920c}},
note = {Snapshot 20260925T034920Z-f70e069d920c, sha256 f70e069d920c8f97a3c982dfbbd698aa9c6afde05307705f36a29921e166642e; accessed 2026-09-25}
}Embed a table or a chart
Paste this into any page. The embed is pinned to the same snapshot, follows the reader's light or dark setting, and always shows the source, license and a link back.
<iframe src="https://datazimuts.com/embed/chart?dataset=llm_benchmark_signals%2Fllm_benchmark_scores_monthly&lang=en&theme=auto&snapshot=20260925T034920Z-f70e069d920c&x=snapshot_date&y=raw_score&agg=avg" title="LLM benchmark scores (monthly)" width="100%" height="380" style="border:0" loading="lazy"></iframe>
Ask about this dataset. Answers come only from its catalog record, measured profile and change history, and list the facts they used.