LLM benchmark scores (monthly)
Monthly cross-leaderboard scoreboard of large language models: every (model, benchmark) score from LMArena's official public leaderboard dataset (Bradley-Terry ratings, text/overall arena, CC-BY-4.0) and Epoch AI's Capabilities & Benchmarking data hub (56 current capability benchmarks, CC-BY), rescaled per benchmark to a comparable 0-100 (min-max within the snapshot; all benchmarks are higher-is-better) with per-benchmark ranks. Models are entity-resolved across the two sources (exact name match, then one arena-variant suffix stripped, unambiguous only; tier recorded per row). Superseded Epoch benchmarks are pruned. Columns: snapshot date (day-granular UTC), source, benchmark, benchmark release date, resolved model key, raw upstream model id, display name, resolution tier, organization, ISO alpha-3 country, raw score, score unit, normalized 0-100 score, rank in benchmark, models in benchmark, vote count (LMArena), source publish date, source URL. Primary key: (snapshot_date, source, benchmark, source_model_id). Cadence: monthly; same-day re-runs are content-hash no-ops. Caveats: the 0-100 scale is within-benchmark relative, not absolute; LMArena measures human preference, Epoch benchmarks measure task accuracy — the composite dataset blends them explicitly. Sample use: filter benchmark = 'GPQA diamond' order by rank_in_benchmark for the current reasoning-benchmark ranking.
Les titres et les descriptions proviennent des sources de données, en anglais.
- Lignes
- 3 121
- Colonnes
- 18
- Cadence de la source
- Mensuelle
- Dernière actualisation
- 25 sept. 2026
- Thème
- technology
| Colonne | Type | Description |
|---|---|---|
| snapshot_date | string | Fetch date at UTC day granularity. Same-day re-runs produce identical snapshots; any upstream change yields a new content-hashed snapshot. (unit: ISO date) |
| source | string | 'lmarena' (LMArena official public leaderboard dataset) or 'epoch_ai' (Epoch AI Capabilities & Benchmarking data hub). |
| benchmark | string | Benchmark name: an Epoch AI benchmark (e.g. 'GPQA diamond') or 'LMArena Text Arena (overall)'. Superseded Epoch benchmarks are pruned. |
| benchmark_release_date | string | Benchmark release date from Epoch AI's benchmark_metadata.csv; empty for LMArena. (unit: ISO date) |
| model_key | string | Resolved canonical model key (separator-stripped lowercase name): the cross-source join key. See resolution_tier for how the join was made. |
| source_model_id | string | Raw upstream model id (LMArena model_name or Epoch AI Model display name); the per-source row identity. |
| display_name | string | Human-readable model name; Epoch AI's display name preferred when the model is resolved, else the LMArena id. |
| resolution_tier | string | Entity-resolution tier: 'exact' (identical normalized names), 'variant' (one arena-variant suffix stripped, unambiguous only), 'epoch_only', 'lmarena_only'. |
| organization | string | Model organization as reported by the source (Epoch ECI metadata preferred). |
| country_code | string | ISO alpha-3 country of the organization (Epoch AI metadata only; empty for LMArena-only models). (unit: ISO alpha-3) |
| raw_score | float | Raw upstream score: Bradley-Terry rating for LMArena, 0-1 performance fraction for Epoch AI. (unit: see score_unit) |
| score_unit | string | 'bradley-terry rating' (LMArena) or 'fraction 0-1' (Epoch AI). |
| normalized_score | float | Per-benchmark min-max rescale to 0-100 within the snapshot: (x - min) / (max - min) * 100. All benchmarks are higher-is-better. Null when undefined (single-observation benchmark). (unit: 0-100) |
| rank_in_benchmark | integer | Rank by raw score descending within the benchmark (1 = best); ties share the rank. (unit: rank) |
| n_models_in_benchmark | integer | Models scored on this benchmark in this snapshot. (unit: count) |
| vote_count | float | LMArena pairwise votes behind the rating; null for Epoch AI rows. (unit: count) |
| source_publish_date | string | Upstream publish date: LMArena leaderboard_publish_date, or the Epoch AI row date. (unit: ISO date) |
| source_url | string | Exact upstream URL behind the row: the LMArena HF dataset page or the Epoch AI benchmark_data.zip download. (unit: url) |
10 premières lignes d’exemple — un aperçu, pas le jeu de données complet.
| snapshot_date | source | benchmark | benchmark_release_date | model_key | source_model_id | display_name | resolution_tier | organization | country_code | raw_score | score_unit | normalized_score | rank_in_benchmark | n_models_in_benchmark | vote_count | source_publish_date | source_url |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | gpt35turbonov2023 | GPT-3.5 Turbo (Nov 2023) | GPT-3.5 Turbo (Nov 2023) | epoch_only | — | — | 0,372 | fraction 0-1 | 100 | 1 | 9 | — | 2023-11-06 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3small74b | phi-3-small 7.4B | phi-3-small 7.4B | epoch_only | — | — | 0,372 | fraction 0-1 | 100 | 1 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | llama38b | Llama 3-8B | Llama 3-8B | variant | — | — | 0,36 | fraction 0-1 | 94,87 | 3 | 9 | — | 2024-04-18 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3medium14b | phi-3-medium 14B | phi-3-medium 14B | epoch_only | — | — | 0,337 | fraction 0-1 | 85,26 | 4 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | mixtral8x7b | Mixtral 8x7B | Mixtral 8x7B | epoch_only | — | — | 0,328 | fraction 0-1 | 81,41 | 5 | 9 | — | 2023-12-11 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi3mini38b | phi-3-mini 3.8B | phi-3-mini 3.8B | epoch_only | — | — | 0,292 | fraction 0-1 | 66,03 | 6 | 9 | — | 2024-04-23 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | gemma7b | Gemma 7B | Gemma 7B | epoch_only | — | — | 0,231 | fraction 0-1 | 39,74 | 7 | 9 | — | 2024-02-21 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | mistral7bv01 | Mistral 7B v0.1 | Mistral 7B v0.1 | epoch_only | — | — | 0,207 | fraction 0-1 | 29,49 | 8 | 9 | — | 2023-09-27 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | ANLI | 2019-10-31 | phi2 | Phi-2 | Phi-2 | epoch_only | — | — | 0,138 | fraction 0-1 | 0 | 9 | 9 | — | 2023-12-12 | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | epoch_ai | APEX-Agents | 2026-01-21 | claudefable51 | Claude Fable 5.1 | Claude Fable 5.1 | variant | — | — | 0,686 | fraction 0-1 | 100 | 1 | 32 | — | 2026-09-01 | https://epoch.ai/data/benchmark_data.zip |
Profilé le 25 sept. 2026 à partir de l’instantané 20260925T034920Z-f70e069d920c
Mesuré- Complétude
- 95,2 %
- Lignes
- 3 121
- Colonnes
- 18
- Colonnes incomplètes
- 1
| Colonne | Manquant | Distinctes | Plage | Distribution |
|---|---|---|---|---|
| snapshot_datevarchar | 0 % | 1 | — |
|
| sourcevarchar | 0 % | 2 | — |
|
| benchmarkvarchar | 0 % | 63 | — |
|
| benchmark_release_datevarchar | 0 % | 52 | — |
|
| model_keyvarchar | 0 % | 582 | — |
|
| source_model_idvarchar | 0 % | 793 | — |
|
| display_namevarchar | 0 % | 578 | — |
|
| resolution_tiervarchar | 0 % | 4 | — |
|
| organizationvarchar | 0 % | 32 | — |
|
| country_codevarchar | 0 % | 1 | — |
|
| raw_scoredouble | 0 % | 2 722 | 0 → 1 508médiane 0,525 | 32 hors du 1er–99e centile |
| score_unitvarchar | 0 % | 2 | — |
|
| normalized_scoredouble | 0 % | 2 070 | 0 → 100médiane 55,85 | |
| rank_in_benchmarkbigint | 0 % | 440 | 1 → 402médiane 35 | 32 hors du 1er–99e centile |
| n_models_in_benchmarkbigint | 0 % | 39 | 6 → 402médiane 76 | 30 hors du 1er–99e centile |
| vote_countdouble | 87,1 % | 409 | 791 → 194 909médiane 21 609 | 10 hors du 1er–99e centile |
| source_publish_datevarchar | 0 % | 206 | — |
|
| source_urlvarchar | 0 % | 2 | — |
|
- Actuelle
20260925T034920Z-f70e069d920c · sha256 f70e069d920c…
3 121 lignes · premier instantané
Dirigez n’importe quel LLM vers le point d’accès des métadonnées — la documentation ci-dessus est aussi lisible par machine (JSON-LD + Croissant).
curl "https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly" | jq '{title, rows, columns_count, license}'import requests
ds = requests.get("https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly").json()
print(ds["title"], ds["rows"], "rows")
# Sample rows for an LLM context window
for row in ds.get("sample_rows", [])[:5]:
print(row)Point d’accès API : https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly
Astuce : récupérez /llms.txt pour le catalogue complet lisible par machine.
D’où viennent ces données et ce qui en a été fait. Le travail des autres apparaît sous forme de décomptes ; seuls les projets partagés sont nommés.
Citer cet instantané
Épinglé à l’instantané 20260925T034920Z-f70e069d920c et à son empreinte, pour que vos lecteurs obtiennent exactement les données utilisées.
LLM benchmark scoreboard (agent-curated). (2026). LLM benchmark scores (monthly) [Data set, snapshot 20260925T034920Z-f70e069d920c, sha256 f70e069d920c]. Datazimuts. Retrieved 2026-09-25, from https://datazimuts.com/fr/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly?snapshot=20260925T034920Z-f70e069d920c
@misc{dz_llm_benchmark_signals_llm_benchmark_scor_f70e069d,
title = {{LLM benchmark scores (monthly)}},
author = {{LLM benchmark scoreboard (agent-curated)}},
year = {2026},
publisher = {Datazimuts},
howpublished = {\url{https://datazimuts.com/fr/datasets/llm_benchmark_signals/llm_benchmark_scores_monthly?snapshot=20260925T034920Z-f70e069d920c}},
note = {Snapshot 20260925T034920Z-f70e069d920c, sha256 f70e069d920c8f97a3c982dfbbd698aa9c6afde05307705f36a29921e166642e; accessed 2026-09-25}
}Intégrer un tableau ou un graphique
Collez ce code dans n’importe quelle page. L’intégration est épinglée au même instantané, suit le thème clair ou sombre du lecteur et affiche toujours la source, la licence et un lien de retour.
<iframe src="https://datazimuts.com/embed/chart?dataset=llm_benchmark_signals%2Fllm_benchmark_scores_monthly&lang=fr&theme=auto&snapshot=20260925T034920Z-f70e069d920c&x=snapshot_date&y=raw_score&agg=avg" title="LLM benchmark scores (monthly)" width="100%" height="380" style="border:0" loading="lazy"></iframe>
Posez une question sur ce jeu de données. Les réponses viennent uniquement de sa fiche, de son profil mesuré et de son historique, et citent les faits utilisés.