LLM composite capability index (monthly)
Monthly per-model composite of the cross-leaderboard scoreboard: for every model appearing in LMArena's official public leaderboard (text/overall, CC-BY-4.0) or Epoch AI's Capabilities & Benchmarking hub (CC-BY), the composite index is the equal-weighted mean of its normalized 0-100 score components (LMArena Bradley-Terry rating rescaled + each Epoch benchmark rescaled), with n_components recorded so single-signal models are distinguishable from broadly measured ones. Epoch's own Capabilities Index (ECI) is carried as a reference column, never a composite input. Columns: snapshot date, composite rank, resolved model key, display name, resolution tier, organization, ISO alpha-3 country, accessibility, component count, Epoch benchmarks covered, LMArena variants merged, Epoch mean normalized, LMArena normalized, LMArena rating/rank/votes, Epoch ECI + CI, composite index, source URLs. Primary key: (snapshot_date, model_key). Cadence: monthly. Caveats: a transparency-first average, not a capability claim; benchmarks differ in difficulty and LMArena measures human preference. Sample use: order by composite_rank for the current cross-source model ranking, or filter n_components >= 5 for broadly-measured models only.
Les titres et les descriptions proviennent des sources de données, en anglais.
- Lignes
- 546
- Colonnes
- 22
- Cadence de la source
- Mensuelle
- Dernière actualisation
- 25 sept. 2026
- Thème
- technology
| Colonne | Type | Description |
|---|---|---|
| snapshot_date | string | Fetch date at UTC day granularity. Same-day re-runs produce identical snapshots; any upstream change yields a new content-hashed snapshot. (unit: ISO date) |
| composite_rank | integer | Rank by composite_index descending (1 = highest). (unit: rank) |
| model_key | string | Resolved canonical model key (separator-stripped lowercase name): the cross-source join key. See resolution_tier for how the join was made. |
| display_name | string | Human-readable model name; Epoch AI's display name preferred when the model is resolved, else the LMArena id. |
| resolution_tier | string | Entity-resolution tier: 'exact' (identical normalized names), 'variant' (one arena-variant suffix stripped, unambiguous only), 'epoch_only', 'lmarena_only'. |
| organization | string | Model organization as reported by the source (Epoch ECI metadata preferred). |
| country_code | string | ISO alpha-3 country of the organization (Epoch AI metadata only; empty for LMArena-only models). (unit: ISO alpha-3) |
| accessibility | string | Model accessibility from Epoch AI metadata (e.g. 'API access'); empty for LMarena-only models. |
| n_components | integer | Number of normalized score components averaged into composite_index (LMArena rating + Epoch benchmarks). (unit: count) |
| n_epoch_benchmarks | integer | Epoch AI benchmarks covering the model. (unit: count) |
| n_lmarena_variants | integer | LMArena arena variants merged into this canonical model (best-rating variant represents it). (unit: count) |
| epoch_mean_normalized | float | Mean of the model's normalized Epoch benchmark scores (0-100). (unit: 0-100) |
| lmarena_normalized | float | LMArena Bradley-Terry rating rescaled 0-100 within the snapshot. (unit: 0-100) |
| lmarena_rating | float | Raw LMArena Bradley-Terry rating of the best variant. (unit: bradley-terry rating) |
| lmarena_rank | float | LMArena overall arena rank of the best variant. (unit: rank) |
| lmarena_vote_count | float | LMArena votes behind the best variant's rating. (unit: count) |
| epoch_eci | float | Epoch AI Capabilities Index for the model (reference column; never a composite input). (unit: index) |
| epoch_eci_ci_low | float | Lower bound of the ECI confidence interval. (unit: index) |
| epoch_eci_ci_high | float | Upper bound of the ECI confidence interval. (unit: index) |
| composite_index | float | Equal-weighted mean of the model's available normalized 0-100 components (LMArena + Epoch benchmarks). A transparency-first average, not a capability claim. (unit: 0-100) |
| source_url_lmarena | string | LMArena dataset URL (empty when the model has no LMArena coverage). (unit: url) |
| source_url_epoch | string | Epoch AI benchmark_data.zip URL (empty when the model has no Epoch coverage). (unit: url) |
10 premières lignes d’exemple — un aperçu, pas le jeu de données complet.
| snapshot_date | composite_rank | model_key | display_name | resolution_tier | organization | country_code | accessibility | n_components | n_epoch_benchmarks | n_lmarena_variants | epoch_mean_normalized | lmarena_normalized | lmarena_rating | lmarena_rank | lmarena_vote_count | epoch_eci | epoch_eci_ci_low | epoch_eci_ci_high | composite_index | source_url_lmarena | source_url_epoch |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-09-25 | 1 | gpt6astra | GPT-6 Astra | variant | OpenAI | USA | API access | 22 | 21 | 1 | 97,5 | 90,53 | 1 443,724 | 59 | 2 693 | 166,6 | 163 | 172,03 | 97,18 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | 2 | qwen35maxpreview | qwen3.5-max-preview | lmarena_only | alibaba | — | — | 1 | 0 | 1 | — | 94,56 | 1 470,932 | 25 | 21 476 | — | — | — | 94,56 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 3 | ernie51 | ernie-5.1 | lmarena_only | baidu | — | — | 1 | 0 | 1 | — | 94,16 | 1 468,195 | 28 | 37 058 | — | — | — | 94,16 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 4 | mimov25pro | mimo-v2.5-pro | lmarena_only | xiaomi | — | — | 1 | 0 | 1 | — | 93,65 | 1 464,749 | 32 | 60 919 | — | — | — | 93,65 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 5 | gemini25pro | gemini-2.5-pro | lmarena_only | — | — | 1 | 0 | 1 | — | 92,61 | 1 457,772 | 36 | 122 554 | — | — | — | 92,61 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — | |
| 2026-09-25 | 6 | gpt55pro | GPT-5.5 Pro | epoch_only | OpenAI | USA | API access | 10 | 10 | 0 | 91,58 | — | — | — | — | 162,45 | 159,23 | 166,64 | 91,58 | — | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | 7 | grok420beta0309reasoning | grok-4.20-beta-0309-reasoning | lmarena_only | xai | — | — | 1 | 0 | 1 | — | 91,57 | 1 450,74 | 42 | 62 168 | — | — | — | 91,57 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 8 | claudeopus4520251101 | claude-opus-4-5-20251101 | lmarena_only | anthropic | — | — | 1 | 0 | 1 | — | 91,53 | 1 450,459 | 44 | 70 013 | — | — | — | 91,53 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 9 | grok420multiagentbeta0309 | grok-4.20-multi-agent-beta-0309 | lmarena_only | xai | — | — | 1 | 0 | 1 | — | 91,45 | 1 449,939 | 46 | 60 777 | — | — | — | 91,45 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 10 | claudefable51 | Claude Fable 5.1 | variant | Anthropic | USA | API access | 19 | 18 | 1 | 90,9 | 100 | 1 507,582 | 1 | 5 783 | 165 | 161,64 | 169,61 | 91,38 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | https://epoch.ai/data/benchmark_data.zip |
Profilé le 25 sept. 2026 à partir de l’instantané 20260925T034926Z-581d6127d4c9
Mesuré- Complétude
- 85,3 %
- Lignes
- 546
- Colonnes
- 22
- Colonnes incomplètes
- 8
| Colonne | Manquant | Distinctes | Plage | Distribution |
|---|---|---|---|---|
| snapshot_datevarchar | 0 % | 1 | — |
|
| composite_rankbigint | 0 % | 607 | 1 → 546médiane 273,5 | 12 hors du 1er–99e centile |
| model_keyvarchar | 0 % | 582 | — |
|
| display_namevarchar | 0 % | 578 | — |
|
| resolution_tiervarchar | 0 % | 4 | — |
|
| organizationvarchar | 0 % | 64 | — |
|
| country_codevarchar | 0 % | 6 | — |
|
| accessibilityvarchar | 0 % | 9 | — |
|
| n_componentsbigint | 0 % | 29 | 1 → 31médiane 1 | 5 hors du 1er–99e centile |
| n_epoch_benchmarksbigint | 0 % | 27 | 0 → 30médiane 0 | 4 hors du 1er–99e centile |
| n_lmarena_variantsbigint | 0 % | 3 | 0 → 2médiane 1 | |
| epoch_mean_normalizeddouble | 50,9 % | 258 | 0,12 → 97,5médiane 45,8 | 6 hors du 1er–99e centile |
| lmarena_normalizeddouble | 29,9 % | 400 | 0 → 100médiane 74,09 | 8 hors du 1er–99e centile |
| lmarena_ratingdouble | 29,9 % | 365 | 833,41 → 1 508médiane 1 333 | 8 hors du 1er–99e centile |
| lmarena_rankdouble | 29,9 % | 402 | 1 → 402médiane 208 | 8 hors du 1er–99e centile |
| lmarena_vote_countdouble | 29,9 % | 389 | 791 → 194 909médiane 20 685 | 8 hors du 1er–99e centile |
| epoch_ecidouble | 50,9 % | 261 | 54,28 → 166,6médiane 135,21 | 6 hors du 1er–99e centile |
| epoch_eci_ci_lowdouble | 51,3 % | 293 | 24,31 → 163médiane 130,08 | 6 hors du 1er–99e centile |
| epoch_eci_ci_highdouble | 51,3 % | 294 | 69,98 → 172,03médiane 137,96 | 6 hors du 1er–99e centile |
| composite_indexdouble | 0 % | 457 | 1,19 → 97,18médiane 57,67 | 12 hors du 1er–99e centile |
| source_url_lmarenavarchar | 0 % | 2 | — |
|
| source_url_epochvarchar | 0 % | 2 | — |
|
- Actuelle
20260925T034926Z-581d6127d4c9 · sha256 581d6127d4c9…
546 lignes · premier instantané
Dirigez n’importe quel LLM vers le point d’accès des métadonnées — la documentation ci-dessus est aussi lisible par machine (JSON-LD + Croissant).
curl "https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly" | jq '{title, rows, columns_count, license}'import requests
ds = requests.get("https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly").json()
print(ds["title"], ds["rows"], "rows")
# Sample rows for an LLM context window
for row in ds.get("sample_rows", [])[:5]:
print(row)Point d’accès API : https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly
Astuce : récupérez /llms.txt pour le catalogue complet lisible par machine.
D’où viennent ces données et ce qui en a été fait. Le travail des autres apparaît sous forme de décomptes ; seuls les projets partagés sont nommés.
Citer cet instantané
Épinglé à l’instantané 20260925T034926Z-581d6127d4c9 et à son empreinte, pour que vos lecteurs obtiennent exactement les données utilisées.
LLM benchmark scoreboard (agent-curated). (2026). LLM composite capability index (monthly) [Data set, snapshot 20260925T034926Z-581d6127d4c9, sha256 581d6127d4c9]. Datazimuts. Retrieved 2026-09-25, from https://datazimuts.com/fr/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly?snapshot=20260925T034926Z-581d6127d4c9
@misc{dz_llm_benchmark_signals_llm_benchmark_comp_581d6127,
title = {{LLM composite capability index (monthly)}},
author = {{LLM benchmark scoreboard (agent-curated)}},
year = {2026},
publisher = {Datazimuts},
howpublished = {\url{https://datazimuts.com/fr/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly?snapshot=20260925T034926Z-581d6127d4c9}},
note = {Snapshot 20260925T034926Z-581d6127d4c9, sha256 581d6127d4c9f260332d9b04ffd9bc17000357e57a9e57d2c2ab9324b6c4b767; accessed 2026-09-25}
}Intégrer un tableau ou un graphique
Collez ce code dans n’importe quelle page. L’intégration est épinglée au même instantané, suit le thème clair ou sombre du lecteur et affiche toujours la source, la licence et un lien de retour.
<iframe src="https://datazimuts.com/embed/chart?dataset=llm_benchmark_signals%2Fllm_benchmark_composite_monthly&lang=fr&theme=auto&snapshot=20260925T034926Z-581d6127d4c9&x=snapshot_date&y=composite_rank&agg=avg" title="LLM composite capability index (monthly)" width="100%" height="380" style="border:0" loading="lazy"></iframe>
Posez une question sur ce jeu de données. Les réponses viennent uniquement de sa fiche, de son profil mesuré et de son historique, et citent les faits utilisés.