LLM composite capability index (monthly)
Monthly per-model composite of the cross-leaderboard scoreboard: for every model appearing in LMArena's official public leaderboard (text/overall, CC-BY-4.0) or Epoch AI's Capabilities & Benchmarking hub (CC-BY), the composite index is the equal-weighted mean of its normalized 0-100 score components (LMArena Bradley-Terry rating rescaled + each Epoch benchmark rescaled), with n_components recorded so single-signal models are distinguishable from broadly measured ones. Epoch's own Capabilities Index (ECI) is carried as a reference column, never a composite input. Columns: snapshot date, composite rank, resolved model key, display name, resolution tier, organization, ISO alpha-3 country, accessibility, component count, Epoch benchmarks covered, LMArena variants merged, Epoch mean normalized, LMArena normalized, LMArena rating/rank/votes, Epoch ECI + CI, composite index, source URLs. Primary key: (snapshot_date, model_key). Cadence: monthly. Caveats: a transparency-first average, not a capability claim; benchmarks differ in difficulty and LMArena measures human preference. Sample use: order by composite_rank for the current cross-source model ranking, or filter n_components >= 5 for broadly-measured models only.
- Rows
- 546
- Columns
- 22
- Source cadence
- Monthly
- Last refreshed
- Sep 25, 2026
- Theme
- technology
| Column | Type | Description |
|---|---|---|
| snapshot_date | string | Fetch date at UTC day granularity. Same-day re-runs produce identical snapshots; any upstream change yields a new content-hashed snapshot. (unit: ISO date) |
| composite_rank | integer | Rank by composite_index descending (1 = highest). (unit: rank) |
| model_key | string | Resolved canonical model key (separator-stripped lowercase name): the cross-source join key. See resolution_tier for how the join was made. |
| display_name | string | Human-readable model name; Epoch AI's display name preferred when the model is resolved, else the LMArena id. |
| resolution_tier | string | Entity-resolution tier: 'exact' (identical normalized names), 'variant' (one arena-variant suffix stripped, unambiguous only), 'epoch_only', 'lmarena_only'. |
| organization | string | Model organization as reported by the source (Epoch ECI metadata preferred). |
| country_code | string | ISO alpha-3 country of the organization (Epoch AI metadata only; empty for LMArena-only models). (unit: ISO alpha-3) |
| accessibility | string | Model accessibility from Epoch AI metadata (e.g. 'API access'); empty for LMarena-only models. |
| n_components | integer | Number of normalized score components averaged into composite_index (LMArena rating + Epoch benchmarks). (unit: count) |
| n_epoch_benchmarks | integer | Epoch AI benchmarks covering the model. (unit: count) |
| n_lmarena_variants | integer | LMArena arena variants merged into this canonical model (best-rating variant represents it). (unit: count) |
| epoch_mean_normalized | float | Mean of the model's normalized Epoch benchmark scores (0-100). (unit: 0-100) |
| lmarena_normalized | float | LMArena Bradley-Terry rating rescaled 0-100 within the snapshot. (unit: 0-100) |
| lmarena_rating | float | Raw LMArena Bradley-Terry rating of the best variant. (unit: bradley-terry rating) |
| lmarena_rank | float | LMArena overall arena rank of the best variant. (unit: rank) |
| lmarena_vote_count | float | LMArena votes behind the best variant's rating. (unit: count) |
| epoch_eci | float | Epoch AI Capabilities Index for the model (reference column; never a composite input). (unit: index) |
| epoch_eci_ci_low | float | Lower bound of the ECI confidence interval. (unit: index) |
| epoch_eci_ci_high | float | Upper bound of the ECI confidence interval. (unit: index) |
| composite_index | float | Equal-weighted mean of the model's available normalized 0-100 components (LMArena + Epoch benchmarks). A transparency-first average, not a capability claim. (unit: 0-100) |
| source_url_lmarena | string | LMArena dataset URL (empty when the model has no LMArena coverage). (unit: url) |
| source_url_epoch | string | Epoch AI benchmark_data.zip URL (empty when the model has no Epoch coverage). (unit: url) |
First 10 sample rows — a preview, not the complete dataset.
| snapshot_date | composite_rank | model_key | display_name | resolution_tier | organization | country_code | accessibility | n_components | n_epoch_benchmarks | n_lmarena_variants | epoch_mean_normalized | lmarena_normalized | lmarena_rating | lmarena_rank | lmarena_vote_count | epoch_eci | epoch_eci_ci_low | epoch_eci_ci_high | composite_index | source_url_lmarena | source_url_epoch |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-09-25 | 1 | gpt6astra | GPT-6 Astra | variant | OpenAI | USA | API access | 22 | 21 | 1 | 97.5 | 90.53 | 1,443.724 | 59 | 2,693 | 166.6 | 163 | 172.03 | 97.18 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | 2 | qwen35maxpreview | qwen3.5-max-preview | lmarena_only | alibaba | — | — | 1 | 0 | 1 | — | 94.56 | 1,470.932 | 25 | 21,476 | — | — | — | 94.56 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 3 | ernie51 | ernie-5.1 | lmarena_only | baidu | — | — | 1 | 0 | 1 | — | 94.16 | 1,468.195 | 28 | 37,058 | — | — | — | 94.16 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 4 | mimov25pro | mimo-v2.5-pro | lmarena_only | xiaomi | — | — | 1 | 0 | 1 | — | 93.65 | 1,464.749 | 32 | 60,919 | — | — | — | 93.65 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 5 | gemini25pro | gemini-2.5-pro | lmarena_only | — | — | 1 | 0 | 1 | — | 92.61 | 1,457.772 | 36 | 122,554 | — | — | — | 92.61 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — | |
| 2026-09-25 | 6 | gpt55pro | GPT-5.5 Pro | epoch_only | OpenAI | USA | API access | 10 | 10 | 0 | 91.58 | — | — | — | — | 162.45 | 159.23 | 166.64 | 91.58 | — | https://epoch.ai/data/benchmark_data.zip |
| 2026-09-25 | 7 | grok420beta0309reasoning | grok-4.20-beta-0309-reasoning | lmarena_only | xai | — | — | 1 | 0 | 1 | — | 91.57 | 1,450.74 | 42 | 62,168 | — | — | — | 91.57 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 8 | claudeopus4520251101 | claude-opus-4-5-20251101 | lmarena_only | anthropic | — | — | 1 | 0 | 1 | — | 91.53 | 1,450.459 | 44 | 70,013 | — | — | — | 91.53 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 9 | grok420multiagentbeta0309 | grok-4.20-multi-agent-beta-0309 | lmarena_only | xai | — | — | 1 | 0 | 1 | — | 91.45 | 1,449.939 | 46 | 60,777 | — | — | — | 91.45 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | — |
| 2026-09-25 | 10 | claudefable51 | Claude Fable 5.1 | variant | Anthropic | USA | API access | 19 | 18 | 1 | 90.9 | 100 | 1,507.582 | 1 | 5,783 | 165 | 161.64 | 169.61 | 91.38 | https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset | https://epoch.ai/data/benchmark_data.zip |
Profiled Sep 25, 2026 from snapshot 20260925T034926Z-581d6127d4c9
Measured- Completeness
- 85.3%
- Rows
- 546
- Columns
- 22
- Columns with gaps
- 8
| Column | Missing | Distinct | Range | Distribution |
|---|---|---|---|---|
| snapshot_datevarchar | 0% | 1 | — |
|
| composite_rankbigint | 0% | 607 | 1 → 546median 273.5 | 12 outside 1st–99th percentile |
| model_keyvarchar | 0% | 582 | — |
|
| display_namevarchar | 0% | 578 | — |
|
| resolution_tiervarchar | 0% | 4 | — |
|
| organizationvarchar | 0% | 64 | — |
|
| country_codevarchar | 0% | 6 | — |
|
| accessibilityvarchar | 0% | 9 | — |
|
| n_componentsbigint | 0% | 29 | 1 → 31median 1 | 5 outside 1st–99th percentile |
| n_epoch_benchmarksbigint | 0% | 27 | 0 → 30median 0 | 4 outside 1st–99th percentile |
| n_lmarena_variantsbigint | 0% | 3 | 0 → 2median 1 | |
| epoch_mean_normalizeddouble | 50.9% | 258 | 0.12 → 97.5median 45.8 | 6 outside 1st–99th percentile |
| lmarena_normalizeddouble | 29.9% | 400 | 0 → 100median 74.09 | 8 outside 1st–99th percentile |
| lmarena_ratingdouble | 29.9% | 365 | 833.41 → 1,508median 1,333 | 8 outside 1st–99th percentile |
| lmarena_rankdouble | 29.9% | 402 | 1 → 402median 208 | 8 outside 1st–99th percentile |
| lmarena_vote_countdouble | 29.9% | 389 | 791 → 194,909median 20,685 | 8 outside 1st–99th percentile |
| epoch_ecidouble | 50.9% | 261 | 54.28 → 166.6median 135.21 | 6 outside 1st–99th percentile |
| epoch_eci_ci_lowdouble | 51.3% | 293 | 24.31 → 163median 130.08 | 6 outside 1st–99th percentile |
| epoch_eci_ci_highdouble | 51.3% | 294 | 69.98 → 172.03median 137.96 | 6 outside 1st–99th percentile |
| composite_indexdouble | 0% | 457 | 1.19 → 97.18median 57.67 | 12 outside 1st–99th percentile |
| source_url_lmarenavarchar | 0% | 2 | — |
|
| source_url_epochvarchar | 0% | 2 | — |
|
- Current
20260925T034926Z-581d6127d4c9 · sha256 581d6127d4c9…
546 rows · first snapshot
Point any LLM at the metadata endpoint — the documentation above is machine-readable too (JSON-LD + Croissant).
curl "https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly" | jq '{title, rows, columns_count, license}'import requests
ds = requests.get("https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly").json()
print(ds["title"], ds["rows"], "rows")
# Sample rows for an LLM context window
for row in ds.get("sample_rows", [])[:5]:
print(row)API endpoint: https://datazimuts.com/v1/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly
Tip: fetch /llms.txt for the full machine-readable catalog.
Where this data comes from and what was made from it. Other people's work shows as counts; only shared projects are named.
Cite this snapshot
Pinned to snapshot 20260925T034926Z-581d6127d4c9 and its content hash, so readers get exactly the data you used.
LLM benchmark scoreboard (agent-curated). (2026). LLM composite capability index (monthly) [Data set, snapshot 20260925T034926Z-581d6127d4c9, sha256 581d6127d4c9]. Datazimuts. Retrieved 2026-09-25, from https://datazimuts.com/en/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly?snapshot=20260925T034926Z-581d6127d4c9
@misc{dz_llm_benchmark_signals_llm_benchmark_comp_581d6127,
title = {{LLM composite capability index (monthly)}},
author = {{LLM benchmark scoreboard (agent-curated)}},
year = {2026},
publisher = {Datazimuts},
howpublished = {\url{https://datazimuts.com/en/datasets/llm_benchmark_signals/llm_benchmark_composite_monthly?snapshot=20260925T034926Z-581d6127d4c9}},
note = {Snapshot 20260925T034926Z-581d6127d4c9, sha256 581d6127d4c9f260332d9b04ffd9bc17000357e57a9e57d2c2ab9324b6c4b767; accessed 2026-09-25}
}Embed a table or a chart
Paste this into any page. The embed is pinned to the same snapshot, follows the reader's light or dark setting, and always shows the source, license and a link back.
<iframe src="https://datazimuts.com/embed/chart?dataset=llm_benchmark_signals%2Fllm_benchmark_composite_monthly&lang=en&theme=auto&snapshot=20260925T034926Z-581d6127d4c9&x=snapshot_date&y=composite_rank&agg=avg" title="LLM composite capability index (monthly)" width="100%" height="380" style="border:0" loading="lazy"></iframe>
Ask about this dataset. Answers come only from its catalog record, measured profile and change history, and list the facts they used.