Skip to content

Open data

Data library

Open datasets, fully documented — searchable here, and readable by any LLM.

  • LLM benchmark scoreboard (agent-curated)

    LLM composite capability index (monthly)

    Monthly per-model composite of the cross-leaderboard scoreboard: for every model appearing in LMArena's official public leaderboard (text/overall, CC-BY-4.0) or Epoch AI's Capabilities & Benchmarking hub (CC-BY), the composite index is the equal-weighted mean of its normalized 0-100 score components (LMArena Bradley-Terry rating rescaled + each Epoch benchmark rescaled), with n_components recorded so single-signal models are distinguishable from broadly measured ones. Epoch's own Capabilities Index (ECI) is carried as a reference column, never a composite input. Columns: snapshot date, composite rank, resolved model key, display name, resolution tier, organization, ISO alpha-3 country, accessibility, component count, Epoch benchmarks covered, LMArena variants merged, Epoch mean normalized, LMArena normalized, LMArena rating/rank/votes, Epoch ECI + CI, composite index, source URLs. Primary key: (snapshot_date, model_key). Cadence: monthly. Caveats: a transparency-first average, not a capability claim; benchmarks differ in difficulty and LMArena measures human preference. Sample use: order by composite_rank for the current cross-source model ranking, or filter n_components >= 5 for broadly-measured models only.

    • ai-research
    • machine-learning
    • technology
    • signals
    rows
    546
    Quality
    96
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK
  • LLM benchmark scoreboard (agent-curated)

    LLM benchmark scores (monthly)

    Monthly cross-leaderboard scoreboard of large language models: every (model, benchmark) score from LMArena's official public leaderboard dataset (Bradley-Terry ratings, text/overall arena, CC-BY-4.0) and Epoch AI's Capabilities & Benchmarking data hub (56 current capability benchmarks, CC-BY), rescaled per benchmark to a comparable 0-100 (min-max within the snapshot; all benchmarks are higher-is-better) with per-benchmark ranks. Models are entity-resolved across the two sources (exact name match, then one arena-variant suffix stripped, unambiguous only; tier recorded per row). Superseded Epoch benchmarks are pruned. Columns: snapshot date (day-granular UTC), source, benchmark, benchmark release date, resolved model key, raw upstream model id, display name, resolution tier, organization, ISO alpha-3 country, raw score, score unit, normalized 0-100 score, rank in benchmark, models in benchmark, vote count (LMArena), source publish date, source URL. Primary key: (snapshot_date, source, benchmark, source_model_id). Cadence: monthly; same-day re-runs are content-hash no-ops. Caveats: the 0-100 scale is within-benchmark relative, not absolute; LMArena measures human preference, Epoch benchmarks measure task accuracy — the composite dataset blends them explicitly. Sample use: filter benchmark = 'GPQA diamond' order by rank_in_benchmark for the current reasoning-benchmark ranking.

    • ai-research
    • machine-learning
    • technology
    • signals
    rows
    3,121
    Quality
    99
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK

Use your own AI key

Once today's free allowance is used up, AI features can run on your own provider account.

Kept in this browser tab only (cleared when you close it) and sent with each AI request. Our servers use it for that request and never store or log it.