LLM benchmark scoreboard (agent-curated)
Monthly cross-leaderboard scoreboard of large language models: every (model, benchmark) score from LMArena's official public leaderboard dataset (Bradley-Terry ratings, text/overall arena, CC-BY-4.0) and Epoch AI's Capabilities & Benchmarking data hub (56 current capability benchmarks, CC-BY), rescaled per benchmark to a comparable 0-100 (min-max within the snapshot; all benchmarks are higher-is-better) with per-benchmark ranks. Models are entity-resolved across the two sources (exact name match, then one arena-variant suffix stripped, unambiguous only; tier recorded per row). Superseded Epoch benchmarks are pruned. Columns: snapshot date (day-granular UTC), source, benchmark, benchmark release date, resolved model key, raw upstream model id, display name, resolution tier, organization, ISO alpha-3 country, raw score, score unit, normalized 0-100 score, rank in benchmark, models in benchmark, vote count (LMArena), source publish date, source URL. Primary key: (snapshot_date, source, benchmark, source_model_id). Cadence: monthly; same-day re-runs are content-hash no-ops. Caveats: the 0-100 scale is within-benchmark relative, not absolute; LMArena measures human preference, Epoch benchmarks measure task accuracy — the composite dataset blends them explicitly. Sample use: filter benchmark = 'GPQA diamond' order by rank_in_benchmark for the current reasoning-benchmark ranking.
- ai-research
- machine-learning
- technology
- signals