Aller au contenu

Données ouvertes

Bibliothèque

Jeux de données ouverts, entièrement documentés — interrogeables ici, et lisibles par n’importe quel LLM.

Les titres et les descriptions proviennent des sources de données, en anglais.

  • LLM benchmark scoreboard (agent-curated)

    LLM composite capability index (monthly)

    Monthly per-model composite of the cross-leaderboard scoreboard: for every model appearing in LMArena's official public leaderboard (text/overall, CC-BY-4.0) or Epoch AI's Capabilities & Benchmarking hub (CC-BY), the composite index is the equal-weighted mean of its normalized 0-100 score components (LMArena Bradley-Terry rating rescaled + each Epoch benchmark rescaled), with n_components recorded so single-signal models are distinguishable from broadly measured ones. Epoch's own Capabilities Index (ECI) is carried as a reference column, never a composite input. Columns: snapshot date, composite rank, resolved model key, display name, resolution tier, organization, ISO alpha-3 country, accessibility, component count, Epoch benchmarks covered, LMArena variants merged, Epoch mean normalized, LMArena normalized, LMArena rating/rank/votes, Epoch ECI + CI, composite index, source URLs. Primary key: (snapshot_date, model_key). Cadence: monthly. Caveats: a transparency-first average, not a capability claim; benchmarks differ in difficulty and LMArena measures human preference. Sample use: order by composite_rank for the current cross-source model ranking, or filter n_components >= 5 for broadly-measured models only.

    • ai-research
    • machine-learning
    • technology
    • signals
    lignes
    546
    Qualité
    96
    Mis à jour
    25 sept. 2026
    À jour
    Licence
    Usage commercial OK
  • LLM benchmark scoreboard (agent-curated)

    LLM benchmark scores (monthly)

    Monthly cross-leaderboard scoreboard of large language models: every (model, benchmark) score from LMArena's official public leaderboard dataset (Bradley-Terry ratings, text/overall arena, CC-BY-4.0) and Epoch AI's Capabilities & Benchmarking data hub (56 current capability benchmarks, CC-BY), rescaled per benchmark to a comparable 0-100 (min-max within the snapshot; all benchmarks are higher-is-better) with per-benchmark ranks. Models are entity-resolved across the two sources (exact name match, then one arena-variant suffix stripped, unambiguous only; tier recorded per row). Superseded Epoch benchmarks are pruned. Columns: snapshot date (day-granular UTC), source, benchmark, benchmark release date, resolved model key, raw upstream model id, display name, resolution tier, organization, ISO alpha-3 country, raw score, score unit, normalized 0-100 score, rank in benchmark, models in benchmark, vote count (LMArena), source publish date, source URL. Primary key: (snapshot_date, source, benchmark, source_model_id). Cadence: monthly; same-day re-runs are content-hash no-ops. Caveats: the 0-100 scale is within-benchmark relative, not absolute; LMArena measures human preference, Epoch benchmarks measure task accuracy — the composite dataset blends them explicitly. Sample use: filter benchmark = 'GPQA diamond' order by rank_in_benchmark for the current reasoning-benchmark ranking.

    • ai-research
    • machine-learning
    • technology
    • signals
    lignes
    3 121
    Qualité
    99
    Mis à jour
    25 sept. 2026
    À jour
    Licence
    Usage commercial OK

Utiliser votre propre clé d’IA

Une fois l’allocation gratuite du jour épuisée, les fonctions d’IA peuvent passer par votre propre compte fournisseur.

Conservée uniquement dans cet onglet (effacée à sa fermeture) et envoyée avec chaque requête d’IA. Nos serveurs l’utilisent pour cette requête et ne la stockent ni ne la journalisent jamais.