Aller au contenu

Données ouvertes

Bibliothèque

Jeux de données ouverts, entièrement documentés — interrogeables ici, et lisibles par n’importe quel LLM.

Les titres et les descriptions proviennent des sources de données, en anglais.

  • AI research vocabulary index (agent-curated)

    AI research vocabulary index (weekly)

    Weekly index of the technical vocabulary driving AI research, mined from the titles and abstracts of every AI paper announced on arXiv in the trailing 7 complete days (cs.AI, cs.CL, cs.CV, cs.LG via the official keyless arXiv query API, 3 s between requests per the API Terms of Use). Deterministic pipeline: LaTeX stripped, hyphenated/spaced spellings merged, plural stemming, stopword + domain-filler screening, document-frequency filter (term must appear in >= 3 papers and <= 50% of the weekly corpus). Each term is scored by TF-IDF-style significance ln(N / df) * mean(1 + ln(tf)) over containing papers, min-max normalized 0-100, and ranked (significance desc, term_id asc). Each term carries a majority-vote AI subtopic (agents, reasoning, evals, llm, finetuning, rl, quantization, multimodal, vision, audio, nlp, robotics, ml-theory, data, education, other), up to 5 sample arXiv ids, and an example paper title. Columns: ISO week, day-granular fetch timestamp, deterministic term_id, display term, token count, paper count, paper share, significance score, emergence rank, AI subtopic, sample arXiv ids, example paper title, first-seen week, week-over-week velocity (attached automatically from the second snapshot on, when a prior snapshot exists in the hub; absent from the baseline snapshot). Primary key: (week, term_id). Cadence: weekly. Nullability: velocity_pct is absent from the first snapshot; every emitted column is otherwise never null. Caveats: subtopic is a keyword-rule majority vote, not a classifier; stemming is a small deterministic rule set, not a lemmatizer; significance measures vocabulary distinctiveness, not research quality. Only descriptive arXiv metadata is mined (CC0 1.0), so commercial_use = yes. Sample use: order by emergence_rank for the week's most distinctive technical vocabulary, or filter ai_subtopic = 'agents'.

    • machine-learning
    • ai-research
    • arxiv
    • technology
    lignes
    11 379
    Qualité
    100
    Mis à jour
    25 sept. 2026
    À jour
    Licence
    Usage commercial OK

Utiliser votre propre clé d’IA

Une fois l’allocation gratuite du jour épuisée, les fonctions d’IA peuvent passer par votre propre compte fournisseur.

Conservée uniquement dans cet onglet (effacée à sa fermeture) et envoyée avec chaque requête d’IA. Nos serveurs l’utilisent pour cette requête et ne la stockent ni ne la journalisent jamais.