Skip to content

Open data

Data library

Open datasets, fully documented — searchable here, and readable by any LLM.

  • AI research vocabulary index (agent-curated)

    AI research vocabulary index (weekly)

    Weekly index of the technical vocabulary driving AI research, mined from the titles and abstracts of every AI paper announced on arXiv in the trailing 7 complete days (cs.AI, cs.CL, cs.CV, cs.LG via the official keyless arXiv query API, 3 s between requests per the API Terms of Use). Deterministic pipeline: LaTeX stripped, hyphenated/spaced spellings merged, plural stemming, stopword + domain-filler screening, document-frequency filter (term must appear in >= 3 papers and <= 50% of the weekly corpus). Each term is scored by TF-IDF-style significance ln(N / df) * mean(1 + ln(tf)) over containing papers, min-max normalized 0-100, and ranked (significance desc, term_id asc). Each term carries a majority-vote AI subtopic (agents, reasoning, evals, llm, finetuning, rl, quantization, multimodal, vision, audio, nlp, robotics, ml-theory, data, education, other), up to 5 sample arXiv ids, and an example paper title. Columns: ISO week, day-granular fetch timestamp, deterministic term_id, display term, token count, paper count, paper share, significance score, emergence rank, AI subtopic, sample arXiv ids, example paper title, first-seen week, week-over-week velocity (attached automatically from the second snapshot on, when a prior snapshot exists in the hub; absent from the baseline snapshot). Primary key: (week, term_id). Cadence: weekly. Nullability: velocity_pct is absent from the first snapshot; every emitted column is otherwise never null. Caveats: subtopic is a keyword-rule majority vote, not a classifier; stemming is a small deterministic rule set, not a lemmatizer; significance measures vocabulary distinctiveness, not research quality. Only descriptive arXiv metadata is mined (CC0 1.0), so commercial_use = yes. Sample use: order by emergence_rank for the week's most distinctive technical vocabulary, or filter ai_subtopic = 'agents'.

    • machine-learning
    • ai-research
    • arxiv
    • technology
    rows
    11,379
    Quality
    100
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK

Use your own AI key

Once today's free allowance is used up, AI features can run on your own provider account.

Kept in this browser tab only (cleared when you close it) and sent with each AI request. Our servers use it for that request and never store or log it.