Skip to content

AI research vocabulary index (weekly)

Commercial use OKFreshLicense: CC0-1.0
Query in workbench

Follow this dataset

Get a notice in your feed when a new snapshot is published. Optionally, we also POST it to your webhook.

Must be a public https address. We never follow redirects.

Sample onlyDownload sample CSVDownload sample JSONSample rows only (up to 20) — not the complete dataset.

Weekly index of the technical vocabulary driving AI research, mined from the titles and abstracts of every AI paper announced on arXiv in the trailing 7 complete days (cs.AI, cs.CL, cs.CV, cs.LG via the official keyless arXiv query API, 3 s between requests per the API Terms of Use). Deterministic pipeline: LaTeX stripped, hyphenated/spaced spellings merged, plural stemming, stopword + domain-filler screening, document-frequency filter (term must appear in >= 3 papers and <= 50% of the weekly corpus). Each term is scored by TF-IDF-style significance ln(N / df) * mean(1 + ln(tf)) over containing papers, min-max normalized 0-100, and ranked (significance desc, term_id asc). Each term carries a majority-vote AI subtopic (agents, reasoning, evals, llm, finetuning, rl, quantization, multimodal, vision, audio, nlp, robotics, ml-theory, data, education, other), up to 5 sample arXiv ids, and an example paper title. Columns: ISO week, day-granular fetch timestamp, deterministic term_id, display term, token count, paper count, paper share, significance score, emergence rank, AI subtopic, sample arXiv ids, example paper title, first-seen week, week-over-week velocity (attached automatically from the second snapshot on, when a prior snapshot exists in the hub; absent from the baseline snapshot). Primary key: (week, term_id). Cadence: weekly. Nullability: velocity_pct is absent from the first snapshot; every emitted column is otherwise never null. Caveats: subtopic is a keyword-rule majority vote, not a classifier; stemming is a small deterministic rule set, not a lemmatizer; significance measures vocabulary distinctiveness, not research quality. Only descriptive arXiv metadata is mined (CC0 1.0), so commercial_use = yes. Sample use: order by emergence_rank for the week's most distinctive technical vocabulary, or filter ai_subtopic = 'agents'.

Rows
11,379
Columns
13
Source cadence
Weekly
Last refreshed
Sep 25, 2026
Theme
technology

Use your own AI key

Once today's free allowance is used up, AI features can run on your own provider account.

Kept in this browser tab only (cleared when you close it) and sent with each AI request. Our servers use it for that request and never store or log it.