Skip to content

Open data

Data library

Open datasets, fully documented — searchable here, and readable by any LLM.

3 datasets

stackoverflow
  • Stack Overflow data/ML question intelligence (agent-curated)

    Stack Overflow data/ML question intelligence (weekly)

    Every distinct Stack Overflow question surfaced by per-tag sweeps of a curated ~38-tag data/ML universe over the trailing 7 complete UTC days, deduplicated on question_id (first sweep wins as primary_tag; all observed tags unioned into tags). The official keyless Stack Exchange /questions API (sort=creation, fixed fromdate/todate window) supplies the raw material; the transform derives day-granular ages from the as-of stamp, scores each question with a documented 0-100 attention composite (60% min-max-normalized log10(1+views_per_day) + 40% min-max-normalized non-negative score_per_day), ranks it as attention_rank (ties: views desc, question_id asc), and flags unanswered_pain_flag when a question has zero answers with at least 2x the snapshot's median view count (visible community demand, no resolution). Titles + metadata only (CC BY-SA 4.0, attribution required) — no bodies, comments, or user data. Primary key: (week, question_id). Cadence: weekly. Nullability: scores and ranks are never null; age is at least 1 day. Caveats: counts tick intraday, so same-day re-runs refresh values; attention scores are within-snapshot relative (recalibrated every week). Sample use: order by attention_rank for the week's hottest data/ML questions, or filter unanswered_pain_flag for under-served demand.

    • technology
    • signals
    • weekly
    • stackoverflow
    rows
    17
    Quality
    100
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK
  • Stack Overflow data/ML question intelligence (agent-curated)

    Stack Overflow data/ML tag heat (weekly)

    Per-tag aggregation of the same weekly data/ML question universe as so_questions_weekly: for each of the ~38 curated sweep tags, the questions carrying it in the trailing 7 complete UTC days, with questions_7d, total_views, answer_rate (answered / total), median_views_per_day, unanswered_count, and pain_count (questions with zero answers at >= 2x the snapshot median views). heat_score is a documented 0-100 composite: 50% min-max-normalized log10(1 + questions_7d) (volume) + 30% min-max-normalized log10(1 + median_views_per_day) (engagement) + 20% min-max-normalized pain_share (unmet demand), ranked as heat_rank (ties: questions_7d desc, tag_name asc). Tags are classified with the shared curated 14-category taxonomy (same as so_tag_signals). Primary key: (week, tag_name). Cadence: weekly. Nullability: scores and ranks never null. Caveats: questions carrying several swept tags count toward each tag's totals; heat scores are within-snapshot relative. Sample use: order by heat_rank for the week's hottest data/ML topics.

    • technology
    • signals
    • weekly
    • stackoverflow
    rows
    38
    Quality
    97
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK
  • Stack Overflow tag interest (agent-curated)

    Stack Overflow tag interest velocity (weekly)

    Weekly developer-interest ranking of the 500 most-asked tags on Stack Overflow. The official keyless Stack Exchange /tags API (sort=popular) supplies the tag universe — top 500 by total question count across 5 polite pages, deduped on tag name — and the /tags/{tags}/synonyms endpoint resolves every synonym-bearing tag to its canonical master with its known aliases (entity resolution). Each tag is classified into a curated 14-category taxonomy (ai-ml, data, languages, web-frontend, web-backend, mobile, devops-cloud, systems, security, testing, game, media, desktop, core-concepts, other) by a deterministic ordered rule set with exact-name pins for short tags (c, r, go, ...). The interest_score is a documented 0-100 composite: 50% min-max-normalized log10(1 + new_questions_7d) (volume) + 50% min-max-normalized week-over-week growth rate clipped to 0..+1000% (growth); interest_rank orders by interest_score desc (ties: new_questions_7d desc, then tag name). breakout_flag marks tags growing >= 50% with >= 20 new questions. Baseline snapshot (first run) is volume-only — growth columns attach from the second weekly snapshot on. fetched_at is the trailing complete UTC day, so a same-day re-run refreshes the snapshot; growth always compares different weeks. Columns: ISO week, fetch date, tag name / Stack Overflow URL / category / total question count / share of the top-500 total synonym count and alias list / volume score / (from snapshot 2: prior total, new questions in 7d, growth rate %, growth score, breakout flag) / interest score / interest rank / deterministic row hash. Primary key: (week, tag_name). Cadence: weekly. Nullability: growth columns are absent from the baseline snapshot; scores and ranks are never null. Caveats: upstream counts exclude deleted questions, so small negative deltas are clamped to zero; counts tick intraday, so same-day re-runs refresh values; the taxonomy is a first-match heuristic, not an authoritative classification. Sample use: order by interest_rank for the week's fastest-rising developer topics, or filter category = 'ai-ml'.

    • technology
    • signals
    • weekly
    • stackoverflow
    rows
    496
    Quality
    100
    Updated
    Sep 25, 2026
    Fresh
    License
    Commercial use OK

Use your own AI key

Once today's free allowance is used up, AI features can run on your own provider account.

Kept in this browser tab only (cleared when you close it) and sent with each AI request. Our servers use it for that request and never store or log it.