Trending datasets on Hugging Face (weekly)
Weekly popularity-velocity ranking of trending datasets on the Hugging Face Hub. Two sweeps of the official keyless Hub datasets API (https://huggingface.co/api/datasets), one sorted by likes and one by downloads (top 2 pages each), are merged and deduped on dataset id; gated, disabled and private datasets are dropped. Each dataset carries its license tag normalized to a canonical id (informational, not legal advice), its task_categories tags normalized to a coarse modality (text, vision, audio, video, multimodal, other, or unknown when no signal) via a documented priority rule set, its task categories, size category, formats and region tags extracted into dedicated columns, and author/name parts. Popularity velocity is likes and downloads per day since dataset creation (age floored at 1 day); the trend_score is a documented 0-100 composite = 50% min-max-normalized likes/day + 50% min-max-normalized downloads/day, and trend_rank orders by trend_score descending (ties: likes desc, then dataset id). Columns: ISO week, fetch timestamp, dataset id / URL / author / name, modality, task categories, size category, formats, region, normalized license id, comma-joined tags, likes, downloads, creation timestamp, age in days, likes/day, downloads/day, trend score (0-100), trend rank. Primary key: (week, dataset_id). Cadence: weekly. Nullability: task categories, size_category, formats, region and license_spdx may be empty when the author supplied none; likes, downloads, trend_score and trend_rank are never null. Caveats: likes/downloads are cumulative totals, so per-day rates are lifetime averages, not trailing-week gains; the rank therefore favors datasets with sustained momentum as well as fast risers; license_spdx comes from author-applied tags and may be missing or wrong — verify before use; popularity is not quality. Sample use: order by trend_rank for the week's hottest datasets, or filter modality = 'vision'.
- Rows
- 250
- Columns
- 21
- Source cadence
- Weekly
- Last refreshed
- Sep 25, 2026
- Theme
- technology
| Column | Type | Description |
|---|---|---|
| week | string | ISO week of the fetch (e.g. 2026-W39). (unit: ISO week) |
| fetched_at | string | — |
| dataset_id | string | — |
| dataset_url | string | Canonical public dataset page URL (https://huggingface.co/datasets/<author>/<name>); never null. (unit: url) |
| author | string | — |
| dataset_name | string | — |
| modality | string | Coarse modality: normalized from the dataset's task_categories tags via a documented priority rule set (text, vision, audio, video, multimodal, other); falls back to the author's modality: tags when no task category is set ('unknown' when neither exists). (unit: category) |
| task_categories | string | Comma-joined distinct task_categories tag values (e.g. text-generation, question-answering); empty when the author set none. (unit: category) |
| size_category | string | Hub size_categories bucket (e.g. 1M<n<10M); empty when the author set none. (unit: bucket) |
| formats | string | Comma-joined distinct format: tag values (e.g. parquet, csv); empty when none set. (unit: format) |
| region | string | Comma-joined distinct region: tag values (e.g. us); empty when none set. (unit: region) |
| license_spdx | string | License id extracted from the dataset's first license:<id> tag, with common values normalized to canonical SPDX spelling. Informational only — verify before use. Empty when the author set no license tag. (unit: license) |
| tags | string | — |
| likes | integer | — |
| downloads | integer | — |
| created_at | string | — |
| age_days | float | Days from dataset creation to fetch, floored at 1.0. (unit: days) |
| likes_per_day | float | Likes divided by age_days (age floored at 1 day). Likes are cumulative, so this is a lifetime average, not a trailing-week gain. (unit: likes/day) |
| downloads_per_day | float | Downloads divided by age_days (age floored at 1 day). Downloads are cumulative, so this is a lifetime average. (unit: downloads/day) |
| trend_score | float | Popularity-velocity composite: 50% min-max-normalized likes_per_day + 50% min-max-normalized downloads_per_day, scaled 0-100 within the snapshot. (unit: 0-100) |
| trend_rank | integer | Rank by trend_score descending (1 = hottest); ties broken by total likes, then dataset id. (unit: rank) |
First 10 sample rows — a preview, not the complete dataset.
| week | fetched_at | dataset_id | dataset_url | author | dataset_name | modality | task_categories | size_category | formats | region | license_spdx | tags | likes | downloads | created_at | age_days | likes_per_day | downloads_per_day | trend_score | trend_rank |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | secemp9/arxiv-complete | https://huggingface.co/datasets/secemp9/arxiv-complete | secemp9 | arxiv-complete | text | text-generation,text-retrieval | 100M<n<1B | parquet | us | other | arxiv,arxiv:2401.18030,format:parquet,full-text,language:en,latex,library:dask,library:datasets,library:mlcroissant,library:polars,license:other,modality:tabular,modality:text,preprints,region:us,scientific-papers,size_categories:100M<n<1B,task_categories:text-generation,task_categories:text-retrieval | 440 | 70,698 | 2026-09-10T21:32:20+0000 | 14.4 | 30.47 | 4,896.03 | 62.98 | 1 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | gfdg34fsd/newe | https://huggingface.co/datasets/gfdg34fsd/newe | gfdg34fsd | newe | unknown | — | — | — | us | — | region:us | 9 | 666,722 | 2026-08-20T22:04:01+0000 | 35.4 | 0.25 | 18,824.46 | 50.32 | 2 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | BuLei/imgbed | https://huggingface.co/datasets/BuLei/imgbed | BuLei | imgbed | image | — | n<1K | imagefolder | us | — | format:imagefolder,library:datasets,library:mlcroissant,modality:image,region:us,size_categories:n<1K | 0 | 446,637 | 2026-09-01T15:39:29+0000 | 23.7 | 0 | 18,857.46 | 50 | 3 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | SKPark1/ngii-map-full-light | https://huggingface.co/datasets/SKPark1/ngii-map-full-light | SKPark1 | ngii-map-full-light | other | other | 1M<n<10M | — | us | other | geospatial,korea,license:other,map,modality:geospatial,ngii,region:us,shp,size_categories:1M<n<10M,task_categories:other | 0 | 273,082 | 2026-09-05T00:25:49+0000 | 20.3 | 0 | 13,439.48 | 35.63 | 4 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | markov-ai/cad-1000-hours | https://huggingface.co/datasets/markov-ai/cad-1000-hours | markov-ai | cad-1000-hours | multimodal | — | — | — | us | — | modality:document,modality:video,region:us | 513 | 113,873 | 2026-08-21T10:31:39+0000 | 34.9 | 14.7 | 3,262.96 | 32.77 | 5 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | xlangai/osworld_v2_assets | https://huggingface.co/datasets/xlangai/osworld_v2_assets | xlangai | osworld_v2_assets | image | — | n<1K | imagefolder | us | — | format:imagefolder,library:datasets,library:mlcroissant,modality:image,region:us,size_categories:n<1K | 16 | 671,793 | 2026-07-23T17:49:13+0000 | 63.6 | 0.25 | 10,563.65 | 28.42 | 6 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | sjkhfuk/fofo | https://huggingface.co/datasets/sjkhfuk/fofo | sjkhfuk | fofo | unknown | — | — | — | us | — | region:us | 0 | 132,573 | 2026-09-12T17:03:49+0000 | 12.6 | 0 | 10,499.73 | 27.84 | 7 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | gfdg34fsd/sh | https://huggingface.co/datasets/gfdg34fsd/sh | gfdg34fsd | sh | unknown | — | — | — | us | — | region:us | 5 | 524,402 | 2026-07-27T19:50:22+0000 | 59.5 | 0.08 | 8,811.9 | 23.5 | 8 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | scorpionjacketguy/physics-course-vids | https://huggingface.co/datasets/scorpionjacketguy/physics-course-vids | scorpionjacketguy | physics-course-vids | multimodal | — | n<1K | — | us | — | modality:document,modality:text,region:us,size_categories:n<1K | 0 | 566,531 | 2026-07-22T05:35:06+0000 | 65.1 | 0 | 8,701.86 | 23.07 | 9 |
| 2026-W39 | 2026-09-25T08:05:43.751189+00:00 | ayuo/hd_tmp | https://huggingface.co/datasets/ayuo/hd_tmp | ayuo | hd_tmp | unknown | — | — | — | us | — | region:us | 44 | 1,448,604 | 2026-04-01T14:22:50+0000 | 176.7 | 0.25 | 8,196.33 | 22.14 | 10 |
Profiled Sep 25, 2026 from snapshot 20260925T080555Z-0a9356470b8b
Measured- Completeness
- 100%
- Rows
- 250
- Columns
- 21
- Columns with gaps
- 0
| Column | Missing | Distinct | Range | Distribution |
|---|---|---|---|---|
| weekvarchar | 0% | 1 | — |
|
| fetched_atvarchar | 0% | 1 | — |
|
| dataset_idvarchar | 0% | 270 | — |
|
| dataset_urlvarchar | 0% | 231 | — |
|
| authorvarchar | 0% | 146 | — |
|
| dataset_namevarchar | 0% | 274 | — |
|
| modalityvarchar | 0% | 12 | — |
|
| task_categoriesvarchar | 0% | 39 | — |
|
| size_categoryvarchar | 0% | 11 | — |
|
| formatsvarchar | 0% | 9 | — |
|
| regionvarchar | 0% | 1 | — |
|
| license_spdxvarchar | 0% | 20 | — |
|
| tagsvarchar | 0% | 215 | — |
|
| likesbigint | 0% | 197 | 0 → 9,845median 284 | 3 outside 1st–99th percentile |
| downloadsbigint | 0% | 250 | 75 → 3,355,072median 178,025 | 6 outside 1st–99th percentile |
| created_atvarchar | 0% | 252 | — |
|
| age_daysdouble | 0% | 256 | 12.6 → 1,667median 557.1 | 3 outside 1st–99th percentile |
| likes_per_daydouble | 0% | 124 | 0 → 30.47median 0.39 | 3 outside 1st–99th percentile |
| downloads_per_daydouble | 0% | 213 | 0.19 → 18,857median 292.64 | 6 outside 1st–99th percentile |
| trend_scoredouble | 0% | 209 | 0.68 → 62.98median 1.95 | 6 outside 1st–99th percentile |
| trend_rankbigint | 0% | 265 | 1 → 250median 125.5 | 6 outside 1st–99th percentile |
Latest change
20260925T080513Z-38b12845bdea → 20260925T080555Z-0a9356470b8b
- Rows now
- 250 (+0)
- New rows
- 250
- Removed rows
- 250
- Unchanged rows
- 0
Rows are compared as whole records over the columns both versions share; an edited row counts as one removed and one new.
Same columns and types as the previous version.
- Current
20260925T080555Z-0a9356470b8b · sha256 0a9356470b8b…
250 rows · +0 rows vs previous
20260925T080513Z-38b12845bdea · sha256 38b12845bdea…
250 rows · first snapshot
Point any LLM at the metadata endpoint — the documentation above is machine-readable too (JSON-LD + Croissant).
curl "https://datazimuts.com/v1/datasets/hf_dataset_trending_signals/hf_trending_datasets_weekly" | jq '{title, rows, columns_count, license}'import requests
ds = requests.get("https://datazimuts.com/v1/datasets/hf_dataset_trending_signals/hf_trending_datasets_weekly").json()
print(ds["title"], ds["rows"], "rows")
# Sample rows for an LLM context window
for row in ds.get("sample_rows", [])[:5]:
print(row)API endpoint: https://datazimuts.com/v1/datasets/hf_dataset_trending_signals/hf_trending_datasets_weekly
Tip: fetch /llms.txt for the full machine-readable catalog.
Where this data comes from and what was made from it. Other people's work shows as counts; only shared projects are named.
Cite this snapshot
Pinned to snapshot 20260925T080555Z-0a9356470b8b and its content hash, so readers get exactly the data you used.
Hugging Face trending datasets (agent-curated). (2026). Trending datasets on Hugging Face (weekly) [Data set, snapshot 20260925T080555Z-0a9356470b8b, sha256 0a9356470b8b]. Datazimuts. Retrieved 2026-09-25, from https://datazimuts.com/en/datasets/hf_dataset_trending_signals/hf_trending_datasets_weekly?snapshot=20260925T080555Z-0a9356470b8b
@misc{dz_hf_dataset_trending_signals_hf_trending__0a935647,
title = {{Trending datasets on Hugging Face (weekly)}},
author = {{Hugging Face trending datasets (agent-curated)}},
year = {2026},
publisher = {Datazimuts},
howpublished = {\url{https://datazimuts.com/en/datasets/hf_dataset_trending_signals/hf_trending_datasets_weekly?snapshot=20260925T080555Z-0a9356470b8b}},
note = {Snapshot 20260925T080555Z-0a9356470b8b, sha256 0a9356470b8be2c970e921fdeaa571fb1876e255c74387c62f70cc4cd56b638f; accessed 2026-09-25}
}Embed a table or a chart
Paste this into any page. The embed is pinned to the same snapshot, follows the reader's light or dark setting, and always shows the source, license and a link back.
<iframe src="https://datazimuts.com/embed/chart?dataset=hf_dataset_trending_signals%2Fhf_trending_datasets_weekly&lang=en&theme=auto&snapshot=20260925T080555Z-0a9356470b8b&x=week&y=likes&agg=avg" title="Trending datasets on Hugging Face (weekly)" width="100%" height="380" style="border:0" loading="lazy"></iframe>
Ask about this dataset. Answers come only from its catalog record, measured profile and change history, and list the facts they used.