|
[SEEKING] Indic Document Dataset (India) — Invoices, Receipts, Utility Bills, Payment Advices, Packing Lists, Commercial Invoices, Credit Notes
|
|
5
|
129
|
August 7, 2026
|
|
Is selling datasets way harder than building them? Or is it just me?
|
|
2
|
89
|
August 7, 2026
|
|
A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on the Hub, showing that 65% declare no license — so teams can avoid licensing traps (and missing-tag repos) before they train on them
|
|
1
|
23
|
August 7, 2026
|
|
Reliability check on my own dataset's annotation layer: five machine raters, one definition, answers from 0 to 78
|
|
0
|
34
|
August 1, 2026
|
|
Same model, up to 4.66x different price — full Inference Providers pricing matrix
|
|
0
|
21
|
August 1, 2026
|
|
cerebras/SlimPajama-627B unexpectedly returns 401/404
|
|
1
|
94
|
August 1, 2026
|
|
Organization Admin needs to urgently delete two datasets, Settings option missing
|
|
1
|
74
|
July 30, 2026
|
|
Fastdedup: Rust-based dataset deduplication — benchmarks on FineWeb sample-10BT
|
|
3
|
152
|
July 28, 2026
|
|
Training LLM model for asking questions
|
|
5
|
442
|
July 28, 2026
|
|
Why is there almost no manipulation data for agriculture?
|
|
1
|
79
|
July 28, 2026
|
|
356,371 unique french races in a database great to feed new frontier models
|
|
0
|
54
|
July 28, 2026
|
|
Claude and ChatGPT vs. a human annotator on "show, don't tell" features one definition, four very different thresholds
|
|
0
|
69
|
July 25, 2026
|
|
Huggy Database — 356,371 French horse races (1996–2026), PMU/PMH, ML-ready
|
|
0
|
16
|
July 23, 2026
|
|
Add IntelligenceLab/Long-Horizon-Terminal-Bench to the Benchmark allow-list
|
|
1
|
57
|
July 20, 2026
|
|
Request to add real5-omnidocbench framework and Benchmark allow list entry
|
|
0
|
36
|
July 20, 2026
|
|
Out-00136.safetensors seems to be corrupted with only 16bytes
|
|
1
|
47
|
July 18, 2026
|
|
Add haifan-gong/CTGroundBench to the Benchmark allow-list
|
|
1
|
51
|
July 15, 2026
|
|
Dataset Viewer API returning 503 Service Temporarily Unavailable for all datasets
|
|
3
|
280
|
July 15, 2026
|
|
Unable to load Hugging Face dataset into Kaggle notebook (Error today)
|
|
1
|
52
|
July 14, 2026
|
|
Url is not fetched from the parquet api
|
|
8
|
268
|
July 14, 2026
|
|
Would a curated dataset of ~4000 social media design layouts be useful for training or fine-tuning design models?
|
|
2
|
62
|
July 13, 2026
|
|
Add thamilvendhan/signalbench to the Benchmark allow-list
|
|
0
|
41
|
July 12, 2026
|
|
Good data to test tensor based soft sparsity and other sparsity models
|
|
5
|
76
|
July 8, 2026
|
|
The Case for an NVC-Annotated Dataset
|
|
1
|
65
|
July 7, 2026
|
|
What are the best practices for detecting and fetching deltas from a dataset?
|
|
3
|
61
|
July 3, 2026
|
|
Add Convence/ParseEmbed as an official benchmark on the Hub (If possible)
|
|
8
|
108
|
July 1, 2026
|
|
Trajlens: a validator for LeRobotDataset, audited 100 Hub datasets
|
|
0
|
40
|
June 30, 2026
|
|
[Concept] Instead of paying for data, we can trade data instead
|
|
1
|
82
|
June 29, 2026
|
|
Dataset Viewer issue: ConfigNamesError
|
|
2
|
74
|
June 21, 2026
|
|
Follow-up: the detector reliability check, now with a second human rater + two LLMs (fresh scenes)
|
|
0
|
41
|
June 18, 2026
|