|
Reliability check on my own dataset's annotation layer: five machine raters, one definition, answers from 0 to 78
|
|
0
|
25
|
August 1, 2026
|
|
Same model, up to 4.66x different price — full Inference Providers pricing matrix
|
|
0
|
16
|
August 1, 2026
|
|
A free, monthly-refreshed compliance matrix of all 4,893 Indic-language datasets on the Hub, showing that 65% declare no license — so teams can avoid licensing traps (and missing-tag repos) before they train on them
|
|
0
|
16
|
August 1, 2026
|
|
cerebras/SlimPajama-627B unexpectedly returns 401/404
|
|
1
|
85
|
August 1, 2026
|
|
Organization Admin needs to urgently delete two datasets, Settings option missing
|
|
1
|
70
|
July 30, 2026
|
|
Fastdedup: Rust-based dataset deduplication — benchmarks on FineWeb sample-10BT
|
|
3
|
145
|
July 28, 2026
|
|
Training LLM model for asking questions
|
|
5
|
439
|
July 28, 2026
|
|
Why is there almost no manipulation data for agriculture?
|
|
1
|
75
|
July 28, 2026
|
|
356,371 unique french races in a database great to feed new frontier models
|
|
0
|
49
|
July 28, 2026
|
|
Claude and ChatGPT vs. a human annotator on "show, don't tell" features one definition, four very different thresholds
|
|
0
|
67
|
July 25, 2026
|
|
Huggy Database — 356,371 French horse races (1996–2026), PMU/PMH, ML-ready
|
|
0
|
10
|
July 23, 2026
|
|
Add IntelligenceLab/Long-Horizon-Terminal-Bench to the Benchmark allow-list
|
|
1
|
55
|
July 20, 2026
|
|
Request to add real5-omnidocbench framework and Benchmark allow list entry
|
|
0
|
34
|
July 20, 2026
|
|
Out-00136.safetensors seems to be corrupted with only 16bytes
|
|
1
|
43
|
July 18, 2026
|
|
Add haifan-gong/CTGroundBench to the Benchmark allow-list
|
|
1
|
48
|
July 15, 2026
|
|
Dataset Viewer API returning 503 Service Temporarily Unavailable for all datasets
|
|
3
|
274
|
July 15, 2026
|
|
Unable to load Hugging Face dataset into Kaggle notebook (Error today)
|
|
1
|
51
|
July 14, 2026
|
|
Url is not fetched from the parquet api
|
|
8
|
263
|
July 14, 2026
|
|
Would a curated dataset of ~4000 social media design layouts be useful for training or fine-tuning design models?
|
|
2
|
59
|
July 13, 2026
|
|
Add thamilvendhan/signalbench to the Benchmark allow-list
|
|
0
|
40
|
July 12, 2026
|
|
Good data to test tensor based soft sparsity and other sparsity models
|
|
5
|
74
|
July 8, 2026
|
|
The Case for an NVC-Annotated Dataset
|
|
1
|
62
|
July 7, 2026
|
|
Is selling datasets way harder than building them? Or is it just me?
|
|
1
|
82
|
July 5, 2026
|
|
What are the best practices for detecting and fetching deltas from a dataset?
|
|
3
|
59
|
July 3, 2026
|
|
Add Convence/ParseEmbed as an official benchmark on the Hub (If possible)
|
|
8
|
102
|
July 1, 2026
|
|
Trajlens: a validator for LeRobotDataset, audited 100 Hub datasets
|
|
0
|
37
|
June 30, 2026
|
|
[Concept] Instead of paying for data, we can trade data instead
|
|
1
|
81
|
June 29, 2026
|
|
[SEEKING] Indic Document Dataset (India) — Invoices, Receipts, Utility Bills, Payment Advices, Packing Lists, Commercial Invoices, Credit Notes
|
|
4
|
121
|
June 25, 2026
|
|
Dataset Viewer issue: ConfigNamesError
|
|
2
|
73
|
June 21, 2026
|
|
Follow-up: the detector reliability check, now with a second human rater + two LLMs (fresh scenes)
|
|
0
|
39
|
June 18, 2026
|