web100k. A compact, no‑frills dataset of 100,000 real homepage HTML documents from popular domains. It’s meant for benchmarking / fuzzing / robustness testing of HTML parsers, link extractors, readability algorithms, ML preprocessing pipelines, etc.

github.com/EmilStenstrom/web100k

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.