nlp-crawler. A horizontally scalable web crawling engine designed for structured content extraction. It performs URL normalization, HTML parsing, Markdown conversion, and LLM post-processing (using xai_sdk with grok). Output is serialized in line-delimited JSON (.jsonl).

github.com/dylan-sutton-chavez/nlp-crawler

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.