WebScraper-Cleaner. A Python script that lists all pages linked from a web address, scrapes an address specified by the user, exports that page as text, then removes script tags from the page and saves the cleaned product as another text file. BeautifulSoup and lxml are used to parse and clean the HTML and text.

github.com/ian-nai/WebScraper-Cleaner

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.