datasetting. Dataset txt files curated and edited by yours truly, primarily for finetuning GPT-Neo family models. Primarily obtained from wikimedia, fandom dumps and other public contributions. Every document is licensed under the GFDL. Most fandoms operate under CC BY-SA 3.0 so citation for both licenses is provided.

github.com/valahraban/datasetting

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.