simpleRL-reason. This is a replicate of DeepSeek-R1-Zero and DeepSeek-R1 training on small models with limited data

github.com/dontriskit/simpleRL-reason

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.