Minimal-GRPO. Fine-tuning Open Language Models (like LlaMa, Qwen) for Tasks with verifiable rewards using Group Relative Policy Optimization (GRPO) or Evolutionary Strategy (ES)

github.com/Bharath2/Minimal-GRPO

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.