TinyLoRA-GRPO-Coder. Inspired by 《Learning to Reason in 13 parameters》, use TinyLoRA+GRPO(32 parameters) to fine-tune Qwen2.5-Coder-3B-Instruct(or other models) to accomplish competitive programming.

github.com/Chi-Shan0707/TinyLoRA-GRPO-Coder

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.