Rare find

Guide-GRPO. Aims for memory-efficient training (24GB VRAM) on consumer GPUs. Optimizing language models through guidance tokens in reasoning chains, based on DeepSeekRL-Extended.

github.com/cnsdqd-dyb/Guide-GRPO

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.