Rare find

R1-VL. R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

github.com/jingyi0000/R1-VL

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.