Rare find

Step-DPO. Implementation for "Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs"

github.com/JIA-Lab-research/Step-DPO

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.