Rare find

PSFT. [ICLR 2026] PSFT is a trust-region–inspired fine-tuning objective that views SFT as a policy gradient method with constant advantages, constraining policy drift to stabilize training and improve generalization.

github.com/zwhong714/PSFT

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.