Rare find

understanding-rlhf. Learning from preferences is a common paradigm for fine-tuning language models. Yet, many algorithmic design decisions come into play. Our new work finds that approaches employing on-policy sampling or negative gradients outperform offline, maximum likelihood objectives.

github.com/Asap7772/understanding-rlhf

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.