AccuracyParadox-RLHF. [EMNLP 2024 Main] Official implementation of the paper "The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models". (by Yanjun Chen)

github.com/EIT-NLP/AccuracyParadox-RLHF

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.