reward-hacking-detector. Reward-Hacking Detection via RL-Post-Training of LLMs - Detecting reward hacking in RL agents using language models as trajectory auditors

github.com/mohammed840/reward-hacking-detector

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.