reward-hacking-oss20b. Reward hacking detection probes for language models - testing bias across multiple domains with structured findings output.

github.com/TakSec/reward-hacking-oss20b

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.