Alpha-RL. On Predictability of Reinforcement Learning Dynamics for Large Language Models (ICLR 2026)
EffOPD. Repository for EffOPD. We are working on polishing the details.
AlphaRL. Python