Rare find

terminal-bench-rl. GRPO training code which scales to 32xH100s for long horizon terminal/coding tasks. Base agent is now the top Qwen3 agent on Stanford's TerminalBench leaderboard.

github.com/Danau5tin/terminal-bench-rl

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.