This is your work, valued
uncalibrated_reasoning. Code repository for "Uncalibrated Reasoning: GRPO Induces Overconfidence for Stochastic Outcomes"