ARBOR. ARBOR(Adaptive Rubric Buffer for Online Reward)是一个面向 search agent RL 训练的可复用 process reward 框架。核心思路是:在 RL 训练过程中按 category 维护一份 rubric memory,从 query-group 内的对比轨迹中诱导 query-local draft rubric,经 admission → consolidation → retirement 生命周期沉淀为 category-level 的 common rubric,在 outcome-only reward 信号坍缩时仍能提供组内过程区分。

github.com/alibaba-multimodal-industrial-ai/ARBOR

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.