防御arXiv 2609.33221
RMB:用奖励模型 boosting 缓解 RLHF 中的奖励作弊
提出 RMB,以多样奖励模型 boosting 缓解 RLHF 奖励作弊
- arXiv
- 2609.33221
- 发表
- 场景
- 模型与 API
- 测试
- Mistral-7B-Instruct-v0.2、Gemma-2B-IT、DeepSeek-R1 等 7 个
摘要RMB 用 HSIC 多样奖励模型加决策树 boosting,提高偏好准确率并缓解 reward hacking。
RMB 是一种奖励建模方法:用多样的代理奖励模型加 boosting 聚合,为 RLHF 提供更稳的奖励信号。。