On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents
AI Digest
多轮代理策略蒸馏课程式指导小模型优化性能提升
文章提出Guided-OPD方法,通过课程式指导减少教师干预,解决多轮代理小模型训练中误差累积问题,显著提升性能。
Guided-OPD reduces teacher intervention via curriculum guidance, addressing error accumulation in multi-turn agent distillation, achieving 21.1% Score and 25.5% Success Rate improvements.
Key points
- 传统OPD在多轮代理训练中易因小错误累积导致教师监督失效 Traditional OPD suffers from error accumulation in multi-turn distillation
- Guided-OPD通过课程式指导动态调整教师干预概率 Guided-OPD dynamically adjusts teacher intervention via curriculum
- 方法在ALFWorld等基准测试中实现显著性能提升 Achieves 21.1% Score and 25.5% Success Rate gains in benchmarks
- 逐步撤除教师监督以恢复纯策略蒸馏模式 Gradually removes teacher supervision to recover pure on-policy mode
Takeaway: Guided-OPD为小模型训练提供有效策略,平衡教师指导与自主学习。 / Guided-OPD offers effective strategy for small model training balancing supervision and autonomy.
Why it matters 为小模型训练提供新范式,可提升复杂任务处理效率。
View original ↗ Back to hot list
This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.