arxiv News score 19

On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

AI Digest

多轮代理策略蒸馏课程式指导小模型优化性能提升

文章提出Guided-OPD方法,通过课程式指导减少教师干预,解决多轮代理小模型训练中误差累积问题,显著提升性能。

Guided-OPD reduces teacher intervention via curriculum guidance, addressing error accumulation in multi-turn agent distillation, achieving 21.1% Score and 25.5% Success Rate improvements.

Key points

  • 传统OPD在多轮代理训练中易因小错误累积导致教师监督失效 Traditional OPD suffers from error accumulation in multi-turn distillation
  • Guided-OPD通过课程式指导动态调整教师干预概率 Guided-OPD dynamically adjusts teacher intervention via curriculum
  • 方法在ALFWorld等基准测试中实现显著性能提升 Achieves 21.1% Score and 25.5% Success Rate gains in benchmarks
  • 逐步撤除教师监督以恢复纯策略蒸馏模式 Gradually removes teacher supervision to recover pure on-policy mode

Takeaway: Guided-OPD为小模型训练提供有效策略,平衡教师指导与自主学习。 / Guided-OPD offers effective strategy for small model training balancing supervision and autonomy.

Why it matters 为小模型训练提供新范式,可提升复杂任务处理效率。

View original ↗ Back to hot list

This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.