Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
AI Digest
Growing Harness任务反馈代码重用部署成本失败修复
该研究提出Growing Harness方法,通过任务反馈将重复控制转化为可重用代码,减少LLM调用并提升效率,实验显示其在多个基准测试中表现最优。
This paper introduces Growing Harness, a method that transforms recurring control into reusable code via task feedback, reducing LLM calls and improving efficiency across benchmarks.
Key points
- 通过任务反馈将控制决策转化为可执行代码,降低LLM调用频率 Converts control decisions into executable code via task feedback to reduce LLM calls
- 实验显示成功率在多数基准测试中领先,部署成本降低超74% Achieves highest success rates in most benchmarks with 74%+ cost reduction
- 代码重用机制使小模型也能保持44.7%的高成功率 Code reuse maintains 44.7% success across small models
- 失败引导训练框架能联合修复多个错误并回滚有害修改 Failure-guided training jointly fixes errors and rolls back harmful changes
- 相比传统工具调用方式,LLM调用减少76-91.8% Reduces LLM calls by 76-91.8% vs traditional tool calling
Takeaway: Growing Harness通过代码重用显著提升LLM代理效率,为小模型应用提供新范式。 / Growing Harness enables efficient LLM agents through code reuse, offering new paradigms for small model applications.
Why it matters 该方法创新性地将控制逻辑转化为代码,为LLM代理优化提供可落地的解决方案。
View original ↗ Back to hot list
This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.