Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark
AI Digest
动态执行推理SWE-Flux基准LLMs表现差异输入扰动生成
研究提出SWE-Flux基准测试,评估LLMs动态执行推理能力,发现模型在复杂场景下表现不佳。
This paper introduces SWE-Flux, a dynamic execution benchmark showing LLMs struggle with complex runtime reasoning tasks.
Key points
- 提出SWE-Flux基准测试,包含真实仓库的动态执行实例 Introduces SWE-Flux benchmark with real-world dynamic execution instances
- 评估结果显示LLMs在跨过程执行等场景准确率仅37% LLMs achieve 37% accuracy in complex runtime reasoning tasks
- 通过输入扰动生成90%实例的挑战性变体 Generates challenging variants for 90% of instances via input perturbation
- 模型在局部控制流等任务表现优于数据流分析 Models perform better on localized control flow than dataflow analysis
Takeaway: LLMs在动态执行推理中存在显著局限,需更复杂基准测试 / LLMs show significant limitations in dynamic execution reasoning tasks
Why it matters 提供新基准测试,揭示LLMs在复杂代码推理中的局限性,对研究者有启发
View original ↗ Back to hot list
This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.