arxiv 资讯 热度 20

Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark

AI 导读

动态执行推理SWE-Flux基准LLMs表现差异输入扰动生成

研究提出SWE-Flux基准测试,评估LLMs动态执行推理能力,发现模型在复杂场景下表现不佳。

This paper introduces SWE-Flux, a dynamic execution benchmark showing LLMs struggle with complex runtime reasoning tasks.

重点速览

  • 提出SWE-Flux基准测试,包含真实仓库的动态执行实例 Introduces SWE-Flux benchmark with real-world dynamic execution instances
  • 评估结果显示LLMs在跨过程执行等场景准确率仅37% LLMs achieve 37% accuracy in complex runtime reasoning tasks
  • 通过输入扰动生成90%实例的挑战性变体 Generates challenging variants for 90% of instances via input perturbation
  • 模型在局部控制流等任务表现优于数据流分析 Models perform better on localized control flow than dataflow analysis

一句话:LLMs在动态执行推理中存在显著局限,需更复杂基准测试 / LLMs show significant limitations in dynamic execution reasoning tasks

推荐理由 提供新基准测试,揭示LLMs在复杂代码推理中的局限性,对研究者有启发

查看原文 ↗ 返回热榜

二次创作声明:本页为 AI 热榜聚合导读,内容与热度数据来自公开来源 (arxiv),版权归原始作者所有;本站仅做转载指引与摘要评述,不复制原文。