arxiv 资讯 热度 11

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

AI 导读

度量失效diff_F1漏洞修复评估

文章指出LLM修复代码漏洞时现有度量标准失效,提出diff_F1作为改进评估方法。

The paper reveals metric failures in LLM-based code repair and proposes diff_F1 as a novel evaluation metric.

重点速览

  • 编译率作为LLM修复效果指标存在系统性偏差 Compilation rate shows systematic bias in LLM repair evaluation
  • diff_F1能识别部分非修复性代码改动 diff_F1 detects non-repairing code modifications
  • 需结合执行分析进行深度漏洞修复评估 Execution analysis is needed for deep repair assessment

一句话:需采用执行导向的评估方法改进LLM漏洞修复效果。 / Execution-grounded evaluation is critical for improving LLM-based vulnerability repair.

推荐理由 揭示现有评估体系缺陷,提出创新指标,对模型优化研究有实践指导价值。

查看原文 ↗ 返回热榜

二次创作声明:本页为 AI 热榜聚合导读,内容与热度数据来自公开来源 (arxiv),版权归原始作者所有;本站仅做转载指引与摘要评述,不复制原文。