Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
AI 导读
度量失效diff_F1漏洞修复评估
文章指出LLM修复代码漏洞时现有度量标准失效,提出diff_F1作为改进评估方法。
The paper reveals metric failures in LLM-based code repair and proposes diff_F1 as a novel evaluation metric.
重点速览
- 编译率作为LLM修复效果指标存在系统性偏差 Compilation rate shows systematic bias in LLM repair evaluation
- diff_F1能识别部分非修复性代码改动 diff_F1 detects non-repairing code modifications
- 需结合执行分析进行深度漏洞修复评估 Execution analysis is needed for deep repair assessment
一句话:需采用执行导向的评估方法改进LLM漏洞修复效果。 / Execution-grounded evaluation is critical for improving LLM-based vulnerability repair.
推荐理由 揭示现有评估体系缺陷,提出创新指标,对模型优化研究有实践指导价值。
二次创作声明:本页为 AI 热榜聚合导读,内容与热度数据来自公开来源 (arxiv),版权归原始作者所有;本站仅做转载指引与摘要评述,不复制原文。