Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen
AI Digest
度量失效diff_F1漏洞修复评估
文章指出LLM修复代码漏洞时现有度量标准失效,提出diff_F1作为改进评估方法。
The paper reveals metric failures in LLM-based code repair and proposes diff_F1 as a novel evaluation metric.
Key points
- 编译率作为LLM修复效果指标存在系统性偏差 Compilation rate shows systematic bias in LLM repair evaluation
- diff_F1能识别部分非修复性代码改动 diff_F1 detects non-repairing code modifications
- 需结合执行分析进行深度漏洞修复评估 Execution analysis is needed for deep repair assessment
Takeaway: 需采用执行导向的评估方法改进LLM漏洞修复效果。 / Execution-grounded evaluation is critical for improving LLM-based vulnerability repair.
Why it matters 揭示现有评估体系缺陷,提出创新指标,对模型优化研究有实践指导价值。
View original ↗ Back to hot list
This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.