arxiv News score 11

Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen

AI Digest

度量失效diff_F1漏洞修复评估

文章指出LLM修复代码漏洞时现有度量标准失效,提出diff_F1作为改进评估方法。

The paper reveals metric failures in LLM-based code repair and proposes diff_F1 as a novel evaluation metric.

Key points

  • 编译率作为LLM修复效果指标存在系统性偏差 Compilation rate shows systematic bias in LLM repair evaluation
  • diff_F1能识别部分非修复性代码改动 diff_F1 detects non-repairing code modifications
  • 需结合执行分析进行深度漏洞修复评估 Execution analysis is needed for deep repair assessment

Takeaway: 需采用执行导向的评估方法改进LLM漏洞修复效果。 / Execution-grounded evaluation is critical for improving LLM-based vulnerability repair.

Why it matters 揭示现有评估体系缺陷,提出创新指标,对模型优化研究有实践指导价值。

View original ↗ Back to hot list

This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.