arxiv News score 13

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

AI Digest

能力误表前沿差距评估透明度模型披露

学术AI评估中存在能力误表问题,论文所用模型普遍落后于前沿模型,差距持续扩大,且披露信息不足。

Academic AI evaluations often use outdated models, with capability gaps widening annually and insufficient disclosure.

Key points

  • 学术论文评估的AI模型普遍落后于前沿模型,中位数差距达+10.85 ECI Median evaluated models lag behind frontier LLMs by +10.85 ECI
  • 能力差距以每年+5.53 ECI的速度持续扩大,与评估日期无关 Capability gap widens at +5.53 ECI/year regardless of evaluation date
  • 仅18.4%论文明确标注评估日期,52.5%摘要仅笼统提及'AI'而非具体模型 Only 18.4% papers explicitly state evaluation dates
  • 推理模型使用情况披露率低,仅3.2%摘要和21.2%正文提及模式状态 Low disclosure rates for reasoning model specifics

Takeaway: 学术AI评估存在系统性能力误表,需建立标准化披露机制 / Systemic capability misrepresentation in academic AI evaluations requires standardized disclosure protocols

Why it matters 揭示学术研究中的评估偏差,为改进AI研究可信度提供关键参考

View original ↗ Back to hot list

This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.