arxiv News score 19

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

AI Digest

临床遗漏基准测试框架依赖隐瞒模型表现差异数据缺失问题

IatroBench基准测试揭示语言模型在医疗场景中存在框架依赖的信息隐瞒现象,不同提问方式导致模型分享信息量差异显著。

IatroBench evaluates language models' clinical omission risks, revealing framing-dependent information withholding across medical scenarios.

Key points

  • IatroBench评估模型在60个预注册临床场景中的信息遗漏风险 IatroBench evaluates clinical omission risks across 60 pre-registered scenarios
  • 不同框架(患者查询/医生咨询)导致模型分享信息量差异显著 Information sharing varies significantly between patient query and doctor consultation frameworks
  • Claude Opus 4.6的隐瞒评分与医生评分高度一致 Claude Opus 4.6's omission scores align closely with physician assessments
  • Llama 4在两种框架下均表现不佳,GPT-5.2数据缺失率达33.2% Llama 4 performs poorly in both framing contexts
  • 标准LLM法官对遗漏危害的评分与结构化评估存在分歧 Standard LLM judges differ from structured evaluations on omission harm

Takeaway: 模型在医疗场景中存在框架依赖的信息隐瞒问题 / Language models exhibit framing-dependent information withholding in clinical contexts

Why it matters 揭示模型在医疗场景中的信息控制机制,对安全应用有指导意义

View original ↗ Back to hot list

This page is an aggregated digest from arxiv; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.