signal新智元zh source2026-09-21
肉眼看不出的幻觉?清华提出视觉源幻觉,仅用0.9%数据实现SOTA
清华大学提出“视觉源幻觉”机制,指出短输出场景下幻觉源于视觉特征提取错位而非仅语言先验,并开发对抗对比微调方法ACFT。该方法仅用COCO数据集0.9%数据、零推理开销,在LLaVA、MiniGPT-4、Qwen2.5-VL上于POPE等基准达SOTA,较基线提升最高5.3%,且能修正图文嵌入错位与注意力分布异常。
- why now
- ACM MM 2026新论文用0.9%数据提升幻觉检测
- topic
- AI Tech & New Models
- source
- 新智元
score
score 8 out of 10. 0-10: how dense the facts are, multiplied by how much you can do with them after reading. 8+ means the topic's evidence bar is met: benchmarks and availability for a new model, amount and investors for a funding round, revenue figures for a solo-money story. Below 5 an item does not enter the digest. A press release scores 3 or less, a reprint loses 2, anything older than 14 days loses 1, a headline that misleads loses 3.
read the source