视觉语言 grounding 中的提示敏感性:细微措辞变化如何影响目标检测
计算机视觉与模式识别
2026-04-21 v1
摘要
视觉语言模型通过自然语言查询实现开词汇目标 grounding,隐式假设等价语义描述会产生一致的输出。我们使用将 DETR 用于目标提案与 CLIP 用于语言条件选择相结合的受控管道,在 263 个 COCO val2017 图像上进行测试。我们发现诸如 "a person"、"a human"、"a pedestrian" 等重叠提示经常选择不同的实例,跨六个提示的平均不稳定性为 2.11 个不同的选择。PCA 分析显示,这种变异是结构化且有方向的,而非随机的。提示集成并未提高质量,反而常常将选择倾向于通用区域。我们进一步表明,文本嵌入的接近性仅解释了 34% 的grounding 不一致 (r = -0.58),确认不稳定性源自 argmax 选择机制而非仅来自文本层面的距离。
引用
@article{arxiv.2604.17126,
title = {Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection},
author = {Dawar Jyoti Deka and Amit Sethi and Syed Mohammad Ali},
journal= {arXiv preprint arXiv:2604.17126},
year = {2026}
}
备注
5 pages, 9 figures, 1 table. Accepted at ICCAI 2026 (The 12th International Conference on Computing and Artificial Intelligence), Okinawa, Japan, April 24-27, 2026