ConVis:基于对比解码与幻觉可视化的多模态大语言模型幻觉消除方法
材料科学
2025-10-10 v1
摘要
多模态大语言模型(MLLMs)中出现的幻觉现象——即生成的响应未能准确反映给定图像——是其可靠性面临的重大挑战。为此,我们提出ConVis,一种新型无训练对比解码方法。ConVis利用文本到图像(T2I)生成模型从幻觉描述语义重构给定图像。通过比较原始和重构图像所产生的对比概率分布,ConVis使MLLMs能够捕获惩罚幻觉生成的视觉对比信号。值得注意的是,该方法完全基于解码过程,不需要额外数据或模型更新。我们在五个流行基准测试上进行了广泛实验,证明ConVis在 various MLLMs中有效减少幻觉,凸显其提升模型可靠性的潜力。
引用
@article{arxiv.2408.13905,
title = {Circularly polarised electroluminescence from chiral excitons in vacuum-sublimed supramolecular semiconductor thin films},
author = {Rituparno Chowdhury and Marco D. Preuss and Hwan-Hee Cho and Joshua J. P. Thompson and Samarpita Sen and Tomi Baikie and Pratyush Ghosh and Yorrick Boeije and Xian-Wei Chua and Kai-Wei Chang and Erjuan Guo and Joost van der Tol and Bart W. L. van den Bersselaar and Andrea Taddeucci and Nicolas Daub and Daphne M. Dekker and Scott T. Keene and Ghislaine Vantomme and Bruno Ehrler and Stefan C. J. Meskers and Akshay Rao and Bartomeu Monserrat and E. W. Meijer and Richard H. Friend},
journal= {arXiv preprint arXiv:2408.13905},
year = {2025}
}