中文

MC-LLaVA:多概念个性化视觉语言模型

计算机视觉与模式识别 2025-03-26 v2 人工智能

摘要

当前视觉语言模型(VLM)在 diverse 任务中表现出卓越的能力,如视觉问答。为提升用户体验,最近的研究探索 VLM 个性化以理解用户提供的概念。然而,这些方法主要关注单概念个性化,忽略了多个概念的存在及其相互作用,这限制了实际应用的可行性。本文提出首个多概念个性化范式,MC-LLaVA。具体而言,MC-LLaVA 采用多概念指令调谋策略,在单个训练步骤中有效整合多个概念。为减少联合训练相关成本,我们提出一种使用视觉 token 信息初始化概念 token 的个性化文本提示。此外,我们在推理过程中引入个性化视觉提示,聚合位置置信度图以增强识别和 grounding 能力。为推动多概念个性化研究,我们进一步贡献了一个高质量的指令调谋数据集。我们仔细从电影中收集具有多个角色和物体的图像,并手动生成用于多概念情景的问答样本,具有优异的多样性。全面的定性和定量实验表明,MC-LLaVA 能够实现惊人的多概念个性化响应,为 VLM 成为更好的用户专属助手铺平了道路。代码和数据集将在 https://github.com/arctanxarc/MC-LLaVA} 上公开。

关键词

引用

@article{arxiv.2503.18854,
  title  = {MC-LLaVA: Multi-Concept Personalized Vision-Language Model},
  author = {Ruichuan An and Sihan Yang and Ming Lu and Renrui Zhang and Kai Zeng and Yulin Luo and Jiajun Cao and Hao Liang and Ying Chen and Qi She and Shanghang Zhang and Wentao Zhang},
  journal= {arXiv preprint arXiv:2503.18854},
  year   = {2025}
}

备注

I sincerely apologize for any inconvenience caused. We actually uploaded this paper to arXiv in November 2024, as arXiv:2411.11706. During this update, we did not consider the replacement operation of arXiv, which led to duplicate submissions. We have made modifications at the original address arXiv:2411.11706