RONA:基于连贯关系的语用多样化图像描述
计算与语言
2025-06-10 v2 人工智能
计算机视觉与模式识别
摘要
写作助手(如 Grammarly、Microsoft Copilot)传统上通过采用句法和语义变化来描述图像组成部分,从而生成多样化的图像描述。然而,人类撰写的描述优先使用语用线索,在传达视觉描述的同时传递核心信息。为了增强描述的多样性,探索结合视觉内容传达这些信息的替代方式至关重要。我们提出了 RONA,一种用于多模态大语言模型(MLLM)的新型提示策略,利用连贯关系作为语用变化的可控轴。我们证明,与多个领域的 MLLM 基线相比,RONA 生成的描述在整体多样性和与真实值的对齐度上均表现更优。我们的代码可在 https://github.com/aashish2000/RONA 获取。
引用
@article{arxiv.2503.10997,
title = {RONA: Pragmatically Diverse Image Captioning with Coherence Relations},
author = {Aashish Anantha Ramakrishnan and Aadarsh Anantha Ramakrishnan and Dongwon Lee},
journal= {arXiv preprint arXiv:2503.10997},
year = {2025}
}
备注
Accepted in the NAACL Fourth Workshop on Intelligent and Interactive Writing Assistants (In2Writing), Albuquerque, New Mexico, May 2025, https://in2writing.glitch.me