English
Related papers

Related papers: ChartCap: Mitigating Hallucination of Dense Chart …

200 papers

Despite continuously improving performance, contemporary image captioning models are prone to "hallucinating" objects that are not actually in a scene. One problem is that standard metrics only measure similarity to ground truth captions…

Computation and Language · Computer Science 2019-04-02 Anna Rohrbach , Lisa Anne Hendricks , Kaylee Burns , Trevor Darrell , Kate Saenko

We argue that generative text-to-image models often struggle with prompt adherence due to the noisy and unstructured nature of large-scale datasets like LAION-5B. This forces users to rely heavily on prompt engineering to elicit desirable…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Nicholas Merchant , Haitz Sáez de Ocáriz Borde , Andrei Cristian Popescu , Carlos Garcia Jurado Suarez

Multimodal large language models (MLLMs) have achieved strong performance on vision-language tasks but still struggle with fine-grained visual differences, leading to hallucinations or missed semantic shifts. We attribute this to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Tianyi Bai , Yuxuan Fan , Jiantao Qiu , Fupeng Sun , Jiayi Song , Junlin Han , Zichen Liu , Conghui He , Wentao Zhang , Binhang Yuan

Large Vision-Language Models (VLMs) now generate highly detailed, paragraphlength image captions, yet evaluating their factual accuracy remains challenging. Current methods often miss fine-grained errors, being designed for shorter texts or…

Computation and Language · Computer Science 2025-06-10 Brian Gordon , Yonatan Bitton , Andreea Marzoca , Yasumasa Onoe , Xiao Wang , Daniel Cohen-Or , Idan Szpektor

The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on manually curated reference texts for evaluation essentially forces models to mimic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Ziyun Chen , Fan Liu , Liang Yao , Chuanyi Zhang , Yuye Ma , Wei Zhou

Image captioning is the process of automatically generating a description of an image in natural language. Image captioning is one of the significant challenges in image understanding since it requires not only recognizing salient objects…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Ghadah Alabduljabbar , Hafida Benhidour , Said Kerrache

Vision language models (VLMs) show strong results on chart understanding, yet existing benchmarks assume clean figures and fact based queries. Real world charts often contain distortions and demand reasoning beyond simple matching. We…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Philip Wootaek Shin , Jack Sampson , Vijaykrishnan Narayanan , Andres Marquez , Mahantesh Halappanavar

Image captioning remains a fundamental task for vision language understanding, yet ground-truth supervision still relies predominantly on human-annotated references. Because human annotations reflect subjective preferences and expertise,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhijiang Tang , Linhua Wang , Jiaxin Qi , Weihao Jiang , Peng Hou , Anxiang Zeng , Jianqiang Huang

Recent advancements in multimodal models highlight the value of rewritten captions for improving performance, yet key challenges remain. For example, while synthetic captions often provide superior quality and image-text alignment, it is…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Zhengfeng Lai , Vasileios Saveris , Chen Chen , Hong-You Chen , Haotian Zhang , Bowen Zhang , Juan Lao Tebar , Wenze Hu , Zhe Gan , Peter Grasch , Meng Cao , Yinfei Yang

Multimodal Large Language Models (MLLMs) have shown impressive capabilities in image understanding and generation. However, current benchmarks fail to accurately evaluate the chart comprehension of MLLMs due to limited chart types and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Zhengzhuo Xu , Sinan Du , Yiyan Qi , Chengjin Xu , Chun Yuan , Jian Guo

Novel view synthesis from images, for example, with 3D Gaussian splatting, has made great progress. Rendering fidelity and speed are now ready even for demanding virtual reality applications. However, the problem of assisting humans in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Ayaka Yasunaga , Hideo Saito , Dieter Schmalstieg , Shohei Mori

The traditional image captioning task uses generic reference captions to provide textual information about images. Different user populations, however, will care about different visual aspects of images. In this paper, we propose a new…

Computation and Language · Computer Science 2020-11-10 Adam Fisch , Kenton Lee , Ming-Wei Chang , Jonathan H. Clark , Regina Barzilay

Synthetic image generation has opened up new opportunities but has also created threats in regard to privacy, authenticity, and security. Detecting fake images is of paramount importance to prevent illegal activities, and previous research…

Computer Vision and Pattern Recognition · Computer Science 2023-02-27 Md Awsafur Rahman , Bishmoy Paul , Najibul Haque Sarker , Zaber Ibn Abdul Hakim , Shaikh Anowarul Fattah

Although Multimodal Large Language Models (MLLMs) have demonstrated proficiency in video captioning, practical applications require captions that follow specific user instructions rather than generating exhaustive, unconstrained…

Figures are essential channels for densely communicating complex ideas in scientific papers. Previous work in automatically generating figure captions has been largely unsuccessful and has defaulted to using single-layer LSTMs, which no…

Computation and Language · Computer Science 2024-07-17 Stanley Cao , Kevin Liu

Image captioning is one of the straightforward tasks that can take advantage of large-scale web-crawled data which provides rich knowledge about the visual world for a captioning model. However, since web-crawled data contains image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Wooyoung Kang , Jonghwan Mun , Sungjun Lee , Byungseok Roh

Video captioning aims to generate comprehensive and coherent descriptions of the video content, contributing to the advancement of both video understanding and generation. However, existing methods often suffer from motion-detail imbalance,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Chunlin Zhong , Qiuxia Hou , Zhangjun Zhou , Shuang Hao , Haonan Lu , Yanhao Zhang , He Tang , Xiang Bai

We propose StyleCap, a method to generate natural language descriptions of speaking styles appearing in speech. Although most of conventional techniques for para-/non-linguistic information recognition focus on the category classification…

Computation and Language · Computer Science 2023-12-29 Kazuki Yamauchi , Yusuke Ijima , Yuki Saito

Visual data storytelling is gaining importance as a means of presenting data-driven information or analysis results, especially to the general public. This has resulted in design principles being proposed for data-driven storytelling, and…

Human-Computer Interaction · Computer Science 2021-05-17 Jian Zhao , Shenyu Xu , Senthil Chandrasegaran , Chris Bryan , Fan Du , Aditi Mishra , Xin Qian , Yiran Li , Kwan-Liu Ma

Image Captioning is a task that requires models to acquire a multi-modal understanding of the world and to express this understanding in natural language text. While the state-of-the-art for this task has rapidly improved in terms of n-gram…

Computer Vision and Pattern Recognition · Computer Science 2018-12-20 Annika Lindh , Robert J. Ross , Abhijit Mahalunkar , Giancarlo Salton , John D. Kelleher