LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles
Abstract
Figure captions are crucial for helping readers understand and remember a figure's key message. Many models have been developed to generate these captions, helping authors compose better quality captions more easily. Yet, authors almost always need to revise generic AI-generated captions to match their writing style and the domain's style, highlighting the need for personalization. Despite language models' personalization (LaMP) advances, these technologies often focus on text-only settings and rarely address scenarios where both inputs and profiles are multimodal. This paper introduces LaMP-Cap, a dataset for personalized figure caption generation with multimodal figure profiles. For each target figure, LaMP-Cap provides not only the needed inputs, such as figure images, but also up to three other figures from the same document--each with its image, caption, and figure-mentioning paragraphs--as a profile to characterize the context. Experiments with four LLMs show that using profile information consistently helps generate captions closer to the original author-written ones. Ablation studies reveal that images in the profile are more helpful than figure-mentioning paragraphs, highlighting the advantage of using multimodal profiles over text-only ones.
Keywords
Cite
@article{arxiv.2506.06561,
title = {LaMP-Cap: Personalized Figure Caption Generation With Multimodal Figure Profiles},
author = {Ho Yin 'Sam' Ng and Ting-Yao Hsu and Aashish Anantha Ramakrishnan and Branislav Kveton and Nedim Lipka and Franck Dernoncourt and Dongwon Lee and Tong Yu and Sungchul Kim and Ryan A. Rossi and Ting-Hao 'Kenneth' Huang},
journal= {arXiv preprint arXiv:2506.06561},
year = {2025}
}
Comments
Accepted to EMNLP 2025 Findings. The LaMP-CAP dataset is publicly available at: https://github.com/Crowd-AI-Lab/lamp-cap