English

TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors

Computer Vision and Pattern Recognition 2026-01-07 v1 Artificial Intelligence Machine Learning

Abstract

Dense video captioning aims to interpret and describe all temporally localized events throughout an input video. Recent state-of-the-art methods leverage large language models (LLMs) to provide detailed moment descriptions for video data. However, existing VideoLLMs remain challenging in identifying precise event boundaries in untrimmed videos, causing the generated captions to be not properly grounded. In this paper, we propose TA-Prompting, which enhances VideoLLMs via Temporal Anchors that learn to precisely localize events and prompt the VideoLLMs to perform temporal-aware video event understanding. During inference, in order to properly determine the output caption sequence from an arbitrary number of events presented within a video, we introduce an event coherent sampling strategy to select event captions with sufficient coherence across temporal events and cross-modal similarity with the given video. Through extensive experiments on benchmark datasets, we show that our TA-Prompting is favorable against state-of-the-art VideoLLMs, yielding superior performance on dense video captioning and temporal understanding tasks including moment retrieval and temporalQA.

Keywords

Cite

@article{arxiv.2601.02908,
  title  = {TA-Prompting: Enhancing Video Large Language Models for Dense Video Captioning via Temporal Anchors},
  author = {Wei-Yuan Cheng and Kai-Po Chang and Chi-Pin Huang and Fu-En Yang and Yu-Chiang Frank Wang},
  journal= {arXiv preprint arXiv:2601.02908},
  year   = {2026}
}

Comments

8 pages for main paper (exclude citation pages), 6 pages for appendix, totally 10 figures 7 tables and 2 algorithms. The paper is accepted by WACV 2026