English
Related papers

Related papers: GOAL: A Challenging Knowledge-grounded Video Capti…

200 papers

Despite the recent emergence of video captioning models, how to generate the text description with specific entity names and fine-grained actions is far from being solved, which however has great applications such as basketball live text…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Zeyu Xi , Ge Shi , Xuefen Li , Junchi Yan , Zun Li , Lifang Wu , Zilin Liu , Liang Wang

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric, overlooking the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Peiran Wu , Yunze Liu , Zhengdong Zhu , Enmin Zhou , Junxiao Shen

Recently, significant advances have been made in Video Large Language Models (Video LLMs) in both academia and industry. However, methods to evaluate and benchmark the performance of different Video LLMs, especially their fine-grained,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Kuangzhi Ge , Lingjun Chen , Kevin Zhang , Yulin Luo , Tianyu Shi , Liaoyuan Fan , Xiang Li , Guanqun Wang , Shanghang Zhang

Recent multimodal large language models (MLLMs) have shown strong capabilities in general video understanding, driving growing interest in automatic sports commentary generation. However, existing benchmarks for this task focus exclusively…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Kaiwen Wang , Kaili Zheng , Rongrong Deng , Yiming Shi , Chenyi Guo , Ji Wu

Video captioning targets interpreting the complex visual contents as text descriptions, which requires the model to fully understand video scenes including objects and their interactions. Prevailing methods adopt off-the-shelf object…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

Cognitive science has shown that humans perceive videos in terms of events separated by the state changes of dominant subjects. State changes trigger new events and are one of the most useful among the large amount of redundant information…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Yuxuan Wang , Difei Gao , Licheng Yu , Stan Weixian Lei , Matt Feiszli , Mike Zheng Shou

Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling object interactions.…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Boxiao Pan , Haoye Cai , De-An Huang , Kuan-Hui Lee , Adrien Gaidon , Ehsan Adeli , Juan Carlos Niebles

Metaphors are a common communication tool used in our day-to-day life. The detection and generation of metaphors in textual form have been studied extensively but metaphors in other forms have been under-explored. Recent studies have shown…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Abisek Rajakumar Kalarani , Pushpak Bhattacharyya , Sumit Shekhar

Soccer is one of the most popular sport worldwide, with live broadcasts frequently available for major matches. However, extracting detailed, frame-by-frame information on player actions from these videos remains a challenge. Utilizing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Shikun Xu , Yandong Zhu , Gen Li , Changhu Wang

Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Xin Wang , Wenhu Chen , Jiawei Wu , Yuan-Fang Wang , William Yang Wang

Video captioning aims to describe the content of videos using natural language. Although significant progress has been made, there is still much room to improve the performance for real-world applications, mainly due to the long-tail words…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Xin Gu , Guang Chen , Yufei Wang , Libo Zhang , Tiejian Luo , Longyin Wen

Tracking objects in soccer videos is extremely important to gather both player and team statistics, whether it is to estimate the total distance run, the ball possession or the team formation. Video processing can help automating the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Anthony Cioppa , Silvio Giancola , Adrien Deliege , Le Kang , Xin Zhou , Zhiyu Cheng , Bernard Ghanem , Marc Van Droogenbroeck

Enhancing the diversity of sentences to describe video contents is an important problem arising in recent video captioning research. In this paper, we explore this problem from a novel perspective of customizing video captions by imitating…

Computer Vision and Pattern Recognition · Computer Science 2021-12-03 Yitian Yuan , Lin Ma , Wenwu Zhu

Text-to-video (T2V) models have shown remarkable performance in generating visually reasonable scenes, while their capability to leverage world knowledge for ensuring semantic consistency and factual accuracy remains largely understudied.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yubin Chen , Xuyang Guo , Zhenmei Shi , Zhao Song , Jiahao Zhang

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Matthew Gwilliam , Michael Cogswell , Meng Ye , Karan Sikka , Abhinav Shrivastava , Ajay Divakaran

Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal grounding and video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Ling You , Wenxuan Huang , Xinni Xie , Xiangyi Wei , Bangyan Li , Shaohui Lin , Yang Li , Changbo Wang

Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory and visual contents are equally important. Although audio…

Computer Vision and Pattern Recognition · Computer Science 2018-12-10 Yapeng Tian , Chenxiao Guan , Justin Goodman , Marc Moore , Chenliang Xu

Generating informative and knowledge-rich image captions remains a challenge for many existing captioning models, which often produce generic descriptions that lack specificity and contextual depth. To address this limitation, we propose…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Reem AlJunaid , Muzammil Behzad

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Andong Wang , Bo Wu , Sunli Chen , Zhenfang Chen , Haotian Guan , Wei-Ning Lee , Li Erran Li , Chuang Gan

Understanding video content and generating caption with context is an important and challenging task. Unlike prior methods that typically attempt to generate generic video captions without context, our architecture contextualizes captioning…

Computer Vision and Pattern Recognition · Computer Science 2020-07-30 Philipp Rimle , Pelin Dogan , Markus Gross