English
Related papers

Related papers: Visual Semantic Role Labeling for Video Understand…

200 papers

With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing real world accidents. Prior work has extensively studied…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Sha Luo , Yogesh Prabhu , Timothy Ossowski , Kaiping Chen , Junjie Hu

Video language continual learning involves continuously adapting to information from video and text inputs, enhancing a model's ability to handle new tasks while retaining prior knowledge. This field is a relatively under-explored area, and…

Artificial Intelligence · Computer Science 2024-12-17 Tianqi Tang , Shohreh Deldari , Hao Xue , Celso De Melo , Flora D. Salim

Existing benchmarks for evaluating long video understanding falls short on two critical aspects, either lacking in scale or quality of annotations. These limitations arise from the difficulty in collecting dense annotations for long videos,…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Aniket Agarwal , Alex Zhang , Karthik Narasimhan , Igor Gilitschenski , Vishvak Murahari , Yash Kant

State-of-the-art vision-language models (VLMs) still have limited performance in structural knowledge extraction, such as relations between objects. In this work, we present ViStruct, a training framework to learn VLMs for effective visual…

Computer Vision and Pattern Recognition · Computer Science 2023-11-23 Yangyi Chen , Xingyao Wang , Manling Li , Derek Hoiem , Heng Ji

Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Huaying Yuan , Jian Ni , Zheng Liu , Yueze Wang , Junjie Zhou , Zhengyang Liang , Bo Zhao , Zhao Cao , Zhicheng Dou , Ji-Rong Wen

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

Semantic Segmentation combines two sub-tasks: the identification of pixel-level image masks and the application of semantic labels to those masks. Recently, so-called Foundation Models have been introduced; general models trained on very…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 David Balaban , Justin Medich , Pranay Gosar , Justin Hart

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jiangtao Wu , Shihao Li , Zhaozhou Bian , Jialu Chen , Runzhe Wen , An Ping , Yiwen He , Jiakai Wang , Yuanxing Zhang , Jiaheng Liu

Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two levels: one is the content directly presented, while the other is the subtext beneath…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Qi Li , Xinchao Wang

Integrating higher level visual and linguistic interpretations is at the heart of human intelligence. As automatic visual category recognition in images is approaching human performance, the high level understanding in the dynamic…

Computer Vision and Pattern Recognition · Computer Science 2015-11-23 Anirudh Goyal , Marius Leordeanu

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription or OCR-based…

Machine Learning · Computer Science 2025-10-29 Arpita Kundu , Joyita Chakraborty , Anindita Desarkar , Aritra Sen , Srushti Anil Patil , Vishwanathan Raman

Large-scale annotated datasets allow AI systems to learn from and build upon the knowledge of the crowd. Many crowdsourcing techniques have been developed for collecting image annotations. These techniques often implicitly rely on the fact…

Human-Computer Interaction · Computer Science 2016-10-07 Gunnar A. Sigurdsson , Olga Russakovsky , Ali Farhadi , Ivan Laptev , Abhinav Gupta

Video moment search, the process of finding relevant moments in a video corpus to match a user's query, is crucial for various applications. Existing solutions, however, often assume a single perfect matching moment, struggle with…

Information Retrieval · Computer Science 2025-01-10 Chongzhi Zhang , Xizhou Zhu , Aixin Sun

Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and serves as a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Po-han Li , Shenghui Chen , Ufuk Topcu , Sandeep Chinchali

In this paper we present VideoSET, a method for Video Summary Evaluation through Text that can evaluate how well a video summary is able to retain the semantic information contained in its original video. We observe that semantics is most…

Computer Vision and Pattern Recognition · Computer Science 2014-06-24 Serena Yeung , Alireza Fathi , Li Fei-Fei

Automatically identifying harmful content in video is an important task with a wide range of applications. However, there is a lack of professionally labeled open datasets available. In this work VidHarm, an open dataset of 3589 video clips…

Computer Vision and Pattern Recognition · Computer Science 2022-09-07 Johan Edstedt , Amanda Berg , Michael Felsberg , Johan Karlsson , Francisca Benavente , Anette Novak , Gustav Grund Pihlgren

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

YouTube users looking for instructions for a specific task may spend a long time browsing content trying to find the right video that matches their needs. Creating a visual summary (abridged version of a video) provides viewers with a quick…

Computer Vision and Pattern Recognition · Computer Science 2022-08-16 Medhini Narasimhan , Arsha Nagrani , Chen Sun , Michael Rubinstein , Trevor Darrell , Anna Rohrbach , Cordelia Schmid

Building models that comprehends videos and responds specific user instructions is a practical and challenging topic, as it requires mastery of both vision understanding and knowledge reasoning. Compared to language and image modalities,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Ji Qi , Kaixuan Ji , Jifan Yu , Duokang Wang , Bin Xu , Lei Hou , Juanzi Li

Video captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level…

Computer Vision and Pattern Recognition · Computer Science 2021-08-11 Yang Bai , Junyan Wang , Yang Long , Bingzhang Hu , Yang Song , Maurice Pagnucco , Yu Guan