English
Related papers

Related papers: Explicit Temporal-Semantic Modeling for Dense Vide…

200 papers

Grounding temporal video segments described in natural language queries effectively and efficiently is a crucial capability needed in vision-and-language fields. In this paper, we deal with the fast video temporal grounding (FVTG) task,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-13 Ziyue Wu , Junyu Gao , Shucheng Huang , Changsheng Xu

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

Computer Vision and Pattern Recognition · Computer Science 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. While prior work on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Jiangtao Wu , Shihao Li , Zhaozhou Bian , Jialu Chen , Runzhe Wen , An Ping , Yiwen He , Jiakai Wang , Yuanxing Zhang , Jiaheng Liu

With the rapid growth of video data on the internet, video summarization is becoming a very important AI technology. However, due to the high labelling cost of video summarization, existing studies have to be conducted on small-scale…

Multimedia · Computer Science 2026-01-13 Cairong Zhao , Chutian Wang , Zifan Song , Guosheng Hu , Haonan Chen , Xiaofan Zhai

Significant progress has been made on visual captioning, largely relying on pre-trained features and later fixed object detectors that serve as rich inputs to auto-regressive models. A key limitation of such methods, however, is that the…

Computer Vision and Pattern Recognition · Computer Science 2022-06-09 Chia-Wen Kuo , Zsolt Kira

Automatically describing videos with natural language is a fundamental challenge for computer vision and natural language processing. Recently, progress in this problem has been achieved through two steps: 1) employing 2-D and/or 3-D…

Computer Vision and Pattern Recognition · Computer Science 2022-02-23 Yuyu Guo , Jingqiu Zhang , Lianli Gao

Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each…

Computer Vision and Pattern Recognition · Computer Science 2022-09-19 Wanrong Zhu , Bo Pang , Ashish V. Thapliyal , William Yang Wang , Radu Soricut

Video grounding aims to locate a moment of interest matching the given query sentence from an untrimmed video. Previous works ignore the {sparsity dilemma} in video annotations, which fails to provide the context information between…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Hongxiang Li , Meng Cao , Xuxin Cheng , Zhihong Zhu , Yaowei Li , Yuexian Zou

The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal videos contain rich and complex dynamic audio-visual components,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-31 Guangyao Li , Henghui Du , Di Hu

Large pre-trained multimodal models have demonstrated significant success in a range of downstream tasks, including image captioning, image-text retrieval, visual question answering (VQA), etc. However, many of these methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Zikang Liu , Sihan Chen , Longteng Guo , Handong Li , Xingjian He , Jing Liu

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models'…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Asmar Nadeem , Faegheh Sardari , Robert Dawes , Syed Sameed Husain , Adrian Hilton , Armin Mustafa

Time awareness is a fundamental ability of omni large language models, especially for understanding long videos and answering complex questions. Previous approaches mainly target vision-language scenarios and focus on the explicit temporal…

Computation and Language · Computer Science 2025-12-11 Yijing Chen , Yihan Wu , Kaisi Guan , Yuchen Ren , Yuyue Wang , Ruihua Song , Liyun Ru

Human perception of events is intrinsically tied to distinguishing between completed (perfect and telic) and ongoing (durative) actions, a process mediated by both linguistic structure and visual cues. In this work, we introduce the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Olga Loginova , Sofía Ortega Loguinova

Query-based moment retrieval aims to localize the most relevant moment in an untrimmed video according to the given natural language query. Existing works often only focus on one aspect of this emerging task, such as the query…

Information Retrieval · Computer Science 2019-07-30 Zhu Zhang , Zhijie Lin , Zhou Zhao , Zhenxin Xiao

We introduce VideoComp, a benchmark and learning framework for advancing video-text compositionality understanding, aimed at improving vision-language models (VLMs) in fine-grained temporal alignment. Unlike existing benchmarks focused on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Dahun Kim , AJ Piergiovanni , Ganesh Mallya , Anelia Angelova

Pretraining general-purpose visual features has become a crucial part of tackling many computer vision tasks. While one can learn such features on the extensively-annotated ImageNet dataset, recent approaches have looked at ways to allow…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Mert Bulent Sariyildiz , Julien Perez , Diane Larlus

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language. Due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Gabriel Fiastre , Antoine Yang , Cordelia Schmid

Pre-trained large language models have recently achieved ground-breaking performance in a wide variety of language understanding tasks. However, the same model can not be applied to multimodal behavior understanding tasks (e.g., video…

Computation and Language · Computer Science 2023-03-30 Md Kamrul Hasan , Md Saiful Islam , Sangwu Lee , Wasifur Rahman , Iftekhar Naim , Mohammed Ibrahim Khan , Ehsan Hoque

Capturing a video's meaning and critical concepts by analyzing the subtle details is a fundamental yet challenging task in video captioning. Identifying the dominant emotional tone in a video significantly enhances the perception of its…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Ehsan Faghihi , Mohammedreza Zarenejad , Ali-Asghar Beheshti Shirazi

Action recognition is a critical task in video understanding, requiring the comprehensive capture of spatio-temporal cues across various scales. However, existing methods often overlook the multi-granularity nature of actions. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Xiaoyang Li , Wenzhu Yang , Kanglin Wang , Tiebiao Wang , Qingsong Fei
‹ Prev 1 3 4 5 6 7 10 Next ›