English
Related papers

Related papers: Team RUC_AIM3 Technical Report at Activitynet 2020…

200 papers

Video description is one of the most challenging problems in vision and language understanding due to the large variability both on the video and language side. Models, hence, typically shortcut the difficulty in recognition and generate…

Computer Vision and Pattern Recognition · Computer Science 2019-05-07 Luowei Zhou , Yannis Kalantidis , Xinlei Chen , Jason J. Corso , Marcus Rohrbach

We present a video generation model that accurately reproduces object motion, changes in camera viewpoint, and new content that arises over time. Existing video generation methods often fail to produce new content as a function of time…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Tim Brooks , Janne Hellsten , Miika Aittala , Ting-Chun Wang , Timo Aila , Jaakko Lehtinen , Ming-Yu Liu , Alexei A. Efros , Tero Karras

Recently video generation has achieved substantial progress with realistic results. Nevertheless, existing AI-generated videos are usually very short clips ("shot-level") depicting a single scene. To deliver a coherent long video…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Xinyuan Chen , Yaohui Wang , Lingjun Zhang , Shaobin Zhuang , Xin Ma , Jiashuo Yu , Yali Wang , Dahua Lin , Yu Qiao , Ziwei Liu

Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Luis Denninger , Sina Mokhtarzadeh Azar , Juergen Gall

We propose a universal video-level modality-awareness tracking model with online dense temporal token learning (called {\modaltracker}). It is designed to support various tracking tasks, including RGB, RGB+Thermal, RGB+Depth, and RGB+Event,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Yaozong Zheng , Bineng Zhong , Qihua Liang , Shengping Zhang , Guorong Li , Xianxian Li , Rongrong Ji

Current image captioning approaches generate descriptions which lack specific information, such as named entities that are involved in the images. In this paper we propose a new task which aims to generate informative image captions, given…

Computation and Language · Computer Science 2018-11-08 Di Lu , Spencer Whitehead , Lifu Huang , Heng Ji , Shih-Fu Chang

In the task of dense video captioning of Soccernet dataset, we propose to generate a video caption of each soccer action and locate the timestamp of the caption. Firstly, we apply Blip as our video caption framework to generate video…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Zheng Ruan , Ruixuan Liu , Shimin Chen , Mengying Zhou , Xinquan Yang , Wei Li , Chen Chen , Wei Shen

The Event Causality Identification Shared Task of CASE 2022 involved two subtasks working on the Causal News Corpus. Subtask 1 required participants to predict if a sentence contains a causal relation or not. This is a supervised binary…

Computation and Language · Computer Science 2022-11-23 Fiona Anting Tan , Hansi Hettiarachchi , Ali Hürriyetoğlu , Tommaso Caselli , Onur Uca , Farhana Ferdousi Liza , Nelleke Oostdijk

Text-to-Video (T2V) models are capable of synthesizing high-quality, temporally coherent dynamic video content, but the diverse generation also inherently introduces critical safety challenges. Existing safety evaluation methods,which focus…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Jiaming He , Guanyu Hou , Hongwei Li , Zhicong Huang , Kangjie Chen , Yi Yu , Wenbo Jiang , Guowen Xu , Tianwei Zhang

Existing video captioning benchmarks and models lack causal-temporal narrative, which is sequences of events linked through cause and effect, unfolding over time and driven by characters or agents. This lack of narrative restricts models'…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Asmar Nadeem , Faegheh Sardari , Robert Dawes , Syed Sameed Husain , Adrian Hilton , Armin Mustafa

In this paper the problem of complex event detection in the continuous domain (i.e. events with unknown starting and ending locations) is addressed. Existing event detection methods are limited to features that are extracted from the local…

Computer Vision and Pattern Recognition · Computer Science 2017-06-14 Iman Abbasnejad , Sridha Sridharan , Simon Denman , Clinton Fookes , Simon Lucey

Dynamic scene graph generation from a video is challenging due to the temporal dynamics of the scene and the inherent temporal fluctuations of predictions. We hypothesize that capturing long-term temporal dependencies is the key to…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Shengyu Feng , Subarna Tripathi , Hesham Mostafa , Marcel Nassar , Somdeb Majumdar

Event cameras offer high temporal resolution and power efficiency, making them well-suited for edge AI applications. However, their high event rates present challenges for data transmission and processing. Subsampling methods provide a…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Hesam Araghi , Jan van Gemert , Nergis Tomen

In this paper, we present Change3D, a framework that reconceptualizes the change detection and captioning tasks through video modeling. Recent methods have achieved remarkable success by regarding each pair of bi-temporal images as separate…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Duowang Zhu , Xiaohu Huang , Haiyan Huang , Hao Zhou , Zhenfeng Shao

Existing video captioning methods struggle to balance visual fidelity and redundancy: holistic captions are compact but lose fine-grained evidence, whereas segment-wise captions improve coverage but introduce heavy redundancy. We propose…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Zihan Lin , Songhe Deng , Shuwei He , Danxiang Zhu , Dan Zhang , Yishu Lei , Xianlong Luo , Shikun Feng , Rui Liu

In problems such as sports video analytics, it is difficult to obtain accurate frame level annotations and exact event duration because of the lengthy videos and sheer volume of video data. This issue is even more pronounced in fast-paced…

Computer Vision and Pattern Recognition · Computer Science 2020-04-15 Kanav Vats , Mehrnaz Fani , Pascale Walters , David A. Clausi , John Zelek

Multimedia information retrieval from videos remains a challenging problem. While recent systems have advanced multimodal search through semantic, object, and OCR queries - and can retrieve temporally consecutive scenes - they often rely on…

Information Retrieval · Computer Science 2025-12-09 Van-Thinh Vo , Minh-Khoi Nguyen , Minh-Huy Tran , Anh-Quan Nguyen-Tran , Duy-Tan Nguyen , Khanh-Loi Nguyen , Anh-Minh Phan

We describe an approach used in the Generic Boundary Event Captioning challenge at the Long-Form Video Understanding Workshop held at CVPR 2022. We designed a Rich Encoder-decoder framework for Video Event CAptioner (REVECA) that utilizes…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Jaehyuk Heo , YongGi Jeong , Sunwoo Kim , Jaehee Kim , Pilsung Kang

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

Generating natural language descriptions for videos, i.e., video captioning, essentially requires step-by-step reasoning along the generation process. For example, to generate the sentence "a man is shooting a basketball", we need to first…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Ganchao Tan , Daqing Liu , Meng Wang , Zheng-Jun Zha
‹ Prev 1 8 9 10 Next ›