English
Related papers

Related papers: A Dataset for Movie Description

200 papers

Remote work and online courses have become important methods of knowledge dissemination, leading to a large number of document-based instructional videos. Unlike traditional video datasets, these videos mainly feature rich-text images and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Haochen Wang , Kai Hu , Liangcai Gao

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains.…

Computation and Language · Computer Science 2025-05-27 Dongqi Liu , Chenxi Whitehouse , Xi Yu , Louis Mahon , Rohit Saxena , Zheng Zhao , Yifu Qiu , Mirella Lapata , Vera Demberg

Event cameras, or Dynamic Vision Sensor (DVS), are very promising sensors which have shown several advantages over frame based cameras. However, most recent work on real applications of these cameras is focused on 3D reconstruction and…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Iñigo Alonso , Ana C. Murillo

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Existing referring…

Computer Vision and Pattern Recognition · Computer Science 2025-12-13 Henghui Ding , Chang Liu , Shuting He , Kaining Ying , Xudong Jiang , Chen Change Loy , Yu-Gang Jiang

Several services for people with visual disabilities have emerged recently due to achievements in Assistive Technologies and Artificial Intelligence areas. Despite the growth in assistive systems availability, there is a lack of services…

Computer Vision and Pattern Recognition · Computer Science 2022-02-17 Daniel Louzada Fernandes , Marcos Henrique Fonseca Ribeiro , Fabio Ribeiro Cerqueira , Michel Melo Silva

The tremendous growth in 3D (stereo) imaging and display technologies has led to stereoscopic content (video and image) becoming increasingly popular. However, both the subjective and the objective evaluation of stereoscopic video content…

Multimedia · Computer Science 2016-04-27 Manasa K , Balasubramanyam Appina , Sumohana S. Channappayya

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes, feature propagation, or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Diandian Guo , Deng-Ping Fan , Tongyu Lu , Christos Sakaridis , Luc Van Gool

Despite progress in vision-based inspection algorithms, real-world industrial challenges -- specifically in data availability, quality, and complex production requirements -- often remain under-addressed. We introduce the VISION Datasets, a…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Haoping Bai , Shancong Mou , Tatiana Likhomanenko , Ramazan Gokberk Cinbis , Oncel Tuzel , Ping Huang , Jiulong Shan , Jianjun Shi , Meng Cao

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli

Translating visual data into natural language is essential for machines to understand the world and interact with humans. In this work, a comprehensive study is conducted on video paragraph captioning, with the goal to generate…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Qinyu Li , Tengpeng Li , Hanli Wang , Chang Wen Chen

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

Computation and Language · Computer Science 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

MPEG is undertaking a new initiative to standardize content description of audio and video data/documents. When it is finalized in 2001, MPEG-7 is expected to provide standardized description schemes for concise and unambiguous content…

Digital Libraries · Computer Science 2007-05-23 Michael J. Hu , Ye Jian

Recent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand…

Multimedia · Computer Science 2023-04-18 Kaiyuan Hu , Yili Jin , Haowen Yang , Junhua Liu , Fangxin Wang

We introduce the first zero-shot approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Qian Wang , Abdelrahman Eldesokey , Mohit Mendiratta , Fangneng Zhan , Adam Kortylewski , Christian Theobalt , Peter Wonka

Video skimming, also known as dynamic video summarization, generates a temporally abridged version of a given video. Skimming can be achieved by identifying significant components either in uni-modal or multi-modal features extracted from…

Computer Vision and Pattern Recognition · Computer Science 2019-10-01 Vivekraj V. K. , Debashis Sen , Balasubramanian Raman

Event cameras, such as dynamic vision sensors (DVS), and dynamic and active-pixel vision sensors (DAVIS) can supplement other autonomous driving sensors by providing a concurrent stream of standard active pixel sensor (APS) images and DVS…

Computer Vision and Pattern Recognition · Computer Science 2017-11-07 Jonathan Binas , Daniel Neil , Shih-Chii Liu , Tobi Delbruck

Video captioning has been attracting broad research attention in multimedia community. However, most existing approaches either ignore temporal information among video frames or just employ local contextual temporal knowledge. In this work,…

Multimedia · Computer Science 2016-06-16 Yi Bin , Yang Yang , Zi Huang , Fumin Shen , Xing Xu , Heng Tao Shen

Automatic video description requires the generation of natural language statements about the actions, events, and objects in the video. An important human trait, when we describe a video, is that we are able to do this with variable levels…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Fatemeh Ziaeetabar , Reza Safabakhsh , Saeedeh Momtazi , Minija Tamosiunaite , Florentin Wörgötter

Our goal is to collect a large-scale audio-visual dataset with low label noise from videos in the wild using computer vision techniques. The resulting dataset can be used for training and evaluating audio recognition models. We make three…

Computer Vision and Pattern Recognition · Computer Science 2020-09-28 Honglie Chen , Weidi Xie , Andrea Vedaldi , Andrew Zisserman

Large vision-language models have achieved remarkable capabilities by training on massive internet-scale data, yet a fundamental asymmetry persists: while LLMs can leverage self-supervised pretraining on abundant text and image data, the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Kidus Zewde , Yuchen Zhou , Dennis Ng , Neo Tiangratanakul , Tommy Duong , Ankit Raj , Yuxin Zhang , Xingyu Shen , Simiao Ren