English
Related papers

Related papers: TVR: A Large-Scale Dataset for Video-Subtitle Mome…

200 papers

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Thanadol Singkhornart , Olarik Surinta

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Yue Feng , Jinwei Hu , Qijia Lu , Jiawei Niu , Li Tan , Shuo Yuan , Ziyi Yan , Yizhen Jia , Qingzhi He , Shiping Ge , Ethan Q. Chen , Wentong Li , Limin Wang , Jie Qin

Text-video retrieval (TVR) has seen substantial advancements in recent years, fueled by the utilization of pre-trained models and large language models (LLMs). Despite these advancements, achieving accurate matching in TVR remains…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Jian Xiao , Zhenzhen Hu , Jia Li , Richang Hong

This paper tackles a recently proposed Video Corpus Moment Retrieval task. This task is essential because advanced video retrieval applications should enable users to retrieve a precise moment from a large video corpus. We propose a novel…

Multimedia · Computer Science 2021-09-22 Zhijian Hou , Chong-Wah Ngo , Wing Kwong Chan

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mohamed Eltahir , Osamah Sarraj , Mohammed Bremoo , Mohammed Khurd , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammad Almatrafi , Tanveer Hussain

Online video web content is richly multimodal: a single video blends vision, speech, ambient audio, and on-screen text. Retrieval systems typically treat these modalities as independent retrieval sources, which can lead to noisy and subpar…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 David Wan , Han Wang , Elias Stengel-Eskin , Jaemin Cho , Mohit Bansal

Existing multimodal machine translation (MMT) datasets consist of images and video captions or general subtitles, which rarely contain linguistic ambiguity, making visual information not so effective to generate appropriate translations. We…

Computation and Language · Computer Science 2022-05-27 Yihang Li , Shuichiro Shimizu , Weiqi Gu , Chenhui Chu , Sadao Kurohashi

True understanding of videos comes from a joint analysis of all its modalities: the video frames, the audio track, and any accompanying text such as closed captions. We present a way to learn a compact multimodal feature representation that…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Vivek Sharma , Makarand Tapaswi , Rainer Stiefelhagen

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation…

Computer Vision and Pattern Recognition · Computer Science 2020-06-18 Xin Wang , Jiawei Wu , Junkun Chen , Lei Li , Yuan-Fang Wang , William Yang Wang

Text-to-Video Retrieval (TVR) aims to retrieve relevant videos based on textual queries. However, as video content evolves continuously, adapting TVR systems to new data remains a critical yet under-explored challenge. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Zecheng Zhao , Zhi Chen , Zi Huang , Shazia Sadiq , Tong Chen

Video retrieval using natural language queries requires learning semantically meaningful joint embeddings between the text and the audio-visual input. Often, such joint embeddings are learnt using pairwise (or triplet) contrastive loss…

Information Retrieval · Computer Science 2021-03-10 Jayaprakash A , Abhishek , Rishabh Dabral , Ganesh Ramakrishnan , Preethi Jyothi

Understanding the content of events occurring in the video and their inherent temporal logic is crucial for video-text retrieval. However, web-crawled pre-training datasets often lack sufficient event information, and the widely adopted…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Zongyang Ma , Ziqi Zhang , Yuxin Chen , Zhongang Qi , Chunfeng Yuan , Bing Li , Yingmin Luo , Xu Li , Xiaojuan Qi , Ying Shan , Weiming Hu

Humans naturally share information with those they are connected to, and video has become one of the dominant mediums for communication and expression on the Internet. To support the creation of high-quality large-scale video content, a…

With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Bo Fang , Wenhao Wu , Chang Liu , Yu Zhou , Yuxin Song , Weiping Wang , Xiangbo Shu , Xiangyang Ji , Jingdong Wang

Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet the demands of…

Computation and Language · Computer Science 2025-05-12 Jinze Lv , Jian Chen , Zi Long , Xianghua Fu , Yin Chen

Text-Video Retrieval (TVR) aims to align relevant video content with natural language queries. To date, most state-of-the-art TVR methods learn image-to-video transfer learning based on large-scale pre-trained visionlanguage models (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Meng Cao , Haoran Tang , Jinfa Huang , Peng Jin , Can Zhang , Ruyang Liu , Long Chen , Xiaodan Liang , Li Yuan , Ge Li

Several large-scale video datasets have been published these years and have advanced the area of video understanding. However, the newly emerged user-generated short-form videos have rarely been studied. This paper presents USV, the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Haoyue Cheng , Su Xu , Liwei Jin , Wayne Wu , Chen Qian , Limin Wang

Video-to-video moment retrieval (Vid2VidMR) is the task of localizing unseen events or moments in a target video using a query video. This task poses several challenges, such as the need for semantic frame-level alignment and modeling…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Yogesh Kumar , Uday Agarwal , Manish Gupta , Anand Mishra

Existing multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating appropriate…

Computation and Language · Computer Science 2023-11-01 Yihang Li , Shuichiro Shimizu , Chenhui Chu , Sadao Kurohashi , Wei Li