English
Related papers

Related papers: Tencent Text-Video Retrieval: Hierarchical Cross-M…

200 papers

The key of the text-to-video retrieval (TVR) task lies in learning the unique similarity between each pair of text (consisting of words) and video (consisting of audio and image frames) representations. However, some problems exist in the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Wenjun Li , Shudong Wang , Dong Zhao , Shenghui Xu , Zhaoming Pan , Zhimin Zhang

Most existing cross-modal language-to-video retrieval (VR) research focuses on single-modal input from video, i.e., visual representation, while the text is omnipresent in human environments and frequently critical to understand video. To…

Computer Vision and Pattern Recognition · Computer Science 2023-05-08 Weijia Wu , Yuzhong Zhao , Zhuang Li , Jiahong Li , Hong Zhou , Mike Zheng Shou , Xiang Bai

Visual data and text data are composed of information at multiple granularities. A video can describe a complex scene that is composed of multiple clips or shots, where each depicts a semantically coherent event or action. Similarly, a…

Computer Vision and Pattern Recognition · Computer Science 2018-10-18 Bowen Zhang , Hexiang Hu , Fei Sha

While hate speech detection (HSD) has been extensively studied in text, existing multi-modal approaches remain limited, particularly in videos. As modalities are not always individually informative, simple fusion methods fail to fully…

Video Moment Retrieval (VMR) aims to retrieve temporal segments in untrimmed videos corresponding to a given language query by constructing cross-modal alignment strategies. However, these existing strategies are often sub-optimal since…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Zhihang Liu , Jun Li , Hongtao Xie , Pandeng Li , Jiannan Ge , Sun-Ao Liu , Guoqing Jin

We address the problem of text-based activity retrieval in video. Given a sentence describing an activity, our task is to retrieve matching clips from an untrimmed video. To capture the inherent structures present in both text and video, we…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Huijuan Xu , Kun He , Bryan A. Plummer , Leonid Sigal , Stan Sclaroff , Kate Saenko

Pre-training a model to learn transferable video-text representation for retrieval has attracted a lot of attention in recent years. Previous dominant works mainly adopt two separate encoders for efficient retrieval, but ignore local…

Computer Vision and Pattern Recognition · Computer Science 2022-03-18 Yuying Ge , Yixiao Ge , Xihui Liu , Dian Li , Ying Shan , Xiaohu Qie , Ping Luo

The user base of short video apps has experienced unprecedented growth in recent years, resulting in a significant demand for video content analysis. In particular, text-video retrieval, which aims to find the top matching videos given text…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Xuzheng Yu , Chen Jiang , Xingning Dong , Tian Gan , Ming Yang , Qingpei Guo

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the challenges posed by…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yichao Cao , Qingfei Tang , Feng Yang , Xiu Su , Shan You , Xiaobo Lu , Chang Xu

Motivated by the success of coarse-grained or fine-grained contrast in text-video retrieval, there emerge multi-grained contrastive learning methods which focus on the integration of contrasts with different granularity. However, due to the…

Information Retrieval · Computer Science 2025-04-08 Xiaolun Jing , Genke Yang , Jian Chu

Long-context video modeling is critical for multimodal large language models (MLLMs), enabling them to process movies, online video streams, and so on. Despite its advances, handling long videos remains challenging due to the difficulty in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Xinhao Li , Yi Wang , Jiashuo Yu , Xiangyu Zeng , Yuhan Zhu , Haian Huang , Jianfei Gao , Kunchang Li , Yinan He , Chenting Wang , Yu Qiao , Yali Wang , Limin Wang

Text-to-image synthesis has progressed to the point where models can generate visually compelling images from natural language prompts. Yet, existing methods often fail to reconcile high-level semantic fidelity with explicit spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-05-16 Hang Wang , Zhi-Qi Cheng , Chenhao Lin , Chao Shen , Lei Zhang

Audio-visual video parsing is the task of categorizing a video at the segment level with weak labels, and predicting them as audible or visible events. Recent methods for this task leverage the attention mechanism to capture the semantic…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Yaru Chen , Ruohao Guo , Xubo Liu , Peipei Wu , Guangyao Li , Zhenbo Li , Wenwu Wang

A more robust and holistic language-video representation is the key to pushing video understanding forward. Despite the improvement in training strategies, the quality of the language-video dataset is less attention to. The current plain…

Multimedia · Computer Science 2024-06-21 Yuchen Yang , Yingxuan Duan

Video-text retrieval (VTR) aims to locate relevant videos using natural language queries. Current methods, often based on pre-trained models like CLIP, are hindered by video's inherent redundancy and their reliance on coarse, final-layer…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Zequn Xie , Boyun Zhang , Yuxiao Lin , Tao Jin

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

A key challenge in video question answering is how to realize the cross-modal semantic alignment between textual concepts and corresponding visual objects. Existing methods mostly seek to align the word representations with the video…

Computer Vision and Pattern Recognition · Computer Science 2022-05-16 Zenan Xu , Wanjun Zhong , Qinliang Su , Zijing Ou , Fuwei Zhang

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

Existing dominant approaches for cross-modal video-text retrieval task are to learn a joint embedding space to measure the cross-modal similarity. However, these methods rarely explore long-range dependency inside video frames or textual…

Multimedia · Computer Science 2020-04-13 Rui Zhao , Kecheng Zheng , Zheng-jun Zha

Multimedia information retrieval from videos remains a challenging problem. While recent systems have advanced multimodal search through semantic, object, and OCR queries - and can retrieve temporally consecutive scenes - they often rely on…

Information Retrieval · Computer Science 2025-12-09 Van-Thinh Vo , Minh-Khoi Nguyen , Minh-Huy Tran , Anh-Quan Nguyen-Tran , Duy-Tan Nguyen , Khanh-Loi Nguyen , Anh-Minh Phan