English
Related papers

Related papers: A Feature-space Multimodal Data Augmentation Techn…

200 papers

Precise video retrieval requires multi-modal correlations to handle unseen vocabulary and scenes, becoming more complex for lengthy videos where models must perform effectively without prior training on a specific dataset. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mohamed Eltahir , Osamah Sarraj , Mohammed Bremoo , Mohammed Khurd , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammad Almatrafi , Tanveer Hussain

In this paper we examine the ability of low-level multimodal features to extract movie similarity, in the context of a content-based movie recommendation approach. In particular, we demonstrate the extraction of multimodal representation…

Information Retrieval · Computer Science 2019-12-19 Konstantinos Bougiatiotis , Theodore Giannakopoulos

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Han Fang , Xianghao Zang , Chao Ban , Zerun Feng , Lanxiang Zhou , Zhongjiang He , Yongxiang Li , Hao Sun

The increasing prevalence of video clips has sparked growing interest in text-video retrieval. Recent advances focus on establishing a joint embedding space for text and video, relying on consistent embedding representations to compute…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Jiamian Wang , Guohao Sun , Pichao Wang , Dongfang Liu , Sohail Dianat , Majid Rabbani , Raghuveer Rao , Zhiqiang Tao

Cross-modal retrieval between videos and texts has gained increasing research interest due to the rapid emergence of videos on the web. Generally, a video contains rich instance and event information and the query text only describes a part…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Chengzhi Lin , Ancong Wu , Junwei Liang , Jun Zhang , Wenhang Ge , Wei-Shi Zheng , Chunhua Shen

Most existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences. However, in real-world scenarios, online videos are often accompanied by relevant text information…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Wenhao Wu , Haipeng Luo , Bo Fang , Jingdong Wang , Wanli Ouyang

In Multimodal Language Models (MLMs), the cost of manually annotating high-quality image-text pair data for fine-tuning and alignment is extremely high. While existing multimodal data augmentation frameworks propose ways to augment…

Artificial Intelligence · Computer Science 2024-08-20 Xiaomeng Jin , Jeonghwan Kim , Yu Zhou , Kuan-Hao Huang , Te-Lin Wu , Nanyun Peng , Heng Ji

We address the problem of data augmentation for video action recognition. Standard augmentation strategies in video are hand-designed and sample the space of possible augmented data points either at random, without knowing which augmented…

Computer Vision and Pattern Recognition · Computer Science 2022-07-26 Shreyank N Gowda , Marcus Rohrbach , Frank Keller , Laura Sevilla-Lara

Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-fitting of the deep…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Li Haopeng , Ke Qiuhong , Gong Mingming , Tom Drummond

Predicting the relevance between two given videos with respect to their visual content is a key component for content-based video recommendation and retrieval. Thanks to the increasing availability of pre-trained image and video…

Computer Vision and Pattern Recognition · Computer Science 2020-04-09 Jianfeng Dong , Xun Wang , Leimin Zhang , Chaoxi Xu , Gang Yang , Xirong Li

The increasing volume of video content in educational, professional, and social domains necessitates effective summarization techniques that go beyond traditional unimodal approaches. This paper proposes a behaviour-aware multimodal video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Md Moinul Islam , Sofoklis Kakouros , Janne Heikkilä , Mourad Oussalah

With the rapid growth of video data, text-video retrieval technology has become increasingly important in numerous application scenarios such as recommendation and search. Early text-video retrieval methods suffer from two critical…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Jiaao Yu , Mingjie Han , Tao Gong , Jian Zhang , Man Lan

This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as instructBLIP and CLIP, to generate text-based product…

Information Retrieval · Computer Science 2024-02-14 Jason Tang , Garrin McGoldrick , Marie Al-Ghossein , Ching-Wei Chen

This article aims to provide the information retrieval community with some reflections on recent advances in retrieval learning by analyzing the reproducibility of image-text retrieval models. Due to the increase of multimodal data over the…

Information Retrieval · Computer Science 2022-08-30 Jun Rao , Fei Wang , Liang Ding , Shuhan Qi , Yibing Zhan , Weifeng Liu , Dacheng Tao

Text-video retrieval is a challenging task that aims to search relevant video contents based on natural language descriptions. The key to this problem is to measure text-video similarities in a joint embedding space. However, most existing…

Computer Vision and Pattern Recognition · Computer Science 2021-04-21 Xiaohan Wang , Linchao Zhu , Yi Yang

State-of-the-art video action classifiers often suffer from overfitting. They tend to be biased towards specific objects and scene cues, rather than the foreground action content, leading to sub-optimal generalization performances. Recent…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Sangdoo Yun , Seong Joon Oh , Byeongho Heo , Dongyoon Han , Jinhyung Kim

Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input audio as a unimodal retrieval query. In contrast, we propose…

Sound · Computer Science 2025-06-11 Choi Changin , Lim Sungjun , Rhee Wonjong

Textual overlays are often used in social media videos as people who watch them without the sound would otherwise miss essential information conveyed in the audio stream. This is why extraction of those overlays can serve as an important…

Computer Vision and Pattern Recognition · Computer Science 2018-05-02 Adam Słucki , Tomasz Trzcinski , Adam Bielski , Paweł Cyrta

Multimodal Person Reidentification is gaining popularity in the research community due to its effectiveness compared to counter-part unimodal frameworks. However, the bottleneck for multimodal deep learning is the need for a large volume of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Mulham Fawakherji , Eduard Vazquez , Pasquale Giampa , Binod Bhattarai