English
Related papers

Related papers: On Metric Learning for Audio-Text Cross-Modal Retr…

200 papers

Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature…

Multimedia · Computer Science 2023-03-07 Zhongweiyang Xu , Xulin Fan , Mark Hasegawa-Johnson

Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most…

Computer Vision and Pattern Recognition · Computer Science 2020-10-08 Jiwei Wei , Xing Xu , Yang Yang , Yanli Ji , Zheng Wang , Heng Tao Shen

Audio captioning aims to generate text descriptions from environmental sounds. One challenge of audio captioning is the difficulty of the generalization due to the lack of audio-text paired training data. In this work, we propose a simple…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-05 Minkyu Kim , Kim Sung-Bin , Tae-Hyun Oh

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

Sound · Computer Science 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

The problem of software artifact retrieval has the goal to effectively locate software artifacts, such as a piece of source code, in a large code repository. This problem has been traditionally addressed through the textual query. In other…

Machine Learning · Computer Science 2016-11-15 Liang Wu , Hui Xiong , Liang Du , Bo Liu , Guandong Xu , Yong Ge , Yanjie Fu , Yuanchun Zhou , Jianhui Li

This paper proposes a new evaluation protocol for cross-media retrieval which better fits the real-word applications. Both image-text and text-image retrieval modes are considered. Traditionally, class labels in the training and testing…

Computer Vision and Pattern Recognition · Computer Science 2017-03-13 Ruoyu Liu , Yao Zhao , Liang Zheng , Shikui Wei , Yi Yang

In this paper, we introduce a novel continual audio-visual sound separation task, aiming to continuously separate sound sources for new classes while preserving performance on previously learned classes, with the aid of visual guidance.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Weiguo Pian , Yiyang Nan , Shijian Deng , Shentong Mo , Yunhui Guo , Yapeng Tian

In this work we introduce a cross modal image retrieval system that allows both text and sketch as input modalities for the query. A cross-modal deep network architecture is formulated to jointly model the sketch and text input modalities…

Computer Vision and Pattern Recognition · Computer Science 2018-05-01 Sounak Dey , Anjan Dutta , Suman K. Ghosh , Ernest Valveny , Josep Lladós , Umapada Pal

Automated Audio captioning (AAC) is a cross-modal task that generates natural language to describe the content of input audio. Most prior works usually extract single-modality acoustic features and are therefore sub-optimal for the…

Sound · Computer Science 2022-04-13 Chen Chen , Nana Hou , Yuchen Hu , Heqing Zou , Xiaofeng Qi , Eng Siong Chng

Different machine learning models can represent the same underlying concept in different ways. This variability is particularly valuable for in-the-wild multimodal retrieval, where the objective is to identify the corresponding…

Information Retrieval · Computer Science 2025-06-11 Fan Xu , Luis A. Leiva

Retrieving target videos based on text descriptions is a task of great practical value and has received increasing attention over the past few years. Despite recent progress, imperfect annotations in existing video retrieval datasets have…

Computer Vision and Pattern Recognition · Computer Science 2022-07-22 Zeyu Wang , Yu Wu , Karthik Narasimhan , Olga Russakovsky

Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in comprehending text information plays a pivotal role in…

Sound · Computer Science 2024-02-28 Liwen Tan , Yin Cao , Yi Zhou

Our objective is to build an embedding model that captures the nuanced relationship between a search query and candidate videos. We cover three aspects of nuanced retrieval: (i) temporal, (ii) negation, and (iii) multimodal. For temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-05-04 Piyush Bagad , Andrew Zisserman

Audio-language pretraining holds promise for general-purpose audio understanding, yet remains underexplored compared to its vision counterpart. While vision-language models like CLIP serve as widely adopted foundations, existing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-24 Wei-Cheng Tseng , Xuanru Zhou , Mingyue Huo , Yiwen Shao , Hao Zhang , Dong Yu

Text-based person retrieval aims to identify specific individuals within an image database using textual descriptions. Due to the high cost of annotation and privacy protection, researchers resort to synthesized data for the paradigm of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Hang Yu , Jiahao Wen , Zhedong Zheng

With the advent of large-scale multimodal video datasets, especially sequences with audio or transcribed speech, there has been a growing interest in self-supervised learning of video representations. Most prior work formulates the…

Computer Vision and Pattern Recognition · Computer Science 2020-09-21 Bruno Korbar , Fabio Petroni , Rohit Girdhar , Lorenzo Torresani

In recent years, datasets of paired audio and captions have enabled remarkable success in automatically generating descriptions for audio clips, namely Automated Audio Captioning (AAC). However, it is labor-intensive and time-consuming to…

Sound · Computer Science 2023-09-22 Theodoros Kouzelis , Vassilis Katsouros

Multi-media communications facilitate global interaction among people. However, despite researchers exploring cross-lingual translation techniques such as machine translation and audio speech translation to overcome language barriers, there…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Xize Cheng , Linjun Li , Tao Jin , Rongjie Huang , Wang Lin , Zehan Wang , Huangdai Liu , Ye Wang , Aoxiong Yin , Zhou Zhao

Cross-modal retrieval aims to bridge the semantic gap between different modalities, such as visual and textual data, enabling accurate retrieval across them. Despite significant advancements with models like CLIP that align cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Zengrong Lin , Zheng Wang , Tianwen Qian , Pan Mu , Sixian Chan , Cong Bai

Text-guided audio editing aims to modify specific acoustic events while strictly preserving non-target content. Despite recent progress, existing approaches remain fundamentally limited. Training-free methods often suffer from signal…

Sound · Computer Science 2026-01-21 Ye Tao , Wen Wu , Chao Zhang , Mengyue Wu , Shuai Wang , Xuenan Xu
‹ Prev 1 8 9 10 Next ›