English
Related papers

Related papers: Bridging the Gap Between Semantic and User Prefere…

200 papers

We present a novel technique for self-supervised video representation learning by: (a) decoupling the learning objective into two contrastive subtasks respectively emphasizing spatial and temporal features, and (b) performing it…

Computer Vision and Pattern Recognition · Computer Science 2021-09-02 Zehua Zhang , David Crandall

Semi-supervised action recognition aims to improve spatio-temporal reasoning ability with a few labeled data in conjunction with a large amount of unlabeled data. Albeit recent advancements, existing powerful methods are still prone to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Yu Wang , Sanping Zhou , Kun Xia , Le Wang

This paper proposes music similarity representation learning (MSRL) based on individual instrument sounds (InMSRL) utilizing music source separation (MSS) and human preference without requiring clean instrument sounds during inference. We…

Sound · Computer Science 2025-03-25 Takehiro Imamura , Yuka Hashizume , Wen-Chin Huang , Tomoki Toda

Cross-modal retrieval has become a highlighted research topic for retrieval across multimedia data such as image and text. A two-stage learning framework is widely adopted by most existing methods based on Deep Neural Network (DNN): The…

Multimedia · Computer Science 2017-08-09 Yuxin Peng , Jinwei Qi , Xin Huang , Yuxin Yuan

Advancements in audio neural networks have established state-of-the-art results on downstream audio tasks. However, the black-box structure of these models makes it difficult to interpret the information encoded in their internal audio…

Sound · Computer Science 2025-04-22 Alice Zhang , Edison Thomaz , Lie Lu

Understanding dark scenes based on multi-modal image data is challenging, as both the visible and auxiliary modalities provide limited semantic information for the task. Previous methods focus on fusing the two modalities but neglect the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 Xiaoyu Dong , Naoto Yokoya

Large language models (LLMs) excel at a range of tasks through in-context learning (ICL), where only a few task examples guide their predictions. However, prior research highlights that LLMs often overlook input-label mapping information in…

Computation and Language · Computer Science 2025-06-10 Keqin Peng , Liang Ding , Yuanxin Ouyang , Meng Fang , Yancheng Yuan , Dacheng Tao

Despite exciting progress in causal language models, the expressiveness of the representations is largely limited due to poor discrimination ability. To remedy this issue, we present ContraCLM, a novel contrastive learning framework at both…

Music representation learning is notoriously difficult for its complex human-related concepts contained in the sequence of numerical signals. To excavate better MUsic SEquence Representation from labeled audio, we propose a novel…

Sound · Computer Science 2023-06-01 Tianyu Chen , Yuan Xie , Shuai Zhang , Shaohan Huang , Haoyi Zhou , Jianxin Li

In sequential recommendation, models recommend items based on user's interaction history. To this end, current models usually incorporate information such as item descriptions and user intent or preferences. User preferences are usually not…

In the context of music information retrieval, similarity-based approaches are useful for a variety of tasks that benefit from a query-by-example scenario. Music however, naturally decomposes into a set of semantically meaningful factors of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-03 Sebastian Ribecky , Jakob Abeßer , Hanna Lukashevich

Currently, learning better unsupervised sentence representations is the pursuit of many natural language processing communities. Lots of approaches based on pre-trained language models (PLMs) and contrastive learning have achieved promising…

Computation and Language · Computer Science 2023-05-11 Nuo Chen , Linjun Shou , Ming Gong , Jian Pei , Bowen Cao , Jianhui Chang , Daxin Jiang , Jia Li

Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform…

Nowadays the measure between heterogeneous data is still an open problem for cross-modal retrieval. The core of cross-modal retrieval is how to measure the similarity between different types of data. Many approaches have been developed to…

Computer Vision and Pattern Recognition · Computer Science 2022-01-31 Haoming Zhang , Xiao-Jun Wu , Tianyang Xu , Donglin Zhang

Modern music streaming services are heavily based on recommendation engines to serve content to users. Sequential recommendation -- continuously providing new items within a single session in a contextually coherent manner -- has been an…

Information Retrieval · Computer Science 2024-09-12 Pavan Seshadri , Shahrzad Shashaani , Peter Knees

Cross-modal retrieval (CMR) typically involves learning common representations to directly measure similarities between multimodal samples. Most existing CMR methods commonly assume multimodal samples in pairs and employ joint training to…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Ruitao Pu , Yang Qin , Dezhong Peng , Xiaomin Song , Huiming Zheng

Human perception and experience of music is highly context-dependent. Contextual variability contributes to differences in how we interpret and interact with music, challenging the design of robust models for information retrieval.…

Sound · Computer Science 2022-10-31 Kleanthis Avramidis , Shanti Stewart , Shrikanth Narayanan

The pre-training for language models captures general language understanding but fails to distinguish the affective impact of a particular context to a specific word. Recent works have sought to introduce contrastive learning (CL) for…

Computation and Language · Computer Science 2024-05-06 Jin Wang , Liang-Chih Yu , Xuejie Zhang

Traditional music search engines rely on retrieval methods that match natural language queries with music metadata. There have been increasing efforts to expand retrieval methods to consider the audio characteristics of music itself, using…

Multimedia · Computer Science 2024-12-10 Shanti Stewart , Kleanthis Avramidis , Tiantian Feng , Shrikanth Narayanan

Audio-visual segmentation (AVS) aims to segment objects in videos based on audio cues. Existing AVS methods are primarily designed to enhance interaction efficiency but pay limited attention to modality representation discrepancies and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Mingfeng Zha , Tianyu Li , Guoqing Wang , Peng Wang , Yangyang Wu , Yang Yang , Heng Tao Shen