English
Related papers

Related papers: CLaMP 2: Multimodal Music Information Retrieval Ac…

200 papers

Contrastive language-audio pretraining (CLAP) has recently emerged as a method for making audio analysis more generalisable. Specifically, CLAP-style models are able to `answer' a diverse set of language queries, extending the capabilities…

Sound · Computer Science 2024-06-12 Xin Jing , Andreas Triantafyllopoulos , Björn Schuller

We discuss a novel task, Chorus Recognition, which could potentially benefit downstream tasks such as song search and music summarization. Different from the existing tasks such as music summarization or lyrics summarization relying on…

Information Retrieval · Computer Science 2021-07-01 Jiaan Wang , Zhixu Li , Binbin Gu , Tingyi Zhang , Qingsheng Liu , Zhigang Chen

Continual learning (CL) empowers pre-trained vision-language models to adapt effectively to novel or previously underrepresented data distributions without comprehensive retraining, enhancing their adaptability and efficiency. While…

Artificial Intelligence · Computer Science 2025-09-04 Zhiyuan Wang , Bokui Chen

Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the symbolic music domain including understanding and generation.…

Sound · Computer Science 2024-08-01 Ziya Zhou , Yuhang Wu , Zhiyue Wu , Xinyue Zhang , Ruibin Yuan , Yinghao Ma , Lu Wang , Emmanouil Benetos , Wei Xue , Yike Guo

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. As the number of reference identities…

Machine Learning · Computer Science 2026-04-10 Yucheng Zhou , Dubing Chen , Huan Zheng , Jianbing Shen

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Efficiently retrieving and synthesizing information from large-scale multimodal collections has become a critical challenge. However, existing video retrieval datasets suffer from scope limitations, primarily focusing on matching…

We consider a specific scenario of text aggregation, in the realm of musical harmonization. Musical harmonization shares similarities with text aggregation, however the language of harmony is more structured than general text. Concretely,…

Sound · Computer Science 2025-09-03 Eyal Briman , Eyal Leizerovich , Nimrod Talmon

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation,…

Machine Learning · Computer Science 2026-01-22 Piyush Singh Pasi

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside world. However, the mixed multiple objects and sounds make it…

Computer Vision and Pattern Recognition · Computer Science 2019-04-22 Di Hu , Feiping Nie , Xuelong Li

Multi-instrument recognition is the task of predicting the presence or absence of different instruments within an audio clip. A considerable challenge in applying deep learning to multi-instrument recognition is the scarcity of labeled…

Sound · Computer Science 2020-01-27 Amir Kenarsari Anhari

The growing availability of music on streaming platforms has led to information overload for users. To address this issue and enhance the user experience, increasingly sophisticated recommendation systems have been proposed. This work…

Information Retrieval · Computer Science 2025-08-19 Ronald Carvalho Boadana , Ademir Guimarães da Costa Junior , Ricardo Rios , Fábio Santos da Silva

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often…

Sound · Computer Science 2024-11-07 Yiming Chen , Xianghu Yue , Xiaoxue Gao , Chen Zhang , Luis Fernando D'Haro , Robby T. Tan , Haizhou Li

We introduce The Benchmark of Linguistic Minimal Pairs (shortened to BLiMP), a challenge set for evaluating what language models (LMs) know about major grammatical phenomena in English. BLiMP consists of 67 sub-datasets, each containing…

Computation and Language · Computer Science 2023-02-15 Alex Warstadt , Alicia Parrish , Haokun Liu , Anhad Mohananey , Wei Peng , Sheng-Fu Wang , Samuel R. Bowman

Tag-based music retrieval is crucial to browse large-scale music libraries efficiently. Hence, automatic music tagging has been actively explored, mostly as a classification task, which has an inherent limitation: a fixed vocabulary. On the…

Information Retrieval · Computer Science 2020-11-02 Minz Won , Sergio Oramas , Oriol Nieto , Fabien Gouyon , Xavier Serra

Audio-text retrieval enables semantic alignment between audio content and natural language queries, supporting applications in multimedia search, accessibility, and surveillance. However, current state-of-the-art approaches struggle with…

Computation and Language · Computer Science 2026-04-28 Meizhu Liu , Matthew Rowe , Amit Agarwal , Michael Avendi , Yassi Abbasi , Hitesh Laxmichand Patel , Paul Li , Kyu J. Han , Tao Sheng , Sujith Ravi , Dan Roth

Recent advances in clinical AI have enabled remarkable progress across many clinical domains. However, existing benchmarks and models are primarily limited to a small set of modalities and tasks, which hinders the development of large-scale…

Machine Learning · Computer Science 2025-03-21 Wei Dai , Peilin Chen , Malinda Lu , Daniel Li , Haowen Wei , Hejie Cui , Paul Pu Liang

Amidst the rapid advances of large language models (LLMs), most LLMs still struggle with mixed-language inputs, limited Codeswitching (CSW) datasets, and evaluation biases, which hinder their deployment in multilingual societies. This…

Computation and Language · Computer Science 2026-04-22 Rajvee Sheth , Samridhi Raj Sinha , Mahavir Patil , Himanshu Beniwal , Mayank Singh

Text-to-music generation (T2M-Gen) faces a major obstacle due to the scarcity of large-scale publicly available music datasets with natural language captions. To address this, we propose the Music Understanding LLaMA (MU-LLaMA), capable of…

Sound · Computer Science 2023-08-23 Shansong Liu , Atin Sakkeer Hussain , Chenshuo Sun , Ying Shan

Extensive works have tackled Language Identification (LID) in the speech domain, however their application to the singing voice trails and performances on Singing Language Identification (SLID) can be improved leveraging recent progresses…

Sound · Computer Science 2021-06-01 Lenny Renault , Andrea Vaglio , Romain Hennequin