English
Related papers

Related papers: Toward Universal Text-to-Music Retrieval

200 papers

We present TALKPLAY, a novel multimodal music recommendation system that reformulates recommendation as a token generation problem using large language models (LLMs). By leveraging the instruction-following and natural language generation…

Information Retrieval · Computer Science 2025-05-27 Seungheon Doh , Keunwoo Choi , Juhan Nam

We report experimental results associated with speech-driven text retrieval, which facilitates retrieving information in multiple domains with spoken queries. Since users speak contents related to a target collection, we produce language…

Computation and Language · Computer Science 2016-11-15 Katunobu Itou , Atsushi Fujii , Tetsuya Ishikawa

Audio-Text retrieval takes a natural language query to retrieve relevant audio files in a database. Conversely, Text-Audio retrieval takes an audio file as a query to retrieve relevant natural language descriptions. Most of the literature…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-29 Soham Deshmukh , Benjamin Elizalde , Huaming Wang

In this paper, we transform tag recommendation into a word-based text generation problem and introduce a sequence-to-sequence model. The model inherits the advantages of LSTM-based encoder for sequential modeling and attention-based decoder…

Computation and Language · Computer Science 2019-12-03 Xuewen Shi , Heyan Huang , Shuyang Zhao , Ping Jian , Yi-Kun Tang

Recently, we proposed a self-attention based music tagging model. Different from most of the conventional deep architectures in music information retrieval, which use stacked 3x3 filters by treating music spectrograms as images, the…

Sound · Computer Science 2019-11-12 Minz Won , Sanghyuk Chun , Xavier Serra

Mood recognition is an important problem in music informatics and has key applications in music discovery and recommendation. These applications have become even more relevant with the rise of music streaming. Our work investigates the…

Sound · Computer Science 2021-10-12 Rajnish Kumar , Manjeet Dahiya

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information Retrieval. In this…

Sound · Computer Science 2022-08-26 Ilaria Manco , Emmanouil Benetos , Elio Quinton , György Fazekas

Neural topic models can successfully find coherent and diverse topics in textual data. However, they are limited in dealing with multimodal datasets (e.g., images and text). This paper presents the first systematic and comprehensive…

Computation and Language · Computer Science 2024-03-27 Felipe González-Pizarro , Giuseppe Carenini

Cross-modal retrieval has become popular in recent years, particularly with the rise of multimedia. Generally, the information from each modality exhibits distinct representations and semantic information, which makes feature tends to be in…

Information Retrieval · Computer Science 2023-08-29 Zichen Yuan , Qi Shen , Bingyi Zheng , Yuting Liu , Linying Jiang , Guibing Guo

Audio-text relevance learning refers to learning the shared semantic properties of audio samples and textual descriptions. The standard approach uses binary relevances derived from pairs of audio samples and their human-provided captions,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-28 Huang Xie , Khazar Khorrami , Okko Räsänen , Tuomas Virtanen

In the context of environmental sound classification, the adaptability of systems is key: which sound classes are interesting depends on the context and the user's needs. Recent advances in text-to-audio retrieval allow for zero-shot audio…

Sound · Computer Science 2023-08-21 Saksham Singh Kushwaha , Magdalena Fuentes

Recent music generation methods based on transformers have a context window of up to a minute. The music generated by these methods is largely unstructured beyond the context window. With a longer context window, learning long-scale…

Sound · Computer Science 2024-10-08 Lilac Atassi

This paper demonstrates the feasibility of learning to retrieve short snippets of sheet music (images) when given a short query excerpt of music (audio) -- and vice versa --, without any symbolic representation of music or scores. This…

Sound · Computer Science 2016-12-16 Matthias Dorfer , Andreas Arzt , Gerhard Widmer

In AI-facilitated teaching, leveraging various query styles to interpret abstract educational content is crucial for delivering effective and accessible learning experiences. However, existing retrieval systems predominantly focus on…

Artificial Intelligence · Computer Science 2025-07-08 Xinyi Wu , Yanhao Jia , Luwei Xiao , Shuai Zhao , Fengkuang Chiang , Erik Cambria

Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing the capabilities of large language models. However, existing RAG evaluation predominantly focuses on text retrieval and relies on opaque, end-to-end…

Information Retrieval · Computer Science 2025-05-19 Chuan Xu , Qiaosheng Chen , Yutong Feng , Gong Cheng

Recent progress in text-to-music generation has enabled models to synthesize high-quality musical segments, full compositions, and even respond to fine-grained control signals, e.g. chord progressions. State-of-the-art (SOTA) systems differ…

Sound · Computer Science 2025-09-05 Or Tal , Felix Kreuk , Yossi Adi

Design sharing sites provide UI designers with a platform to share their works and also an opportunity to get inspiration from others' designs. To facilitate management and search of millions of UI design images, many design sharing sites…

Human-Computer Interaction · Computer Science 2020-10-06 Chunyang Chen , Sidong Feng , Zhengyang Liu , Zhenchang Xing , Shengdong Zhao

Vision-language alignment learning for video-text retrieval arouses a lot of attention in recent years. Most of the existing methods either transfer the knowledge of image-text pretraining model to video-text retrieval task without fully…

Computer Vision and Pattern Recognition · Computer Science 2023-01-31 Yizhen Chen , Jie Wang , Lijian Lin , Zhongang Qi , Jin Ma , Ying Shan

Audio carries richer information than text, including emotion, speaker traits, and environmental context, while also enabling lower-latency processing compared to speech-to-text pipelines. However, recent multimodal information retrieval…

Sound · Computer Science 2026-04-23 Tong Zhao , Chenghao Zhang , Yutao Zhu , Zhicheng Dou