中文
相关论文

相关论文: Universal Music Representations? Evaluating Founda…

200 篇论文

While both the data volume and heterogeneity of the digital music content is huge, it has become increasingly important and convenient to build a recommendation or search system to facilitate surfacing these content to the user or consumer…

Musical instruments recognition is a widely used application for music information retrieval. As most of previous musical instruments recognition dataset focus on western musical instruments, it is difficult for researcher to study and…

音频与语音处理 · 电气工程与系统科学 2021-12-14 Xia Gong , Yuxiang Zhu , Haidi Zhu , Haoran Wei

Music performances are representative scenarios for audio-visual modeling. Unlike common scenarios with sparse audio, music performances continuously involve dense audio signals throughout. While existing multimodal learning methods on the…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Xingjian Diao , Chunhui Zhang , Tingxuan Wu , Ming Cheng , Zhongyu Ouyang , Weiyi Wu , Jiang Gui

The number of possible melodies is unfathomably large, yet despite this virtually unlimited potential for melodic variation, melodies from different societies can be surprisingly similar. The motor constraint hypothesis accounts for certain…

声音 · 计算机科学 2025-07-25 John M McBride , Nahie Kim , Yuri Nishikawa , Mekhmed Saadakeev , Marcus T Pearce , Tsvi Tlusty

A range of applications of multi-modal music information retrieval is centred around the problem of connecting large collections of sheet music (images) to corresponding audio recordings, that is, identifying pairs of audio and score…

声音 · 计算机科学 2023-09-22 Luis Carvalho , Gerhard Widmer

We apply foundation models to data discovery and exploration tasks. Foundation models include large language models (LLMs) that show promising performance on a range of diverse tasks unrelated to their training. We show that these models…

数据库 · 计算机科学 2024-04-09 Moe Kayali , Anton Lykov , Ilias Fountalis , Nikolaos Vasiloglou , Dan Olteanu , Dan Suciu

A music piece is both comprehended hierarchically, from sonic events to melodies, and sequentially, in the form of repetition and variation. Music from different cultures establish different aesthetics by having different style conventions…

声音 · 计算机科学 2021-11-25 Shlomo Dubnov , Kevin Huang , Cheng-i Wang

Music similarity is an essential aspect of music retrieval, recommendation systems, and music analysis. Moreover, similarity is of vital interest for music experts, as it allows studying analogies and influences among composers and…

声音 · 计算机科学 2023-06-22 Andrea Poltronieri

The rapid development of generative audio raises ethical and security concerns stemming from forged data, making deepfake sound detection an important safeguard against the malicious use of such technologies. Although prior studies have…

声音 · 计算机科学 2025-09-29 Zeyu Xie , Yaoyun Zhang , Xuenan Xu , Yongkang Yin , Chenxing Li , Mengyue Wu , Yuexian Zou

Music preferences are strongly shaped by the cultural and socio-economic background of the listener, which is reflected, to a considerable extent, in country-specific music listening profiles. Previous work has already identified several…

信息检索 · 计算机科学 2021-02-08 Markus Schedl , Christine Bauer , Wolfgang Reisinger , Dominik Kowald , Elisabeth Lex

Compositional generalization, the ability to recognize familiar parts in novel contexts, is a defining property of intelligent systems. Although modern models are trained on massive datasets, they still cover only a tiny fraction of the…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Arnas Uselis , Andrea Dittadi , Seong Joon Oh

CLaMP 3 is a unified framework developed to address challenges of cross-modal and cross-lingual generalization in music information retrieval. Using contrastive learning, it aligns all major music modalities--including sheet music,…

Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose…

Music is a structured and perceptually rich sequence of sounds in time, whose perception is shaped by the interplay of expectation and uncertainty about what comes next. Yet the uncertainty we infer from music depends on how the musical…

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural…

人工智能 · 计算机科学 2025-10-23 Brandon James Carone , Iran R. Roman , Pablo Ripollés

At present, neural network-based models, including transformers, struggle to generate memorable and readily comprehensible music from unified and repetitive musical material due to a lack of understanding of musical structure. Consequently,…

声音 · 计算机科学 2026-01-21 Shangxuan Luo , Joshua Reiss

Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced…

声音 · 计算机科学 2026-05-27 Abhinaba Roy , Junyi Liang , Dorien Herremans

Recently, research on audio foundation models has witnessed notable advances, as illustrated by the ever improving results on complex downstream tasks. Subsequently, those pretrained networks have quickly been used for various audio…

声音 · 计算机科学 2025-02-19 David Genova , Philippe Esling , Tom Hurlin

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive…

There has been a rapid growth of digitally available music data, including audio recordings, digitized images of sheet music, album covers and liner notes, and video clips. This huge amount of data calls for retrieval strategies that allow…

信息检索 · 计算机科学 2019-02-13 Meinard Müller , Andreas Arzt , Stefan Balke , Matthias Dorfer , Gerhard Widmer