English
Related papers

Related papers: Music Foundation Model as Generic Booster for Musi…

200 papers

This paper investigates foundation models tailored for music informatics, a domain currently challenged by the scarcity of labeled data and generalization issues. To this end, we conduct an in-depth comparative study among various…

Sound · Computer Science 2023-11-07 Minz Won , Yun-Ning Hung , Duc Le

Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles…

Sound · Computer Science 2025-09-10 Ganghui Ru , Jieying Wang , Jiahao Zhao , Yulun Wu , Yi Yu , Nannan Jiang , Wei Wang , Wei Li

In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-24 Orchid Chetia Phukan , Sarthak Jain , Swarup Ranjan Behera , Arun Balaji Buduru , Rajesh Sharma , S. R Mahadeva Prasanna

Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections.…

In recent years, foundation models (FMs) such as large language models (LLMs) and latent diffusion models (LDMs) have profoundly impacted diverse sectors, including music. This comprehensive review examines state-of-the-art (SOTA)…

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

Sound · Computer Science 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

Music foundation models possess impressive music generation capabilities. When people compose music, they may infuse their understanding of music into their work, by using notes and intervals to craft melodies, chords to build progressions,…

Sound · Computer Science 2024-10-02 Megan Wei , Michael Freeman , Chris Donahue , Chen Sun

Music Information Retrieval (MIR) research is increasingly leveraging representation learning to obtain more compact, powerful music audio representations for various downstream MIR tasks. However, current representation evaluation methods…

Sound · Computer Science 2023-12-13 Christos Plachouras , Pablo Alonso-Jiménez , Dmitry Bogdanov

In recent years, foundation models have become very popular due to their exceptional performance, mainly in natural language (NLP) tasks where they were first introduced. These models usually consist of hundreds of millions, or even…

Sound · Computer Science 2026-01-15 Petros Vavaroutsos , Theodoros Palamas , Pantelis Vikatos

Moonbeam is a transformer-based foundation model for symbolic music, pretrained on a large and diverse collection of MIDI data totaling 81.6K hours of music and 18 billion tokens. Moonbeam incorporates music-domain inductive biases by…

Sound · Computer Science 2025-05-22 Zixun Guo , Simon Dixon

Recently, research on audio foundation models has witnessed notable advances, as illustrated by the ever improving results on complex downstream tasks. Subsequently, those pretrained networks have quickly been used for various audio…

Sound · Computer Science 2025-02-19 David Genova , Philippe Esling , Tom Hurlin

Foundation models have revolutionized music information retrieval, but questions remain about their ability to generalize across diverse musical traditions. This paper presents a comprehensive evaluation of five state-of-the-art audio…

Sound · Computer Science 2025-06-23 Charilaos Papaioannou , Emmanouil Benetos , Alexandros Potamianos

In this work, we introduce the task of singing voice deepfake source attribution (SVDSA). We hypothesize that multimodal foundation models (MMFMs) such as ImageBind, LanguageBind will be most effective for SVDSA as they are better equipped…

Music source separation (MSS) is a task that involves isolating individual sound sources, or stems, from mixed audio signals. This paper presents an ensemble approach to MSS, combining several state-of-the-art architectures to achieve…

Sound · Computer Science 2024-10-29 Saarth Vardhan , Pavani R Acharya , Samarth S Rao , Oorjitha Ratna Jasthi , S Natarajan

The field of Music Information Retrieval (MIR) is fragmented, with specialized models excelling at isolated tasks. In this work, we challenge this paradigm by introducing a unified foundation model named MuFun for holistic music…

Sound · Computer Science 2025-08-05 Yi Jiang , Wei Wang , Xianwen Guo , Huiyun Liu , Hanrui Wang , Youri Xu , Haoqi Gu , Zhongqian Xie , Chuanjiang Luo

Foundation models promise to unify multiple clinical tasks within a single framework, but recent ultrasound studies report that unified models can underperform task-specific baselines. We hypothesize that this degradation arises not from…

Image and Video Processing · Electrical Eng. & Systems 2026-05-25 Fangyijie Wang , Tanya Akumu , Vien Ngoc Dang , Amelia Jiménez-Sánchez , Jieyun Bai , Guénolé Silvestre , Karim Lekadir , Kathleen M. Curran

Foundation models (FMs) unlock unprecedented multimodal and multitask intelligence, yet their cloud-centric deployment precludes real-time responsiveness and compromises user privacy. Meanwhile, monolithic execution at the edge remains…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-01-28 Juan Zhu , Zixin Wang , Shenghui Song , Jun Zhang , Khaled Ben Letaief

Symbolic Music Generation relies on the contextual representation capabilities of the generative model, where the most prevalent approach is the Transformer-based model. The learning of musical context is also related to the structural…

Sound · Computer Science 2022-07-12 Guowei Wu , Shipei Liu , Xiaoya Fan

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

Sound · Computer Science 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music audios, we construct two flow matching networks to model the…

‹ Prev 1 2 3 10 Next ›