中文
相关论文

相关论文: SongFormer: Scaling Music Structure Analysis with …

200 篇论文

Music structure analysis (MSA) systems aim to segment a song recording into non-overlapping sections with useful labels. Previous MSA systems typically predict abstract labels in a post-processing step and require the full context of the…

声音 · 计算机科学 2022-11-30 Ju-Chiang Wang , Jordan B. L. Smith , Yun-Ning Hung

Music structure analysis (MSA) methods traditionally search for musically meaningful patterns in audio: homogeneity, repetition, novelty, and segment-length regularity. Hand-crafted audio features such as MFCCs or chromagrams are often used…

音频与语音处理 · 电气工程与系统科学 2022-05-03 Ju-Chiang Wang , Jordan B. L. Smith , Wei-Tsung Lu , Xuchen Song

Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on supervised deep learning, but these methods are bottlenecked by the need for heavily annotated data…

声音 · 计算机科学 2026-03-31 Axel Marmoret

Music Structure Analysis (MSA) consists in segmenting a music piece in several distinct sections. We approach MSA within a compression framework, under the hypothesis that the structure is more easily revealed by a simplified representation…

声音 · 计算机科学 2022-04-18 Axel Marmoret , Jérémy E. Cohen , Frédéric Bimbot

The deep learning community has witnessed an exponentially growing interest in self-supervised learning (SSL). However, it still remains unexplored how to build a framework for learning useful representations of raw music waveforms in a…

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

音频与语音处理 · 电气工程与系统科学 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

Creating lyrics and melodies for the vocal track in a symbolic format, known as song composition, demands expert musical knowledge of melody, an advanced understanding of lyrics, and precise alignment between them. Despite achievements in…

声音 · 计算机科学 2025-06-03 Shuangrui Ding , Zihan Liu , Xiaoyi Dong , Pan Zhang , Rui Qian , Junhao Huang , Conghui He , Dahua Lin , Jiaqi Wang

We present Music Tagging Transformer that is trained with a semi-supervised approach. The proposed model captures local acoustic characteristics in shallow convolutional layers, then temporally summarizes the sequence of the extracted…

声音 · 计算机科学 2021-11-29 Minz Won , Keunwoo Choi , Xavier Serra

Chord recognition serves as a critical task in music information retrieval due to the abstract and descriptive nature of chords in music analysis. While audio chord recognition systems have achieved significant accuracy for small…

Music Structure Analysis (MSA) is the task aiming at identifying musical segments that compose a music track and possibly label them based on their similarity. In this paper we propose a supervised approach for the task of music boundary…

声音 · 计算机科学 2023-09-06 Geoffroy Peeters

Audio-based music structure analysis (MSA) is an essential task in Music Information Retrieval that remains challenging due to the complexity and variability of musical form. Recent advances highlight the potential of fine-tuning…

声音 · 计算机科学 2025-07-21 Yixiao Zhang , Haonan Chen , Ju-Chiang Wang , Jitong Chen

Music generative artificial intelligence (AI) is rapidly expanding music content, necessitating automated song aesthetics evaluation. However, existing studies largely focus on speech, audio or singing quality, leaving song aesthetics…

声音 · 计算机科学 2026-01-21 Yishan Lv , Jing Luo , Boyuan Ju , Yang Zhang , Xinda Wu , Bo Yuan , Xinyu Yang

Recent commercial systems such as Suno demonstrate strong capabilities in long-form song generation, while academic research remains largely non-reproducible due to the lack of publicly available training data, hindering fair comparison and…

Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source…

声音 · 计算机科学 2026-01-01 Shuyang Liu , Yuan Jin , Rui Lin , Shizhe Chen , Junyu Dai , Tao Jiang

In this paper, we propose GaMMA, a state-of-the-art (SoTA) large multimodal model (LMM) designed to achieve comprehensive musical content understanding. GaMMA inherits the streamlined encoder-decoder design of LLaVA, enabling effective…

声音 · 计算机科学 2026-05-04 Zuyao You , Zhesong Yu , Mingyu Liu , Bilei Zhu , Yuan Wan , Zuxuan Wu

Music Structure Analysis (MSA) is a Music Information Retrieval task consisting of representing a song in a simplified, organized manner by breaking it down into sections typically corresponding to ``chorus'', ``verse'', ``solo'', etc. In…

声音 · 计算机科学 2023-12-01 Axel Marmoret , Jérémy E. Cohen , Frédéric Bimbot

The lack of data tends to limit the outcomes of deep learning research, particularly when dealing with end-to-end learning stacks processing raw data such as waveforms. In this study, 1.2M tracks annotated with musical labels are available…

声音 · 计算机科学 2018-06-18 Jordi Pons , Oriol Nieto , Matthew Prockup , Erik Schmidt , Andreas Ehmann , Xavier Serra

Large language models perform strongly on general tasks but remain constrained in specialized settings such as music, particularly in the music-entertainment domain, where corpus scale, purity, and the match between data and training…

计算与语言 · 计算机科学 2025-11-19 Kai Tian , Yirong Mao , Wendong Bi , Hanjie Wang , Que Wenhui

We investigate the feasibility of a singing voice synthesis (SVS) system by using a decomposed framework to improve flexibility in generating singing voices. Due to data-driven approaches, SVS performs a music score-to-waveform mapping;…

声音 · 计算机科学 2024-07-15 Lester Phillip Violeta , Taketo Akama

We present a family of open-source Music Foundation Models designed to advance large-scale music understanding and generation across diverse tasks and modalities. Our framework consists of four major components: (1) HeartCLAP, an audio-text…

‹ 上一页 1 2 3 10 下一页 ›