English
Related papers

Related papers: Mel-Band RoFormer for Music Source Separation

200 papers

Music source separation is an audio-to-audio retrieval task of extracting one or more constituent components, or composites thereof, from a musical audio mixture. Each of these constituent components is often referred to as a "stem" in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-28 Karn N. Watcharasupat , Alexander Lerch

In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument…

Sound · Computer Science 2025-12-19 Ken O'Hanlon , Basil Woods , Lin Wang , Mark Sandler

Despite phenomenal progress in recent years, state-of-the-art music separation systems produce source estimates with significant perceptual shortcomings, such as adding extraneous noise or removing harmonics. We propose a post-processing…

Sound · Computer Science 2022-08-29 Noah Schaffer , Boaz Cogan , Ethan Manilow , Max Morrison , Prem Seetharaman , Bryan Pardo

Although the conventional mask-based minimum variance distortionless response (MVDR) could reduce the non-linear distortion, the residual noise level of the MVDR separated speech is still high. In this paper, we propose a spatio-temporal…

Sound · Computer Science 2021-04-06 Yong Xu , Zhuohuang Zhang , Meng Yu , Shi-Xiong Zhang , Dong Yu

Music source separation (MSS) aims to extract 'vocals', 'drums', 'bass' and 'other' tracks from a piece of mixed music. While deep learning methods have shown impressive results, there is a trend toward larger models. In our paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-20 Junyu Chen , Susmitha Vekkot , Pancham Shukla

In this paper, we propose a simple yet effective method for multiple music source separation using convolutional neural networks. Stacked hourglass network, which was originally designed for human pose estimation in natural images, is…

Sound · Computer Science 2018-06-25 Sungheon Park , Taehoon Kim , Kyogu Lee , Nojun Kwak

Autoregressive music generation depends strongly on the audio tokenizer. Existing high-fidelity codecs often use residual multi-codebook quantization, which preserves reconstruction quality but complicates language modeling after sequence…

Sound · Computer Science 2026-05-18 Yuqing Cheng , Xingyu Ma , Guochen Yu , Xiaotao Gu

Music source separation (MSS) aims to extract individual instrument sources from their mixture. While most existing methods focus on the widely adopted four-stem separation setup (vocals, bass, drums, and other instruments), this approach…

Sound · Computer Science 2025-08-06 Yutong Wen , Minje Kim , Paris Smaragdis

Music structure analysis (MSA) underpins music understanding and controllable generation, yet progress has been limited by small, inconsistent corpora. We present SongFormer, a scalable framework that learns from heterogeneous supervision.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-09 Chunbo Hao , Ruibin Yuan , Jixun Yao , Qixin Deng , Xinyi Bai , Yanbo Wang , Wei Xue , Lei Xie

Deep neural networks with convolutional layers usually process the entire spectrogram of an audio signal with the same time-frequency resolutions, number of filters, and dimensionality reduction scale. According to the constant-Q transform,…

Sound · Computer Science 2019-10-22 Emad M. Grais , Fei Zhao , Mark D. Plumbley

With the recent advancements of data driven approaches using deep neural networks, music source separation has been formulated as an instrument-specific supervised problem. While existing deep learning models implicitly absorb the spatial…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-16 Darius Petermann , Minje Kim

This research paper presents a novel deep learning-based neural network architecture, named Y-Net, for achieving music source separation. The proposed architecture performs end-to-end hybrid source separation by extracting features from…

Can we perform an end-to-end music source separation with a variable number of sources using a deep learning model? We present an extension of the Wave-U-Net model which allows end-to-end monaural source separation with a non-fixed number…

Sound · Computer Science 2019-05-10 Olga Slizovskaia , Leo Kim , Gloria Haro , Emilia Gomez

With the increasing usage of radiograph images as a most common medical imaging system for diagnosis, treatment planning, and clinical studies, it is increasingly becoming a vital factor to use machine learning-based systems to provide…

Computer Vision and Pattern Recognition · Computer Science 2020-01-01 Ata Jodeiri , Reza A. Zoroofi , Yuta Hiasa , Masaki Takao , Nobuhiko Sugano , Yoshinobu Sato , Yoshito Otake

Rate splitting multiple access (RSMA) is a promising non-orthogonal transmission strategy for next-generation wireless networks. It has been shown to outperform existing multiple access schemes in terms of spectral and energy efficiency…

Information Theory · Computer Science 2022-11-04 Bho Matthiesen , Yijie Mao , Armin Dekorsy , Petar Popovski , Bruno Clerckx

The task of radio map estimation aims to generate a dense representation of electromagnetic spectrum quantities, such as the received signal strength at each grid point within a geographic region, based on measurements from a subset of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Zheng Fang , Kangjun Liu , Ke Chen , Qingyu Liu , Jianguo Zhang , Lingyang Song , Yaowei Wang

Rate splitting multiple access (RSMA) relies on beamforming design for attaining spectral efficiency and energy efficiency gains over traditional multiple access schemes. While conventional optimization approaches such as weighted minimum…

Signal Processing · Electrical Eng. & Systems 2024-07-10 Yiwen Wang , Yijie Mao , Sijie Ji

Most of the currently successful source separation techniques use the magnitude spectrogram as input, and are therefore by default omitting part of the signal: the phase. To avoid omitting potentially useful information, we study the…

Sound · Computer Science 2019-07-01 Francesc Lluís , Jordi Pons , Xavier Serra

Music source separation represents the task of extracting all the instruments from a given song. Recent breakthroughs on this challenge have gravitated around a single dataset, MUSDB, only limited to four instrument classes. Larger datasets…

Sound · Computer Science 2021-12-02 Alexandru Mocanu , Benjamin Ricaud , Milos Cernak

Tandem mass spectra capture fragmentation patterns that provide key structural information about a molecule. Although mass spectrometry is applied in many areas, the vast majority of small molecules lack experimental reference spectra. For…

Machine Learning · Computer Science 2023-05-03 Adamo Young , Bo Wang , Hannes Röst