English
Related papers

Related papers: The SJTU X-LANCE Lab System for MSR Challenge 2025

200 papers

In this paper we propose a conditioned UNet for Music Source Separation (MSS). MSS is generally performed by multi-output neural networks, typically UNets, with each output representing a particular stem from a predefined instrument…

Sound · Computer Science 2025-12-19 Ken O'Hanlon , Basil Woods , Lin Wang , Mark Sandler

We consider the multiuser successive refinement (MSR) problem, where the users are connected to a central server via links with different noiseless capacities, and each user wishes to reconstruct in a successive-refinement fashion. An…

Information Theory · Computer Science 2007-11-13 Chao Tian , Jun Chen , Suhas Diggavi

Music source separation (MSS) faces challenges due to the limited availability of correctly-labeled individual instrument tracks. With the push to acquire larger datasets to improve MSS performance, the inevitability of encountering…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-25 Junghyun Koo , Yunkee Chae , Chang-Bin Jeon , Kyogu Lee

Multimodal Large Language Models (MLLMs) have demonstrated capabilities in audio understanding, but current evaluations may obscure fundamental weaknesses in relational reasoning. We introduce the Music Understanding and Structural…

Artificial Intelligence · Computer Science 2025-10-23 Brandon James Carone , Iran R. Roman , Pablo Ripollés

In this paper, we present MuLanTTS, the Microsoft end-to-end neural text-to-speech (TTS) system designed for the Blizzard Challenge 2023. About 50 hours of audiobook corpus for French TTS as hub task and another 2 hours of speaker…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-13 Zhihang Xu , Shaofei Zhang , Xi Wang , Jiajun Zhang , Wenning Wei , Lei He , Sheng Zhao

We propose a generative framework for multi-track music source separation (MSS) that reformulates the task as conditional discrete token generation. Unlike conventional approaches that directly estimate continuous signals in the time or…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-20 Pengbo Lyu , Xiangyu Zhao , Chengwei Liu , Haoyin Yan , Xiaotao Liang , Hongyu Wang , Shaofei Xue

Self-supervised cross-modal super-resolution (SR) can overcome the difficulty of acquiring paired training data, but is challenging because only low-resolution (LR) source and high-resolution (HR) guide images from different modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Xiaoyu Dong , Naoto Yokoya , Longguang Wang , Tatsumi Uezato

Machine unlearning (MU) enables the removal of selected training data from trained models, to address privacy compliance, security, and liability issues in recommender systems. Existing MU benchmarks poorly reflect real-world recommender…

Information Retrieval · Computer Science 2026-03-10 Pierre Lubitzsch , Maarten de Rijke , Sebastian Schelter

In this paper, we describe the top-scoring submissions for team RTZR VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22) in the closed dataset, speaker verification Track 1. The top performed system is a fusion of 7 models, which…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-22 Sangwon Suh , Sunjong Park

In this technical report, the systems we submitted for subtask 1B of the DCASE 2021 challenge, regarding audiovisual scene classification, are described in detail. They are essentially multi-source transformers employing a combination of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-20 Wim Boes , Hugo Van hamme

Despite phenomenal progress in recent years, state-of-the-art music separation systems produce source estimates with significant perceptual shortcomings, such as adding extraneous noise or removing harmonics. We propose a post-processing…

Sound · Computer Science 2022-08-29 Noah Schaffer , Boaz Cogan , Ethan Manilow , Max Morrison , Prem Seetharaman , Bryan Pardo

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any other…

Sound · Computer Science 2021-04-29 Alexandre Défossez , Nicolas Usunier , Léon Bottou , Francis Bach

Over the years, Music Information Retrieval (MIR) has proposed various models pretrained on large amounts of music data. Transfer learning showcases the proven effectiveness of pretrained backend models with a broad spectrum of downstream…

Information Retrieval · Computer Science 2024-09-16 Yan-Martin Tamm , Anna Aljanaki

While deep neural network-based music source separation (MSS) is very effective and achieves high performance, its model size is often a problem for practical deployment. Deep implicit architectures such as deep equilibrium models (DEQ)…

This paper provides a comprehensive review of the NTIRE 2024 challenge, focusing on efficient single-image super-resolution (ESR) solutions and their outcomes. The task of this challenge is to super-resolve an input image with a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Bin Ren , Yawei Li , Nancy Mehta , Radu Timofte , Hongyuan Yu , Cheng Wan , Yuxin Hong , Bingnan Han , Zhuoyuan Wu , Yajun Zou , Yuqing Liu , Jizhe Li , Keji He , Chao Fan , Heng Zhang , Xiaolin Zhang , Xuanwu Yin , Kunlong Zuo , Bohao Liao , Peizhe Xia , Long Peng , Zhibo Du , Xin Di , Wangkai Li , Yang Wang , Wei Zhai , Renjing Pei , Jiaming Guo , Songcen Xu , Yang Cao , Zhengjun Zha , Yan Wang , Yi Liu , Qing Wang , Gang Zhang , Liou Zhang , Shijie Zhao , Long Sun , Jinshan Pan , Jiangxin Dong , Jinhui Tang , Xin Liu , Min Yan , Qian Wang , Menghan Zhou , Yiqiang Yan , Yixuan Liu , Wensong Chan , Dehua Tang , Dong Zhou , Li Wang , Lu Tian , Barsoum Emad , Bohan Jia , Junbo Qiao , Yunshuai Zhou , Yun Zhang , Wei Li , Shaohui Lin , Shenglong Zhou , Binbin Chen , Jincheng Liao , Suiyi Zhao , Zhao Zhang , Bo Wang , Yan Luo , Yanyan Wei , Feng Li , Mingshen Wang , Yawei Li , Jinhan Guan , Dehua Hu , Jiawei Yu , Qisheng Xu , Tao Sun , Long Lan , Kele Xu , Xin Lin , Jingtong Yue , Lehan Yang , Shiyi Du , Lu Qi , Chao Ren , Zeyu Han , Yuhan Wang , Chaolin Chen , Haobo Li , Mingjun Zheng , Zhongbao Yang , Lianhong Song , Xingzhuo Yan , Minghan Fu , Jingyi Zhang , Baiang Li , Qi Zhu , Xiaogang Xu , Dan Guo , Chunle Guo , Jiadi Chen , Huanhuan Long , Chunjiang Duanmu , Xiaoyan Lei , Jie Liu , Weilin Jia , Weifeng Cao , Wenlong Zhang , Yanyu Mao , Ruilong Guo , Nihao Zhang , Qian Wang , Manoj Pandey , Maksym Chernozhukov , Giang Le , Shuli Cheng , Hongyuan Wang , Ziyan Wei , Qingting Tang , Liejun Wang , Yongming Li , Yanhui Guo , Hao Xu , Akram Khatami-Rizi , Ahmad Mahmoudi-Aznaveh , Chih-Chung Hsu , Chia-Ming Lee , Yi-Shiuan Chou , Amogh Joshi , Nikhil Akalwadi , Sampada Malagi , Palani Yashaswini , Chaitra Desai , Ramesh Ashok Tabib , Ujwala Patil , Uma Mudenagudi

Music source separation is the task of isolating the instrumental tracks from a music song. Despite its spectacular recent progress, the trend towards more complex architectures and training protocols exacerbates reproducibility issues. The…

Sound · Computer Science 2026-03-11 Paul Magron , Romain Serizel , Constance Douwes

Background and Objective: Processing electrophysiological signals often requires blind source separation (BSS) due to the nature of mixing source signals. However, its complex computational demands make real-time BSS challenging. The…

Human-Computer Interaction · Computer Science 2024-11-28 Yao Li , Haowen Zhao , Yunfei Liu , Xu Zhang

Recently, pre-trained models for music information retrieval based on self-supervised learning (SSL) are becoming popular, showing success in various downstream tasks. However, there is limited research on the specific meanings of the…

Sound · Computer Science 2025-05-23 Yizhi Zhou , Haina Zhu , Hangting Chen

Recently, many methods based on deep learning have been proposed for music source separation. Some state-of-the-art methods have shown that stacking many layers with many skip connections improve the SDR performance. Although such a deep…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-25 Minseok Kim , Woosung Choi , Jaehwa Chung , Daewon Lee , Soonyoung Jung

Vocal recordings on consumer devices commonly suffer from multiple concurrent degradations: noise, reverberation, band-limiting, and clipping. We present Smule Renaissance Small (SRS), a compact single-stage model that performs end-to-end…