English
Related papers

Related papers: The SJTU X-LANCE Lab System for MSR Challenge 2025

200 papers

A main challenge in applying deep learning to music processing is the availability of training data. One potential solution is Multi-task Learning, in which the model also learns to solve related auxiliary tasks on additional datasets to…

Sound · Computer Science 2018-04-06 Daniel Stoller , Sebastian Ewert , Simon Dixon

Recently, multi-band spectrogram-based approaches such as Band-Split RNN (BSRNN) have demonstrated promising results for music source separation. In our recent work, we introduce the BS-RoFormer model which inherits the idea of band-split…

Sound · Computer Science 2023-10-04 Ju-Chiang Wang , Wei-Tsung Lu , Minz Won

This paper describes our audio-quality-based multi-strategy approach for the audio-visual target speaker extraction (AVTSE) task in the Multi-modal Information based Speech Processing (MISP) 2023 Challenge. Specifically, our approach adopts…

Sound · Computer Science 2024-03-08 Runduo Han , Xiaopeng Yan , Weiming Xu , Pengcheng Guo , Jiayao Sun , He Wang , Quan Lu , Ning Jiang , Lei Xie

The aim of this study is to implement a method to remove ambient noise in biomedical sounds captured in auscultation. We propose an incremental approach based on multichannel non-negative matrix partial co-factorization (NMPCF) for ambient…

We investigate which loss functions provide better separations via benchmarking an extensive set of those for music source separation. To that end, we first survey the most representative audio source separation losses we identified, to…

Sound · Computer Science 2022-02-17 Enric Gusó , Jordi Pons , Santiago Pascual , Joan Serrà

Music source separation and pitch estimation are two vital tasks in music information retrieval. Typically, the input of pitch estimation is obtained from the output of music source separation. Therefore, existing methods have tried to…

Sound · Computer Science 2025-01-08 Haojie Wei , Jun Yuan , Rui Zhang , Quanyu Dai , Yueguo Chen

Supervised deep learning approaches to underdetermined audio source separation achieve state-of-the-art performance but require a dataset of mixtures along with their corresponding isolated source signals. Such datasets can be extremely…

This paper reports on the design and results of the 2024 ICASSP SP Cadenza Challenge: Music Demixing/Remixing for Hearing Aids. The Cadenza project is working to enhance the audio quality of music for those with a hearing loss. The scenario…

Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline -…

Sound · Computer Science 2025-07-24 Tobias Morocutti , Jonathan Greif , Paul Primus , Florian Schmid , Gerhard Widmer

Self-supervised learning (SSL) has shown promising results in various speech and natural language processing applications. However, its efficacy in music information retrieval (MIR) still remains largely unexplored. While previous SSL…

This paper presents the NTIRE 2025 image super-resolution ($\times$4) challenge, one of the associated competitions of the 10th NTIRE Workshop at CVPR 2025. The challenge aims to recover high-resolution (HR) images from low-resolution (LR)…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Zheng Chen , Kai Liu , Jue Gong , Jingkai Wang , Lei Sun , Zongwei Wu , Radu Timofte , Yulun Zhang , Xiangyu Kong , Xiaoxuan Yu , Hyunhee Park , Suejin Han , Hakjae Jeon , Dafeng Zhang , Hyung-Ju Chun , Donghun Ryou , Inju Ha , Bohyung Han , Lu Zhao , Yuyi Zhang , Pengyu Yan , Jiawei Hu , Pengwei Liu , Fengjun Guo , Hongyuan Yu , Pufan Xu , Zhijuan Huang , Shuyuan Cui , Peng Guo , Jiahui Liu , Dongkai Zhang , Heng Zhang , Huiyuan Fu , Huadong Ma , Yanhui Guo , Sisi Tian , Xin Liu , Jinwen Liang , Jie Liu , Jie Tang , Gangshan Wu , Zeyu Xiao , Zhuoyuan Li , Yinxiang Zhang , Wenxuan Cai , Vijayalaxmi Ashok Aralikatti , Nikhil Akalwadi , G Gyaneshwar Rao , Chaitra Desai , Ramesh Ashok Tabib , Uma Mudenagudi , Marcos V. Conde , Alejandro Merino , Bruno Longarela , Javier Abad , Weijun Yuan , Zhan Li , Zhanglu Chen , Boyang Yao , Aagam Jain , Milan Kumar Singh , Ankit Kumar , Shubh Kawa , Divyavardhan Singh , Anjali Sarvaiya , Kishor Upla , Raghavendra Ramachandra , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Risheek V Hiremath , Yashaswini Palani , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Jingwei Liao , Yuqing Yang , Wenda Shao , Junyi Zhao , Qisheng Xu , Kele Xu , Sunder Ali Khowaja , Ik Hyun Lee , Snehal Singh Tomar , Rajarshi Ray , Klaus Mueller , Sachin Chaudhary , Surya Vashisth , Akshay Dudhane , Praful Hambarde , Satya Naryan Tazi , Prashant Patil , Santosh Kumar Vipparthi , Subrahmanyam Murala , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Zahra Moammeri , Ahmad Mahmoudi-Aznaveh , Ali Karbasi , Hossein Motamednia , Liangyan Li , Guanhua Zhao , Kevin Le , Yimo Ning , Haoxuan Huang , Jun Chen

This paper summarizes the 3rd NTIRE challenge on stereo image super-resolution (SR) with a focus on new solutions and results. The task of this challenge is to super-resolve a low-resolution stereo image pair to a high-resolution one with a…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Longguang Wang , Yulan Guo , Juncheng Li , Hongda Liu , Yang Zhao , Yingqian Wang , Zhi Jin , Shuhang Gu , Radu Timofte

Blind source separation (BSS) techniques aims at joint estimation of source signals and a mixing matrix from observations of mixtures. This paper addresses a doubly nonstationary BSS problem, where the mixing matrix is time dependent and…

Signal Processing · Electrical Eng. & Systems 2019-06-25 Adrien Meynard

Musical source separation (MSS) has recently seen a big breakthrough in separating instruments from a mixture in the context of Western music, but research on non-Western instruments is still limited due to a lack of data. In this demo, we…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Richa Namballa , Giovana Morais , Magdalena Fuentes

Figure skating scoring is challenging because it requires judging the technical moves of the players as well as their coordination with the background music. Most learning-based methods cannot solve it well for two reasons: 1) each move in…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Jingfei Xia , Mingchen Zhuge , Tiantian Geng , Shun Fan , Yuantai Wei , Zhenyu He , Feng Zheng

In recent years, automated radiology report generation has experienced significant growth. This paper introduces MRScore, an automatic evaluation metric tailored for radiology report generation by leveraging Large Language Models (LLMs).…

Computation and Language · Computer Science 2024-04-30 Yunyi Liu , Zhanyu Wang , Yingshu Li , Xinyu Liang , Lingqiao Liu , Lei Wang , Luping Zhou

Environmental sound analysis is currently getting more and more attentions. In the domain, acoustic scene classification and acoustic event classification are two closely related tasks. In this letter, a two-stage method is proposed for the…

Sound · Computer Science 2021-03-31 Weiping Zheng , Dacan Jiang , Gansen Zhao

Environmental sound recognition (ESR) is an emerging research topic in audio pattern recognition. Many tasks are presented to resort to computational models for ESR in real-life applications. However, current models are usually designed for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-22 Jisheng Bai , Jianfeng Chen , Mou Wang , Muhammad Saad Ayub

Spatial reasoning has emerged as a critical capability for Multimodal Large Language Models (MLLMs), drawing increasing attention and rapid advancement. However, existing benchmarks primarily focus on single-step perception-to-judgment…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Rui Zhu , Xin Shen , Shuchen Wu , Chenxi Miao , Xin Yu , Yang Li , Weikang Li , Deguo Xia , Jizhou Huang

Retrieval-augmented systems are typically evaluated in settings where information required to answer the query can be found within a single source or the answer is short-form or factoid-based. However, many real-world applications demand…

Computation and Language · Computer Science 2025-08-29 Rohan Phanse , Yijie Zhou , Kejian Shi , Wencai Zhang , Yixin Liu , Yilun Zhao , Arman Cohan