English
Related papers

Related papers: The SJTU X-LANCE Lab System for MSR Challenge 2025

200 papers

This paper presents the results of the 2025 Automatic Music Transcription (AMT) Challenge, an online competition to benchmark progress in multi-instrument transcription. Eight teams submitted valid solutions; two outperformed the baseline…

Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and…

Sound · Computer Science 2021-02-08 Ho-Hsiang Wu , Chieh-Chi Kao , Qingming Tang , Ming Sun , Brian McFee , Juan Pablo Bello , Chao Wang

Universal sound separation faces a fundamental misalignment: models optimized for low-level signal metrics often produce semantically contaminated outputs, failing to suppress perceptually salient interference from acoustically similar…

Sound · Computer Science 2026-02-18 Zihan Zhang , Xize Cheng , Zhennan Jiang , Dongjie Fu , Jingyuan Chen , Zhou Zhao , Tao Jin

We present the Open ASR Leaderboard, a reproducible benchmarking platform with community contributions from academia and industry. It compares 86 open-source and proprietary systems across 12 datasets, with English short- and long-form and…

This study introduces a novel formulation to enhance Support Vector Machines (SVMs) in handling class imbalance and noise. Unlike the conventional Soft Margin SVM, which penalizes the magnitude of constraint violations, the proposed model…

Machine Learning · Computer Science 2025-03-20 Seyed Mojtaba Mohasel , Hamidreza Koosha

We present the UTokyo-SaruLab mean opinion score (MOS) prediction system submitted to VoiceMOS Challenge 2022. The challenge is to predict the MOS values of speech samples collected from previous Blizzard Challenges and Voice Conversion…

Music source separation is an audio-to-audio retrieval task of extracting one or more constituent components, or composites thereof, from a musical audio mixture. Each of these constituent components is often referred to as a "stem" in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-28 Karn N. Watcharasupat , Alexander Lerch

With the growing amount of musical data available, automatic instrument recognition, one of the essential problems in Music Information Retrieval (MIR), is drawing more and more attention. While automatic recognition of single instruments…

Sound · Computer Science 2023-06-16 Lifan Zhong , Erica Cooper , Junichi Yamagishi , Nobuaki Minematsu

This paper presents the crossing scheme (X-scheme) for improving the performance of deep neural network (DNN)-based music source separation (MSS) with almost no increasing calculation cost. It consists of three components: (i) multi-domain…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-07 Ryosuke Sawata , Naoya Takahashi , Stefan Uhlich , Shusuke Takahashi , Yuki Mitsufuji

Music Recommendation Systems (MRSs) are a cornerstone of modern streaming platforms. Existing recommendation models, spanning both recall and ranking stages, predominantly rely on collaborative filtering, which fails to exploit the…

Information Retrieval · Computer Science 2026-04-24 Yizhi Zhou , Jia-Qi Yang , De-Chuan Zhan , Da-Wei Zhou

Large language models perform strongly on general tasks but remain constrained in specialized settings such as music, particularly in the music-entertainment domain, where corpus scale, purity, and the match between data and training…

Computation and Language · Computer Science 2025-11-19 Kai Tian , Yirong Mao , Wendong Bi , Hanjie Wang , Que Wenhui

Sequential recommender systems (SRSs) aim to suggest next item for a user based on her historical interaction sequences. Recently, many research efforts have been devoted to attenuate the influence of noisy items in sequences by either…

Information Retrieval · Computer Science 2024-06-21 Xiaofei Zhu , Liang Li , Weidong Liu , Xin Luo

When Large Language Models produce structured outputs such as travel plans, code solutions, or multi-step proofs, individual reasoning steps may appear correct while the output as a whole violates budgets, fails test cases, or contradicts…

Machine Learning · Computer Science 2026-05-20 Shireen Kudukkil Manchingal , Abhey Kalia , Fernanda Gonçalves , Shebin Rawther

This technical report details our systems submitted for Task 3 of the DCASE 2024 Challenge: Audio and Audiovisual Sound Event Localization and Detection (SELD) with Source Distance Estimation (SDE). We address only the audio-only SELD with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Jun Wei Yeow , Ee-Leng Tan , Jisheng Bai , Santi Peksi , Woon-Seng Gan

Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Language Models to interpret full musical notation remains…

This study presents an exploratory evaluation of Music Generation Systems (MGS) within contemporary music production workflows by examining eight open-source systems. The evaluation framework combines technical insights with practical…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-03 Shayan Dadman , Bernt Arild Bremdal , Andreas Bergsland

We introduce our submission to the AudioMOS Challenge (AMC) 2025 Track 3: mean opinion score (MOS) prediction for speech with multiple sampling frequencies (SFs). Our submitted model integrates an SF-independent (SFI) convolutional layer…

Objective: To develop and evaluate a scalable methodology for harmonizing inconsistent units in large-scale clinical datasets, addressing a key barrier to data interoperability. Materials and Methods: We designed a novel unit harmonization…

Machine Learning · Computer Science 2025-11-18 Jordi de la Torre

This paper studies stable recovery of a collection of point sources from its noisy $M+1$ low-frequency Fourier coefficients. We focus on the super-resolution regime where the minimum separation of the point sources is below $1/M$. We…

Information Theory · Computer Science 2019-05-03 Weilin Li , Wenjing Liao

Dereverberation of a moving speech source in the presence of other directional interferers, is a harder problem than that of stationary source and interference cancellation. We explore joint multi channel linear prediction (MCLP) and…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-23 Srikanth Raj Chetupalli , Thippur V. Sreenivas
‹ Prev 1 3 4 5 6 7 10 Next ›