中文
相关论文

相关论文: Mel-Band RoFormer for Music Source Separation

200 篇论文

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack…

音频与语音处理 · 电气工程与系统科学 2025-02-18 Jaime Garcia-Martinez , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen , Julio J. Carabias-Orti , Pedro Vera-Candeas

Recently, many deep learning based beamformers have been proposed for multi-channel speech separation. Nevertheless, most of them rely on extra cues known in advance, such as speaker feature, face image or directional information. In this…

音频与语音处理 · 电气工程与系统科学 2022-12-08 Yanjie Fu , Haoran Yin , Meng Ge , Longbiao Wang , Gaoyan Zhang , Jianwu Dang , Chengyun Deng , Fei Wang

We present a new table structure recognition (TSR) approach, called TSRFormer, to robustly recognizing the structures of complex tables with geometrical distortions from various table images. Unlike previous methods, we formulate table…

计算机视觉与模式识别 · 计算机科学 2022-08-10 Weihong Lin , Zheng Sun , Chixiang Ma , Mingze Li , Jiawei Wang , Lei Sun , Qiang Huo

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recently, such models have been used to synthesize audio waveforms…

Phase-Based Ranging (PBR) offers several advantages for estimating distances between wirelessly connected devices, including high accuracy over large distances and the removal of the need for antenna arrays at each transceiver. This study…

信号处理 · 电气工程与系统科学 2025-11-26 Pantelis Stefanakis , Ming Shen

Mel-scale spectrum features are used in various recognition and classification tasks on speech signals. There is no reason to expect that these features are optimal for all different tasks, including speaker verification (SV). This paper…

音频与语音处理 · 电气工程与系统科学 2022-06-16 Jingyu Li , Yusheng Tian , Tan Lee

While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises with other dominant yet less intrusive sounds. However,…

声音 · 计算机科学 2026-02-02 Chi Zuo , Martin B. Møller , Pablo Martínez-Nuevo , Huayang Huang , Yu Wu , Ye Zhu

In this paper, a block-based inter-band predictor (BIP) with multilayer propagation neural network model (MLPNN) is presented by a completely new framework. This predictor can combine with diversity entropy coding methods. Hyperspectral…

多媒体 · 计算机科学 2019-02-13 Rui Dusselaar , Manoranjan Paul

Recurrent Neural Networks (RNNs) have long been the dominant architecture in sequence-to-sequence learning. RNNs, however, are inherently sequential models that do not allow parallelization of their computations. Transformers are emerging…

音频与语音处理 · 电气工程与系统科学 2021-03-10 Cem Subakan , Mirco Ravanelli , Samuele Cornell , Mirko Bronzi , Jianyuan Zhong

The Inaugural Music Source Restoration (MSR) Challenge targets the recovery of original, unprocessed stems from fully mixed and mastered music. Unlike conventional music source separation, MSR requires reversing complex production processes…

声音 · 计算机科学 2026-03-19 Xinlong Deng , Yu Xia , Jie Jiang

We introduce a framework for audio source separation using embeddings on a hyperbolic manifold that compactly represent the hierarchical relationship between sound sources and time-frequency features. Inspired by recent successes modeling…

音频与语音处理 · 电气工程与系统科学 2022-12-12 Darius Petermann , Gordon Wichern , Aswin Subramanian , Jonathan Le Roux

We introduce a new method for robust beamforming, where the goal is to estimate a signal from array samples when there is uncertainty in the angle of arrival. Our method offers state-of-the-art performance on narrowband signals and is…

信号处理 · 电气工程与系统科学 2024-06-25 Nakul Singh , Coleman DeLude , Mark A. Davenport , Justin Romberg

Performance of multicell systems is inevitably limited by interference and available resources. Although intercell interference can be mitigated by Base Station (BS) Coordination, the demand on inter-BS information exchange and…

信息论 · 计算机科学 2014-07-10 Mohammad Hossein Akbari , Vahid Tabataba Vakili

Singing Voice Separation (SVS) tries to separate singing voice from a given mixed musical signal. Recently, many U-Net-based models have been proposed for the SVS task, but there were no existing works that evaluate and compare various…

音频与语音处理 · 电气工程与系统科学 2020-10-09 Woosung Choi , Minseok Kim , Jaehwa Chung , Daewon Lee , Soonyoung Jung

We propose a novel Neural Steering technique that adapts the target area of a spatial-aware multi-microphone sound source separation algorithm during inference without the necessity of retraining the deep neural network (DNN). To achieve…

音频与语音处理 · 电气工程与系统科学 2024-10-23 Martin Strauss , Wolfgang Mack , María Luis Valero , Okan Köpüklü

In this paper we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel non-negative matrix factorization (NMF) and extracting…

声音 · 计算机科学 2017-10-30 Joonas Nikunen , Aleksandr Diment , Tuomas Virtanen

This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare the rank characteristics of Mel-spectrum with other common…

声音 · 计算机科学 2025-11-11 Andong Li , Tong Lei , Rilin Chen , Kai Li , Meng Yu , Xiaodong Li , Dong Yu , Chengshi Zheng

Audio source separation aims to separate a mixture into target sources. Previous audio source separation systems usually conduct one-step inference, which does not fully explore the separation ability of models. In this work, we reveal that…

声音 · 计算机科学 2025-05-27 Yongyi Zang , Jingyi Li , Qiuqiang Kong

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal statistics for time-invariant minimum variance…

声音 · 计算机科学 2021-10-04 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

Consumer-grade music recordings such as those captured by mobile devices typically contain distortions in the form of background noise, reverb, and microphone-induced EQ. This paper presents a deep learning approach to enhance low-quality…

声音 · 计算机科学 2022-04-29 Nikhil Kandpal , Oriol Nieto , Zeyu Jin