English
Related papers

Related papers: X-SepFormer: End-to-end Speaker Extraction Network…

200 papers

In this work, we introduce S4M, a new efficient speech separation framework based on neural state-space models (SSM). Motivated by linear time-invariant systems for sequence modeling, our SSM-based approach can efficiently model input…

Sound · Computer Science 2023-05-29 Chen Chen , Chao-Han Huck Yang , Kai Li , Yuchen Hu , Pin-Jui Ku , Eng Siong Chng

End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially.…

Computation and Language · Computer Science 2022-06-10 Biao Zhang , Barry Haddow , Rico Sennrich

In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T based SCD model loss optimizes all output tokens equally. Due to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-06 Guanlong Zhao , Quan Wang , Han Lu , Yiling Huang , Ignacio Lopez Moreno

Supervised learning based on a deep neural network recently has achieved substantial improvement on speech enhancement. Denoising networks learn mapping from noisy speech to clean one directly, or to a spectrum mask which is the ratio…

Sound · Computer Science 2023-03-10 Jaeyoung Kim , Mostafa El-Khamy , Jungwon Lee

Speech translation models are unable to directly process long audios, like TED talks, which have to be split into shorter segments. Speech translation datasets provide manual segmentations of the audios, which are not available in…

The previous SpEx+ has yielded outstanding performance in speaker extraction and attracted much attention. However, it still encounters inadequate utilization of multi-scale information and speaker embedding. To this end, this paper…

Sound · Computer Science 2023-06-29 Jun Chen , Wei Rao , Zilin Wang , Jiuxin Lin , Yukai Ju , Shulin He , Yannan Wang , Zhiyong Wu

This paper proposes a generative pretraining foundation model for high-quality speech restoration tasks. By directly operating on complex-valued short-time Fourier transform coefficients, our model does not rely on any vocoders for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-26 Pin-Jui Ku , Alexander H. Liu , Roman Korostik , Sung-Feng Huang , Szu-Wei Fu , Ante Jukić

Speaker extraction is to extract a target speaker's voice from multi-talker speech. It simulates humans' cocktail party effect or the selective listening ability. The prior work mostly performs speaker extraction in frequency domain, then…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Chenglin Xu , Wei Rao , Eng Siong Chng , Haizhou Li

This paper investigates the use of relative cues for text-based target speech extraction (TSE). We first provide a theoretical justification for relative cues from the perspectives of human perception and label quantization, showing that…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-03 Wang Dai , Archontis Politis , Tuomas Virtanen

This paper addresses end-to-end automatic speech recognition (ASR) for long audio recordings such as lecture and conversational speeches. Most end-to-end ASR models are designed to recognize independent utterances, but contextual…

Computation and Language · Computer Science 2021-04-20 Takaaki Hori , Niko Moritz , Chiori Hori , Jonathan Le Roux

Modern speaker verification systems primarily rely on speaker embeddings, followed by verification based on cosine similarity between the embedding vectors of the enrollment and test utterances. While effective, these methods struggle with…

Sound · Computer Science 2025-07-04 Wan Lin , Junhui Chen , Tianhao Wang , Zhenyu Zhou , Lantian Li , Dong Wang

While existing end-to-end beamformers achieve impressive performance in various front-end speech processing tasks, they usually encapsulate the whole process into a black box and thus lack adequate interpretability. As an attempt to fill…

Sound · Computer Science 2022-03-17 Andong Li , Guochen Yu , Chengshi Zheng , Xiaodong Li

Text-based speech editing (TSE) techniques are designed to enable users to edit the output audio by modifying the input text transcript instead of the audio itself. Despite much progress in neural network-based TSE techniques, the current…

Sound · Computer Science 2023-09-25 Rui Liu , Jiatian Xi , Ziyue Jiang , Haizhou Li

This paper introduces a practical approach for leveraging a real-time deep learning model to alternate between speech enhancement and joint speech enhancement and separation depending on whether the input mixture contains one or two active…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-17 Kashyap Patel , Anton Kovalyov , Issa Panahi

Determining 'who spoke what and when' remains challenging in real-world applications. In typical scenarios, Speaker Diarization (SD) is employed to address the problem of 'who spoke when,' while Target Speaker Extraction (TSE) or Target…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-20 Bang Zeng , Ming Li

Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Hao Ma , Zhiyuan Peng , Xu Li , Yukai Li , Mingjie Shao , Qiuqiang Kong , Ju Liu

Speech separation and enhancement (SSE) has advanced remarkably and achieved promising results in controlled settings, such as a fixed number of speakers and a fixed array configuration. Towards a universal SSE system, single-channel…

Deep learning has become a de facto method of choice for speech enhancement tasks with significant improvements in speech quality. However, real-time processing with reduced size and computations for low-power edge devices drastically…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-28 Monisankha Pal , Arvind Ramanathan , Ted Wada , Ashutosh Pandey

Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec 2.0 deep speech…

We present ESPnet-SE, which is designed for the quick development of speech enhancement and speech separation systems in a single framework, along with the optional downstream speech recognition module. ESPnet-SE is a new project which…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-18 Chenda Li , Jing Shi , Wangyou Zhang , Aswin Shanmugam Subramanian , Xuankai Chang , Naoyuki Kamo , Moto Hira , Tomoki Hayashi , Christoph Boeddeker , Zhuo Chen , Shinji Watanabe