English
Related papers

Related papers: Dual-path Transformer Based Neural Beamformer for …

200 papers

Speech separation (SS) seeks to disentangle a multi-talker speech mixture into single-talker speech streams. Although SS can be generally achieved using offline methods, such a processing paradigm is not suitable for real-time streaming…

Sound · Computer Science 2025-04-04 Wupeng Wang , Zexu Pan , Xinke Li , Shuai Wang , Haizhou Li

This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction, the BiVocoder takes…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

Target speaker extraction is to extract the target speaker, specified by enrollment utterance, in an environment with other competing speakers. Therefore, the task needs to solve two problems, speaker identification and separation, at the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Pengjie Shen , Shulin He , Xueliang Zhang

Speech-to-speech models handle turn-taking naturally but offer limited support for tool-calling or complex reasoning, while production ASR-LLM-TTS voice pipelines offer these capabilities but rely on silence timeouts, which lead to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Shangeth Rajaa

We propose a speaker selection mechanism (SSM) for the training of an end-to-end beamforming neural network, based on recent findings that a listener usually looks to the target speaker with a certain undershot angle. The mechanism allows…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-25 Luan Vinícius Fiorio , Bruno Defraene , Johan David , Alex Young , Frans Widdershoven , Wim van Houtum , Ronald M. Aarts

Recent speech enhancement methods based on convolutional neural networks (CNNs) and transformer have been demonstrated to efficaciously capture time-frequency (T-F) information on spectrogram. However, the correlation of each channels of…

Sound · Computer Science 2024-07-16 Jizhen Li , Xinmeng Xu , Weiping Tu , Yuhong Yang , Rong Zhu

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech…

Sound · Computer Science 2021-03-01 Jiatong Shi , Chunlei Zhang , Chao Weng , Shinji Watanabe , Meng Yu , Dong Yu

In this paper, we propose a speech enhancement method us ing dual-path Multi-Channel Linear Prediction (MCLP) filters and multi-norm beamforming. Specifically, the MCLP part in the proposed method is designed with dual-path filters in both…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-25 Chengyuan Qin , Wenmeng Xiong , Jing Zhou , Maoshen Jia , Changchun Bao

We propose a cross-modal transformer-based neural correction models that refines the output of an automatic speech recognition (ASR) system so as to exclude ASR errors. Generally, neural correction models are composed of encoder-decoder…

Computation and Language · Computer Science 2021-07-06 Tomohiro Tanaka , Ryo Masumura , Mana Ihori , Akihiko Takashima , Takafumi Moriya , Takanori Ashihara , Shota Orihashi , Naoki Makishima

Neural end-to-end text-to-speech (TTS) , which adopts either a recurrent model, e.g. Tacotron, or an attention one, e.g. Transformer, to characterize a speech utterance, has achieved significant improvement of speech synthesis. However, it…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-18 Xi Wang , Huaiping Ming , Lei He , Frank K. Soong

Conformer has proven to be effective in many speech processing tasks. It combines the benefits of extracting local dependencies using convolutions and global dependencies using self-attention. Inspired by this, we propose a more flexible,…

Computation and Language · Computer Science 2022-07-08 Yifan Peng , Siddharth Dalmia , Ian Lane , Shinji Watanabe

Using more test-time computation during language model inference, such as generating more intermediate thoughts or sampling multiple candidate answers, has proven effective in significantly improving model performance. This paper takes an…

Machine Learning · Computer Science 2025-08-20 Xingwu Chen , Miao Lu , Beining Wu , Difan Zou

Recent studies in deep learning-based speech separation have proven the superiority of time-domain approaches to conventional time-frequency-based methods. Unlike the time-frequency domain approaches, the time-domain separation systems…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-30 Yi Luo , Zhuo Chen , Takuya Yoshioka

Latest development of neural models has connected the encoder and decoder through a self-attention mechanism. In particular, Transformer, which is solely based on self-attention, has led to breakthroughs in Natural Language Processing (NLP)…

Computation and Language · Computer Science 2019-11-07 Xindian Ma , Peng Zhang , Shuai Zhang , Nan Duan , Yuexian Hou , Dawei Song , Ming Zhou

The human auditory system has the ability to selectively focus on key speech elements in an audio stream while giving secondary attention to less relevant areas such as noise or distortion within the background, dynamically adjusting its…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-09 Nursadul Mamun , John H. L. Hansen

Visual information can serve as an effective cue for target speaker extraction (TSE) and is vital to improving extraction performance. In this paper, we propose AV-SepFormer, a SepFormer-based attention dual-scale model that utilizes cross-…

Binaural speech separation in real-world scenarios often involves moving speakers. Most current speech separation methods use utterance-level permutation invariant training (u-PIT) for training. In inference time, however, the order of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-15 Cong Han , Nima Mesgarani

During the Covid, online meetings have become an indispensable part of our lives. This trend is likely to continue due to their convenience and broad reach. However, background noise from other family members, roommates, office-mates not…

Sound · Computer Science 2022-07-22 Wei Sun , Mei Wang , Lili Qiu

Transformer based models have provided significant performance improvements in monaural speech separation. However, there is still a performance gap compared to a recent proposed upper bound. The major limitation of the current dual-path…

Sound · Computer Science 2023-02-24 Shengkui Zhao , Bin Ma

Extracting the speech of participants in a conversation amidst interfering speakers and noise presents a challenging problem. In this paper, we introduce the novel task of target conversation extraction, where the goal is to extract the…

Computation and Language · Computer Science 2024-09-26 Tuochao Chen , Qirui Wang , Bohan Wu , Malek Itani , Sefik Emre Eskimez , Takuya Yoshioka , Shyamnath Gollakota
‹ Prev 1 8 9 10 Next ›