English
Related papers

Related papers: Reading to Listen at the Cocktail Party: Multi-Mod…

200 papers

Speech data collected in real-world scenarios often encounters two issues. First, multiple sources may exist simultaneously, and the number of sources may vary with time. Second, the existence of background noise in recording is inevitable.…

Sound · Computer Science 2020-05-21 Yuan-Kuei Wu , Chao-I Tuan , Hung-yi Lee , Yu Tsao

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-17 Jisi Zhang , Catalin Zorila , Rama Doddipatla , Jon Barker

Accurate recognition of cocktail party speech containing overlapping speakers, noise and reverberation remains a highly challenging task to date. Motivated by the invariance of visual modality to acoustic signal corruption, an audio-visual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-07 Guinan Li , Jiajun Deng , Mengzhe Geng , Zengrui Jin , Tianzi Wang , Shujie Hu , Mingyu Cui , Helen Meng , Xunying Liu

Recent studies in neural network-based monaural speech separation (SS) have achieved a remarkable success thanks to increasing ability of long sequence modeling. However, they would degrade significantly when put under realistic noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Yuchen Hu , Chen Chen , Heqing Zou , Xionghu Zhong , Eng Siong Chng

The problem of speech separation, also known as the cocktail party problem, refers to the task of isolating a single speech signal from a mixture of speech signals. Previous work on source separation derived an upper bound for the source…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-27 Shahar Lutati , Eliya Nachmani , Lior Wolf

This work introduces sequential neural beamforming, which alternates between neural network based spectral separation and beamforming based spatial separation. Our neural networks for separation use an advanced convolutional architecture…

Deep learning based models have significantly improved the performance of speech separation with input mixtures like the cocktail party. Prominent methods (e.g., frequency-domain and time-domain speech separation) usually build regression…

Sound · Computer Science 2022-01-11 Jing Shi , Xuankai Chang , Tomoki Hayashi , Yen-Ju Lu , Shinji Watanabe , Bo Xu

In this paper, we analyzed how audio-visual speech enhancement can help to perform the ASR task in a cocktail party scenario. Therefore we considered two simple end-to-end LSTM-based models that perform single-channel audio-visual speech…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-28 Luca Pasa , Giovanni Morrone , Leonardo Badino

Target speech separation refers to isolating target speech from a multi-speaker mixture signal by conditioning on auxiliary information about the target speaker. Different from the mainstream audio-visual approaches which usually require…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Leyuan Qu , Cornelius Weber , Stefan Wermter

Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have…

Sound · Computer Science 2018-05-16 Hiroshi Seki , Takaaki Hori , Shinji Watanabe , Jonathan Le Roux , John R. Hershey

In this paper, we address the problem of enhancing the speech of a speaker of interest in a cocktail party scenario when visual information of the speaker of interest is available. Contrary to most previous studies, we do not learn visual…

Computation and Language · Computer Science 2021-02-04 Giovanni Morrone , Luca Pasa , Vadim Tikhanoff , Sonia Bergamaschi , Luciano Fadiga , Leonardo Badino

Visual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The…

Sound · Computer Science 2023-02-23 Meng Liu , Kong Aik Lee , Longbiao Wang , Hanyi Zhang , Chang Zeng , Jianwu Dang

Recent progress in separating the speech signals from multiple overlapping speakers using a single audio channel has brought us closer to solving the cocktail party problem. However, most studies in this area use a constrained problem…

Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning network that…

Sound · Computer Science 2022-10-13 Tao Sun , Nidal Abuhajar , Shuyu Gong , Zhewei Wang , Charles D. Smith , Xianhui Wang , Li Xu , Jundong Liu

One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-19 Joe Caroselli , Arun Narayanan , Yiteng Huang

In this paper, a novel architecture for speaker recognition is proposed by cascading speech enhancement and speaker processing. Its aim is to improve speaker recognition performance when speech signals are corrupted by noise. Instead of…

Computation and Language · Computer Science 2020-05-25 Yanpei Shi , Qiang Huang , Thomas Hain

Our goal is to isolate individual speakers from multi-talker simultaneous speech in videos. Existing works in this area have focussed on trying to separate utterances from known speakers in controlled environments. In this paper, we propose…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

Although fully end-to-end speaker diarization systems have made significant progress in recent years, modular systems often achieve superior results in real-world scenarios due to their greater adaptability and robustness. Historically,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-26 Ruoyu Wang , Shutong Niu , Gaobin Yang , Jun Du , Shuangqing Qian , Tian Gao , Jia Pan

The field of speech separation, addressing the "cocktail party problem", has seen revolutionary advances with DNNs. Speech separation enhances clarity in complex acoustic environments and serves as crucial pre-processing for speech…

Sound · Computer Science 2025-08-15 Kai Li , Guo Chen , Wendi Sang , Yi Luo , Zhuo Chen , Shuai Wang , Shulin He , Zhong-Qiu Wang , Andong Li , Zhiyong Wu , Xiaolin Hu