English
Related papers

Related papers: Reading to Listen at the Cocktail Party: Multi-Mod…

200 papers

Speech separation is very important in real-world applications such as human-machine interaction, hearing aids devices, and automatic meeting transcription. In recent years, a significant improvement occurred towards the solution based on…

Sound · Computer Science 2024-08-29 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

Recent advances in multi-modal large language models (MLLMs) have opened new possibilities for unified modeling of speech, text, images, and other modalities. Building on our prior work, this paper examines the conditions and model…

Sound · Computer Science 2025-07-28 Yiwen Guan , Viet Anh Trinh , Vivek Voleti , Jacob Whitehill

Speech separation with several speakers is a challenging task because of the non-stationarity of the speech and the strong signal similarity between interferent sources. Current state-of-the-art solutions can separate well the different…

Signal Processing · Electrical Eng. & Systems 2021-02-09 Nicolas Furnon , Romain Serizel , Irina Illina , Slim Essid

Speech recognition in cocktail-party environments remains a significant challenge for state-of-the-art speech recognition systems, as it is extremely difficult to extract an acoustic signal of an individual speaker from a background of…

Audio and Speech Processing · Electrical Eng. & Systems 2019-06-17 Guan-Lin Chao , William Chan , Ian Lane

Multimodal speech recognition aims to improve the performance of automatic speech recognition (ASR) systems by leveraging additional visual information that is usually associated to the audio input. While previous approaches make crucial…

Sound · Computer Science 2022-04-29 Dan Oneata , Horia Cucu

While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios.…

Sound · Computer Science 2024-07-31 Tianrui Pan , Jie Liu , Bohan Wang , Jie Tang , Gangshan Wu

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-23 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Target speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-03 Hiroshi Sato , Tsubasa Ochiai , Keisuke Kinoshita , Marc Delcroix , Tomohiro Nakatani , Shoko Araki

Modern smart glasses leverage advanced audio sensing and machine learning technologies to offer real-time transcribing and captioning services, considerably enriching human experiences in daily communications. However, such systems…

Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information of audio sequences in speech separation is equally important.…

Sound · Computer Science 2023-03-08 Zhaoxi Mu , Xinyu Yang , Wenjing Zhu

Target speech separation refers to extracting the target speaker's speech from mixed signals. Despite the recent advances in deep learning based close-talk speech separation, the applications to real-world are still an open issue. Two main…

Sound · Computer Science 2020-01-03 Rongzhi Gu , Yuexian Zou

Vision is often used as a complementary modality for audio speech recognition (ASR), especially in the noisy environment where performance of solo audio modality significantly deteriorates. After combining visual modality, ASR is upgraded…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Bo Xu , Cheng Lu , Yandong Guo , Jacob Wang

Since the first speech recognition systems were built more than 30 years ago, improvement in voice technology has enabled applications such as smart assistants and automated customer support. However, conversation intelligence of the future…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-15 Desh Raj

We introduce the task of Multi-Modal Context-Aware Recognition (MCoRec) in the ninth CHiME Challenge, which addresses the cocktail-party problem of overlapping conversations in a single-room setting using audio, visual, and contextual cues.…

Computation and Language · Computer Science 2026-02-13 Thai-Binh Nguyen , Katerina Zmolikova , Pingchuan Ma , Ngoc Quan Pham , Christian Fuegen , Alexander Waibel

Robust selective auditory attention under multilingual interference is critical for reliable deployment of Large Audio Language Models (LALMs). We introduce MUSA, a cocktail party-inspired multilingual benchmark for source-grounded…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Heejoon Koo

The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality…

Sound · Computer Science 2024-05-07 Zhaoxi Mu , Xinyu Yang

Speech separation refers to extracting each individual speech source in a given mixed signal. Recent advancements in speech separation and ongoing research in this area, have made these approaches as promising techniques for pre-processing…

Machine Learning · Computer Science 2019-12-18 Fahimeh Bahmaninezhad , Shi-Xiong Zhang , Yong Xu , Meng Yu , John H. L. Hansen , Dong Yu

Deep clustering is a recently introduced deep learning architecture that uses discriminatively trained embeddings as the basis for clustering. It was recently applied to spectrogram segmentation, resulting in impressive results on…

Machine Learning · Computer Science 2016-07-11 Yusuf Isik , Jonathan Le Roux , Zhuo Chen , Shinji Watanabe , John R. Hershey

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-enrolled speaker…

Sound · Computer Science 2020-11-05 Soo-Whan Chung , Soyeon Choe , Joon Son Chung , Hong-Goo Kang

In this work, we propose a technique to transfer speech recognition capabilities from audio speech recognition systems to visual speech recognizers, where our goal is to utilize audio data during lipreading model training. Impressive…

Multimedia · Computer Science 2022-07-13 Hadeel Mabrouk , Omar Abugabal , Nourhan Sakr , Hesham M. Eraqi