English
Related papers

Related papers: D{\'e}tection de locuteurs dans les s{\'e}ries TV

200 papers

Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning network that…

Sound · Computer Science 2022-10-13 Tao Sun , Nidal Abuhajar , Shuyu Gong , Zhewei Wang , Charles D. Smith , Xianhui Wang , Li Xu , Jundong Liu

The vast majority of speech separation methods assume that the number of speakers is known in advance, hence they are specific to the number of speakers. By contrast, a more realistic and challenging task is to separate a mixture in which…

Sound · Computer Science 2022-03-31 Zhenhao Jin , Xiang Hao , Xiangdong Su

In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Matteo Torcoli , Emanuël A. P. Habets

Speaker verification is the process by which a speakers claim of identity is tested against a claimed speaker by his or her voice. Speaker verification is done by the use of some parameters (features) from the speakers voice which can be…

Sound · Computer Science 2019-08-16 Bhavana V. S , Pradip K. Das

Overlapped speech is notoriously problematic for speaker diarization systems. Consequently, the use of speech separation has recently been proposed to improve their performance. Although promising, speech separation models struggle with…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-02 Elio Gruttadauria , Mathieu Fontaine , Slim Essid

Speaker diarization in real-world videos presents significant challenges due to varying acoustic conditions, diverse scenes, the presence of off-screen speakers, etc. This paper builds upon a previous study (AVR-Net) and introduces a novel…

Multimedia · Computer Science 2024-03-15 Yongkang Yin , Xu Li , Ying Shan , Yuexian Zou

This paper describes an effective unsupervised speaker indexing approach. We suggest a two stage algorithm to speed-up the state-of-the-art algorithm based on the Bayesian Information Criterion (BIC). In the first stage of the merging…

Sound · Computer Science 2010-09-27 Konstantin Biatov

A speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track.…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Zexu Pan , Ruijie Tao , Chenglin Xu , Haizhou Li

Speaker segmentation consists in partitioning a conversation between one or more speakers into speaker turns. Usually addressed as the late combination of three sub-tasks (voice activity detection, speaker change detection, and overlapped…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-11 Hervé Bredin , Antoine Laurent

This paper presents a robust multi-channel speaker extraction algorithm designed to handle inaccuracies in reference information. While existing approaches often rely solely on either spatial or spectral cues to identify the target speaker,…

Sound · Computer Science 2025-12-24 Aviad Eisenberg , Sharon Gannot , Shlomo E. Chazan

This paper proposes a guided speaker embedding extraction system, which extracts speaker embeddings of the target speaker using speech activities of target and interference speakers as clues. Several methods for long-form overlapped…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-03 Shota Horiguchi , Takafumi Moriya , Atsushi Ando , Takanori Ashihara , Hiroshi Sato , Naohiro Tawara , Marc Delcroix

Speaker embeddings extracted with deep 2D convolutional neural networks are typically modeled as projections of first and second order statistics of channel-frequency pairs onto a linear layer, using either average or attentive pooling…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-08 Themos Stafylakis , Johan Rohdin , Lukas Burget

Recently, we proposed a novel speaker diarization method called End-to-End-Neural-Diarization-vector clustering (EEND-vector clustering) that integrates clustering-based and end-to-end neural network-based diarization approaches into one…

Audio and Speech Processing · Electrical Eng. & Systems 2021-09-01 Keisuke Kinoshita , Marc Delcroix , Naohiro Tawara

The x-vector architecture has recently achieved state-of-the-art results on the speaker verification task. This architecture incorporates a central layer, referred to as temporal pooling, which stacks statistical parameters of the acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-11 Mickael Rouvier , Pierre-Michel Bousquet , Jarod Duret

The performance of speaker verification degrades significantly when the test speech is corrupted by interference speakers. Speaker diarization does well to separate speakers if the speakers are temporally overlapped. However, if…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-08 Wei Rao , Chenglin Xu , Eng Siong Chng , Haizhou Li

The paper introduces Diff-Filter, a multichannel speech enhancement approach based on the diffusion probabilistic model, for improving speaker verification performance under noisy and reverberant conditions. It also presents a new two-step…

Sound · Computer Science 2023-07-06 Sandipana Dowerah , Ajinkya Kulkarni , Romain Serizel , Denis Jouvet

Since diarization and source separation of meeting data are closely related tasks, we here propose an approach to perform the two objectives jointly. It builds upon the target-speaker voice activity detection (TS-VAD) diarization approach,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-23 Christoph Boeddeker , Aswin Shanmugam Subramanian , Gordon Wichern , Reinhold Haeb-Umbach , Jonathan Le Roux

Turn-taking has played an essential role in structuring the regulation of a conversation. The task of identifying the main speaker (who is properly taking his/her turn of speaking) and the interrupters (who are interrupting or reacting to…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Thanh-Dat Truong , Chi Nhan Duong , The De Vu , Hoang Anh Pham , Bhiksha Raj , Ngan Le , Khoa Luu

Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD)…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-21 Ming Cheng , Weiqing Wang , Yucong Zhang , Xiaoyi Qin , Ming Li

Identification and localization of sounds are both integral parts of computational auditory scene analysis. Although each can be solved separately, the goal of forming coherent auditory objects and achieving a comprehensive spatial scene…

Sound · Computer Science 2019-12-24 Ivo Trowitzsch , Christopher Schymura , Dorothea Kolossa , Klaus Obermayer