中文
相关论文

相关论文: EML System Description for VoxCeleb Speaker Diariz…

200 篇论文

Recently, x-vector has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances.…

声音 · 计算机科学 2022-01-02 Wentao Zhu , Tianlong Kong , Shun Lu , Jixiang Li , Dawei Zhang , Feng Deng , Xiaorui Wang , Sen Yang , Ji Liu

Speaker identification in noisy audio recordings, specifically those from collaborative learning environments, can be extremely challenging. There is a need to identify individual students talking in small groups from other students talking…

音频与语音处理 · 电气工程与系统科学 2022-07-05 Antonio Gomez

This report describes our submission to the VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020. We perform a careful analysis of speaker recognition models based on the popular ResNet architecture, and train a number of…

音频与语音处理 · 电气工程与系统科学 2020-09-30 Hee Soo Heo , Bong-Jin Lee , Jaesung Huh , Joon Son Chung

The INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Xiaoyi Qin , Ming Li , Hui Bu , Wei Rao , Rohan Kumar Das , Shrikanth Narayanan , Haizhou Li

Neural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance.…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Ruijie Tao , Kong Aik Lee , Zhan Shi , Haizhou Li

Distance Metric Learning (DML) has typically dominated the audio-visual speaker verification problem space, owing to strong performance in new and unseen classes. In our work, we explored multitask learning techniques to further enhance…

声音 · 计算机科学 2024-09-25 Anith Selvakumar , Homa Fashandi

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

声音 · 计算机科学 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

Spoken language diarization (LD) and related tasks are mostly explored using the phonotactic approach. Phonotactic approaches mostly use explicit way of language modeling, hence requiring intermediate phoneme modeling and transcribed data.…

音频与语音处理 · 电气工程与系统科学 2023-06-23 Jagabandhu Mishra , Amartya Chowdhury , S. R. Mahadeva Prasanna

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to empower the clustering-based speaker diarization system to handle overlapped…

声音 · 计算机科学 2022-02-11 Chen Shen , Yi Liu , Wenzhi Fan , Bin Wang , Shixue Wen , Yao Tian , Jun Zhang , Jingsheng Yang , Zejun Ma

In this technical report we describe the IDLAB top-scoring submissions for the VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20) in the supervised and unsupervised speaker verification tracks. For the supervised verification tracks we…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

Diarization partitions an audio stream into segments based on the voices of the speakers. Real-time diarization systems that include an enrollment step should limit enrollment training samples to reduce user interaction time. Although…

声音 · 计算机科学 2022-08-09 Dirk Padfield , Daniel J. Liebling

This paper details our speaker diarization system designed for multi-domain, multi-microphone casual conversations. The proposed diarization pipeline uses weighted prediction error (WPE)-based dereverberation as a front end, then applies…

音频与语音处理 · 电气工程与系统科学 2023-09-25 Naohiro Tawara , Marc Delcroix , Atsushi Ando , Atsunori Ogawa

We introduce DIVE, an end-to-end speaker diarization algorithm. Our neural algorithm presents the diarization task as an iterative process: it repeatedly builds a representation for each speaker before predicting the voice activity of each…

声音 · 计算机科学 2021-05-31 Neil Zeghidour , Olivier Teboul , David Grangier

This paper presents a system for the 2024 Text-Dependent Speaker Verification (TdSV) Challenge. The system achieved a Minimum Detection Cost Function (MinDCF) of 0.0461 and an Equal Error Rate (EER) of 1.3\%. Our approach focused on…

声音 · 计算机科学 2026-05-15 Amir Mohammad Rostami , Pourya Jafarzadeh

This paper describes the systems submitted by team HCCL to the Far-Field Speaker Verification Challenge. Our previous work in the AIshell Speaker Verification Challenge 2019 shows that the powerful modeling abilities of Neural Network…

声音 · 计算机科学 2021-07-06 Zhuo Li , Ce Fang , Runqiu Xiao , Zhigao Chen , Wenchao Wang , Yonghong Yan

The VoxCeleb Speaker Recognition Challenge 2019 aimed to assess how well current speaker recognition technology is able to identify speakers in unconstrained or `in the wild' data. It consisted of: (i) a publicly available speaker…

The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each person…

Speaker verification (SV) systems are currently being used to make sensitive decisions like giving access to bank accounts or deciding whether the voice of a suspect coincides with that of the perpetrator of a crime. Ensuring that these…

音频与语音处理 · 电气工程与系统科学 2025-11-18 Mariel Estevez , Luciana Ferrer

This work considers training neural networks for speaker recognition with a much smaller dataset size compared to contemporary work. We artificially restrict the amount of data by proposing three subsets of the popular VoxCeleb2 dataset.…

声音 · 计算机科学 2023-02-28 Nik Vaessen , David A. van Leeuwen

Speaker verification (SV) utilizing features obtained from models pre-trained via self-supervised learning has recently demonstrated impressive performances. However, these pre-trained models (PTMs) usually have a temporal resolution of 20…

音频与语音处理 · 电气工程与系统科学 2026-01-28 Jisoo Myoung , Sangwook Han , Kihyuk Kim , Jong Won Shin