中文
相关论文

相关论文: Build a SRE Challenge System: Lessons from VoxSRC …

200 篇论文

Despite its broad practical applications such as in fraud prevention, open-set speaker identification (OSI) has received less attention in the speaker recognition community compared to speaker verification (SV). OSI deals with determining…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Raghuveer Peri , Seyed Omid Sadjadi , Daniel Garcia-Romero

We present a cost-effective two-step authentication system that integrates face identification and speaker verification using only a camera and microphone available on common devices. The pipeline first performs face recognition to identify…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Kuan Wei Chen , Ting Yi Lin , Wen Ren Yang , Aryan Kesarwani , Riya Singh

Speaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker. This paper demonstrates the effectiveness of meta-learning and…

音频与语音处理 · 电气工程与系统科学 2022-07-25 Ashutosh Chaubey , Sparsh Sinha , Susmita Ghose

This document outlines the Text-dependent Speaker Verification (TdSV) Challenge 2024, which centers on analyzing and exploring novel approaches for text-dependent speaker verification. The primary goal of this challenge is to motive…

声音 · 计算机科学 2024-04-23 Zeinali Hossein , Lee Kong Aik , Alam Jahangir , Burget Lukas

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been…

计算机视觉与模式识别 · 计算机科学 2023-06-29 Pingchuan Ma , Alexandros Haliassos , Adriana Fernandez-Lopez , Honglie Chen , Stavros Petridis , Maja Pantic

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

音频与语音处理 · 电气工程与系统科学 2022-05-13 Otavio Braga , Olivier Siohan

This paper describes the DKU-MSXF submission to track 4 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). Our system pipeline contains voice activity detection, clustering-based diarization, overlapped speech detection, and…

音频与语音处理 · 电气工程与系统科学 2023-08-21 Ming Cheng , Weiqing Wang , Xiaoyi Qin , Yuke Lin , Ning Jiang , Guoqing Zhao , Ming Li

Spoken language recognition (SLR) is the task of automatically identifying the language present in a speech signal. Existing SLR models are either too computationally expensive or too large to run effectively on devices with limited…

计算与语言 · 计算机科学 2023-06-06 Oriol Nieto , Zeyu Jin , Franck Dernoncourt , Justin Salamon

Deep learning approaches are still not very common in the speaker verification field. We investigate the possibility of using deep residual convolutional neural network with spectrograms as an input features in the text-dependent speaker…

声音 · 计算机科学 2017-05-31 Egor Malykh , Sergey Novoselov , Oleg Kudashev

Audio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE). Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Ruijie Tao , Rohan Kumar Das , Haizhou Li

Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under…

音频与语音处理 · 电气工程与系统科学 2024-02-06 Kerlos Atia Abdalmalak , Ascensión Gallardo-Antol'in

We introduce a method to identify speakers by computing with high-dimensional random vectors. Its strengths are simplicity and speed. With only 1.02k active parameters and a 128-minute pass through the training data we achieve Top-1 and…

声音 · 计算机科学 2022-08-30 Ping-Chen Huang , Denis Kleyko , Jan M. Rabaey , Bruno A. Olshausen , Pentti Kanerva

This paper describes the TSUP team's submission to the ISCSLP 2022 conversational short-phrase speaker diarization (CSSD) challenge which particularly focuses on short-phrase conversations with a new evaluation metric called conversational…

声音 · 计算机科学 2023-10-26 Bowen Pang , Huan Zhao , Gaosheng Zhang , Xiaoyue Yang , Yang Sun , Li Zhang , Qing Wang , Lei Xie

The CL-UZH team submitted one system each for the fixed and open conditions of the NIST SRE 2024 challenge. For the closed-set condition, results for the audio-only trials were achieved using the X-vector system developed with Kaldi. For…

音频与语音处理 · 电气工程与系统科学 2025-10-08 Aref Farhadipour , Shiran Liu , Masoumeh Chapariniya , Valeriia Vyshnevetska , Srikanth Madikeri , Teodora Vukovic , Volker Dellwo

In neural network based speaker verification, speaker embedding is expected to be discriminative between speakers while the intra-speaker distance should remain small. A variety of loss functions have been proposed to achieve this goal. In…

声音 · 计算机科学 2019-04-09 Yi Liu , Liang He , Jia Liu

The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organized the third edition of the challenge and constructed and…

音频与语音处理 · 电气工程与系统科学 2020-08-31 Yi Zhao , Wen-Chin Huang , Xiaohai Tian , Junichi Yamagishi , Rohan Kumar Das , Tomi Kinnunen , Zhenhua Ling , Tomoki Toda

We present the latest iteration of the voice conversion challenge (VCC) series, a bi-annual scientific event aiming to compare and understand different voice conversion (VC) systems based on a common dataset. This year we shifted our focus…

声音 · 计算机科学 2023-07-07 Wen-Chin Huang , Lester Phillip Violeta , Songxiang Liu , Jiatong Shi , Tomoki Toda

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

This document provides a brief description of the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) conversational telephone speech (CTS) Superset. The CTS Superset has been created in an attempt to…

声音 · 计算机科学 2021-08-17 Seyed Omid Sadjadi

This paper introduces SpoofCeleb, a dataset designed for Speech Deepfake Detection (SDD) and Spoofing-robust Automatic Speaker Verification (SASV), utilizing source data from real-world conditions and spoofing attacks generated by…