中文
相关论文

相关论文: Wespeaker baselines for VoxSRC2023

200 篇论文

This technical report describes the SJTU X-LANCE Lab system for the three tracks in CNSRC 2022. In this challenge, we explored the speaker embedding modeling ability of deep ResNet (Deeper r-vector). All the systems are only trained on the…

声音 · 计算机科学 2023-05-16 Zhengyang Chen , Bei Liu , Bing Han , Leying Zhang , Yanmin Qian

In this report, we describe our submitted system for track 2 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). We fuse a variety of good-performing models ranging from supervised models to self-supervised learning(SSL)…

声音 · 计算机科学 2022-09-26 Gang Liu , Tianyan Zhou , Yong Zhao , Yu Wu , Zhuo Chen , Yao Qian , Jian Wu

This paper is the system description of the DKU-MSXF System for the track1, track2 and track3 of the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23). For Track 1, we utilize a network structure based on ResNet for training. By…

音频与语音处理 · 电气工程与系统科学 2023-08-21 Ze Li , Yuke Lin , Xiaoyi Qin , Ning Jiang , Guoqing Zhao , Ming Li

The VoxCeleb Speaker Recognition Challenge 2019 aimed to assess how well current speaker recognition technology is able to identify speakers in unconstrained or `in the wild' data. It consisted of: (i) a publicly available speaker…

This paper describes the ByteDance speaker diarization system for the fourth track of the VoxCeleb Speaker Recognition Challenge 2021 (VoxSRC-21). The VoxSRC-21 provides both the dev set and test set of VoxConverse for use in validation and…

声音 · 计算机科学 2021-09-07 Keke Wang , Xudong Mao , Hao Wu , Chen Ding , Chuxiang Shang , Rui Xia , Yuxuan Wang

This paper summarizes the outcomes from the ISCSLP 2022 Intelligent Cockpit Speech Recognition Challenge (ICSRC). We first address the necessity of the challenge and then introduce the associated dataset collected from a new-energy vehicle…

声音 · 计算机科学 2022-11-04 Ao Zhang , Fan Yu , Kaixun Huang , Lei Xie , Longbiao Wang , Eng Siong Chng , Hui Bu , Binbin Zhang , Wei Chen , Xin Xu

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expansion over the…

音频与语音处理 · 电气工程与系统科学 2024-07-17 Yuke Lin , Ming Cheng , Fulin Zhang , Yingying Gao , Shilei Zhang , Ming Li

We describe the system used by our team for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC 2022) in the speaker diarization track. Our solution was designed around a new combination of voice activity detection algorithms that uses…

声音 · 计算机科学 2023-01-19 Yannis Tevissen , Jérôme Boudy , Frédéric Petitpont

This paper introduces ESPnet-SPK, a toolkit designed with several objectives for training speaker embedding extractors. First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build…

We introduce a new zero resource code-switched speech benchmark designed to directly assess the code-switching capabilities of self-supervised speech encoders. We showcase a baseline system of language modeling on discrete units to…

音频与语音处理 · 电气工程与系统科学 2024-03-19 Kuan-Po Huang , Chih-Kai Yang , Yu-Kuan Fu , Ewan Dunbar , Hung-yi Lee

The first Chinese Continuous Visual Speech Recognition Challenge aimed to probe the performance of Large Vocabulary Continuous Visual Speech Recognition (LVC-VSR) on two tasks: (1) Single-speaker VSR for a particular speaker and (2)…

计算与语言 · 计算机科学 2024-06-18 Chen Chen , Zehua Liu , Xiaolou Li , Lantian Li , Dong Wang

This report describes the submission system by the GIST-AiTeR team for the VoxCeleb Speaker Recognition Challenge 2023 (VoxSRC-23) Track 4. Our submission system focuses on implementing diverse speaker diarization (SD) techniques, including…

音频与语音处理 · 电气工程与系统科学 2023-08-28 Dongkeon Park , Ji Won Kim , Kang Ryeol Kim , Do Hyun Lee , Hong Kook Kim

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneous behaviors in conversational speech. The challenge…

This paper describes the BUCEA speaker diarization system for the 2022 VoxCeleb Speaker Recognition Challenge. Voxsrc-22 provides the development set and test set of VoxConverse, and we mainly use the test set of VoxConverse for parameter…

声音 · 计算机科学 2022-09-21 Ruohua Zhou , Yuxuan Du , Chenlei Hu

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge(VoxSRC) 2020. We will first explain our system…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Xiong Xiao , Naoyuki Kanda , Zhuo Chen , Tianyan Zhou , Takuya Yoshioka , Sanyuan Chen , Yong Zhao , Gang Liu , Yu Wu , Jian Wu , Shujie Liu , Jinyu Li , Yifan Gong

Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors. Numerous studies have highlighted the efficacy of discrete units in various applications such as…

This report describes the submission from Technical University of Catalonia (UPC) to the VoxCeleb Speaker Recognition Challenge (VoxSRC-20) at Interspeech 2020. The final submission is a combination of three systems. System-1 is an…

音频与语音处理 · 电气工程与系统科学 2020-10-28 Umair Khan , Javier Hernando

The primary goal of the L3DAS23 Signal Processing Grand Challenge at ICASSP 2023 is to promote and support collaborative research on machine learning for 3D audio signal processing, with a specific emphasis on 3D speech enhancement and 3D…

音频与语音处理 · 电气工程与系统科学 2024-02-15 Christian Marinoni , Riccardo Fosco Gramaccioni , Changan Chen , Aurelio Uncini , Danilo Comminiello

In this technical report, we describe the Royalflush submissions for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our submissions contain track 1, which is for supervised speaker verification and track 3, which is for…

声音 · 计算机科学 2022-09-21 Jingguang Tian , Xinhui Hu , Xinkang Xu

Speaker verification (SV) provides billions of voice-enabled devices with access control, and ensures the security of voice-driven technologies. As a type of biometrics, it is necessary that SV is unbiased, with consistent and reliable…

音频与语音处理 · 电气工程与系统科学 2022-09-14 Wiebke Toussaint Hutiri , Lauriane Gorce , Aaron Yi Ding