English
Related papers

Related papers: VoxBlink: A Large Scale Speaker Verification Datas…

200 papers

In this work, we introduce X-FACT: the largest publicly available multilingual dataset for factual verification of naturally existing real-world claims. The dataset contains short statements in 25 languages and is labeled for veracity by…

Computation and Language · Computer Science 2021-06-18 Ashim Gupta , Vivek Srikumar

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge(VoxSRC) 2020. We will first explain our system…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Xiong Xiao , Naoyuki Kanda , Zhuo Chen , Tianyan Zhou , Takuya Yoshioka , Sanyuan Chen , Yong Zhao , Gang Liu , Yu Wu , Jian Wu , Shujie Liu , Jinyu Li , Yifan Gong

We introduce a method to identify speakers by computing with high-dimensional random vectors. Its strengths are simplicity and speed. With only 1.02k active parameters and a 128-minute pass through the training data we achieve Top-1 and…

Sound · Computer Science 2022-08-30 Ping-Chen Huang , Denis Kleyko , Jan M. Rabaey , Bruno A. Olshausen , Pentti Kanerva

Large-scale self-supervised Pre-Trained Models (PTMs) have shown significant improvements in the speaker verification (SV) task by providing rich feature representations. In this paper, we utilize w2v-BERT 2.0, a model with approximately…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Ze Li , Ming Cheng , Ming Li

Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack generalization to the full spectrum of human diversity in ethnicity, language, and age…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shunian Chen , Hejin Huang , Yexin Liu , Zihan Ye , Pengcheng Chen , Chenghao Zhu , Michael Guan , Rongsheng Wang , Junying Chen , Guanbin Li , Ser-Nam Lim , Harry Yang , Benyou Wang

Speaker recognition systems based on deep speaker embeddings have achieved significant performance in controlled conditions according to the results obtained for early NIST SRE (Speaker Recognition Evaluation) datasets. From the practical…

Recent research in speaker verification has increasingly focused on achieving robust and reliable recognition under challenging channel conditions and noisy environments. Identifying speakers in radio communications is particularly…

Sound · Computer Science 2024-06-18 Wenhao Yang , Jianguo Wei , Wenhuan Lu , Lei Li , Xugang Lu

Voice recognition and speaker identification are vital for applications in security and personal assistants. This paper presents a lightweight 1D-Convolutional Neural Network (1D-CNN) designed to perform speaker identification on minimal…

Sound · Computer Science 2024-11-25 Irfan Nafiz Shahan , Pulok Ahmed Auvi

As Speech Language Models (SLMs) transition from personal devices to shared, multi-user environments such as smart homes, a new challenge emerges: the model is expected to distinguish between users to manage information flow appropriately.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-29 Yuxiang Wang , Hongyu Liu , Dekun Chen , Xueyao Zhang , Zhizheng Wu

Voice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can…

Sound · Computer Science 2023-07-27 Yanzhen Ren , Hongcheng Zhu , Liming Zhai , Zongkun Sun , Rubing Shen , Lina Wang

The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker…

Sound · Computer Science 2020-10-30 Yoohwan Kwon , Hee-Soo Heo , Bong-Jin Lee , Joon Son Chung

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Jeong Hun Yeo , Minsu Kim , Shinji Watanabe , Yong Man Ro

As audio-first agents become increasingly common in physical AI, conversational robots, and screenless wearables, audio large language models (audio-LLMs) must integrate speaker-specific understanding to support user authorization,…

Sound · Computer Science 2026-05-15 KiHyun Nam , Jungwoo Heo , Siu Bae , Ha-Jin Yu , Joon Son Chung

The growing influence of video content as a medium for communication and misinformation underscores the urgent need for effective tools to analyze claims in multilingual and multi-topic settings. Existing efforts in misinformation detection…

Computation and Language · Computer Science 2025-10-13 Patrick Giedemann , Pius von Däniken , Jan Deriu , Alvaro Rodrigo , Anselmo Peñas , Mark Cieliebak

This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-10 Jenthe Thienpondt , Nilesh Madhu , Kris Demuynck

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

The SAFE Challenge evaluates synthetic speech detection across three tasks: unmodified audio, processed audio with compression artifacts, and laundered audio designed to evade detection. We systematically explore self-supervised learning…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Hashim Ali , Surya Subramani , Lekha Bollinani , Nithin Sai Adupa , Sali El-Loh , Hafiz Malik

The rapid development of large-scale models has catalyzed significant breakthroughs in the digital human domain. These advanced methodologies offer high-fidelity solutions for avatar driving and rendering, leading academia to focus on the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Youliang Zhang , Zhaoyang Li , Duomin Wang , Jiahe Zhang , Deyu Zhou , Zixin Yin , Xili Dai , Gang Yu , Xiu Li

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is provided only at the…

Sound · Computer Science 2018-06-25 Martin Karu , Tanel Alumäe

Modern speaker recognition system relies on abundant and balanced datasets for classification training. However, diverse defective datasets, such as partially-labelled, small-scale, and imbalanced datasets, are common in real-world…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-03 Ruijie Tao , Zhan Shi , Yidi Jiang , Tianchi Liu , Haizhou Li