中文
相关论文

相关论文: CN-Celeb: multi-genre speaker recognition

200 篇论文

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under…

音频与语音处理 · 电气工程与系统科学 2019-11-06 Yue Fan , Jiawen Kang , Lantian Li , Kaicheng Li , Haolin Chen , Sitong Cheng , Pengyuan Zhang , Ziya Zhou , Yunqi Cai , Dong Wang

Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the utterances are in different speech genres. Previous…

声音 · 计算机科学 2025-01-03 Hoang Long Vu , Phuong Tuan Dat , Pham Thao Nhi , Nguyen Song Hao , Nguyen Thi Thu Trang

Multi-genre speaker recognition is becoming increasingly popular due to its ability to better represent the complexities of real-world applications. However, a major challenge is the significant shift in the distribution of speaker vectors…

声音 · 计算机科学 2023-09-26 Zhenyu Zhou , Junhui Chen , Namin Wang , Lantian Li , Dong Wang

The objective of this paper is speaker recognition under noisy and unconstrained conditions. We make two key contributions. First, we introduce a very large-scale audio-visual speaker recognition dataset collected from open-source media.…

声音 · 计算机科学 2020-11-05 Joon Son Chung , Arsha Nagrani , Andrew Zisserman

Audio-visual person recognition (AVPR) has received extensive attention. However, most datasets used for AVPR research so far are collected in constrained environments, and thus cannot reflect the true performance of AVPR systems in…

计算机视觉与模式识别 · 计算机科学 2023-07-31 Lantian Li , Xiaolou Li , Haoyu Jiang , Chen Chen , Ruihai Hou , Dong Wang

In this paper, we propose a multi-label classification framework to detect multiple speaking styles in a speech sample. Unlike previous studies that have primarily focused on identifying a single target style, our framework effectively…

音频与语音处理 · 电气工程与系统科学 2025-09-19 Miseul Kim , Seyun Um , Hyeonjin Cha , Hong-goo Kang

The goal of this paper is to learn robust speaker representation for bilingual speaking scenario. The majority of the world's population speak at least two languages; however, most speaker recognition systems fail to recognise the same…

音频与语音处理 · 电气工程与系统科学 2023-06-08 Kihyun Nam , Youkyum Kim , Jaesung Huh , Hee Soo Heo , Jee-weon Jung , Joon Son Chung

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Many speaker recognition challenges have been held to assess the speaker verification system in the wild and probe the performance limit. Voxceleb Speaker Recognition Challenge (VoxSRC), based on the voxceleb, is the most popular. Besides,…

声音 · 计算机科学 2023-06-02 Zhengyang Chen , Bing Han , Xu Xiang , Houjun Huang , Bei Liu , Yanmin Qian

The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identification. The challenge…

As speech generation technology advances, the risk of misuse through deepfake audio has become a pressing concern, which underscores the critical need for robust detection systems. However, many existing speech deepfake datasets are limited…

声音 · 计算机科学 2025-07-30 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

Current audio classification models have small class vocabularies relative to the large number of sound event classes of interest in the real world. Thus, they provide a limited view of the world that may miss important yet unexpected or…

声音 · 计算机科学 2023-10-24 Sripathi Sridhar , Mark Cartwright

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

声音 · 计算机科学 2024-02-08 Dan Lyth , Simon King

Motivated by unconsolidated data situation and the lack of a standard benchmark in the field, we complement our previous efforts and present a comprehensive corpus designed for training and evaluating text-independent multi-channel speaker…

音频与语音处理 · 电气工程与系统科学 2021-11-15 Ladislav Mošner , Oldřich Plchot , Lukáš Burget , Jan Černocký

In recent years, an association is established between faces and voices of celebrities leveraging large scale audio-visual information from YouTube. The availability of large scale audio-visual datasets is instrumental in developing speaker…

声音 · 计算机科学 2023-02-28 Saqlain Hussain Shah , Muhammad Saad Saeed , Shah Nawaz , Muhammad Haroon Yousaf

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass

Many commercial and forensic applications of speech demand the extraction of information about the speaker characteristics, which falls into the broad category of speaker profiling. The speaker characteristics needed for profiling include…

音频与语音处理 · 电气工程与系统科学 2020-07-14 Shareef Babu Kalluri , Deepu Vijayasenan , Sriram Ganapathy , Ragesh Rajan M , Prashant Krishnan

The ability of countermeasure models to generalize from seen speech synthesis methods to unseen ones has been investigated in the ASVspoof challenge. However, a new mismatch scenario in which fake audio may be generated from real audio with…

音频与语音处理 · 电气工程与系统科学 2023-05-19 Chang Zeng , Xin Wang , Xiaoxiao Miao , Erica Cooper , Junichi Yamagishi

Research in speaker recognition has recently seen significant progress due to the application of neural network models and the availability of new large-scale datasets. There has been a plethora of work in search for more powerful…

声音 · 计算机科学 2020-02-04 Joon Son Chung , Jaesung Huh , Seongkyu Mun
‹ 上一页 1 2 3 10 下一页 ›