English
Related papers

Related papers: VoxKnesset: A Large-Scale Longitudinal Hebrew Spee…

200 papers

This paper introduces Mixat: a dataset of Emirati speech code-mixed with English. Mixat was developed to address the shortcomings of current speech recognition resources when applied to Emirati speech, and in particular, to bilignual…

Computation and Language · Computer Science 2024-05-07 Maryam Al Ali , Hanan Aldarmaki

An utterance-level speaker embedding is typically obtained by aggregating a sequence of frame-level representations. However, in real-world scenarios, individual frames encode not only speaker-relevant information but also various nuisance…

Sound · Computer Science 2026-03-25 Junjie Li , Kong Aik Lee

Speech enhancement has recently achieved great success with various deep learning methods. However, most conventional speech enhancement systems are trained with supervised methods that impose two significant challenges. First, a majority…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Viet Anh Trinh , Sebastian Braun

This paper describes the Microsoft speaker diarization system for monaural multi-talker recordings in the wild, evaluated at the diarization track of the VoxCeleb Speaker Recognition Challenge(VoxSRC) 2020. We will first explain our system…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-26 Xiong Xiao , Naoyuki Kanda , Zhuo Chen , Tianyan Zhou , Takuya Yoshioka , Sanyuan Chen , Yong Zhao , Gang Liu , Yu Wu , Jian Wu , Shujie Liu , Jinyu Li , Yifan Gong

Creating universal speaker encoders which are robust for different acoustic and speech duration conditions is a big challenge today. According to our observations systems trained on short speech segments are optimal for short phrase speaker…

Sound · Computer Science 2022-10-31 Sergey Novoselov , Vladimir Volokhov , Galina Lavrentyeva

Speaker diarization is usually referred to as the task that determines ``who spoke when'' in a recording. Until a few years ago, all competitive approaches were modular. Systems based on this framework reached state-of-the-art performance…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-15 Federico Landini

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

Automatic height and age estimation of speakers using acoustic features is widely used for the purpose of human-computer interaction, forensics, etc. In this work, we propose a novel approach of using attention mechanism to build an…

Sound · Computer Science 2021-01-14 Manav Kaushik , Van Tung Pham , Eng Siong Chng

This paper introduces a new multi-speaker English dataset for training text-to-speech models. The dataset is based on LibriVox audiobooks and Project Gutenberg texts, both in the public domain. The new dataset contains about 292 hours of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Evelina Bakhturina , Vitaly Lavrukhin , Boris Ginsburg , Yang Zhang

Speech-to-text (STT) systems have a wide range of applications. They are available in many languages, albeit at different quality levels. Although Kurdish is considered a less-resourced language from a processing perspective, SST is…

Computation and Language · Computer Science 2025-08-14 Renas Adnan , Hossein Hassani

In this paper, we demonstrate a method for training speaker embedding extractors using weak annotation. More specifically, we are using the full VoxCeleb recordings and the name of the celebrities appearing on each video without knowledge…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-10 Themos Stafylakis , Ladislav Mošner , Oldřich Plchot , Johan Rohdin , Anna Silnova , Lukáš Burget , Jan "Honza'' Černocký

With the development of large text-to-speech (TTS) models and scale-up of the training data, state-of-the-art TTS systems have achieved impressive performance. In this paper, we present WenetSpeech4TTS, a multi-domain Mandarin corpus…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-21 Linhan Ma , Dake Guo , Kun Song , Yuepeng Jiang , Shuai Wang , Liumeng Xue , Weiming Xu , Huan Zhao , Binbin Zhang , Lei Xie

The short duration of an input utterance is one of the most critical threats that degrade the performance of speaker verification systems. This study aimed to develop an integrated text-independent speaker verification system that inputs…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-11 Jee-weon Jung , Hee-soo Heo , Hye-jin Shim , Ha-jin Yu

Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian…

High-fidelity speech can be synthesized by end-to-end text-to-speech models in recent years. However, accessing and controlling speech attributes such as speaker identity, prosody, and emotion in a text-to-speech system remains a challenge.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-05 Zexin Cai , Chuxiong Zhang , Ming Li

In this technical report, we describe the Royalflush submissions for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). Our submissions contain track 1, which is for supervised speaker verification and track 3, which is for…

Sound · Computer Science 2022-09-21 Jingguang Tian , Xinhui Hu , Xinkang Xu

When it comes to authentication in speaker verification systems, not all utterances are created equal. It is essential to estimate the quality of test utterances in order to account for varying acoustic conditions. In addition to the…

Audio and Speech Processing · Electrical Eng. & Systems 2024-07-12 Nicholas Klein , Ganesh Sivaraman , Elie Khoury

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

The convolutional neural network (CNN) based approaches have shown great success for speaker verification (SV) tasks, where modeling long temporal context and reducing information loss of speaker characteristics are two important challenges…

Sound · Computer Science 2021-08-31 Yanfeng Wu , Chenkai Guo , Junan Zhao , Xiao Jin , Jing Xu

Recent breakthroughs in intelligent speech and digital human technologies have primarily targeted mainstream adult users, often overlooking the distinct vocal patterns and interaction styles of seniors and children. These demographics…

Sound · Computer Science 2025-07-22 Haiying Xu , Haoze Liu , Mingshi Li , Siyu Cai , Guangxuan Zheng , Yuhuang Jia , Jinghua Zhao , Yong Qin