English
Related papers

Related papers: The VoiceMOS Challenge 2022

200 papers

The scientific community is increasingly aware of the necessity to embrace pluralism and consistently represent major and minor social groups. Currently, there are no standard evaluation techniques for different types of biases.…

Computation and Language · Computer Science 2022-05-17 Marta R. Costa-jussà , Christine Basta , Gerard I. Gállego

Automatic mean opinion score (MOS) prediction provides a more perceptual alternative to objective metrics, offering deeper insights into the evaluated models. With the rapid progress of multimodal large language models (MLLMs), their…

Sound · Computer Science 2025-09-23 Yuhang Jia , Xu Zhang , Yang Chen , Hui Wang , Enzhi Wang , Yong Qin

Speech emotion recognition (SER), particularly for naturally expressed emotions, remains a challenging computational task. Key challenges include the inherent subjectivity in emotion annotation and the imbalanced distribution of emotion…

Sound · Computer Science 2025-06-03 Tiantian Feng , Thanathai Lertpetchpun , Dani Byrd , Shrikanth Narayanan

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captioning. Our submission focuses on solving two indeterminacy…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Yuma Koizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Kunio Kashino

Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to new speakers with only several minutes of speech or…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Yan Deng , Lei He , Frank Soong

In this paper, we study the task of subjective speech quality assessment (SSQA), which refers to predicting the perceptual quality of speech. Owing to the development of deep neural network models, SSQA has greatly advanced and has been…

Sound · Computer Science 2026-04-27 Wen-Chin Huang , Erica Cooper , Tomoki Toda

Automatic speech quality assessment is essential for audio researchers, developers, speech and language pathologists, and system quality engineers. The current state-of-the-art systems are based on framewise speech features (hand-engineered…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Karl El Hajal , Zihan Wu , Neil Scheidwasser-Clow , Gasser Elbanna , Milos Cernak

The "VOiCES from a Distance Challenge 2019" is designed to foster research in the area of speaker recognition and automatic speech recognition (ASR) with the special focus on single channel distant/far-field audio, under noisy conditions.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-03-01 Mahesh Kumar Nandwana , Julien van Hout , Mitchell McLaren , Colleen Richey , Aaron Lawson , Maria Alejandra Barrios

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

Computer Vision and Pattern Recognition · Computer Science 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS),…

Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled…

We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for…

The ICASSP 2024 Speech Signal Improvement Grand Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. This marks our second challenge, building upon the success from the…

Representing speech and audio signals in discrete units has become a compelling alternative to traditional high-dimensional feature vectors. Numerous studies have highlighted the efficacy of discrete units in various applications such as…

Mean Opinion Score (MOS) prediction has made significant progress in specific domains. However, the unstable performance of MOS prediction models across diverse samples presents ongoing challenges in the practical application of these…

Machine Learning · Computer Science 2024-08-26 Hui Wang , Shiwan Zhao , Jiaming Zhou , Xiguang Zheng , Haoqin Sun , Xuechen Wang , Yong Qin

This paper reports on the second GENEA Challenge to benchmark data-driven automatic co-speech gesture generation. Participating teams used the same speech and motion dataset to build gesture-generation systems. Motion generated by all these…

Human-Computer Interaction · Computer Science 2022-08-23 Youngwoo Yoon , Pieter Wolfert , Taras Kucherenko , Carla Viegas , Teodor Nikolov , Mihail Tsakov , Gustav Eje Henter

The Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-06 Qicong Xie , Xiaohai Tian , Guanghou Liu , Kun Song , Lei Xie , Zhiyong Wu , Hai Li , Song Shi , Haizhou Li , Fen Hong , Hui Bu , Xin Xu

The quality of the speech communication systems, which include noise suppression algorithms, are typically evaluated in laboratory experiments according to the ITU-T Rec. P.835, in which participants rate background noise, speech signal,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-19 Babak Naderi , Ross Cutler

The ground truth used for training image, video, or speech quality prediction models is based on the Mean Opinion Scores (MOS) obtained from subjective experiments. Usually, it is necessary to conduct multiple experiments, mostly with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-15 Gabriel Mittag , Saman Zadtootaghaj , Thilo Michael , Babak Naderi , Sebastian Möller

Automatic speech recognition systems are part of people's daily lives, embedded in personal assistants and mobile phones, helping as a facilitator for human-machine interaction while allowing access to information in a practically intuitive…

Sound · Computer Science 2021-10-05 Julio Cesar Duarte , Sérgio Colcher