English
Related papers

Related papers: Evaluation of Speech Representations for MOS predi…

200 papers

As a practical alternative of speech separation, target speaker extraction (TSE) aims to extract the speech from the desired speaker using additional speaker cue extracted from the speaker. Its main challenge lies in how to properly extract…

Sound · Computer Science 2023-01-18 Kai Liu , Xucheng Wan , Ziqing Du , Huan Zhou

Speaker verification is an established yet challenging task in speech processing and a very vibrant research area. Recent speaker verification (SV) systems rely on deep neural networks to extract high-level embeddings which are able to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-23 Fei Tao , Gokhan Tur

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual…

Computation and Language · Computer Science 2025-08-15 Leonora Vesterbacka , Faton Rekathati , Robin Kurtz , Justyna Sikora , Agnes Toftgård

WaveNet is a state-of-the-art text-to-speech vocoder that remains challenging to deploy due to its autoregressive loop. In this work we focus on ways to accelerate the original WaveNet architecture directly, as opposed to modifying the…

Machine Learning · Computer Science 2020-11-23 Sam Davis , Giuseppe Coccia , Sam Gooch , Julian Mack

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource Speech Benchmark 2021: a suite of 4 black-box, zero-shot…

Computation and Language · Computer Science 2020-12-02 Tu Anh Nguyen , Maureen de Seyssel , Patricia Rozé , Morgane Rivière , Evgeny Kharitonov , Alexei Baevski , Ewan Dunbar , Emmanuel Dupoux

Self-supervised learned models have been found to be very effective for certain speech tasks such as automatic speech recognition, speaker identification, keyword spotting and others. While the features are undeniably useful in speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-05 Ravi Shankar , Ke Tan , Buye Xu , Anurag Kumar

Second-pass rescoring is employed in most state-of-the-art speech recognition systems. Recently, BERT based models have gained popularity for re-ranking the n-best hypothesis by exploiting the knowledge from masked language model…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-19 Prashanth Gurunath Shivakumar , Jari Kolehmainen , Yile Gu , Ankur Gandhe , Ariya Rastrow , Ivan Bulyko

In this paper, we focus on Whisper, a recent automatic speech recognition model trained with a massive 680k hour labeled speech corpus recorded in diverse conditions. We first show an interesting finding that while Whisper is very robust…

Sound · Computer Science 2023-10-10 Yuan Gong , Sameer Khurana , Leonid Karlinsky , James Glass

Personalized speech enhancement (PSE) models can improve the audio quality of teleconferencing systems by adapting to the characteristics of a speaker's voice. However, most existing methods require a separate speaker embedding model to…

Sound · Computer Science 2024-06-17 Tanel Pärnamaa , Ando Saabas

Recent techniques for speech deepfake detection often rely on pre-trained self-supervised models. These systems, initially developed for Automatic Speech Recognition (ASR), have proved their ability to offer a meaningful representation of…

Over the recent years, various deep learning-based methods were proposed for extracting a fixed-dimensional embedding vector from speech signals. Although the deep learning-based embedding extraction methods have shown good performance in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-08 Woo Hyun Kang , Jahangir Alam , Abderrahim Fathan

Whisper, as a form of speech, is not sufficiently addressed by mainstream speech applications. This is due to the fact that systems built for normal speech do not work as expected for whispered speech. A first step to building a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under…

Computation and Language · Computer Science 2023-06-06 Bashar Talafha , Abdul Waheed , Muhammad Abdul-Mageed

Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-13 Sewade Ogun , Vincent Colotte , Emmanuel Vincent

In recent years, much speech separation research has focused primarily on improving model performance. However, for low-latency speech processing systems, high efficiency is equally important. Therefore, we propose a speech separation model…

Sound · Computer Science 2026-03-02 Mohan Xu , Kai Li , Guo Chen , Xiaolin Hu

Large-scale end-to-end models such as Whisper have shown strong performance on diverse speech tasks, but their internal behavior on pathological speech remains poorly understood. Understanding how dysarthric speech is represented across…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Zhengjun Yue , Devendra Kayande , Zoran Cvetkovic , Erfan Loweimi

Neural speaker embeddings encode the speaker's speech characteristics through a DNN model and are prevalent for speaker verification tasks. However, few studies have investigated the usage of neural speaker embeddings for an ASR system. In…

Computation and Language · Computer Science 2023-09-21 Christoph Lüscher , Jingjing Xu , Mohammad Zeineldeen , Ralf Schlüter , Hermann Ney

Speech recognition performance varies by language, domain, and speaker characteristics such as accent, but fine-tuning a model on any of these categories may lead to catastrophic forgetting. Token-level $k$ nearest neighbor search ($k$NN),…

Computation and Language · Computer Science 2025-02-12 Maya K. Nachesa , Vlad Niculae

Assessing the perceptual quality of synthetic speech is crucial for guiding the development and refinement of speech generation models. However, it has traditionally relied on human subjective ratings such as the Mean Opinion Score (MOS),…

Although deep learning (DL) has achieved notable progress in speech enhancement (SE), further research is still required for a DL-based SE system to adapt effectively and efficiently to particular speakers. In this study, we propose a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-11 Cheng Yu , Szu-Wei Fu , Tsun-An Hsieh , Yu Tsao , Mirco Ravanelli