中文
相关论文

相关论文: Comprehensive evaluation of statistical speech wav…

200 篇论文

Text-to-Speech synthesis systems are generally evaluated using Mean Opinion Score (MOS) tests, where listeners score samples of synthetic speech on a Likert scale. A major drawback of MOS tests is that they only offer a general measure of…

音频与语音处理 · 电气工程与系统科学 2021-07-07 Elijah Gutierrez , Pilar Oplustil-Gallegos , Catherine Lai

Designing a speech quality assessment (SQA) system for estimating mean-opinion-score (MOS) of multi-rate speech with varying sampling frequency (16-48 kHz) is a challenging task. The challenge arises due to the limited availability of a…

音频与语音处理 · 电气工程与系统科学 2026-02-17 Fengyuan Cao , Xinyu Liang , Fredrik Cumlin , Victor Ungureanu , Chandan K. A. Reddy , Christian Schuldt , Saikat Chatterjee

Text-to-Speech (TTS) synthesis using deep learning relies on voice quality. Modern TTS models are advanced, but they need large amount of data. Given the growing computational complexity of these models and the scarcity of large,…

声音 · 计算机科学 2023-10-10 Ze Liu

In this study, we present an innovative technique for speaker adaptation in order to improve the accuracy of segmentation with application to unit-selection Text-To-Speech (TTS) systems. Unlike conventional techniques for speaker…

音频与语音处理 · 电气工程与系统科学 2020-05-01 Claudio Zito , Fabio Tesser , Mauro Nicolao , Piero Cosi

Modern autoregressive speech synthesis models leveraging language models have demonstrated remarkable performance. However, the sequential nature of next token prediction in these models leads to significant latency, hindering their…

声音 · 计算机科学 2025-06-04 Zijian Lin , Yang Zhang , Yougen Yuan , Yuming Yan , Jinjiang Liu , Zhiyong Wu , Pengfei Hu , Qun Yu

Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This…

音频与语音处理 · 电气工程与系统科学 2024-02-19 Vinotha R , Hepsiba D , L. D. Vijay Anand , Deepak John Reji

In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available. Conventional…

声音 · 计算机科学 2022-07-12 Naoki Makishima , Satoshi Suzuki , Atsushi Ando , Ryo Masumura

Speech synthesis has come a long way as current text-to-speech (TTS) models can now generate natural human-sounding speech. However, most of the TTS research focuses on using adult speech data and there has been very limited work done on…

声音 · 计算机科学 2022-04-05 Rishabh Jain , Mariam Yiwere , Dan Bigioi , Peter Corcoran , Horia Cucu

This paper proposes a zero-shot text-to-speech (TTS) conditioned by a self-supervised speech-representation model acquired through self-supervised learning (SSL). Conventional methods with embedding vectors from x-vector or global style…

声音 · 计算机科学 2023-12-19 Kenichi Fujita , Takanori Ashihara , Hiroki Kanagawa , Takafumi Moriya , Yusuke Ijima

With the rapid advancement in synthetic speech generation technologies, great interest in differentiating spoof speech from the natural speech is emerging in the research community. The identification of these synthetic signals is a…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Tadipatri Uday Kiran Reddy , Sahukari Chaitanya Varun , Kota Pranav Kumar Sankala Sreekanth , Kodukula Sri Rama Murty

Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a…

While supervised quality predictors for synthesized speech have demonstrated strong correlations with human ratings, their requirement for in-domain labeled training data hinders their generalization ability to new domains. Unsupervised…

音频与语音处理 · 电气工程与系统科学 2025-09-08 Erica Cooper , Takuma Okamoto , Yamato Ohtani , Tomoki Toda , Hisashi Kawai

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction…

The keyword spotting (KWS) problem requires large amounts of real speech training data to achieve high accuracy across diverse populations. Utilizing large amounts of text-to-speech (TTS) synthesized data can reduce the cost and time…

Advances in automatic speaker verification (ASV) promote research into the formulation of spoofing detection systems for real-world applications. The performance of ASV systems can be degraded severely by multiple types of spoofing attacks,…

声音 · 计算机科学 2024-08-27 Zhenyu Wang , John H. L. Hansen

Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models. These networks are often employed in a frozen state, under the assumption that the SSL…

声音 · 计算机科学 2025-06-17 Tony Alex , Sara Ahmed , Armin Mustafa , Muhammad Awais , Philip JB Jackson

This study investigates voice mapping as an evaluation framework for text-to-speech (TTS) synthesis quality. The study analyzes six TTS models, including historical and recent ones. The metrics are crest factor, spectrum balance, and…

音频与语音处理 · 电气工程与系统科学 2026-05-05 Huanchen Cai , Sten Ternström

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

声音 · 计算机科学 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

Assessing the naturalness of speech using mean opinion score (MOS) prediction models has positive implications for the automatic evaluation of speech synthesis systems. Early MOS prediction models took the raw waveform or amplitude spectrum…

声音 · 计算机科学 2024-11-19 Yu-Fei Shi , Yang Ai , Ye-Xin Lu , Hui-Peng Du , Zhen-Hua Ling

In this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the…