English
Related papers

Related papers: VANI: Very-lightweight Accent-controllable TTS for…

200 papers

This paper reports our work on building up a Cantonese Speech-to-Text (STT) system with a syllable based acoustic model. This is a part of an effort in building a STT system to aid dyslexic students who have cognitive deficiency in writing…

Computation and Language · Computer Science 2024-02-15 Timothy Wong , Claire Li , Sam Lam , Billy Chiu , Qin Lu , Minglei Li , Dan Xiong , Roy Shing Yu , Vincent T. Y. Ng

In this paper, we introduce the Variational Autoencoder (VAE) to an end-to-end speech synthesis model, to learn the latent representation of speaking styles in an unsupervised manner. The style representation learned through VAE shows good…

Computation and Language · Computer Science 2019-02-15 Ya-Jie Zhang , Shifeng Pan , Lei He , Zhen-Hua Ling

We introduce a text-to-speech(TTS) framework based on a neural transducer. We use discretized semantic tokens acquired from wav2vec2.0 embeddings, which makes it easy to adopt a neural transducer for the TTS framework enjoying its monotonic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Minchan Kim , Myeonghun Jeong , Byoung Jin Choi , Dongjune Lee , Nam Soo Kim

Currently, many multi-speaker speech synthesis and voice conversion systems address speaker variations with an embedding vector. Modeling it directly allows new voices outside of training data to be synthesized. GMM based approaches such as…

Sound · Computer Science 2023-09-26 Yao Shi , Ming Li

We release 70 small and discriminative test sets for machine translation (MT) evaluation called variance-aware test sets (VAT), covering 35 translation directions from WMT16 to WMT20 competitions. VAT is automatically created by a novel…

Computation and Language · Computer Science 2021-11-09 Runzhe Zhan , Xuebo Liu , Derek F. Wong , Lidia S. Chao

Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of…

Sound · Computer Science 2024-12-30 Yunyi Liu , Craig Jin

Recently, large language model (LLM) based text-to-speech (TTS) systems have gradually become the mainstream in the industry due to their high naturalness and powerful zero-shot voice cloning capabilities.Here, we introduce the IndexTTS…

Sound · Computer Science 2025-02-11 Wei Deng , Siyi Zhou , Jingchen Shu , Jinchao Wang , Lu Wang

We propose PromptTTS++, a prompt-based text-to-speech (TTS) synthesis system that allows control over speaker identity using natural language descriptions. To control speaker identity within the prompt-based TTS framework, we introduce the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-29 Reo Shimizu , Ryuichi Yamamoto , Masaya Kawamura , Yuma Shirahata , Hironori Doi , Tatsuya Komatsu , Kentaro Tachibana

The vanilla LSTM has become one of the most potential architectures in word-level language modeling, like other recurrent neural networks, overfitting is always a key barrier for its effectiveness. The existing noise-injected…

Computation and Language · Computer Science 2021-03-18 Rui Li , Kai Shuang , Mengyu Gu , Sen Su

This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…

Sound · Computer Science 2022-07-20 Juan F. Montesinos , Venkatesh S. Kadandale , Gloria Haro

In this paper we propose a new cross-lingual Voice Conversion (VC) approach which can generate all speech parameters (MCEP, LF0, BAP) from one DNN model using PPGs (Phonetic PosteriorGrams) extracted from inputted speech using several ASR…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-29 Qinghua Sun , Kenji Nagamatsu

With the demand for autonomous control and personalized speech generation, the style control and transfer in Text-to-Speech (TTS) is becoming more and more important. In this paper, we propose a new TTS system that can perform style…

Sound · Computer Science 2023-07-12 Wenhao Guan , Tao Li , Yishuang Li , Hukai Huang , Qingyang Hong , Lin Li

The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity well, or used…

Sound · Computer Science 2024-10-08 Tuan Nam Nguyen , Ngoc Quan Pham , Alexander Waibel

Hidden-Markov-model (HMM) based text-to-speech (HTS) offers flexibility in speaking styles along with fast training and synthesis while being computationally less intense. HTS performs well even in low-resource scenarios. The primary…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-14 Sudhanshu Srivastava , Ishika Gupta , Anusha Prakash , Jom Kuriakose , Hema A. Murthy

This paper introduces the T23 team's system submitted to the Singing Voice Conversion Challenge 2023. Following the recognition-synthesis framework, our singing conversion model is based on VITS, incorporating four key modules: a prior…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-05 Ziqian Ning , Yuepeng Jiang , Zhichao Wang , Bin Zhang , Lei Xie

With the continued proliferation of Large Language Model (LLM) based chatbots, there is a growing demand for generating responses that are not only linguistically fluent but also consistently aligned with persona-specific traits in…

Computation and Language · Computer Science 2025-12-12 Qi Lin , Weikai Xu , Lisi Chen , Bin Dai

Existing Large Language Model (LLM) based autoregressive (AR) text-to-speech (TTS) systems, while achieving state-of-the-art quality, still face critical challenges. The foundation of this LLM-based paradigm is the discretization of the…

Existing text-to-image (T2I) diffusion models usually struggle in interpreting complex prompts, especially those with quantity, object-attribute binding, and multi-subject descriptions. In this work, we introduce a semantic panel as the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Yutong Feng , Biao Gong , Di Chen , Yujun Shen , Yu Liu , Jingren Zhou

Text-to-speech (TTS) technology has achieved impressive results for widely spoken languages, yet many under-resourced languages remain challenged by limited data and linguistic complexities. In this paper, we present a novel methodology…

Sound · Computer Science 2025-04-11 Yizhong Geng , Jizhuo Xu , Zeyu Liang , Jinghan Yang , Xiaoyi Shi , Xiaoyu Shen

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement.…

‹ Prev 1 3 4 5 6 7 10 Next ›