English
Related papers

Related papers: VoiceSculptor: Your Voice, Designed By You

200 papers

In this paper, we present ControlSpeech, a text-to-speech (TTS) system capable of fully cloning the speaker's voice and enabling arbitrary control and adjustment of speaking style. Prior zero-shot TTS models only mimic the speaker's voice…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-05 Shengpeng Ji , Qian Chen , Wen Wang , Jialong Zuo , Minghui Fang , Ziyue Jiang , Hai Huang , Zehan Wang , Xize Cheng , Siqi Zheng , Zhou Zhao

Voice design from natural language aims to generate speaker timbres directly from free-form textual descriptions, allowing users to create voices tailored to specific roles, personalities, and emotions. Such controllable voice creation…

We introduce MiniMax-Speech, an autoregressive Transformer-based Text-to-Speech (TTS) model that generates high-quality speech. A key innovation is our learnable speaker encoder, which extracts timbre features from a reference audio without…

This study proposes FlexiVoice, a text-to-speech (TTS) synthesis system capable of flexible style control with zero-shot voice cloning. The speaking style is controlled by a natural-language instruction and the voice timbre is provided by a…

Sound · Computer Science 2026-01-09 Dekun Chen , Xueyao Zhang , Yuancheng Wang , Kenan Dai , Li Ma , Zhizheng Wu

Voice design from natural language descriptions is emerging as a new task in text-to-speech multimodal generation, aiming to synthesize speech with target timbre and speaking style without relying on reference audio. However, existing…

Sound · Computer Science 2026-04-10 Xiaosu Su , Zihan Sun , Peilei Jia , Jun Gao

We present an open-source system designed for multilingual translation and speech regeneration, addressing challenges in communication and accessibility across diverse linguistic contexts. The system integrates Whisper for speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-04 Mateo Cámara , Juan Gutiérrez , María Pilar Daza , José Luis Blanco

We propose a novel description-based controllable text-to-speech (TTS) method with cross-lingual control capability. To address the lack of audio-description paired data in the target language, we combine a TTS model trained on the target…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-27 Ryuichi Yamamoto , Yuma Shirahata , Masaya Kawamura , Kentaro Tachibana

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Minsu Kim , Pingchuan Ma , Honglie Chen , Stavros Petridis , Maja Pantic

We propose Cotatron, a transcription-guided speech encoder for speaker-independent linguistic representation. Cotatron is based on the multispeaker TTS architecture and can be trained with conventional TTS datasets. We train a voice…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-17 Seung-won Park , Doo-young Kim , Myun-chul Joe

We introduce OpenVoice, a versatile voice cloning approach that requires only a short audio clip from the reference speaker to replicate their voice and generate speech in multiple languages. OpenVoice represents a significant advancement…

Sound · Computer Science 2024-08-20 Zengyi Qin , Wenliang Zhao , Xumin Yu , Xin Sun

The imitation of voice, targeted on specific speech attributes such as timbre and speaking style, is crucial in speech generation. However, existing methods rely heavily on annotated data, and struggle with effectively disentangling timbre…

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-08 Akshay Anand , Chenxu Guo , Cheol Jun Cho , Jiachen Lian , Gopala Anumanchipalli

Personalized TTS is an exciting and highly desired application that allows users to train their TTS voice using only a few recordings. However, TTS training typically requires many hours of recording and a large model, making it unsuitable…

Sound · Computer Science 2023-03-22 Sung-Feng Huang , Chia-ping Chen , Zhi-Sheng Chen , Yu-Pao Tsai , Hung-yi Lee

Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g.,…

Computation and Language · Computer Science 2025-08-26 Tianxin Xie , Yan Rong , Pengfei Zhang , Wenwu Wang , Li Liu

Controllable speech generation methods typically rely on single or fixed prompts, hindering creativity and flexibility. These limitations make it difficult to meet specific user needs in certain scenarios, such as adjusting the style while…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-01 Hanzhao Li , Yuke Li , Xinsheng Wang , Jingbin Hu , Qicong Xie , Shan Yang , Lei Xie

Text-to-speech (TTS) and text-to-music (TTM) models face significant limitations in instruction-based control. TTS systems usually depend on reference audio for timbre, offer only limited text-level attribute control, and rarely support…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-25 Chunyu Qiang , Kang Yin , Xiaopeng Wang , Yuzhe Liang , Jiahui Zhao , Ruibo Fu , Tianrui Wang , Cheng Gong , Chen Zhang , Longbiao Wang , Jianwu Dang

Voice cloning is the task of learning to synthesize the voice of an unseen speaker from a few samples. While current voice cloning methods achieve promising results in Text-to-Speech (TTS) synthesis for a new voice, these approaches lack…

Sound · Computer Science 2021-02-02 Paarth Neekhara , Shehzeen Hussain , Shlomo Dubnov , Farinaz Koushanfar , Julian McAuley

Modern TTS systems are capable of creating highly realistic and natural-sounding speech. Despite these developments, the process of customizing TTS voices remains a complex task, mostly requiring the expertise of specialists within the…

Human-Computer Interaction · Computer Science 2024-08-23 Silvan Mertes , Daksitha Withanage Don , Otto Grothe , Johanna Kuch , Ruben Schlagowski , Elisabeth André

Recently, there has been a growing interest in the field of controllable Text-to-Speech (TTS). While previous studies have relied on users providing specific style factor values based on acoustic knowledge or selecting reference speeches…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-26 Shengpeng Ji , Jialong Zuo , Minghui Fang , Ziyue Jiang , Feiyang Chen , Xinyu Duan , Baoxing Huai , Zhou Zhao
‹ Prev 1 2 3 10 Next ›