English
Related papers

Related papers: EmoSphere++: Emotion-Controllable Zero-Shot Text-t…

200 papers

Despite rapid advances in the field of emotional text-to-speech (TTS), recent studies primarily focus on mimicking the average style of a particular emotion. As a result, the ability to manipulate speech emotion remains constrained to…

Sound · Computer Science 2024-11-06 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Sang-Hoon Lee , Seong-Whan Lee

People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-18 Haibin Wu , Xiaofei Wang , Sefik Emre Eskimez , Manthan Thakker , Daniel Tompkins , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Sheng Zhao , Jinyu Li , Naoyuki Kanda

This paper introduces EmoSSLSphere, a novel framework for multilingual emotional text-to-speech (TTS) synthesis that combines spherical emotion vectors with discrete token features derived from self-supervised learning (SSL). By encoding…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-07 Joonyong Park , Kenichi Nakamura

Emotional Text-To-Speech (TTS) is an important task in the development of systems (e.g., human-like dialogue agents) that require natural and emotional speech. Existing approaches, however, only aim to produce emotional TTS for seen…

Sound · Computer Science 2023-05-24 Minki Kang , Wooseok Han , Sung Ju Hwang , Eunho Yang

Text-to-speech (TTS) has shown great progress in recent years. However, most existing TTS systems offer only coarse and rigid emotion control, typically via discrete emotion labels or a carefully crafted and detailed emotional text prompt,…

Sound · Computer Science 2025-10-28 Tianxin Xie , Shan Yang , Chenxing Li , Dong Yu , Li Liu

Human speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-14 Guanrou Yang , Chen Yang , Qian Chen , Ziyang Ma , Wenxi Chen , Wen Wang , Tianrui Wang , Yifan Yang , Zhikang Niu , Wenrui Liu , Fan Yu , Zhihao Du , Zhifu Gao , ShiLiang Zhang , Xie Chen

Zero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker's voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language…

While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-26 Wangzixi Zhou , Bagus Tris Atmaja , Sakriani Sakti

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach for emotional TTS…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-19 Xiong Cai , Dongyang Dai , Zhiyong Wu , Xiang Li , Jingbei Li , Helen Meng

Advances in text-to-speech (TTS) technology have significantly improved the quality of generated speech, closely matching the timbre and intonation of the target speaker. However, due to the inherent complexity of human emotional…

Sound · Computer Science 2024-12-13 Weizhen Bian , Yubo Zhou , Kaitai Zhang , Xiaohan Gu

Achieving precise and controllable emotional expression is crucial for producing natural and context-appropriate speech in text-to-speech (TTS) synthesis. However, many emotion-aware TTS systems, including large language model (LLM)-based…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-02 Li Zhou , Hao Jiang , Junjie Li , Tianrui Wang , Haizhou Li

This paper proposes an effective emotion control method for an end-to-end text-to-speech (TTS) system. To flexibly control the distinct characteristic of a target emotion category, it is essential to determine embedding vectors representing…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-07 Se-Yun Um , Sangshin Oh , Kyungguen Byun , Inseon Jang , Chunghyun Ahn , Hong-Goo Kang

Many frameworks for emotional text-to-speech (E-TTS) rely on human-annotated emotion labels that are often inaccurate and difficult to obtain. Learning emotional prosody implicitly presents a tough challenge due to the subjective nature of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Shreeram Suresh Chandra , Zongyang Du , Berrak Sisman

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Kun Zhou , You Zhang , Dianwen Ng , Shengkui Zhao , Hao Wang , Bin Ma

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Tianrui Wang , Haoyu Wang , Meng Ge , Cheng Gong , Chunyu Qiang , Ziyang Ma , Zikang Huang , Guanrou Yang , Xiaobao Wang , Eng Siong Chng , Xie Chen , Longbiao Wang , Jianwu Dang

We propose FEIM-TTS, an innovative zero-shot text-to-speech (TTS) model that synthesizes emotionally expressive speech, aligned with facial images and modulated by emotion intensity. Leveraging deep learning, FEIM-TTS transcends traditional…

Sound · Computer Science 2024-09-25 Yunji Chu , Yunseob Shim , Unsang Park

Current emotional Text-To-Speech (TTS) and style transfer methods rely on reference encoders to control global style or emotion vectors, but do not capture nuanced acoustic details of the reference speech. To this end, we propose a novel…

Sound · Computer Science 2025-10-03 Jianing Yang , Sheng Li , Takahiro Shinozaki , Yuki Saito , Hiroshi Saruwatari

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

While controllable Text-to-Speech (TTS) has achieved notable progress, most existing methods remain limited to inter-utterance-level control, making fine-grained intra-utterance expression challenging due to their reliance on non-public…

Sound · Computer Science 2026-05-19 Qifan Liang , Yuansen Liu , Ruixin Wei , Nan Lu , Junchuan Zhao , Ye Wang

Existing autoregressive large-scale text-to-speech (TTS) models have advantages in speech naturalness, but their token-by-token generation mechanism makes it difficult to precisely control the duration of synthesized speech. This becomes a…

Computation and Language · Computer Science 2025-09-04 Siyi Zhou , Yiquan Zhou , Yi He , Xun Zhou , Jinchao Wang , Wei Deng , Jingchen Shu
‹ Prev 1 2 3 10 Next ›