中文
相关论文

相关论文: EmoMix: Emotion Mixing via Diffusion Models for Em…

200 篇论文

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

声音 · 计算机科学 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang

While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have…

音频与语音处理 · 电气工程与系统科学 2026-05-26 Wangzixi Zhou , Bagus Tris Atmaja , Sakriani Sakti

In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Bingyuan Zhang , Xulong Zhang , Ning Cheng , Jun Yu , Jing Xiao , Jianzong Wang

Text-based speech editing (TSE) modifies speech using only text, eliminating re-recording. However, existing TSE methods, mainly focus on the content accuracy and acoustic consistency of synthetic speech segments, and often overlook the…

音频与语音处理 · 电气工程与系统科学 2025-05-28 Rui Liu , Pu Gao , Jiatian Xi , Berrak Sisman , Carlos Busso , Haizhou Li

This paper aims to build a multi-speaker expressive TTS system, synthesizing a target speaker's speech with multiple styles and emotions. To this end, we propose a novel contrastive learning-based TTS approach to transfer style and emotion…

音频与语音处理 · 电气工程与系统科学 2024-04-26 Xinfa Zhu , Yuke Li , Yi Lei , Ning Jiang , Guoqing Zhao , Lei Xie

Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the…

计算与语言 · 计算机科学 2022-07-14 Yookyung Shin , Younggun Lee , Suhee Jo , Yeongtae Hwang , Taesu Kim

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to…

声音 · 计算机科学 2023-03-15 Haobin Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Emotional expression in human speech is nuanced and compositional, often involving multiple, sometimes conflicting, affective cues that may diverge from linguistic content. In contrast, most expressive text-to-speech systems enforce a…

声音 · 计算机科学 2026-02-04 Siyi Wang , Shihong Tan , Siyi Liu , Hong Jia , Gongping Huang , James Bailey , Ting Dang

Emotional text-to-speech (TTS) technology has achieved significant progress in recent years; however, challenges remain owing to the inherent complexity of emotions and limitations of the available emotional speech datasets and models.…

声音 · 计算机科学 2025-04-18 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Seong-Whan Lee

This paper proposes a unified model to conduct emotion transfer, control and prediction for sequence-to-sequence based fine-grained emotional speech synthesis. Conventional emotional speech synthesis often needs manual labels or reference…

声音 · 计算机科学 2020-11-18 Yi Lei , Shan Yang , Lei Xie

Traditional anti-spoofing focuses on models and datasets built on synthetic speech with mostly neutral state, neglecting diverse emotional variations. As a result, their robustness against high-quality, emotionally expressive synthetic…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Aurosweta Mahapatra , Ismail Rasim Ulgen , Abinay Reddy Naini , Carlos Busso , Berrak Sisman

Despite the rapid progress in image generation, emotional image editing remains under-explored. The semantics, context, and structure of an image can evoke emotional responses, making emotional image editing techniques valuable for various…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Qing Lin , Jingfeng Zhang , Yew-Soon Ong , Mengmi Zhang

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion…

音频与语音处理 · 电气工程与系统科学 2023-01-05 Guangyan Zhang , Ying Qin , Wenjie Zhang , Jialun Wu , Mei Li , Yutao Gai , Feijun Jiang , Tan Lee

Adding an emotions using prosody manipulation method for Indonesian text to speech system. Text To Speech (TTS) is a system that can convert text in one language into speech, accordance with the reading of the text in the language used. The…

声音 · 计算机科学 2016-06-30 Salita Ulitia Prini , Ary Setijadi Prihatmanto

Speech Emotion Recognition is a crucial area of research in human-computer interaction. While significant work has been done in this field, many state-of-the-art networks struggle to accurately recognize emotions in speech when the data is…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Rashedul Hasan , Meher Nigar , Nursadul Mamun , Sayan Paul

While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce…

声音 · 计算机科学 2024-09-11 Xin Jing , Kun Zhou , Andreas Triantafyllopoulos , Björn W. Schuller

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Emotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech…

声音 · 计算机科学 2022-07-19 Kun Zhou , Berrak Sisman , Rajib Rana , Björn W. Schuller , Haizhou Li

Emotion recognition is a critical task in human-computer interaction, enabling more intuitive and responsive systems. This study presents a multimodal emotion recognition system that combines low-level information from audio and text,…

音频与语音处理 · 电气工程与系统科学 2025-01-23 Shamin Bin Habib Avro , Taieba Taher , Nursadul Mamun

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

音频与语音处理 · 电气工程与系统科学 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter