English
Related papers

Related papers: Mixed-EVC: Mixed Emotion Synthesis and Control in …

200 papers

Previous studies on visual customization primarily rely on the objective alignment between various control signals (e.g., language, layout and canny) and the edited images, which largely ignore the subjective emotional contents, and more…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Jiamin Luo , Xuqian Gu , Jingjing Wang , Jiahong Lu

Emotion recognition (ER) from speech signals is a robust approach since it cannot be imitated like facial expression or text based sentiment analysis. Valuable information underlying the emotions are significant for human-computer…

Sound · Computer Science 2023-12-19 David Hason Rudd , Huan Huo , Guandong Xu

It remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with…

Sound · Computer Science 2024-05-16 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Emotional text-to-speech synthesis (ETTS) has seen much progress in recent years. However, the generated voice is often not perceptually identifiable by its intended emotion category. To address this problem, we propose a new interactive…

Computation and Language · Computer Science 2021-06-15 Rui Liu , Berrak Sisman , Haizhou Li

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Soumya Dutta

Recent text-to-speech models have reached the level of generating natural speech similar to what humans say. But there still have limitations in terms of expressiveness. The existing emotional speech synthesis models have shown…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-30 Yoori Oh , Juheon Lee , Yoseob Han , Kyogu Lee

An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional Image Content Generation (C-EICG), which aims to generate images that remain faithful to a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Jingyuan Yang , Weibin Luo , Hui Huang

Emotion Recognition in Conversation (ERC) is critical for enabling natural human-machine interactions. However, existing methods predominantly employ categorical or dimensional emotion annotations, which often fail to adequately represent…

Computation and Language · Computer Science 2026-03-10 Yoshiki Tanaka , Ryuichi Uehara , Koji Inoue , Michimasa Inaba

Emotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based…

Sound · Computer Science 2025-06-24 Sho Inoue , Kun Zhou , Shuai Wang , Haizhou Li

Current Emotion Recognition in Conversation (ERC) research follows a closed-domain assumption. However, there is no clear consensus on emotion classification in psychology, which presents a challenge for models when it comes to recognizing…

Computation and Language · Computer Science 2025-08-28 Kun Peng , Cong Cao , Hao Peng , Guanlin Wu , Zhifeng Hao , Lei Jiang , Yanbing Liu , Philip S. Yu

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

Sound · Computer Science 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang

Several works have developed end-to-end pipelines for generating lip-synced talking faces with various real-world applications, such as teaching and language translation in videos. However, these prior works fail to create realistic-looking…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Sahil Goyal , Shagun Uppal , Sarthak Bhagat , Yi Yu , Yifang Yin , Rajiv Ratn Shah

Emotion Recognition in Conversations (ERC) presents unique challenges, requiring models to capture the temporal flow of multi-turn dialogues and to effectively integrate cues from multiple modalities. We propose Mixture of Speech-Text…

Computation and Language · Computer Science 2026-02-27 Soumya Dutta , Smruthi Balaji , Sriram Ganapathy

Emotion recognition in conversation (ERC) aims to detect the emotion label for each utterance. Motivated by recent studies which have proven that feeding training examples in a meaningful order rather than considering them randomly can…

Computation and Language · Computer Science 2022-04-22 Lin Yang , Yi Shen , Yue Mao , Longjun Cai

Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of generated speech, such…

Sound · Computer Science 2025-12-23 Pengchao Feng , Yao Xiao , Ziyang Ma , Zhikang Niu , Shuai Fan , Yao Li , Sheng Wang , Xie Chen

This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Puneet Kumar , Sarthak Malik , Balasubramanian Raman

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-05 Guangyan Zhang , Ying Qin , Wenjie Zhang , Jialun Wu , Mei Li , Yutao Gai , Feijun Jiang , Tan Lee

While expressive speech synthesis or voice conversion systems mainly focus on controlling or manipulating abstract prosodic characteristics of speech, such as emotion or accent, we here address the control of perceptual voice qualities…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Frederik Rautenberg , Michael Kuhlmann , Fritz Seebauer , Jana Wiechmann , Petra Wagner , Reinhold Haeb-Umbach

We propose a multi-singer emotional singing voice synthesizer, Muse-SVS, that expresses emotion at various intensity levels by controlling subtle changes in pitch, energy, and phoneme duration while accurately following the score. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Sungjae Kim , Yewon Kim , Jewoo Jun , Injung Kim

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zongyang Qiu , Bingyuan Wang , Xingbei Chen , Yingqing He , Zeyu Wang
‹ Prev 1 4 5 6 7 8 10 Next ›