English
Related papers

Related papers: Seeing What You Say: Expressive Image Generation f…

200 papers

Despite the rapid progress in image generation, emotional image editing remains under-explored. The semantics, context, and structure of an image can evoke emotional responses, making emotional image editing techniques valuable for various…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Qing Lin , Jingfeng Zhang , Yew-Soon Ong , Mengmi Zhang

While current emotional text-to-speech (TTS) systems can generate highly intelligible emotional speech, achieving fine control over emotion rendering of the output speech still remains a significant challenge. In this paper, we introduce…

Sound · Computer Science 2024-09-11 Xin Jing , Kun Zhou , Andreas Triantafyllopoulos , Björn W. Schuller

We present a methodology to train our multi-speaker emotional text-to-speech synthesizer that can express speech for 10 speakers' 7 different emotions. All silences from audio samples are removed prior to learning. This results in fast…

Computation and Language · Computer Science 2021-12-08 Sungjae Cho , Soo-Young Lee

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by…

Sound · Computer Science 2024-10-01 Jingyi Xu , Hieu Le , Zhixin Shu , Yang Wang , Yi-Hsuan Tsai , Dimitris Samaras

Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, generating effective…

Sound · Computer Science 2020-07-29 Siddique Latif , Rajib Rana , Junaid Qadir , Julien Epps

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

Existing expressive text-to-speech (TTS) systems primarily model a limited set of categorical emotions, whereas human conversations extend far beyond these predefined emotions, making it essential to explore more diverse emotional speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Xiaoxue Gao , Huayun Zhang , Nancy F. Chen

Although significant progress has been made to audio-driven talking face generation, existing methods either neglect facial emotion or cannot be applied to arbitrary subjects. In this paper, we propose the Emotion-Aware Motion Model (EAMM)…

Computer Vision and Pattern Recognition · Computer Science 2022-09-26 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Qianyi Wu , Wayne Wu , Feng Xu , Xun Cao

Spoken dialogue systems that assist users to solve complex tasks such as movie ticket booking have become an emerging research topic in artificial intelligence and natural language processing areas. With a well-designed dialogue system as…

Computation and Language · Computer Science 2021-09-01 Shang-Yu Su , Po-Wei Lin , Yun-Nung Chen

Affect is an emotional characteristic encompassing valence, arousal, and intensity, and is a crucial attribute for enabling authentic conversations. While existing text-to-speech (TTS) and speech-to-speech systems rely on strength embedding…

Visual-textual sentiment analysis aims to predict sentiment with the input of a pair of image and text, which poses a challenge in learning effective features for diverse input images. To address this, we propose a holistic method that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Junyu Chen , Jie An , Hanjia Lyu , Christopher Kanan , Jiebo Luo

Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting…

Multimedia · Computer Science 2026-01-16 Diqiong Jiang , Kai Zhu , Dan Song , Jian Chang , Chenglizhao Chen , Zhenyu Wu

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generation begins to attract the attention of related research…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Junhao Cheng , Xi Lu , Hanhui Li , Khun Loun Zai , Baiqiao Yin , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

We propose a text-to-talking-face synthesis framework leveraging latent speech representations from HierSpeech++. A Text-to-Vec module generates Wav2Vec2 embeddings from text, which jointly condition speech and face generation. To handle…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Dogucan Yaman , Seymanur Akti , Fevziye Irem Eyiokur , Alexander Waibel

Visually-grounded spoken language datasets can enable models to learn cross-modal correspondences with very weak supervision. However, modern audio-visual datasets contain biases that undermine the real-world performance of models trained…

Computation and Language · Computer Science 2021-10-15 Ian Palmer , Andrew Rouditchenko , Andrei Barbu , Boris Katz , James Glass

Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fashion and can therefore capture expressive aspects of speech…

Visual speech recognition models traditionally consist of two stages, feature extraction and classification. Several deep learning approaches have been recently presented aiming to replace the feature extraction stage by automatically…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Stavros Petridis , Yujiang Wang , Pingchuan Ma , Zuwei Li , Maja Pantic

Existing 3D facial emotion modeling have been constrained by limited emotion classes and insufficient datasets. This paper introduces "Emo3D", an extensive "Text-Image-Expression dataset" spanning a wide spectrum of human emotions, each…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Mahshid Dehghani , Amirahmad Shafiee , Ali Shafiei , Neda Fallah , Farahmand Alizadeh , Mohammad Mehdi Gholinejad , Hamid Behroozi , Jafar Habibi , Ehsaneddin Asgari

We present a general theory and corresponding declarative model for the embodied grounding and natural language based analytical summarisation of dynamic visuo-spatial imagery. The declarative model ---ecompassing spatio-linguistic…

Artificial Intelligence · Computer Science 2015-08-14 Jakob Suchan , Mehul Bhatt , Harshita Jhavar

We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss…

Computation and Language · Computer Science 2023-12-27 Ziyang Ma , Zhisheng Zheng , Jiaxin Ye , Jinchao Li , Zhifu Gao , Shiliang Zhang , Xie Chen