English
Related papers

Related papers: High-fidelity Generalized Emotional Talking Face G…

200 papers

Affective computing stands at the forefront of artificial intelligence (AI), seeking to imbue machines with the ability to comprehend and respond to human emotions. Central to this field is emotion recognition, which endeavors to identify…

Machine Learning · Computer Science 2024-07-08 Fei Ma , Yucheng Yuan , Yifan Xie , Hongwei Ren , Ivan Liu , Ying He , Fuji Ren , Fei Richard Yu , Shiguang Ni

Emotion recognition from facial videos enables non-contact inference of human emotional states. Although facial expressions are widely used cues, they cannot fully reflect intrinsic affective states. Remote photoplethysmography (rPPG)…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Xiwen Luo , Jia Li , Rencheng Song , Yu Liu , Juan Cheng

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

A 3D avatar typically has one of six cardinal facial expressions. To simulate realistic emotional variation, we should be able to render a facial transition between two arbitrary expressions. This study presents a new framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-01-14 Anh H. Vo , Tae-Seok Kim , Hulin Jin , Soo-Mi Choi , Yong-Guk Kim

Accurate recognition of human emotions is a crucial challenge in affective computing and human-robot interaction (HRI). Emotional states play a vital role in shaping behaviors, decisions, and social interactions. However, emotional…

Robotics · Computer Science 2024-09-19 Youssef Mohamed , Severin Lemaignan , Arzu Guneysu , Patric Jensfelt , Christian Smith

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Haozhe Wu , Jia Jia , Haoyu Wang , Yishun Dou , Chao Duan , Qingshan Deng

Listener head generation centers on generating non-verbal behaviors (e.g., smile) of a listener in reference to the information delivered by a speaker. A significant challenge when generating such responses is the non-deterministic nature…

Graphics · Computer Science 2023-10-10 Luchuan Song , Guojun Yin , Zhenchao Jin , Xiaoyi Dong , Chenliang Xu

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroundings, is present.…

Computer Vision and Pattern Recognition · Computer Science 2024-02-29 Meidai Xuanyuan , Yuwang Wang , Honglei Guo , Qionghai Dai

Emotion recognition and sentiment analysis are pivotal tasks in speech and language processing, particularly in real-world scenarios involving multi-party, conversational data. This paper presents a multimodal approach to tackle these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Aref Farhadipour , Hossein Ranjbar , Masoumeh Chapariniya , Teodora Vukovic , Sarah Ebling , Volker Dellwo

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Virtual humans have gained considerable attention in numerous industries, e.g., entertainment and e-commerce. As a core technology, synthesizing photorealistic face frames from target speech and facial identity has been actively studied…

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by…

Sound · Computer Science 2024-10-01 Jingyi Xu , Hieu Le , Zhixin Shu , Yang Wang , Yi-Hsuan Tsai , Dimitris Samaras

In this paper, we propose a multi-speaker face-to-speech waveform generation model that also works for unseen speaker conditions. Using a generative adversarial network (GAN) with linguistic and speaker characteristic features as auxiliary…

Computer Vision and Pattern Recognition · Computer Science 2023-03-16 Se-Yun Um , Jihyun Kim , Jihyun Lee , Hong-Goo Kang

Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate emotionally expressive portraits while preserving speaker…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Weipeng Tan , Chuming Lin , Chengming Xu , FeiFan Xu , Xiaobin Hu , Xiaozhong Ji , Junwei Zhu , Chengjie Wang , Yanwei Fu

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Speech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Hejia Chen , Haoxian Zhang , Shoulong Zhang , Xiaoqiang Liu , Sisi Zhuang , Yuan Zhang , Pengfei Wan , Di Zhang , Shuai Li

Generating vivid and diverse 3D co-speech gestures is crucial for various applications in animating virtual avatars. While most existing methods can generate gestures from audio directly, they usually overlook that emotion is one of the key…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Xingqun Qi , Chen Liu , Lincheng Li , Jie Hou , Haoran Xin , Xin Yu

The goal of this paper is to synthesise talking faces with controllable facial motions. To achieve this goal, we propose two key ideas. The first is to establish a canonical space where every face has the same motion patterns but different…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Youngjoon Jang , Kyeongha Rho , Jong-Bin Woo , Hyeongkeun Lee , Jihwan Park , Youshin Lim , Byeong-Yeol Kim , Joon Son Chung

The existing text-guided image synthesis methods can only produce limited quality results with at most \mbox{$\text{256}^2$} resolution and the textual instructions are constrained in a small Corpus. In this work, we propose a unified…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Weihao Xia , Yujiu Yang , Jing-Hao Xue , Baoyuan Wu

Cross-modality generation is an emerging topic that aims to synthesize data in one modality based on information in a different modality. In this paper, we consider a task of such: given an arbitrary audio speech and one lip image of…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Lele Chen , Zhiheng Li , Ross K. Maddox , Zhiyao Duan , Chenliang Xu
‹ Prev 1 4 5 6 7 8 10 Next ›