English
Related papers

Related papers: Disentangle Identity, Cooperate Emotion: Correlati…

200 papers

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are present in the entire data, which is essential for speaker…

Sound · Computer Science 2025-09-30 Chaohao Lin , Xu Zheng , Kaida Wu , Peihao Xiang , Ou Bai

Cross-speaker emotion transfer in speech synthesis relies on extracting speaker-independent emotion embeddings for accurate emotion modeling without retaining speaker traits. However, existing timbre compression methods fail to fully…

Sound · Computer Science 2025-10-20 Deok-Hyeon Cho , Hyung-Seok Oh , Seung-Bin Kim , Seong-Whan Lee

We propose a novel talking head synthesis pipeline called "DiT-Head", which is based on diffusion transformers and uses audio as a condition to drive the denoising process of a diffusion model. Our method is scalable and can generalise to…

Artificial Intelligence · Computer Science 2023-12-12 Aaron Mir , Eduardo Alonso , Esther Mondragón

Human facial images encode a rich spectrum of information, encompassing both stable identity-related traits and mutable attributes such as pose, expression, and emotion. While recent advances in image generation have enabled high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Kazuaki Mishima , Antoni Bigata Casademunt , Stavros Petridis , Maja Pantic , Kenji Suzuki

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need to overcome two…

Computer Vision and Pattern Recognition · Computer Science 2025-08-21 Ziqi Zhang , Cheng Deng

Emotionally talking head video generation aims to generate expressive portrait videos with accurate lip synchronization and emotional facial expressions. Current methods rely on simple emotional labels, leading to insufficient semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yahui Li , Yinfeng Yu , Liejun Wang , Shengjie Shen

Drawing on recent advancements in diffusion models for text-to-image generation, identity-preserved personalization has made significant progress in accurately capturing specific identities with just a single reference image. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yi Wu , Ziqiang Li , Heliang Zheng , Chaoyue Wang , Bin Li

Emotional Talking Face synthesis is pivotal in multimedia and signal processing, yet existing 3D methods suffer from two critical challenges: poor audio-vision emotion alignment, manifested as difficult audio emotion extraction and…

Artificial Intelligence · Computer Science 2026-01-28 Nanhan Shen , Zhilei Liu

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Wenqing Wang , Yun Fu

Audio-driven 3D facial animation aims to generate synchronized lip movements and vivid facial expressions from arbitrary audio clips. While existing methods can produce synchronized lip motions, they often rely on predefined identity or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xuangeng Chu , Yuan Gan , Ziteng Cui , Shuhong Liu , Jian Wang , Bing Zhou , Tatsuya Harada

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

Emotion detection is a critical technology extensively employed in diverse fields. While the incorporation of commonsense knowledge has proven beneficial for existing emotion detection methods, dialogue-based emotion detection encounters…

Computation and Language · Computer Science 2023-09-14 Yuting Su , Yichen Wei , Weizhi Nie , Sicheng Zhao , Anan Liu

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

Audio-driven talking head generation is a significant and challenging task applicable to various fields such as virtual avatars, film production, and online conferences. However, the existing GAN-based models emphasize generating…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Jintao Tan , Xize Cheng , Lingyu Xiong , Lei Zhu , Xiandong Li , Xianjia Wu , Kai Gong , Minglei Li , Yi Cai

Recent photo-realistic 3D talking head via 3D Gaussian Splatting still has significant shortcoming in emotional expression manipulation, especially for fine-grained and expansive dynamics emotional editing using multi-modal control. This…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Chang Liu , Tianjiao Jing , Chengcheng Ma , Xuanqi Zhou , Zhengxuan Lian , Qin Jin , Hongliang Yuan , Shi-Sheng Huang

Diffusion-based talking head generation has achieved remarkable visual quality, yet scaling it to long-term videos remains challenging. The widely adopted chunk-wise paradigm introduces two fundamental failures: (1) temporal-spatial…

Machine Learning · Computer Science 2026-05-12 Yuxin Lu , Jiayang Sun , Guibo Zhu , Min Cao

Recent advances in generative modeling have enabled the generation of high-quality synthetic data that is applicable in a variety of domains, including face recognition. Here, state-of-the-art generative models typically rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Darian Tomašević , Fadi Boutros , Chenhao Lin , Naser Damer , Vitomir Štruc , Peter Peer

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu