English
Related papers

Related papers: PGSTalker: Real-Time Audio-Driven Talking Head Gen…

200 papers

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise to the utterance or…

Computer Vision and Pattern Recognition · Computer Science 2019-10-03 Gaurav Mittal , Baoyuan Wang

By equipping the most recent 3D Gaussian Splatting representation with head 3D morphable models (3DMM), existing methods manage to create head avatars with high fidelity. However, most existing methods only reconstruct a head without the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Tianhao Wu , Jing Yang , Zhilin Guo , Jingyi Wan , Fangcheng Zhong , Cengiz Oztireli

Although neural rendering has made significant advances in creating lifelike, animatable full-body and head avatars, incorporating detailed expressions into full-body avatars remains largely unexplored. We present DEGAS, the first 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Zhijing Shao , Duotun Wang , Qing-Yao Tian , Yao-Dong Yang , Hengyu Meng , Zeyu Cai , Bo Dong , Yu Zhang , Kang Zhang , Zeyu Wang

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper…

Sound · Computer Science 2026-02-03 Zhipeng Chen , Xinheng Wang , Lun Xie , Haijie Yuan , Hang Pan

Despite much progress, achieving real-time high-fidelity head avatar animation is still difficult and existing methods have to trade-off between speed and quality. 3DMM based methods often fail to model non-facial structures such as…

Graphics · Computer Science 2024-06-25 Zhongyuan Zhao , Zhenyu Bao , Qing Li , Guoping Qiu , Kanglin Liu

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

Audio-driven talking head synthesis is a promising topic with wide applications in digital human, film making and virtual reality. Recent NeRF-based approaches have shown superiority in quality and fidelity compared to previous studies.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Dongze Li , Kang Zhao , Wei Wang , Bo Peng , Yingya Zhang , Jing Dong , Tieniu Tan

NeRF-based 3D-aware Generative Adversarial Networks (GANs) like EG3D or GIRAFFE have shown very high rendering quality under large representational variety. However, rendering with Neural Radiance Fields poses challenges for 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Florian Barthel , Arian Beckmann , Wieland Morgenstern , Anna Hilsmann , Peter Eisert

Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy…

Computation and Language · Computer Science 2024-08-02 Rui Liu , Yifan Hu , Yi Ren , Xiang Yin , Haizhou Li

Novel view synthesis via Neural Radiance Fields (NeRFs) or 3D Gaussian Splatting (3DGS) typically necessitates dense observations with hundreds of input images to circumvent artifacts. We introduce Deceptive-NeRF/3DGS to enhance sparse-view…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Xinhang Liu , Jiaben Chen , Shiu-hong Kao , Yu-Wing Tai , Chi-Keung Tang

3D scene reconstruction and novel-view synthesis are fundamental for VR, robotics, and content creation. However, most NeRF and 3D Gaussian Splatting pipelines assume clean inputs and degrade under real noise and artifacts. We therefore…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Fuzhen Jiang , Zhuoran Li , Yinlin Zhang

Although 3D Gaussian Splatting (3D-GS) achieves efficient rendering for novel view synthesis, extending it to dynamic scenes still results in substantial memory overhead from replicating Gaussians across frames. To address this challenge,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Chun-Tin Wu , Jun-Cheng Chen

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Hoang-Son Vo , Quang-Vinh Nguyen , Seungwon Kim , Hyung-Jeong Yang , Soonja Yeom , Soo-Hyung Kim

This work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh, recent research has…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Arthur Moreau , Jifei Song , Helisa Dhamo , Richard Shaw , Yiren Zhou , Eduardo Pérez-Pellitero

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

Computer Vision and Pattern Recognition · Computer Science 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

In this work, we introduce Monocular and Generalizable Gaussian Talking Head Animation (MGGTalk), which requires monocular datasets and generalizes to unseen identities without personalized re-training. Compared with previous 3D Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Shengjie Gong , Haojie Li , Jiapeng Tang , Dongming Hu , Shuangping Huang , Hao Chen , Tianshui Chen , Zhuoman Liu

Talking head generation based on the neural radiation fields model has shown promising visual effects. However, the slow rendering speed of NeRF seriously limits its application, due to the burdensome calculation process over hundreds of…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Niu Guanchen

Spatial audio is fundamental to immersive virtual experiences, yet synthesizing high-fidelity binaural audio from sparse observations remains a significant challenge. Existing methods typically rely on implicit neural representations…

Sound · Computer Science 2026-04-13 Chunhao Bi , Houqiang Zhong , Zhixin Xu , Li Song , Zhengxue Cheng

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion models for 3D facial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yifan Yang , Zhi Cen , Sida Peng , Xiangwei Chen , Yifu Deng , Xinyu Zhu , Fan Jia , Xiaowei Zhou , Hujun Bao