中文
相关论文

相关论文: EGSTalker: Real-Time Audio-Driven Talking Head Gen…

200 篇论文

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Since the beginning of the COVID-19 pandemic, remote conferencing and school-teaching have become important tools. The previous applications aim to save the commuting cost with real-time interactions. However, our application is going to…

人工智能 · 计算机科学 2022-10-14 Aolan Sun , Xulong Zhang , Tiandong Ling , Jianzong Wang , Ning Cheng , Jing Xiao

Speech-driven facial animation methods usually contain two main classes, 3D and 2D talking face, both of which attract considerable research attention in recent years. However, to the best of our knowledge, the research on 3D talking face…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Yixiang Zhuang , Baoping Cheng , Yao Cheng , Yuntao Jin , Renshuai Liu , Chengyang Li , Xuan Cheng , Jing Liao , Juncong Lin

Real-time speech-driven 3D facial animation has been attractive in academia and industry. Traditional methods mainly focus on learning a deterministic mapping from speech to animation. Recent approaches start to consider the…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Peng Chen , Xiaobao Wei , Ming Lu , Hui Chen , Feng Tian

Modeling open-vocabulary language fields in 3D is essential for intuitive human-AI interaction and querying within physical environments. State-of-the-art approaches, such as LangSplat, leverage 3D Gaussian Splatting to efficiently…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Pranav Saxena

We present an audio-driven real-time system for animating photorealistic 3D facial avatars with minimal latency, designed for social interactions in virtual reality for anyone. Central to our approach is an encoder model that transforms…

图形学 · 计算机科学 2025-11-04 Jiye Lee , Chenghui Li , Linh Tran , Shih-En Wei , Jason Saragih , Alexander Richard , Hanbyul Joo , Shaojie Bai

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D…

计算机视觉与模式识别 · 计算机科学 2025-08-05 Anh Thai , Songyou Peng , Kyle Genova , Leonidas Guibas , Thomas Funkhouser

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Real-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Kartik Teotia , Hyeongwoo Kim , Pablo Garrido , Marc Habermann , Mohamed Elgharib , Christian Theobalt

Audio-driven talking-head generation is a crucial and useful technology for virtual human interaction and film-making. While recent advances have focused on improving image fidelity and lip synchronization, generating accurate emotional…

计算机视觉与模式识别 · 计算机科学 2025-05-05 Wenqing Wang , Yun Fu

Audio-driven single-image talking portrait generation plays a crucial role in virtual reality, digital human creation, and filmmaking. Existing approaches are generally categorized into keypoint-based and image-based methods. Keypoint-based…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Chaolong Yang , Kai Yao , Yuyao Yan , Chenru Jiang , Weiguang Zhao , Jie Sun , Guangliang Cheng , Yifei Zhang , Bin Dong , Kaizhu Huang

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic method to generate…

计算机视觉与模式识别 · 计算机科学 2021-08-29 Xinsheng Wang , Qicong Xie , Jihua Zhu , Lei Xie , Scharenborg

The creation of lifelike speech-driven 3D facial animation requires a natural and precise synchronization between audio input and facial expressions. However, existing works still fail to render shapes with flexible head poses and natural…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Wei Zhao , Yijun Wang , Tianyu He , Lianying Yin , Jianxin Lin , Xin Jin

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework.…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Haiyang Liu , Xiaolin Hong , Xuancheng Yang , Yudi Ruan , Xiang Lian , Michael Lingelbach , Hongwei Yi , Wei Li

In this work, we introduce the first autoregressive framework for real-time, audio-driven portrait animation, a.k.a, talking head. Beyond the challenge of lengthy animation times, a critical challenge in realistic talking head generation…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Dingcheng Zhen , Shunshun Yin , Shiyang Qin , Hou Yi , Ziwei Zhang , Siyuan Liu , Gan Qi , Ming Tao

We present EmbodiedHead, a speech-driven talking-head framework that equips LLMs with real-time visual avatars for conversation. A practical embodied avatar must achieve real-time generation, unified listening-speaking behavior, and high…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yu Zhang , Kaiyuan Shen , Yang Li

Text-to-3D synthesis has recently seen intriguing advances by combining the text-to-image priors with 3D representation methods, e.g., 3D Gaussian Splatting (3D GS), via Score Distillation Sampling (SDS). However, a hurdle of existing…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Lutao Jiang , Xu Zheng , Yuanhuiyi Lyu , Jiazhou Zhou , Lin Wang