English
Related papers

Related papers: Towards High-fidelity 3D Talking Avatar with Perso…

200 papers

This paper presents FaceXHuBERT, a text-less speech-driven 3D facial animation generation method that allows to capture personalized and subtle cues in speech (e.g. identity, emotion and hesitation). It is also very robust to background…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Kazi Injamamul Haque , Zerrin Yumak

Recent advances in deep learning for sequential data have given rise to fast and powerful models that produce realistic videos of talking humans. The state of the art in talking face generation focuses mainly on lip-syncing, being…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Georgios Milis , Panagiotis P. Filntisis , Anastasios Roussos , Petros Maragos

3D meshes are widely used in computer vision and graphics for their efficiency in animation and minimal memory use, playing a crucial role in movies, games, AR, and VR. However, creating temporally consistent and realistic textures for mesh…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Jingzhi Bao , Xueting Li , Ming-Hsuan Yang

Realistic 3D full-body talking avatars hold great potential in AR, with applications ranging from e-commerce live streaming to holographic communication. Despite advances in 3D Gaussian Splatting (3DGS) for lifelike avatar creation,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Jianchuan Chen , Jingchuan Hu , Gaige Wang , Zhonghua Jiang , Tiansong Zhou , Zhiwen Chen , Chengfei Lv

Audio-driven emotional 3D face animation aims to generate emotionally expressive talking heads with synchronized lip movements. However, previous research has often overlooked the influence of diverse emotions on facial expressions or…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Chang Liu , Qunfen Lin , Zijiao Zeng , Ye Pan

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

Talking face generation (TFG) aims to animate a target identity's face to create realistic talking videos. Personalized TFG is a variant that emphasizes the perceptual identity similarity of the synthesized result (from the perspective of…

Computer Vision and Pattern Recognition · Computer Science 2024-10-16 Zhenhui Ye , Tianyun Zhong , Yi Ren , Ziyue Jiang , Jiawei Huang , Rongjie Huang , Jinglin Liu , Jinzheng He , Chen Zhang , Zehan Wang , Xize Chen , Xiang Yin , Zhou Zhao

To adequately utilize the available image evidence in multi-view video-based avatar modeling, we propose TexVocab, a novel avatar representation that constructs a texture vocabulary and associates body poses with texture maps for animation.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yuxiao Liu , Zhe Li , Yebin Liu , Haoqian Wang

We are interested in a novel task, namely low-resource text-to-talking avatar. Given only a few-minute-long talking person video with the audio track as the training data and arbitrary texts as the driving input, we aim to synthesize…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Zhenhui Ye , Ziyue Jiang , Yi Ren , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

Audio-driven talking head generation is crucial for applications in virtual reality, digital avatars, and film production. While NeRF-based methods enable high-fidelity reconstruction, they suffer from low rendering efficiency and…

Sound · Computer Science 2025-09-23 Tianheng Zhu , Yinfeng Yu , Liejun Wang , Fuchun Sun , Wendong Zheng

Current text-to-avatar methods often rely on implicit representations (e.g., NeRF, SDF, and DMTet), leading to 3D content that artists cannot easily edit and animate in graphics software. This paper introduces a novel framework for…

Graphics · Computer Science 2025-05-01 Duotun Wang , Hengyu Meng , Zeyu Cai , Zhijing Shao , Qianxi Liu , Lin Wang , Mingming Fan , Xiaohang Zhan , Zeyu Wang

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Renjie Lu , Xulong Zhang , Xiaoyang Qu , Jianzong Wang , Shangfei Wang

Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and average facial shape…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ziyu Yao , Xuxin Cheng , Zhiqi Huang

We propose TextToon, a method to generate a drivable toonified avatar. Given a short monocular video sequence and a written instruction about the avatar style, our model can generate a high-fidelity toonified avatar that can be driven in…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Luchuan Song , Lele Chen , Celong Liu , Pinxin Liu , Chenliang Xu

To be widely adopted, 3D facial avatars must be animated easily, realistically, and directly from speech signals. While the best recent methods generate 3D animations that are synchronized with the input audio, they largely ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Radek Daněček , Kiran Chhatre , Shashank Tripathi , Yandong Wen , Michael J. Black , Timo Bolkart

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

Although existing speech-driven talking face generation methods achieve significant progress, they are far from real-world application due to the avatar-specific training demand and unstable lip movements. To address the above issues, we…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Haiming Zhang , Zhihao Yuan , Chaoda Zheng , Xu Yan , Baoyuan Wang , Guanbin Li , Song Wu , Shuguang Cui , Zhen Li

Although significant progress has been made in audio-driven talking head generation, text-driven methods remain underexplored. In this work, we present OmniTalker, a unified framework that jointly generates synchronized talking audio-video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Zhongjian Wang , Peng Zhang , Jinwei Qi , Guangyuan Wang , Chaonan Ji , Sheng Xu , Bang Zhang , Liefeng Bo

We introduce AvatarBooth, a novel method for generating high-quality 3D avatars using text prompts or specific images. Unlike previous approaches that can only synthesize avatars based on simple text descriptions, our method enables the…

Computer Vision and Pattern Recognition · Computer Science 2023-06-19 Yifei Zeng , Yuanxun Lu , Xinya Ji , Yao Yao , Hao Zhu , Xun Cao
‹ Prev 1 4 5 6 7 8 10 Next ›