English
Related papers

Related papers: SentiAvatar: Towards Expressive and Interactive Di…

200 papers

Dynamic facial emotion is essential for believable AI-generated avatars, yet most systems remain visually static, limiting their use in simulations like virtual training for investigative interviews with abused children. We present a…

Human-Computer Interaction · Computer Science 2025-07-09 Pegah Salehi , Sajad Amouei Sheshkal , Vajira Thambawita , Michael A. Riegler , Pål Halvorsen

Controllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Zhiyuan Ma , Xiangyu Zhu , Guojun Qi , Zhen Lei , Lei Zhang

Recent years have witnessed significant progress in audio-driven human animation. However, critical challenges remain in (i) generating highly dynamic videos while preserving character consistency, (ii) achieving precise emotion alignment…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Yi Chen , Sen Liang , Zixiang Zhou , Ziyao Huang , Yifeng Ma , Junshu Tang , Qin Lin , Yuan Zhou , Qinglin Lu

Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Chang Liu , Ye Pan , Chenyang Ding , Susanto Rahardja , Xiaokang Yang

We propose VASA-3D, an audio-driven, single-shot 3D head avatar generator. This research tackles two major challenges: capturing the subtle expression details present in real human faces, and reconstructing an intricate 3D head avatar from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Sicheng Xu , Guojun Chen , Jiaolong Yang , Yizhong Zhang , Yu Deng , Steve Lin , Baining Guo

Building realistic and animatable avatars still requires minutes of multi-view or monocular self-rotating videos, and most methods lack precise control over gestures and expressions. To push this boundary, we address the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jun Xiang , Yudong Guo , Leipeng Hu , Boyang Guo , Yancheng Yuan , Juyong Zhang

Generating high-fidelity 3D avatars from text or image prompts is highly sought after in virtual reality and human-computer interaction. However, existing text-driven methods often rely on iterative Score Distillation Sampling (SDS) or CLIP…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Hong Li , Yutang Feng , Minqi Meng , Yichen Yang , Xuhui Liu , Baochang Zhang

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech,…

Artificial Intelligence · Computer Science 2024-10-23 Saif Punjwani , Larry Heck

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

With the rising interest from the community in digital avatars coupled with the importance of expressions and gestures in communication, modeling natural avatar behavior remains an important challenge across many industries such as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Kefan Chen , Sergiu Oprea , Justin Theiss , Sreyas Mohan , Srinath Sridhar , Aayush Prakash

Human communication involves a complex interplay of verbal and nonverbal signals, essential for conveying meaning and achieving interpersonal goals. To develop socially intelligent AI technologies, it is crucial to develop models that can…

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

Sound · Computer Science 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Zhenglin Zhou , Fan Ma , Hehe Fan , Zongxin Yang , Yi Yang

Text-to-Avatar generation has recently made significant strides due to advancements in diffusion models. However, most existing work remains constrained by limited diversity, producing avatars with subtle differences in appearance for a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Weijing Tao , Biwen Lei , Kunhao Liu , Shijian Lu , Miaomiao Cui , Xuansong Xie , Chunyan Miao

Conversation is an essential component of virtual avatar activities in the metaverse. With the development of natural language processing, textual and vocal conversation generation has achieved a significant breakthrough. However,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Yichao Yan , Zanwei Zhou , Zi Wang , Jingnan Gao , Xiaokang Yang

Synthesizing high-fidelity head avatars is a central problem for computer vision and graphics. While head avatar synthesis algorithms have advanced rapidly, the best ones still face great obstacles in real-world scenarios. One of the vital…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Dongwei Pan , Long Zhuo , Jingtan Piao , Huiwen Luo , Wei Cheng , Yuxin Wang , Siming Fan , Shengqi Liu , Lei Yang , Bo Dai , Ziwei Liu , Chen Change Loy , Chen Qian , Wayne Wu , Dahua Lin , Kwan-Yee Lin

In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body or half-body poses,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Linrui Tian , Siqi Hu , Qi Wang , Bang Zhang , Liefeng Bo

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

We propose the first approach to automatically and jointly synthesize both the synchronous 3D conversational body and hand gestures, as well as 3D face and head animations, of a virtual character from speech input. Our algorithm uses a CNN…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Ikhsanul Habibie , Weipeng Xu , Dushyant Mehta , Lingjie Liu , Hans-Peter Seidel , Gerard Pons-Moll , Mohamed Elgharib , Christian Theobalt