English
Related papers

Related papers: SentiAvatar: Towards Expressive and Interactive Di…

200 papers

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

This paper introduces a framework, called EMOTION, for generating expressive motion sequences in humanoid robots, enhancing their ability to engage in humanlike non-verbal communication. Non-verbal cues such as facial expressions, gestures,…

Robotics · Computer Science 2024-10-31 Peide Huang , Yuhan Hu , Nataliya Nechyporenko , Daehwa Kim , Walter Talbott , Jian Zhang

To adequately utilize the available image evidence in multi-view video-based avatar modeling, we propose TexVocab, a novel avatar representation that constructs a texture vocabulary and associates body poses with texture maps for animation.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Yuxiao Liu , Zhe Li , Yebin Liu , Haoqian Wang

Audio-driven 3D facial animation aims to map input audio to realistic facial motion. Despite significant progress, limitations arise from inconsistent 3D annotations, restricting previous models to training on specific annotations and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Xiangyu Fan , Jiaqi Li , Zhiqian Lin , Weiye Xiao , Lei Yang

We propose GaussianTalker, a novel framework for real-time generation of pose-controllable talking heads. It leverages the fast rendering capabilities of 3D Gaussian Splatting (3DGS) while addressing the challenges of directly controlling…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Kyusun Cho , Joungbin Lee , Heeji Yoon , Yeobin Hong , Jaehoon Ko , Sangjun Ahn , Seungryong Kim

Creating photorealistic 3D head avatars from limited input has become increasingly important for applications in virtual reality, telepresence, and digital entertainment. While recent advances like neural rendering and 3D Gaussian splatting…

Graphics · Computer Science 2026-03-12 Chen Guo , Zhuo Su , Liao Wang , Jian Wang , Shuang Li , Xu Chang , Zhaohu Li , Yang Zhao , Guidong Wang , Yebin Liu , Ruqi Huang

We introduce FaceTalk, a novel generative approach designed for synthesizing high-fidelity 3D motion sequences of talking human heads from input audio signal. To capture the expressive, detailed nature of human heads, including hair, ears,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Shivangi Aneja , Justus Thies , Angela Dai , Matthias Nießner

Large Language Models (LLMs) have revolutionized various industries by harnessing their power to improve productivity and facilitate learning across different fields. One intriguing application involves combining LLMs with visual models to…

Human-Computer Interaction · Computer Science 2023-12-14 Hussam Azzuni , Sharim Jamal , Abdulmotaleb Elsaddik

Cartoon avatars have been widely used in various applications, including social media, online tutoring, and gaming. However, existing cartoon avatar datasets and generation methods struggle to present highly expressive avatars with…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Hao Yu , Rupayan Mallick , Margrit Betke , Sarah Adel Bargal

Creating high-fidelity, animatable 3D talking heads is crucial for immersive applications, yet often hindered by the prevalence of low-quality image or video sources, which yield poor 3D reconstructions. In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ding-Jiun Huang , Yuanhao Wang , Shao-Ji Yuan , Albert Mosella-Montoro , Francisco Vicente Carrasco , Cheng Zhang , Fernando De la Torre

High-quality, long-horizon demonstrations are essential for embodied AI, yet acquiring such data for tightly coupled wheeled mobile manipulators remains a fundamental bottleneck. Unlike fixed-base systems, mobile manipulators require…

Robotics · Computer Science 2026-03-09 Tongqing Chen , Hang Wu , Jiasen Wang , Xiaotao Li , Zhu Jin , Lu Fang

High-fidelity head avatar reconstruction plays a crucial role in AR/VR, gaming, and multimedia content creation. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated effectiveness in modeling complex geometry with real-time…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Shikun Zhang , Cunjian Chen , Yiqun Wang , Qiuhong Ke , Yong Li

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed $\textbf{Soul}$, which generates semantically coherent videos from a single-frame portrait image, text prompts, and audio, achieving precise…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jiangning Zhang , Junwei Zhu , Zhenye Gan , Donghao Luo , Chuming Lin , Feifan Xu , Xu Peng , Jianlong Hu , Yuansen Liu , Yijia Hong , Weijian Cao , Han Feng , Xu Chen , Chencan Fu , Keke He , Xiaobin Hu , Chengjie Wang

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

In face-to-face conversations, individuals need to switch between speaking and listening roles seamlessly. Existing 3D talking head generation models focus solely on speaking or listening, neglecting the natural dynamics of interactive…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ziqiao Peng , Yanbo Fan , Haoyu Wu , Xuan Wang , Hongyan Liu , Jun He , Zhaoxin Fan

We present a CloseUpAvatar - a novel approach for articulated human avatar representation dealing with more general camera motions, while preserving rendering quality for close-up views. CloseUpAvatar represents an avatar as a set of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 David Svitov , Pietro Morerio , Lourdes Agapito , Alessio Del Bue

Modeling and rendering photorealistic avatars is of crucial importance in many applications. Existing methods that build a 3D avatar from visual observations, however, struggle to reconstruct clothed humans. We introduce PhysAvatar, a novel…

Existing methods for synthesizing 3D human gestures from speech have shown promising results, but they do not explicitly model the impact of emotions on the generated gestures. Instead, these methods directly output animations from speech…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Kiran Chhatre , Radek Daněček , Nikos Athanasiou , Giorgio Becherini , Christopher Peters , Michael J. Black , Timo Bolkart

We introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP). UHAP is trained on cross-identity multi-view videos. In…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Kartik Teotia , Helge Rhodin , Mohit Mendiratta , Hyeongwoo Kim , Marc Habermann , Christian Theobalt

Recent advancements in human motion synthesis have focused on specific types of motions, such as human-scene interaction, locomotion or human-human interaction, however, there is a lack of a unified system capable of generating a diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-02-14 Jianqi Chen , Panwen Hu , Xiaojun Chang , Zhenwei Shi , Michael Kampffmeyer , Xiaodan Liang
‹ Prev 1 8 9 10 Next ›