English
Related papers

Related papers: Monocular and Generalizable Gaussian Talking Head …

200 papers

Recent advancements in Gaussian Splatting have enabled increasingly accurate reconstruction of photorealistic head avatars, opening the door to numerous applications in visual effects, videoconferencing, and virtual reality. This, however,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Kelian Baert , Mae Younes , Francois Bourel , Marc Christie , Adnane Boukhayma

This work addresses the problem of real-time rendering of photorealistic human body avatars learned from multi-view videos. While the classical approaches to model and render virtual humans generally use a textured mesh, recent research has…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Arthur Moreau , Jifei Song , Helisa Dhamo , Richard Shaw , Yiren Zhou , Eduardo Pérez-Pellitero

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Anh Thai , Songyou Peng , Kyle Genova , Leonidas Guibas , Thomas Funkhouser

Animatable 3D reconstruction has significant applications across various fields, primarily relying on artists' handcraft creation. Recently, some studies have successfully constructed animatable 3D models from monocular videos. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Tingyang Zhang , Qingzhe Gao , Weiyu Li , Libin Liu , Baoquan Chen

Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Peng Chen , Xiaobao Wei , Yi Yang , Naiming Yao , Hui Chen , Feng Tian

Recent advancements in radiance field rendering show promising results in 3D scene representation, where Gaussian splatting-based techniques emerge as state-of-the-art due to their quality and efficiency. Gaussian splatting is widely used…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Arnab Dey , Cheng-You Lu , Andrew I. Comport , Srinath Sridhar , Chin-Teng Lin , Jean Martinet

Reconstructing dynamic humans together with static scenes from monocular videos remains difficult, especially under fast motion, where RGB frames suffer from motion blur. Event cameras exhibit distinct advantages, e.g., microsecond temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Xiaoting Yin , Hao Shi , Kailun Yang , Jiajun Zhai , Shangwei Guo , Lin Wang , Kaiwei Wang

We present a novel approach, termed ADGaussian, for generalizable street scene reconstruction. The proposed method enables high-quality rendering from merely single-view input. Unlike prior Gaussian Splatting methods that primarily focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Qi Song , Chenghong Li , Haotong Lin , Sida Peng , Rui Huang

Recently, Gaussian Splatting has sparked a new trend in the field of computer vision. Apart from novel view synthesis, it has also been extended to the area of multi-view reconstruction. The latest methods facilitate complete, detailed…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Han Huang , Yulun Wu , Chao Deng , Ge Gao , Ming Gu , Yu-Shen Liu

Most current audio-driven facial animation research primarily focuses on generating videos with neutral emotions. While some studies have addressed the generation of facial videos driven by emotional audio, efficiently generating…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Chuhang Ma , Shuai Tan , Ye Pan , Jiaolong Yang , Xin Tong

Tracking the 6DoF pose of unknown objects in monocular RGB video sequences is crucial for robotic manipulation. However, existing approaches typically rely on accurate depth information, which is non-trivial to obtain in real-world…

Computer Vision and Pattern Recognition · Computer Science 2024-12-04 Zhiyuan Chen , Fan Lu , Guo Yu , Bin Li , Sanqing Qu , Yuan Huang , Changhong Fu , Guang Chen

Recently, 3D Gaussian Splatting (3DGS) has emerged as an efficient approach for accurately representing scenes. However, despite its superior novel view synthesis capabilities, extracting the geometry of the scene directly from the Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Yaniv Wolf , Amit Bracha , Ron Kimmel

Gaussian splatting has become a popular representation for novel-view synthesis, exhibiting clear strengths in efficiency, photometric quality, and compositional edibility. Following its success, many works have extended Gaussians to 4D,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Colton Stearns , Adam Harley , Mikaela Uy , Florian Dubost , Federico Tombari , Gordon Wetzstein , Leonidas Guibas

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Daniel Rho , Jun Myeong Choi , Matthew Thornton , Biswadip Dey , Roni Sengupta

Many works have succeeded in reconstructing Gaussian human avatars from multi-view videos. However, they either struggle to capture pose-dependent appearance details with a single MLP, or rely on a computationally intensive neural network…

Graphics · Computer Science 2025-04-29 Youyi Zhan , Tianjia Shao , Yin Yang , Kun Zhou

Achieving high synchronization in the synthesis of realistic, speech-driven talking head videos presents a significant challenge. A lifelike talking head requires synchronized coordination of subject identity, lip movements, facial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ziqiao Peng , Wentao Hu , Junyuan Ma , Xiangyu Zhu , Xiaomei Zhang , Hao Zhao , Hui Tian , Jun He , Hongyan Liu , Zhaoxin Fan

Understanding open-vocabulary 3D scenes with Gaussian-based representations remains challenging due to fragmented and spatially inconsistent semantic predictions across multi-view observations. In this paper, we present OpenGaFF, a novel…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kunyi Li , Michael Niemeyer , Sen Wang , Stefano Gasperini , Nassir Navab , Federico Tombari

We introduce RMAvatar, a novel human avatar representation with Gaussian splatting embedded on mesh to learn clothed avatar from a monocular video. We utilize the explicit mesh geometry to represent motion and shape of a virtual human and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Sen Peng , Weixing Xie , Zilong Wang , Xiaohu Guo , Zhonggui Chen , Baorong Yang , Xiao Dong

We present HAHA - a novel approach for animatable human avatar generation from monocular input videos. The proposed method relies on learning the trade-off between the use of Gaussian splatting and a textured mesh for efficient and high…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 David Svitov , Pietro Morerio , Lourdes Agapito , Alessio Del Bue

In-the-wild photo collections often contain limited volumes of imagery and exhibit multiple appearances, e.g., taken at different times of day or seasons, posing significant challenges to scene reconstruction and novel view synthesis.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Deming Li , Kaiwen Jiang , Yutao Tang , Ravi Ramamoorthi , Rama Chellappa , Cheng Peng
‹ Prev 1 4 5 6 7 8 10 Next ›