中文
相关论文

相关论文: Geometry Forcing: Marrying Video Diffusion and 3D …

200 篇论文

Camera-controlled video generation has achieved remarkable progress in recent years. However, existing video-to-video re-rendering methods primarily rely on Supervised Fine-Tuning using synthetic datasets. At present, there is an extreme…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Zizun Li , Haoyu Guo , Runzhe Teng , Chunhua Shen , Tong He

Diffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos into overlapping…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Yue-Jiang Dong , Wang Zhao , Jiale Xu , Ying Shan , Song-Hai Zhang

Recent video foundation models demonstrate impressive visual synthesis but frequently suffer from geometric inconsistencies. While existing methods attempt to inject 3D priors via architectural modifications, they often incur high…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Weijie Wang , Xiaoxuan He , Youping Gu , Yifan Yang , Zeyu Zhang , Yefei He , Yanbo Ding , Xirui Hu , Donny Y. Chen , Zhiyuan He , Yuqing Yang , Bohan Zhuang

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Duolikun Danier , Ge Gao , Steven McDonagh , Changjian Li , Hakan Bilen , Oisin Mac Aodha

Object-level manipulation, relocating or reorienting objects in images or videos while preserving scene realism, is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core…

计算机视觉与模式识别 · 计算机科学 2026-02-13 Penghui Ruan , Bojia Zi , Xianbiao Qi , Youze Huang , Rong Xiao , Pichao Wang , Jiannong Cao , Yuhui Shi

Precise geometric control in image generation is essential for engineering \& product design and creative industries to control 3D object features accurately in image space. Traditional 3D editing approaches are time-consuming and demand…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Phillip Mueller , Talip Uenlue , Sebastian Schmidt , Marcel Kollovieh , Jiajie Fan , Stephan Guennemann , Lars Mikelsons

Generating realistic 3D objects from single-view images requires natural appearance, 3D consistency, and the ability to capture multiple plausible interpretations of unseen regions. Existing approaches often rely on fine-tuning pretrained…

计算机视觉与模式识别 · 计算机科学 2025-10-21 Pufan Li , Bi'an Du , Wei Hu

A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D space, making it challenging for existing models to establish…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Junyi Zhang , Yiming Wang , Yunhong Lu , Qichao Wang , Wenzhe Qian , Xiaoyin Xu , David Gu , Min Zhang

It is inherently ambiguous to lift 2D results from pre-trained diffusion models to a 3D world for text-to-3D generation. 2D diffusion models solely learn view-agnostic priors and thus lack 3D knowledge during the lifting, leading to the…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Weiyu Li , Rui Chen , Xuelin Chen , Ping Tan

In recent years, video generation has seen significant advancements. However, challenges still persist in generating complex motions and interactions. To address these challenges, we introduce ReVision, a plug-and-play framework that…

计算机视觉与模式识别 · 计算机科学 2026-01-12 Qihao Liu , Ju He , Qihang Yu , Liang-Chieh Chen , Alan Yuille

Reconstructing photorealistic and animatable 4D head avatars from a single portrait image remains a fundamental challenge in computer vision. While diffusion models have enabled remarkable progress in image and video generation for avatar…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Chao Xu , Xiaochen Zhao , Xiang Deng , Jingxiang Sun , Donglin Di , Zhuo Su , Yebin Liu

Video diffusion models generate high-quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we propose a new method…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Lukas Höllein , Matthias Nießner

Manual modeling of material parameters and 3D geometry is a time consuming yet essential task in the gaming and film industries. While recent advances in 3D reconstruction have enabled accurate approximations of scene geometry and…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Philipp Langsteiner , Jan-Niklas Dihlmann , Hendrik P. A. Lensch

Using image models naively for solving inverse video problems often suffers from flickering, texture-sticking, and temporal inconsistency in generated videos. To tackle these problems, in this paper, we view frames as continuous functions…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Giannis Daras , Weili Nie , Karsten Kreis , Alex Dimakis , Morteza Mardani , Nikola Borislavov Kovachki , Arash Vahdat

3D object generation from a single image involves estimating the full 3D geometry and texture of unseen views from an unposed RGB image captured in the wild. Accurately reconstructing an object's complete 3D structure and texture has…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Hritam Basak , Hadi Tabatabaee , Shreekant Gayaka , Ming-Feng Li , Xin Yang , Cheng-Hao Kuo , Arnie Sen , Min Sun , Zhaozheng Yin

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

While recent video diffusion models (VDMs) produce visually impressive results, they fundamentally struggle to maintain 3D structural consistency, often resulting in object deformation or spatial drift. We hypothesize that these failures…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Hongyang Du , Junjie Ye , Xiaoyan Cong , Runhao Li , Jingcheng Ni , Aman Agarwal , Zeqi Zhou , Zekun Li , Randall Balestriero , Yue Wang

Grasping objects of different shapes and sizes - a foundational, effortless skill for humans - remains a challenging task in robotics. Although model-based approaches can predict stable grasp configurations for known object models, they…

机器人学 · 计算机科学 2022-11-22 Malte Mosbach , Sven Behnke

Neural representations of 3D data have been widely adopted across various applications, particularly in recent work leveraging coordinate-based networks to model scalar or vector fields. However, these approaches face inherent challenges,…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Biao Zhang , Jing Ren , Peter Wonka

Estimating accurate and temporally consistent 3D human geometry from videos is a challenging problem in computer vision. Existing methods, primarily optimized for single images, often suffer from temporal inconsistencies and fail to capture…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Gwanghyun Kim , Xueting Li , Ye Yuan , Koki Nagano , Tianye Li , Jan Kautz , Se Young Chun , Umar Iqbal