中文
相关论文

相关论文: ViDAR: Video Diffusion-Aware 4D Reconstruction Fro…

200 篇论文

View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Wen-Hsuan Chu , Lei Ke , Katerina Fragkiadaki

We propose 4Real-Video, a novel framework for generating 4D videos, organized as a grid of video frames with both time and viewpoint axes. In this grid, each row contains frames sharing the same timestep, while each column contains frames…

Diffusion models currently achieve state-of-the-art performance for both conditional and unconditional image generation. However, so far, image diffusion models do not support tasks required for 3D understanding, such as view-consistent 3D…

计算机视觉与模式识别 · 计算机科学 2024-02-22 Titas Anciukevičius , Zexiang Xu , Matthew Fisher , Paul Henderson , Hakan Bilen , Niloy J. Mitra , Paul Guerrero

Reconstructing 3D scenes and synthesizing novel views from sparse input views is a highly challenging task. Recent advances in video diffusion models have demonstrated strong temporal reasoning capabilities, making them a promising tool for…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yuqi Zhang , Guanying Chen , Jiaxing Chen , Chuanyu Fu , Chuan Huang , Shuguang Cui

We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expressions). Our method generates Multiplane Images (MPIs) that…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Yuan Li , Ziqian Bai , Feitong Tan , Zhaopeng Cui , Sean Fanello , Yinda Zhang

Directly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D CT volume. In this work, we facilitate complex 2D X-ray…

图像与视频处理 · 电气工程与系统科学 2025-03-25 Xing Xie , Jiawei Liu , Huijie Fan , Zhi Han , Yandong Tang , Liangqiong Qu

Active object reconstruction is crucial for many robotic applications. A key aspect in these scenarios is generating object-specific view configurations to obtain informative measurements for reconstruction. One-shot view planning enables…

机器人学 · 计算机科学 2025-04-17 Sicong Pan , Liren Jin , Xuying Huang , Cyrill Stachniss , Marija Popović , Maren Bennewitz

Humans can infer 3D structure from 2D images of an object based on past experience and improve their 3D understanding as they see more images. Inspired by this behavior, we introduce SAP3D, a system for 3D reconstruction and novel view…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Xinyang Han , Zelin Gao , Angjoo Kanazawa , Shubham Goel , Yossi Gandelsman

Recent 3D novel view synthesis (NVS) methods often require extensive 3D data for training, and also typically lack generalization beyond the training distribution. Moreover, they tend to be object centric and struggle with complex and…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Taewon Kang , Divya Kothandaraman , Dinesh Manocha , Ming C. Lin

In this work, we address a challenge in video inpainting: reconstructing occluded regions in dynamic, real-world scenarios. Motivated by the need for continuous human motion monitoring in healthcare settings, where facial features are…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Zheyan Zhang , Diego Klabjan , Renee CB Manworren

Multi-modal 3D object detection is important for reliable perception in robotics and autonomous driving. However, its effectiveness remains limited under adverse weather conditions due to weather-induced distortions and misalignment between…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Zhijian He , Feifei Liu , Yuwei Li , Zhanpeng Luo , Jintao Cheng , Xieyuanli Chen , Xiaoyu Tang

We introduce a novel, training-free system for reconstructing, understanding, and rendering 3D indoor scenes from a sparse set of unposed RGB images. Unlike traditional radiance field approaches that require dense views and per-scene…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jiatong Xia , Lingqiao Liu

We present Farm3D, a method for learning category-specific 3D reconstructors for articulated objects, relying solely on "free" virtual supervision from a pre-trained 2D diffusion-based image generator. Recent approaches can learn a…

计算机视觉与模式识别 · 计算机科学 2024-05-15 Tomas Jakab , Ruining Li , Shangzhe Wu , Christian Rupprecht , Andrea Vedaldi

We present SetDiff, a geometry-grounded multi-view diffusion framework that enhances novel-view renderings produced by 3D Gaussian Splatting. Our method integrates explicit 3D priors, pixel-aligned coordinate maps and pose-aware Plucker ray…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Farhad G. Zanjani , Hong Cai , Amirhossein Habibian

Recent developments in 3D Gaussian Splatting have significantly enhanced novel view synthesis, yet generating high-quality renderings from extreme novel viewpoints or partially observed regions remains challenging. Meanwhile, diffusion…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Jiaxin Wei , Stefan Leutenegger , Simon Schaefer

We present 3DiM, a diffusion model for 3D novel view synthesis, which is able to translate a single input view into consistent and sharp completions across many views. The core component of 3DiM is a pose-conditional image-to-image…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Daniel Watson , William Chan , Ricardo Martin-Brualla , Jonathan Ho , Andrea Tagliasacchi , Mohammad Norouzi

In dynamic Neural Radiance Fields (NeRF) systems, state-of-the-art novel view synthesis methods often fail under significant viewpoint deviations, producing unstable and unrealistic renderings. To address this, we introduce Expanded Dynamic…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Le Jiang , Shaotong Zhu , Yedi Luo , Shayda Moezzi , Sarah Ostadabbas

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Morris Alper , David Novotny , Filippos Kokkinos , Hadar Averbuch-Elor , Tom Monnier

We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Bharath Raj Nagoor Kani , Hsin-Ying Lee , Sergey Tulyakov , Shubham Tulsiani

We present Vista4D, a robust and flexible video reshooting framework that grounds the input video and target cameras in a 4D point cloud. Specifically, given an input video, our method re-synthesizes the scene with the same dynamics from a…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Kuan Heng Lin , Zhizheng Liu , Pablo Salamanca , Yash Kant , Ryan Burgert , Yuancheng Xu , Koichi Namekata , Yiwei Zhao , Bolei Zhou , Micah Goldblum , Paul Debevec , Ning Yu