中文
相关论文

相关论文: EditCast3D: Single-Frame-Guided 3D Editing with Vi…

200 篇论文

Propagation-based video editing enables precise user control by propagating a single edited frame into following frames while maintaining the original context such as motion and structures. However, training such models requires…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Wonyong Seo , Jaeho Moon , Jaehyup Lee , Soo Ye Kim , Munchurl Kim

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Generating consistent multiple views for 3D reconstruction tasks is still a challenge to existing image-to-3D diffusion models. Generally, incorporating 3D representations into diffusion model decrease the model's speed as well as…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Emmanuelle Bourigault , Pauline Bourigault

With the rapid advancement of 3D representation techniques and generative models, substantial progress has been made in reconstructing full-body 3D avatars from a single image. However, this task remains fundamentally ill-posedness due to…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Gaofeng Liu , Hengsen Li , Ruoyu Gao , Xuetong Li , Zhiyuan Ma , Tao Fang

Recent advancements in 3D foundation models have enabled the generation of high-fidelity assets, yet precise 3D manipulation remains a significant challenge. Existing 3D editing frameworks often face a difficult trade-off between visual…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Inbar Gat , Dana Cohen-Bar , Guy Levy , Elad Richardson , Daniel Cohen-Or

The problem of 3D reconstruction from posed images is undergoing a fundamental transformation, driven by continuous advances in 3D Gaussian Splatting (3DGS). By modeling scenes explicitly as collections of 3D Gaussians, 3DGS enables…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Vitor Pereira Matias , Daniel Perazzo , Vinicius Silva , Alberto Raposo , Luiz Velho , Afonso Paiva , Tiago Novello

We introduce 3DEgo to address a novel problem of directly synthesizing photorealistic 3D scenes from monocular videos guided by textual prompts. Conventional methods construct a text-conditioned 3D scene through a three-stage process,…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Umar Khalid , Hasan Iqbal , Azib Farooq , Jing Hua , Chen Chen

Multi-view 3D reconstruction remains a core challenge in computer vision. Recent methods, such as DUST3R and its successors, directly regress pointmaps from image pairs without relying on known scene geometry or camera parameters. However,…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Jiahao Ma , Lei Wang , Miaomiao liu , David Ahmedt-Aristizabal , Chuong Nguyen

The task of 3D shape captioning occupies a significant place within the domain of computer graphics and has garnered considerable interest in recent years. Traditional approaches to this challenge frequently depend on the utilization of…

图形学 · 计算机科学 2025-09-30 Zhenyu Shu , Jiawei Wen , Shiyang Li , Shiqing Xin , Ligang Liu

Distilling 3D representations from pretrained 2D diffusion models is essential for 3D creative applications across gaming, film, and interior design. Current SDS-based methods are hindered by inefficient information distillation from…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Haoran Li , Yuli Tian , Yonghui Wang , Yong Liao , Lin Wang , Yuyang Wang , Peng Yuan Zhou

We propose Flash3D, a method for scene reconstruction and novel view synthesis from a single image which is both very generalisable and efficient. For generalisability, we start from a "foundation" model for monocular depth estimation and…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Stanislaw Szymanowicz , Eldar Insafutdinov , Chuanxia Zheng , Dylan Campbell , João F. Henriques , Christian Rupprecht , Andrea Vedaldi

World building with 3D scene representations is increasingly important for content creation, simulation, and interactive experiences, yet real workflows are inherently iterative: creators must repeatedly extend an existing scene under user…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Zijian He , Renjie Liu , Yihao Wang , Weizhi Zhong , Huan Yuan , Kun Gai , Guangrun Wang , Guanbin Li

Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer promising 3D priors…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Yabo Chen , Chen Yang , Jiemin Fang , Xiaopeng Zhang , Lingxi Xie , Wei Shen , Wenrui Dai , Hongkai Xiong , Qi Tian

3D scene reconstruction is essential for applications in virtual reality, robotics, and autonomous driving, enabling machines to understand and interact with complex environments. Traditional 3D Gaussian Splatting techniques rely on images…

图形学 · 计算机科学 2025-03-04 Changlin Song , Jiaqi Wang , Liyun Zhu , He Weng

Endovascular surgical tool reconstruction represents an important factor in advancing endovascular tool navigation, which is an important step in endovascular surgery. However, the lack of publicly available datasets significantly restricts…

图像与视频处理 · 电气工程与系统科学 2024-10-30 Tudor Jianu , Baoru Huang , Hoan Nguyen , Binod Bhattarai , Tuong Do , Erman Tjiputra , Quang Tran , Pierre Berthet-Rayne , Ngan Le , Sebastiano Fichera , Anh Nguyen

Despite advances in neural rendering, due to the scarcity of high-quality 3D datasets and the inherent limitations of multi-view diffusion models, view synthesis and 3D model generation are restricted to low resolutions with suboptimal…

计算机视觉与模式识别 · 计算机科学 2025-04-30 Yihang Luo , Shangchen Zhou , Yushi Lan , Xingang Pan , Chen Change Loy

Automatic 3D content creation seeks to replace labor-intensive modeling and scanning pipelines with systems that can synthesize or recover 3D assets directly from text or images. Its applications span video games, virtual reality, robotics,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Jiahao Li

Diffusion-based image editing is a composite process of preserving the source image content and generating new content or applying modifications. While current editing approaches have made improvements under text guidance, most of them have…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Tianrui Huang , Pu Cao , Lu Yang , Chun Liu , Mengjie Hu , Zhiwei Liu , Qing Song

Recovering 3D scenes from sparse views is a challenging task due to its inherent ill-posed problem. Conventional methods have developed specialized solutions (e.g., geometry regularization or feed-forward deterministic model) to mitigate…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Hanyang Wang , Fangfu Liu , Jiawei Chi , Yueqi Duan

This paper presents a novel method for building scalable 3D generative models utilizing pre-trained video diffusion models. The primary obstacle in developing foundation 3D generative models is the limited availability of 3D data. Unlike…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Junlin Han , Filippos Kokkinos , Philip Torr