中文
相关论文

相关论文: UniRecGen: Unifying Multi-View 3D Reconstruction a…

200 篇论文

Open-world 3D reconstruction models have recently garnered significant attention. However, without sufficient 3D inductive bias, existing methods typically entail expensive training costs and struggle to extract high-quality 3D meshes. In…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Minghua Liu , Chong Zeng , Xinyue Wei , Ruoxi Shi , Linghao Chen , Chao Xu , Mengqi Zhang , Zhaoning Wang , Xiaoshuai Zhang , Isabella Liu , Hongzhi Wu , Hao Su

A fundamental problem in the texturing of 3D meshes using pre-trained text-to-image models is to ensure multi-view consistency. State-of-the-art approaches typically use diffusion models to aggregate multi-view inputs, where common issues…

计算机视觉与模式识别 · 计算机科学 2024-08-05 Zhengyi Zhao , Chen Song , Xiaodong Gu , Yuan Dong , Qi Zuo , Weihao Yuan , Liefeng Bo , Zilong Dong , Qixing Huang

3D human mesh recovery from monocular RGB images aims to estimate anatomically plausible 3D human models for downstream applications, but remains challenging under partial or severe occlusions. Regression-based methods are efficient yet…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Yang Liu , Zhiyong Zhang

Reconstructing 3D scenes from a single image is a fundamentally ill-posed task due to the severely under-constrained nature of the problem. Consequently, when the scene is rendered from novel camera views, existing single image to 3D…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Sarosij Bose , Arindam Dutta , Sayak Nag , Junge Zhang , Jiachen Li , Konstantinos Karydis , Amit K. Roy Chowdhury

Single-image 3D generation has emerged as a prominent research topic, playing a vital role in virtual reality, 3D modeling, and digital content creation. However, existing methods face challenges such as a lack of multi-view geometric…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Jinbo Yan , Alan Zhao , Yixin Hu

Transforming casually captured, monocular videos into fully immersive dynamic experiences is a highly ill-posed task, and comes with significant challenges, e.g., reconstructing unseen regions, and dealing with the ambiguity in monocular…

图形学 · 计算机科学 2026-04-08 Denis Rozumny , Jonathon Luiten , Numair Khan , Johannes Schönberger , Peter Kontschieder

The past few years have witnessed the rapid development of vision-centric 3D perception in autonomous driving. Although the 3D perception models share many structural and conceptual similarities, there still exist gaps in their feature…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yu Hong , Qian Liu , Huayuan Cheng , Danjiao Ma , Hang Dai , Yu Wang , Guangzhi Cao , Yong Ding

Diffusion-based models demonstrate impressive generation capabilities. However, they also have a massive number of parameters, resulting in enormous model sizes, thus making them unsuitable for deployment on resource-constraint devices.…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Avideep Mukherjee , Soumya Banerjee , Piyush Rai , Vinay P. Namboodiri

While many unsupervised learning models focus on one family of tasks, either generative or discriminative, we explore the possibility of a unified representation learner: a model which uses a single pre-training stage to address both…

Semantic-aware 3D reconstruction from sparse, unposed images remains challenging for feed-forward 3D Gaussian Splatting (3DGS). Existing methods often predict an over-complete set of Gaussian primitives under sparse-view supervision,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Guibiao Liao , Qian Ren , Kaimin Liao , Hua Wang , Zhi Chen , Luchao Wang , Yaohua Tang

We propose UpFusion, a system that can perform novel view synthesis and infer 3D representations for an object given a sparse set of reference images without corresponding pose information. Current sparse-view 3D inference methods typically…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Bharath Raj Nagoor Kani , Hsin-Ying Lee , Sergey Tulyakov , Shubham Tulsiani

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

We present UniRef-Image-Edit, a high-performance multi-modal generation system that unifies single-image editing and multi-image composition within a single framework. Existing diffusion-based editing methods often struggle to maintain…

Sparse-view novel view synthesis is fundamentally ill-posed due to severe geometric ambiguity. Current methods are caught in a trade-off: regressive models are geometrically faithful but incomplete, whereas generative models can complete…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Atakan Topaloglu , Kunyi Li , Michael Niemeyer , Nassir Navab , A. Murat Tekalp , Federico Tombari

Low-field to high-field MRI synthesis has emerged as a cost-effective strategy to enhance image quality under hardware and acquisition constraints, particularly in scenarios where access to high-field scanners is limited or impractical.…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Zhenxuan Zhang , Peiyuan Jing , Ruicheng Yuan , Liwei Hu , Anbang Wang , Fanwen Wang , Yinzhe Wu , Kh Tohidul Islam , Zhaolin Chen , Zi Wang , Peter Lally , Guang Yang

Although recent advances in visual generation have been remarkable, most existing architectures still depend on distinct encoders for images and text. This separation constrains diffusion models' ability to perform cross-modal reasoning and…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Kevin Li , Manuel Brack , Sudeep Katakol , Hareesh Ravi , Ajinkya Kale

Effectively designing molecular geometries is essential to advancing pharmaceutical innovations, a domain, which has experienced great attention through the success of generative models and, in particular, diffusion models. However, current…

生物大分子 · 定量生物学 2025-01-07 Sirine Ayadi , Leon Hetzel , Johanna Sommer , Fabian Theis , Stephan Günnemann

We introduce a new diffusion-based approach for shape completion on 3D range scans. Compared with prior deterministic and probabilistic methods, we strike a balance between realism, multi-modality, and high fidelity. We propose DiffComplete…

计算机视觉与模式识别 · 计算机科学 2023-06-29 Ruihang Chu , Enze Xie , Shentong Mo , Zhenguo Li , Matthias Nießner , Chi-Wing Fu , Jiaya Jia

Recent advances in video generation have made it possible to produce visually compelling videos, with wide-ranging applications in content creation, entertainment, and virtual reality. However, most existing diffusion transformer based…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Teng Hu , Jiangning Zhang , Zihan Su , Ran Yi

We present a diffusion-based model for 3D-aware generative novel view synthesis from as few as a single input image. Our model samples from the distribution of possible renderings consistent with the input and, even in the presence of…