中文
相关论文

相关论文: DL3DV-10K: A Large-Scale Scene Dataset for Deep Le…

200 篇论文

In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, high-quality training data in indoor layout design.…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Wenxu Zhou , Kaixuan Nie , Hang Du , Dong Yin , Wei Huang , Siqiang Guo , Xiaobo Zhang , Pengbo Hu

The development of generalizable Novel View Synthesis (NVS) models is critically limited by the scarcity of large-scale training data featuring diverse and precise camera trajectories. While real-world captures are photorealistic, they are…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Chenhan Jiang , Yu Chen , Qingwen Zhang , Jifei Song , Songcen Xu , Dit-Yan Yeung , Jiankang Deng

Neural Rendering representations have significantly contributed to the field of 3D computer vision. Given their potential, considerable efforts have been invested to improve their performance. Nonetheless, the essential question of…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Wenhui Xiao , Rodrigo Santa Cruz , David Ahmedt-Aristizabal , Olivier Salvado , Clinton Fookes , Leo Lebrat

Recent advances in scene understanding have leveraged multimodal large language models (MLLMs) for 3D reasoning by capitalizing on their strong 2D pretraining. However, the lack of explicit 3D data during MLLM pretraining limits 3D…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Xiaohu Huang , Jingjing Wu , Qunyi Xie , Kai Han

Neural rendering methods such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have achieved significant progress in photorealistic 3D scene reconstruction and novel view synthesis. However, most existing models assume…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Weeyoung Kwon , Jeahun Sung , Minkyu Jeon , Chanho Eom , Jihyong Oh

Implicit neural representation has demonstrated promising results in 3D reconstruction on various scenes. However, existing approaches either struggle to model fast-moving objects or are incapable of handling large-scale camera ego-motions…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Tianchen Deng , Yanbo Wang , Yejia Liu , Chenpeng Su , Jingchuan Wang , Danwei Wang , Shao-Yuan Lo , Weidong Chen

We tackle the problem of retrieving high-resolution (HR) texture maps of objects that are captured from multiple view points. In the multi-view case, model-based super-resolution (SR) methods have been recently proved to recover high…

计算机视觉与模式识别 · 计算机科学 2019-06-05 Yawei Li , Vagia Tsiminaki , Radu Timofte , Marc Pollefeys , Luc van Gool

We present DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications. Our approach is structured around three key dimensions: We strive to ensure our data is diverse,…

Modern deep learning developments create new opportunities for 3D mapping technology, scene reconstruction pipelines, and virtual reality development. Despite advances in 3D deep learning technology, direct training of deep learning models…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Xueyang Kang

We cast multiview reconstruction from unknown pose as a generative modeling problem. From a collection of unannotated 2D images of a scene, our approach simultaneously learns both a network to predict camera pose from 2D image input, as…

计算机视觉与模式识别 · 计算机科学 2024-06-12 Xin Yuan , Rana Hanocka , Michael Maire

In recent years, the field of implicit neural representation has progressed significantly. Models such as neural radiance fields (NeRF), which uses relatively small neural networks, can represent high-quality scenes and achieve…

计算机视觉与模式识别 · 计算机科学 2022-04-01 David Dadon , Ohad Fried , Yacov Hel-Or

Novel view synthesis (NVS) of multi-human scenes imposes challenges due to the complex inter-human occlusions. Layered representations handle the complexities by dividing the scene into multi-layered radiance fields, however, they are…

计算机视觉与模式识别 · 计算机科学 2023-09-22 Youssef Abdelkareem , Shady Shehata , Fakhri Karray

Realtime 4D reconstruction for dynamic scenes remains a crucial challenge for autonomous driving perception. Most existing methods rely on depth estimation through self-supervision or multi-modality sensor fusion. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Xin Fei , Wenzhao Zheng , Yueqi Duan , Wei Zhan , Masayoshi Tomizuka , Kurt Keutzer , Jiwen Lu

Inferring a meaningful geometric scene representation from a single image is a fundamental problem in computer vision. Approaches based on traditional depth map prediction can only reason about areas that are visible in the image.…

计算机视觉与模式识别 · 计算机科学 2023-04-20 Felix Wimbauer , Nan Yang , Christian Rupprecht , Daniel Cremers

Developing vision-language models (VLMs) capable of understanding 3D scenes has been a longstanding research goal. Despite recent progress, 3D VLMs still struggle with spatial reasoning and robustness. We identify three key obstacles…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Jiangyong Huang , Xiaojian Ma , Xiongkun Linghu , Junchao He , Qing Li , Song-Chun Zhu , Yixin Chen , Baoxiong Jia , Siyuan Huang

Progress in 3D computer vision tasks demands a huge amount of data, yet annotating multi-view images with 3D-consistent annotations, or point clouds with part segmentation is both time-consuming and challenging. This paper introduces…

计算机视觉与模式识别 · 计算机科学 2024-08-20 Yu Chi , Fangneng Zhan , Sibo Wu , Christian Theobalt , Adam Kortylewski

Neural Radiance Fields (NeRFs), despite their outstanding performance on novel view synthesis, often need dense input views. Many papers train one model for each scene respectively and few of them explore incorporating multi-modal data into…

计算机视觉与模式识别 · 计算机科学 2022-10-12 Haoyi Zhu , Hao-Shu Fang , Cewu Lu

Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean multi-view training data, beyond manually curated datasets…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Morris Alper , David Novotny , Filippos Kokkinos , Hadar Averbuch-Elor , Tom Monnier

In this paper, we provide a comprehensive overview of existing scene representation methods for robotics, covering traditional representations such as point clouds, voxels, signed distance functions (SDF), and scene graphs, as well as more…

We propose a learning-based approach for novel view synthesis for multi-camera 360$^{\circ}$ panorama capture rigs. Previous work constructs RGBD panoramas from such data, allowing for view synthesis with small amounts of translation, but…

计算机视觉与模式识别 · 计算机科学 2020-08-06 Kai-En Lin , Zexiang Xu , Ben Mildenhall , Pratul P. Srinivasan , Yannick Hold-Geoffroy , Stephen DiVerdi , Qi Sun , Kalyan Sunkavalli , Ravi Ramamoorthi