English
Related papers

Related papers: 3D sans 3D Scans: Scalable Pre-training from Video…

200 papers

3D object detection is an important task in computer vision. Most existing methods require a large number of high-quality 3D annotations, which are expensive to collect. Especially for outdoor scenes, the problem becomes more severe due to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Hongyi Xu , Fengqi Liu , Qianyu Zhou , Jinkun Hao , Zhijie Cao , Zhengyang Feng , Lizhuang Ma

The significant achievements of pre-trained models leveraging large volumes of data in the field of NLP and 2D vision inspire us to explore the potential of extensive data pre-training for 3D perception in autonomous driving. Toward this…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Shumin Wang , Zhuoran Yang , Lidian Wang , Zhipeng Tang , Heng Li , Lehan Pan , Sha Zhang , Jie Peng , Jianmin Ji , Yanyong Zhang

Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use varied tokenization approaches. This work presents a rigorous…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Hugues Thomas , Chen Chen , Jian Zhang

Geometric feature extraction is a crucial component of point cloud registration pipelines. Recent work has demonstrated how supervised learning can be leveraged to learn better and more compact 3D features. However, those approaches'…

Computer Vision and Pattern Recognition · Computer Science 2021-06-02 Mohamed El Banani , Justin Johnson

Acquiring detailed 3D scenes typically demands costly equipment, multi-view data, or labor-intensive modeling. Therefore, a lightweight alternative, generating complex 3D scenes from a single top-down image, plays an essential role in…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Kaizhi Zheng , Ruijian Zha , Zishuo Xu , Jing Gu , Jie Yang , Xin Eric Wang

Shape priors learned from data are commonly used to reconstruct 3D objects from partial or noisy data. Yet no such shape priors are available for indoor scenes, since typical 3D autoencoders cannot handle their scale, complexity, or…

Computer Vision and Pattern Recognition · Computer Science 2020-03-23 Chiyu Max Jiang , Avneesh Sud , Ameesh Makadia , Jingwei Huang , Matthias Nießner , Thomas Funkhouser

In this paper we set out to solve the task of 6-DOF 3D object detection from 2D images, where the only supervision is a geometric representation of the objects we aim to find. In doing so, we remove the need for 6-DOF labels (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-19 David Griffiths , Jan Boehm , Tobias Ritschel

3D object detection is fundamental for spatial understanding. Real-world environments demand models capable of recognizing diverse, previously unseen objects, which remains a major limitation of closed-set methods. Existing open-vocabulary…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Andrey Lemeshko , Bulat Gabdullin , Nikita Drozdov , Anton Konushin , Danila Rukhovich , Maksim Kolodiazhnyi

Large Reconstruction Models (LRMs) have recently become a popular method for creating 3D foundational models. Training 3D reconstruction models with 2D visual data traditionally requires prior knowledge of camera poses for the training…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Shiu-hong Kao , Xiao Li , Jinglu Wang , Yang Li , Chi-Keung Tang , Yu-Wing Tai , Yan Lu

In this work, we propose to learn local descriptors for point clouds in a self-supervised manner. In each iteration of the training, the input of the network is merely one unlabeled point cloud. On top of our previous work, that directly…

Robotics · Computer Science 2020-03-12 Yijun Yuan , Jiawei Hou , Andreas Nüchter , Sören Schwertfeger

3D object detection is a core perceptual challenge for robotics and autonomous driving. However, the class-taxonomies in modern autonomous driving datasets are significantly smaller than many influential 2D detection datasets. In this work,…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Benjamin Wilson , Zsolt Kira , James Hays

We propose a training-free and robust solution to offer camera movement control for off-the-shelf video diffusion models. Unlike previous work, our method does not require any supervised finetuning on camera-annotated datasets or…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Chen Hou , Zhibo Chen

We propose an unsupervised method for 3D geometry-aware representation learning of articulated objects, in which no image-pose pairs or foreground masks are used for training. Though photorealistic images of articulated objects can be…

Computer Vision and Pattern Recognition · Computer Science 2022-09-28 Atsuhiro Noguchi , Xiao Sun , Stephen Lin , Tatsuya Harada

Point cloud registration is the process of aligning a pair of point sets via searching for a geometric transformation. Recent works leverage the power of deep learning for registering a pair of point sets. However, unfortunately, deep…

Computational Geometry · Computer Science 2020-06-12 Lingjing Wang , Xiang Li , Yi Fang

Understanding the 3D world without supervision is currently a major challenge in computer vision as the annotations required to supervise deep networks for tasks in this domain are expensive to obtain on a large scale. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Octave Mariotti , Oisin Mac Aodha , Hakan Bilen

Recent years have seen flourishing research on both semi-supervised learning and 3D room layout reconstruction. In this work, we explore the intersection of these two fields to advance the research objective of enabling more accurate 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-05-18 Phi Vu Tran

Graph Convolutional Networks(GCNs) play a crucial role in graph learning tasks, however, learning graph embedding with few supervised signals is still a difficult problem. In this paper, we propose a novel training algorithm for Graph…

Machine Learning · Computer Science 2020-02-21 Ke Sun , Zhouchen Lin , Zhanxing Zhu

Large Multimodal Models (LMMs) that process 3D data typically rely on heavy, pre-trained visual encoders to extract geometric features. While recent 2D LMMs have begun to eliminate such encoders for efficiency and scalability, extending…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Guofeng Mei , Wei Lin , Luigi Riz , Yujiao Wu , Yiming Wang , Fabio Poiesi

We present a novel approach to the generation of static and articulated 3D assets that has a 3D autodecoder at its core. The 3D autodecoder framework embeds properties learned from the target dataset in the latent space, which can then be…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Evangelos Ntavelis , Aliaksandr Siarohin , Kyle Olszewski , Chaoyang Wang , Luc Van Gool , Sergey Tulyakov

Modeling the 3D world from sensor data for simulation is a scalable way of developing testing and validation environments for robotic learning problems such as autonomous driving. However, manually creating or re-creating real-world-like…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Bokui Shen , Xinchen Yan , Charles R. Qi , Mahyar Najibi , Boyang Deng , Leonidas Guibas , Yin Zhou , Dragomir Anguelov