English
Related papers

Related papers: Physically Guided Visual Mass Estimation from a Si…

200 papers

We show that generative models can be used to capture visual geometry constraints statistically. We use this fact to infer the 3D shape of object categories from raw single-view images. Differently from prior work, we use no external…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Shangzhe Wu , Christian Rupprecht , Andrea Vedaldi

We present a method to infer 3D pose and shape of vehicles from a single image. To tackle this ill-posed problem, we optimize two-scale projection consistency between the generated 3D hypotheses and their 2D pseudo-measurements.…

Computer Vision and Pattern Recognition · Computer Science 2019-01-14 Tong He , Stefano Soatto

Semantic aware reconstruction is more advantageous than geometric-only reconstruction for future robotic and AR/VR applications because it represents not only where things are, but also what things are. Object-centric mapping is a task to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Kejie Li , Hamid Rezatofighi , Ian Reid

Gaze object prediction (GOP) aims to predict the category and location of the object that a human is looking at. Previous methods utilized box-level supervision to identify the object that a person is looking at, but struggled with semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Yang Jin , Lei Zhang , Shi Yan , Bin Fan , Binglu Wang

Accurate 6D object pose estimation is a prerequisite for successfully completing robotic prehensile and non-prehensile manipulation tasks. At present, 6D pose estimation for robotic manipulation generally relies on depth sensors based on,…

Robotics · Computer Science 2025-06-23 Teng Guo , Baichuan Huang , Jingjin Yu

Previous image based relighting methods require capturing multiple images to acquire high frequency lighting effect under different lighting conditions, which needs nontrivial effort and may be unrealistic in certain practical use…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Di Qiu , Jin Zeng , Zhanghan Ke , Wenxiu Sun , Chengxi Yang

The goal of this paper is to compare surface-based and volumetric 3D object shape representations, as well as viewer-centered and object-centered reference frames for single-view 3D shape prediction. We propose a new algorithm for…

Computer Vision and Pattern Recognition · Computer Science 2018-06-13 Daeyun Shin , Charless C. Fowlkes , Derek Hoiem

Most self-supervised 6D object pose estimation methods can only work with additional depth information or rely on the accurate annotation of 2D segmentation masks, limiting their application range. In this paper, we propose a 6D object pose…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Yang Hai , Rui Song , Jiaojiao Li , David Ferstl , Yinlin Hu

A single color image can contain many cues informative towards different aspects of local geometric structure. We approach the problem of monocular depth estimation by using a neural network to produce a mid-level representation that…

Computer Vision and Pattern Recognition · Computer Science 2016-09-08 Ayan Chakrabarti , Jingyu Shao , Gregory Shakhnarovich

Perceiving 3D structures from RGB images based on CAD model primitives can enable an effective, efficient 3D object-based representation of scenes. However, current approaches rely on supervision from expensive annotations of CAD models…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Daoyi Gao , Dávid Rozenberszki , Stefan Leutenegger , Angela Dai

Monocular normal estimation aims to estimate the normal map from a single RGB image of an object under arbitrary lights. Existing methods rely on deep models to directly predict normal maps. However, they often suffer from 3D misalignment:…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zongrui Li , Xinhua Ma , Minghui Hu , Yunqing Zhao , Yingchen Yu , Qian Zheng , Chang Liu , Xudong Jiang , Song Bai

Detecting and localizing glass in 3D environments poses significant challenges for visual perception systems, as the optical properties of glass often hinder conventional sensors from accurately distinguishing glass surfaces. The lack of…

Robotics · Computer Science 2025-09-09 Kai Zhang , Guoyang Zhao , Jianxing Shi , Bonan Liu , Weiqing Qi , Jun Ma

Reliance on images for dietary assessment is an important strategy to accurately and conveniently monitor an individual's health, making it a vital mechanism in the prevention and care of chronic diseases and obesity. However, image-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Gautham Vinod , Fengqing Zhu

Metric depth estimation from visual sensors is crucial for robots to perceive, navigate, and interact with their environment. Traditional range imaging setups, such as stereo or structured light cameras, face hassles including calibration,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Blanca Lasheras-Hernandez , Klaus H. Strobl , Sergio Izquierdo , Tim Bodenmüller , Rudolph Triebel , Javier Civera

In recent years, 3D generation has made great strides in both academia and industry. However, generating 3D scenes from a single RGB image remains a significant challenge, as current approaches often struggle to ensure both object…

Graphics · Computer Science 2026-02-18 Xiang Tang , Ruotong Li , Xiaopeng Fan

Monocular depth estimation is very challenging because clues to the exact depth are incomplete in a single RGB image. To overcome the limitation, deep neural networks rely on various visual hints such as size, shade, and texture extracted…

Computer Vision and Pattern Recognition · Computer Science 2023-04-26 Kyuhong Shim , Jiyoung Kim , Gusang Lee , Byonghyo Shim

We consider the problem of next frame prediction from video input. A recurrent convolutional neural network is trained to predict depth from monocular video input, which, along with the current video image and the camera trajectory, can…

Machine Learning · Computer Science 2017-06-14 Reza Mahjourian , Martin Wicke , Anelia Angelova

We consider the problem of vision-based pose estimation for autonomous systems. While deep neural networks have been successfully used for vision-based tasks, they inherently lack provable guarantees on the correctness of their output,…

Robotics · Computer Science 2026-01-27 Ulices Santa Cruz , Mahmoud Elfar , Yasser Shoukry

Monocular 3D object detection task aims to predict the 3D bounding boxes of objects based on monocular RGB images. Since the location recovery in 3D space is quite difficult on account of absence of depth information, this paper proposes a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-10 Yingjie Cai , Buyu Li , Zeyu Jiao , Hongsheng Li , Xingyu Zeng , Xiaogang Wang

In this paper, we propose a monocular 3D object detection framework in the domain of autonomous driving. Unlike previous image-based methods which focus on RGB feature extracted from 2D images, our method solves this problem in the…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Xinzhu Ma , Zhihui Wang , Haojie Li , Pengbo Zhang , Xin Fan , Wanli Ouyang