English
Related papers

Related papers: RGBD Datasets: Past, Present and Future

200 papers

3D scene understanding is important for robots to interact with the 3D world in a meaningful way. Most previous works on 3D scene understanding focus on recognizing geometrical or semantic properties of the scene independently. In this…

Computer Vision and Pattern Recognition · Computer Science 2017-06-01 Yu Xiang , Dieter Fox

6D object pose estimation aims at determining an object's translation, rotation, and scale, typically from a single RGBD image. Recent advancements have expanded this estimation from instance-level to category-level, allowing models to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Mengchen Zhang , Tong Wu , Tai Wang , Tengfei Wang , Ziwei Liu , Dahua Lin

Estimating depth from RGB images is a long-standing ill-posed problem, which has been explored for decades by the computer vision, graphics, and machine learning communities. In this article, we provide a comprehensive survey of the recent…

Computer Vision and Pattern Recognition · Computer Science 2019-06-17 Hamid Laga

Recognizing objects in images is a fundamental problem in computer vision. Although detecting objects in 2D images is common, many applications require determining their pose in 3D space. Traditional category-level methods rely on RGB-D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Tom Fischer , Xiaojie Zhang , Eddy Ilg

We present the first publicly-available RGB-thermal dataset designed for aerial robotics operating in natural environments. Our dataset captures a variety of terrain across the United States, including rivers, lakes, coastlines, deserts,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-02 Connor Lee , Matthew Anderson , Nikhil Raganathan , Xingxing Zuo , Kevin Do , Georgia Gkioxari , Soon-Jo Chung

Augmenting RGB data with measured depth has been shown to improve the performance of a range of tasks in computer vision including object detection and semantic segmentation. Although depth sensors such as the Microsoft Kinect have…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 Yuanzhouhan Cao , Chunhua Shen , Heng Tao Shen

3D hand pose estimation based on RGB images has been studied for a long time. Most of the studies, however, have performed frame-by-frame estimation based on independent static images. In this paper, we attempt to not only consider the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-13 John Yang , Hyung Jin Chang , Seungeui Lee , Nojun Kwak

Over the last decade, Computer Vision, the branch of Artificial Intelligence aimed at understanding the visual world, has evolved from simply recognizing objects in images to describing pictures, answering questions about images, aiding…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Ranjay Krishna , Mitchell Gordon , Li Fei-Fei , Michael Bernstein

In this paper, we present a dataset capturing diverse visual data formats that target varying luminance conditions. While RGB cameras provide nourishing and intuitive information, changes in lighting conditions potentially result in…

Robotics · Computer Science 2022-04-15 Alex Junho Lee , Younggun Cho , Young-sik Shin , Ayoung Kim , Hyun Myung

Image semantic segmentation is more and more being of interest for computer vision and machine learning researchers. Many applications on the rise need accurate and efficient segmentation mechanisms: autonomous driving, indoor navigation,…

Computer Vision and Pattern Recognition · Computer Science 2017-04-25 Alberto Garcia-Garcia , Sergio Orts-Escolano , Sergiu Oprea , Victor Villena-Martinez , Jose Garcia-Rodriguez

Most vision models are trained on RGB images processed through ISP pipelines optimized for human perception, which can discard sensor-level information useful for machine reasoning. RAW images preserve unprocessed scene data, enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Mishal Fatima , Shashank Agnihotri , Kanchana Vaishnavi Gandikota , Michael Moeller , Margret Keuper

We present the first real-time method to capture the full global 3D skeletal pose of a human in a stable, temporally consistent manner using a single RGB camera. Our method combines a new convolutional neural network (CNN) based pose…

Computer Vision and Pattern Recognition · Computer Science 2018-11-26 Dushyant Mehta , Srinath Sridhar , Oleksandr Sotnychenko , Helge Rhodin , Mohammad Shafiei , Hans-Peter Seidel , Weipeng Xu , Dan Casas , Christian Theobalt

Tracking and reconstructing the 3D pose and geometry of two hands in interaction is a challenging problem that has a high relevance for several human-computer interaction applications, including AR/VR, robotics, or sign language…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Jiayi Wang , Franziska Mueller , Florian Bernard , Suzanne Sorli , Oleksandr Sotnychenko , Neng Qian , Miguel A. Otaduy , Dan Casas , Christian Theobalt

Current approaches to 3D scene graph generation rely on dedicated depth sensors, such as LiDAR or RGB-D cameras, for metric 3D reconstruction. This limits deployment to specialized robotic platforms and excludes settings where only RGB…

Robotics · Computer Science 2026-05-19 Giorgia Modi , Davide Buoso , Giuseppe Averta , Daniele De Martini

Over the past decade deep learning has driven progress in 2D image understanding. Despite these advancements, techniques for automatic 3D sensed data understanding, such as point clouds, is comparatively immature. However, with a range of…

Computer Vision and Pattern Recognition · Computer Science 2019-07-11 David Griffiths , Jan Boehm

We address the problem of registering synchronized color (RGB) and multi-spectral (MS) images featuring very different resolution by solving stereo matching correspondences. Purposely, we introduce a novel RGB-MS dataset framing 13…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Fabio Tosi , Pierluigi Zama Ramirez , Matteo Poggi , Samuele Salti , Stefano Mattoccia , Luigi Di Stefano

Referring Multi-Object Tracking (RMOT) aims to track specific targets based on language descriptions and is vital for interactive AI systems such as robotics and autonomous driving. However, existing RMOT models rely solely on 2D RGB data,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Sijia Chen , Lijuan Ma , Yanqiu Yu , En Yu , Liman Liu , Wenbing Tao

We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuki Kawana , Tatsuya Harada

In this paper, we introduce a new benchmark dataset named IPN Hand with sufficient size, variety, and real-world elements able to train and evaluate deep neural networks. This dataset contains more than 4,000 gesture samples and 800,000 RGB…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Gibran Benitez-Garcia , Jesus Olivares-Mercado , Gabriel Sanchez-Perez , Keiji Yanai

Gender is an important demographic attribute of people. This paper provides a survey of human gender recognition in computer vision. A review of approaches exploiting information from face and whole body (either from a still image or gait…

Computer Vision and Pattern Recognition · Computer Science 2012-04-10 Choon Boon Ng , Yong Haur Tay , Bok Min Goi