English
Related papers

Related papers: Rel3D: A Minimally Contrastive Benchmark for Groun…

200 papers

We present a simple unsupervised method for learning an encoder mapping short 3D pose sequences into embedding vectors suitable for sequence-to-sequence alignment by dynamic time warping. Training samples consist of temporal windows of…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Robert T. Collins

Contrastive Language-Image Pre-training (CLIP) starts to emerge in many computer vision tasks and has achieved promising performance. However, it remains underexplored whether CLIP can be generalized to 3D hand pose estimation, as bridging…

Multimedia · Computer Science 2023-09-29 Shaoxiang Guo , Qing Cai , Lin Qi , Junyu Dong

Annotation of large-scale 3D data is notoriously cumbersome and costly. As an alternative, weakly-supervised learning alleviates such a need by reducing the annotation by several order of magnitudes. We propose COARSE3D, a novel…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Rong Li , Anh-Quan Cao , Raoul de Charette

3D pose estimation from a single image is a challenging task in computer vision. We present a weakly supervised approach to estimate 3D pose points, given only 2D pose landmarks. Our method does not require correspondences between 2D and 3D…

Computer Vision and Pattern Recognition · Computer Science 2018-08-23 Dylan Drover , Rohith MV , Ching-Hang Chen , Amit Agrawal , Ambrish Tyagi , Cong Phuoc Huynh

We study 3D shape modeling from a single image and make contributions to it in three aspects. First, we present Pix3D, a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment. Pix3D has wide applications in…

Computer Vision and Pattern Recognition · Computer Science 2018-04-13 Xingyuan Sun , Jiajun Wu , Xiuming Zhang , Zhoutong Zhang , Chengkai Zhang , Tianfan Xue , Joshua B. Tenenbaum , William T. Freeman

Image Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D…

Computer Vision and Pattern Recognition · Computer Science 2024-06-17 Vincent Leroy , Yohann Cabon , Jérôme Revaud

Intelligent robots require object-level scene understanding to reason about possible tasks and interactions with the environment. Moreover, many perception tasks such as scene reconstruction, image retrieval, or place recognition can…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Cathrin Elich , Iro Armeni , Martin R. Oswald , Marc Pollefeys , Joerg Stueckler

We explore the task of geometric reconstruction of images captured from a mixture of ground and aerial views. Current state-of-the-art learning-based approaches fail to handle the extreme viewpoint variation between aerial-ground image…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Khiem Vuong , Anurag Ghosh , Deva Ramanan , Srinivasa Narasimhan , Shubham Tulsiani

Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scene-language models often struggle with this relational understanding,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jintang Xue , Ganning Zhao , Jie-En Yao , Hong-En Chen , Yue Hu , Meida Chen , Suya You , C. -C. Jay Kuo

Understanding social interactions from egocentric views is crucial for many applications, ranging from assistive robotics to AR/VR. Key to reasoning about interactions is to understand the body pose and motion of the interaction partner…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Siwei Zhang , Qianli Ma , Yan Zhang , Zhiyin Qian , Taein Kwon , Marc Pollefeys , Federica Bogo , Siyu Tang

Leveraging pre-trained 2D image representations in behavior cloning policies has achieved great success and has become a standard approach for robotic manipulation. However, such representations fail to capture the 3D spatial information…

Robotics · Computer Science 2026-05-07 I-Chun Arthur Liu , Krzysztof Choromanski , Sandy Huang , Connor Schenck

Detecting object-level changes between two images across possibly different views is a core task in many applications that involve visual inspection or camera surveillance. Existing change-detection approaches suffer from three major…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Hung Huy Nguyen , Pooyan Rahmanzadehgervi , Long Mai , Anh Totti Nguyen

Estimating the geometry level of human-scene contact aims to ground specific contact surface points at 3D human geometries, which provides a spatial prior and bridges the interaction between human and scene, supporting applications such as…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Chengfeng Wang , Wei Zhai , Yuhang Yang , Yang Cao , Zhengjun Zha

Recent advances in large multimodal models suggest that explicit reasoning mechanisms play a critical role in improving model reliability, interpretability, and cross-modal alignment. While such reasoning-centric approaches have been proven…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Tianjiao Yu , Xinzhuo Li , Yifan Shen , Yuanzhe Liu , Ismini Lourentzou

The default strategy for training single-view Large Reconstruction Models (LRMs) follows the fully supervised route using large-scale datasets of synthetic 3D assets or multi-view captures. Although these resources simplify the training…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Hanwen Jiang , Qixing Huang , Georgios Pavlakos

We consider the problem of learning object arrangements in a 3D scene. The key idea here is to learn how objects relate to human poses based on their affordances, ease of use and reachability. In contrast to modeling object-object…

Machine Learning · Computer Science 2012-07-03 Yun Jiang , Marcus Lim , Ashutosh Saxena

Capturing the interactions between humans and their environment in 3D is important for many applications in robotics, graphics, and vision. Recent works to reconstruct the 3D human and object from a single RGB image do not have consistent…

Computer Vision and Pattern Recognition · Computer Science 2023-11-01 Xianghui Xie , Bharat Lal Bhatnagar , Gerard Pons-Moll

To achieve dexterity comparable to that of humans, robots must intelligently process tactile sensor data. Taxel-based tactile signals often have low spatial-resolution, with non-standardized representations. In this paper, we propose a…

Robotics · Computer Science 2024-08-16 Hongyu Li , Snehal Dikhale , Jinda Cui , Soshi Iba , Nawid Jamali

Humans can naturally identify and mentally complete occluded objects in cluttered environments. However, imparting similar cognitive ability to robotics remains challenging even with advanced reconstruction techniques, which models scenes…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zesong Yang , Bangbang Yang , Wenqi Dong , Chenxuan Cao , Liyuan Cui , Yuewen Ma , Zhaopeng Cui , Hujun Bao

Video provides us with the spatio-temporal consistency needed for visual learning. Recent approaches have utilized this signal to learn correspondence estimation from close-by frame pairs. However, by only relying on close-by frame pairs,…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Mohamed El Banani , Ignacio Rocco , David Novotny , Andrea Vedaldi , Natalia Neverova , Justin Johnson , Benjamin Graham
‹ Prev 1 8 9 10 Next ›