English
Related papers

Related papers: Self-supervised Learning of Interpretable Keypoint…

200 papers

Recovering 3D human pose from 2D joints is still a challenging problem, especially without any 3D annotation, video information, or multi-view information. In this paper, we present an unsupervised GAN-based model consisting of multiple…

Computer Vision and Pattern Recognition · Computer Science 2022-04-14 Yicheng Deng , Cheng Sun , Jiahui Zhu , Yongqi Sun

Video-based person re-identification deals with the inherent difficulty of matching unregulated sequences with different length and with incomplete target pose/viewpoint structure. Common approaches operate either by reducing the problem to…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Alessandro Borgia , Yang Hua , Elyor Kodirov , Neil M. Robertson

Visual scenes are extremely rich in diversity, not only because there are infinite combinations of objects and background, but also because the observations of the same scene may vary greatly with the change of viewpoints. When observing a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Jinyang Yuan , Bin Li , Xiangyang Xue

Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Fa-Ting Hong , Li Shen , Dan Xu

We present 6-PACK, a deep learning approach to category-level 6D object pose tracking on RGB-D data. Our method tracks in real-time novel object instances of known object categories such as bowls, laptops, and mugs. 6-PACK learns to…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Chen Wang , Roberto Martín-Martín , Danfei Xu , Jun Lv , Cewu Lu , Li Fei-Fei , Silvio Savarese , Yuke Zhu

We propose a technique for learning single-view 3D object pose estimation models by utilizing a new source of data -- in-the-wild videos where objects turn. Such videos are prevalent in practice (e.g., cars in roundabouts, airplanes near…

Computer Vision and Pattern Recognition · Computer Science 2022-12-14 Zezhou Cheng , Matheus Gadelha , Subhransu Maji

The practicality of 3D object pose estimation remains limited for many applications due to the need for prior knowledge of a 3D model and a training period for new objects. To address this limitation, we propose an approach that takes a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Van Nguyen Nguyen , Thibault Groueix , Yinlin Hu , Mathieu Salzmann , Vincent Lepetit

We show that generative models can be used to capture visual geometry constraints statistically. We use this fact to infer the 3D shape of object categories from raw single-view images. Differently from prior work, we use no external…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Shangzhe Wu , Christian Rupprecht , Andrea Vedaldi

Category-level object pose estimation aims to predict the 6D pose and size of previously unseen instances from predefined categories, requiring strong generalization across diverse object instances. Although many previous methods attempt to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Xiao Zhang , Lu Zou , Tao Lu , Yuan Yao , Zhangjin Huang , Guoping Wang

We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-10-28 Angtian Wang , Shenxiao Mei , Alan Yuille , Adam Kortylewski

Accurate 3D human pose estimation from single images is possible with sophisticated deep-net architectures that have been trained on very large datasets. However, this still leaves open the problem of capturing motions for which no such…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Helge Rhodin , Jörg Spörri , Isinsu Katircioglu , Victor Constantin , Frédéric Meyer , Erich Müller , Mathieu Salzmann , Pascal Fua

Segment Anything (SAM) provides an unprecedented foundation for human segmentation, but may struggle under occlusion, where keypoints may be partially or fully invisible. We adapt SAM 2.1 for pose-guided segmentation with minimal encoder…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Constantin Kolomiiets , Miroslav Purkrabek , Jiri Matas

Current best local descriptors are learned on a large dataset of matching and non-matching keypoint pairs. However, data of this kind is not always available since detailed keypoint correspondences can be hard to establish. On the other…

Computer Vision and Pattern Recognition · Computer Science 2019-05-08 Nenad Markuš , Igor S. Pandžić , Jörgen Ahlberg

We introduce an unsupervised feature learning approach that embeds 3D shape information into a single-view image representation. The main idea is a self-supervised training objective that, given only a single 2D image, requires all unseen…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Dinesh Jayaraman , Ruohan Gao , Kristen Grauman

Human pose estimation from single images is a challenging problem in computer vision that requires large amounts of labeled training data to be solved accurately. Unfortunately, for many human activities (\eg outdoor sports) such training…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Bastian Wandt , Marco Rudolph , Petrissa Zell , Helge Rhodin , Bodo Rosenhahn

A video autoencoder is proposed for learning disentan- gled representations of 3D structure and camera pose from videos in a self-supervised manner. Relying on temporal continuity in videos, our work assumes that the 3D scene structure in…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Zihang Lai , Sifei Liu , Alexei A. Efros , Xiaolong Wang

This paper introduces KeyDiff3D, a framework for unsupervised monocular 3D keypoints estimation that accurately predicts 3D keypoints from a single image. While previous methods rely on manual annotations or calibrated multi-view images,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Subin Jeon , In Cho , Junyoung Hong , Seon Joo Kim

Learning neural implicit fields of 3D shapes is a rapidly emerging field that enables shape representation at arbitrary resolutions. Due to the flexibility, neural implicit fields have succeeded in many research areas, including shape…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Yifei Shi , Boyan Wan , Xin Xu , Kai Xu

This paper introduces self-taught object localization, a novel approach that leverages deep convolutional networks trained for whole-image recognition to localize objects in images without additional human supervision, i.e., without using…

Computer Vision and Pattern Recognition · Computer Science 2016-02-03 Loris Bazzani , Alessandro Bergamo , Dragomir Anguelov , Lorenzo Torresani

We propose a new method for estimating the relative pose between two images, where we jointly learn keypoint detection, description extraction, matching and robust pose estimation. While our architecture follows the traditional pipeline for…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Antoine Fond , Luca Del Pero , Nikola Sivacki , Marco Paladini