English
Related papers

Related papers: EquiCaps: Predictor-Free Pose-Aware Pre-Trained Ca…

200 papers

Vision Transformer (ViT) has achieved remarkable performance in computer vision. However, positional encoding in ViT makes it substantially difficult to learn the intrinsic equivariance in data. Initial attempts have been made on designing…

Computer Vision and Pattern Recognition · Computer Science 2023-07-10 Renjun Xu , Kaifan Yang , Ke Liu , Fengxiang He

This work introduces a novel approach to achieving architecture-agnostic equivariance in deep learning, particularly addressing the limitations of traditional layerwise equivariant architectures and the inefficiencies of the existing…

Machine Learning · Computer Science 2024-11-18 Siba Smarak Panigrahi , Arnab Kumar Mondal

For human pose estimation in monocular images, joint occlusions and overlapping upon human bodies often result in deviated pose predictions. Under these circumstances, biologically implausible pose predictions may be produced. In contrast,…

Computer Vision and Pattern Recognition · Computer Science 2017-05-03 Yu Chen , Chunhua Shen , Xiu-Shen Wei , Lingqiao Liu , Jian Yang

Unsupervised representation learning holds the promise of exploiting large amounts of unlabeled data to learn general representations. A promising technique for unsupervised learning is the framework of Variational Auto-encoders (VAEs).…

Computer Vision and Pattern Recognition · Computer Science 2020-04-09 Kamal Gupta , Saurabh Singh , Abhinav Shrivastava

This paper proposes a convolution structure for learning SE(3)-equivariant features from 3D point clouds. It can be viewed as an equivariant version of kernel point convolutions (KPConv), a widely used convolution form to process point…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Minghan Zhu , Maani Ghaffari , William A. Clark , Huei Peng

Capsule network is a type of neural network that uses the spatial relationship between features to classify images. By capturing the poses and relative positions between features, its ability to recognize affine transformation is improved,…

Machine Learning · Computer Science 2021-12-21 Jiazhu Dai , Siwei Xiong

Accurately modeling agent behaviors is an important task in self-driving. It is also a task with many symmetries, such as equivariance to the order of agents and objects in the scene or equivariance to arbitrary roto-translations of the…

Robotics · Computer Science 2026-04-03 Scott Xu , Dian Chen , Kelvin Wong , Chris Zhang , Kion Fallah , Raquel Urtasun

Self-supervised learning for inverse problems allows to train a reconstruction network from noise and/or incomplete data alone. These methods have the potential of enabling learning-based solutions when obtaining ground-truth references for…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Victor Sechaud , Jérémy Scanvic , Quentin Barthélemy , Patrice Abry , Julián Tachella

Training accurate 3D human pose estimators requires large amount of 3D ground-truth data which is costly to collect. Various weakly or self supervised pose estimation methods have been proposed due to lack of 3D data. Nevertheless, these…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Muhammed Kocabas , Salih Karagoz , Emre Akbas

Estimating 3D from 2D is one of the central tasks in computer vision. In this work, we consider the monocular setting, i.e. single-view input, for 3D human pose estimation (HPE). Here, the task is to predict a 3D point set of human skeletal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Pavlo Melnyk , Cuong Le , Urs Waldmann , Per-Erik Forssén , Bastian Wandt

Capsule Networks, as alternatives to Convolutional Neural Networks, have been proposed to recognize objects from images. The current literature demonstrates many advantages of CapsNets over CNNs. However, how to create explanations for…

Computer Vision and Pattern Recognition · Computer Science 2021-03-09 Jindong Gu , Volker Tresp

Video Capsule Endoscopy (VCE) has become an indispensable diagnostic tool for gastrointestinal (GI) disorders due to its non-invasive nature and ability to capture high-resolution images of the small intestine. However, the enormous volume…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Vamshi Krishna Kancharla , Pavan Kumar Kaveti , Dasari Naga Raju

Conventional 2D pose estimation models are constrained by their design to specific object categories. This limits their applicability to predefined objects. To overcome these limitations, category-agnostic pose estimation (CAPE) emerged as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Matan Rusanovsky , Or Hirschorn , Shai Avidan

3D point cloud is an efficient and flexible representation of 3D structures. Recently, neural networks operating on point clouds have shown superior performance on 3D understanding tasks such as shape classification and part segmentation.…

Computer Vision and Pattern Recognition · Computer Science 2019-10-21 Wentao Yuan , David Held , Christoph Mertz , Martial Hebert

We propose a method to learn object representations from 3D point clouds using bundles of geometrically interpretable hidden units, which we call geometric capsules. Each geometric capsule represents a visual entity, such as an object or a…

Machine Learning · Computer Science 2019-12-10 Nitish Srivastava , Hanlin Goh , Ruslan Salakhutdinov

Neural architecture search has proven to be highly effective in the design of efficient convolutional neural networks that are better suited for mobile deployment than hand-designed networks. Hypothesizing that neural architecture search…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 William McNally , Kanav Vats , Alexander Wong , John McPhee

Human pose estimation in complicated situations has always been a challenging task. Many Transformer-based pose networks have been proposed recently, achieving encouraging progress in improving performance. However, the remarkable…

Computer Vision and Pattern Recognition · Computer Science 2023-11-27 Chengpeng Wu , Guangxing Tan , Chunyu Li

This work seeks to improve the generalization and robustness of existing neural networks for 3D point clouds by inducing group equivariance under general group transformations. The main challenge when designing equivariant models for point…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Thuan N. A. Trang , Thieu N. Vo , Khuong D. Nguyen

Point cloud registration is a foundational task for 3D alignment and reconstruction applications. While both traditional and learning-based registration approaches have succeeded, leveraging the intrinsic symmetry of point cloud data,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Xueyang Kang , Zhaoliang Luan , Kourosh Khoshelham , Bing Wang

Recently proposed Capsule Network is a brain inspired architecture that brings a new paradigm to deep learning by modelling input domain variations through vector based representations. Despite being a seminal contribution, CapsNet does not…

Computer Vision and Pattern Recognition · Computer Science 2018-10-17 Sameera Ramasinghe , C. D. Athuralya , Salman Khan