English
Related papers

Related papers: Adaptive Geodesic Conformal Prediction for Egocent…

200 papers

We propose the first real-time approach for the egocentric estimation of 3D human body pose in a wide range of unconstrained everyday activities. This setting has a unique set of challenges, such as mobility of the hardware setup, and…

Computer Vision and Pattern Recognition · Computer Science 2019-01-24 Weipeng Xu , Avishek Chatterjee , Michael Zollhoefer , Helge Rhodin , Pascal Fua , Hans-Peter Seidel , Christian Theobalt

The recently released Ego4D dataset and benchmark significantly scales and diversifies the first-person visual perception data. In Ego4D, the Visual Queries 2D Localization task aims to retrieve objects appeared in the past from the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Mengmeng Xu , Cheng-Yang Fu , Yanghao Li , Bernard Ghanem , Juan-Manuel Perez-Rua , Tao Xiang

This paper presents SIM-Sync, a certifiably optimal algorithm that estimates camera trajectory and 3D scene structure directly from multiview image keypoints. SIM-Sync fills the gap between pose graph optimization and bundle adjustment; the…

Robotics · Computer Science 2023-09-12 Xihang Yu , Heng Yang

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Estimating relative camera poses from consecutive frames is a fundamental problem in visual odometry (VO) and simultaneous localization and mapping (SLAM), where classic methods consisting of hand-crafted features and sampling-based outlier…

Computer Vision and Pattern Recognition · Computer Science 2020-07-31 You-Yi Jau , Rui Zhu , Hao Su , Manmohan Chandraker

Large-scale colored point clouds have many advantages in navigation or scene display. Relying on cameras and LiDARs, which are now widely used in reconstruction tasks, it is possible to obtain such colored point clouds. However, the…

Robotics · Computer Science 2023-02-28 Jiadi Cui , Sören Schwertfeger

Estimating the 3D hand articulation from a single color image is an important problem with applications in Augmented Reality (AR), Virtual Reality (VR), Human-Computer Interaction (HCI), and robotics. Apart from the absence of depth…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Christos Pantazopoulos , Spyridon Thermos , Gerasimos Potamianos

We present a method for simultaneously estimating 3D human pose and body shape from a sparse set of wide-baseline camera views. We train a symmetric convolutional autoencoder with a dual loss that enforces learning of a latent…

Computer Vision and Pattern Recognition · Computer Science 2018-07-05 Matthew Trumble , Andrew Gilbert , Adrian Hilton , John Collomosse

Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Yamei Chen , Yan Di , Guangyao Zhai , Fabian Manhardt , Chenyangguang Zhang , Ruida Zhang , Federico Tombari , Nassir Navab , Benjamin Busam

Partial point cloud registration is essential for autonomous perception and 3D scene understanding, yet it remains challenging owing to structural ambiguity, partial visibility, and noise. We address these issues by proposing Confidence…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yongqiang Wang , Weigang Li , Wenping Liu , Zhe Xu , Zhiqiang Tian

In this paper, we present Adaptive Computation Steps (ACS) algo-rithm, which enables end-to-end speech recognition models to dy-namically decide how many frames should be processed to predict a linguistic output. The model that applies ACS…

Audio and Speech Processing · Electrical Eng. & Systems 2018-09-27 Mohan Li , Min Liu , Masanori Hattori

We consider the problem of estimating object pose and shape from an RGB-D image. Our first contribution is to introduce CRISP, a category-agnostic object pose and shape estimation pipeline. The pipeline implements an encoder-decoder model…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Jingnan Shi , Rajat Talak , Harry Zhang , David Jin , Luca Carlone

Estimating rigid objects' poses is one of the fundamental problems in computer vision, with a range of applications across automation and augmented reality. Most existing approaches adopt one network per object class strategy, depend…

Computer Vision and Pattern Recognition · Computer Science 2024-10-14 Jianyu Zhao , Wei Quan , Bogdan J. Matuszewski

Category-level articulated object pose estimation focuses on the pose estimation of unknown articulated objects within known categories. Despite its significance, this task remains challenging due to the varying shapes and poses of objects,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Yuchen Che , Ryo Furukawa , Asako Kanezaki

Camera, and associated with its objects within the field of view, localization could benefit many computer vision fields, such as autonomous driving, robot navigation, and augmented reality (AR). In this survey, we first introduce specific…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Meng Xu , Youchen Wang , Bin Xu , Jun Zhang , Jian Ren , Stefan Poslad , Pengfei Xu

Category-level 6D object pose estimation aims to estimate the rotation, translation and size of unseen instances within specific categories. In this area, dense correspondence-based methods have achieved leading performance. However, they…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Xiao Lin , Wenfei Yang , Yuan Gao , Tianzhu Zhang

Hand pose represents key information for action recognition in the egocentric perspective, where the user is interacting with objects. We propose to improve egocentric 3D hand pose estimation based on RGB frames only by using pseudo-depth…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Wiktor Mucha , Michael Wray , Martin Kampel

In multimodal perception systems, achieving precise extrinsic calibration between LiDAR and camera is of critical importance. Previous calibration methods often required specific targets or manual adjustments, making them both…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Xingchen Li , Yifan Duan , Beibei Wang , Haojie Ren , Guoliang You , Yu Sheng , Jianmin Ji , Yanyong Zhang

Human gaze is essential for various appealing applications. Aiming at more accurate gaze estimation, a series of recent works propose to utilize face and eye images simultaneously. Nevertheless, face and eye images only serve as independent…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Yihua Cheng , Shiyao Huang , Fei Wang , Chen Qian , Feng Lu

Understanding the camera wearer's activity is central to egocentric vision, yet one key facet of that activity is inherently invisible to the camera--the wearer's body pose. Prior work focuses on estimating the pose of hands and arms when…

Computer Vision and Pattern Recognition · Computer Science 2016-03-28 Hao Jiang , Kristen Grauman