English
Related papers

Related papers: WorldCam: Interactive Autoregressive 3D Gaming Wor…

200 papers

Most model-free visual object tracking methods formulate the tracking task as object location estimation given by a 2D segmentation or a bounding box in each video frame. We argue that this representation is limited and instead propose to…

Computer Vision and Pattern Recognition · Computer Science 2023-04-14 Denys Rozumnyi , Jiri Matas , Marc Pollefeys , Vittorio Ferrari , Martin R. Oswald

Prior panorama stitching approaches heavily rely on pairwise feature correspondences and are unable to leverage geometric consistency across multiple views. This leads to severe distortion and misalignment, especially in challenging scenes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Zhengdong Zhu , Weiyi Xue , Zuyuan Yang , Wenlve Zhou , Zhiheng Zhou

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Jiahao Lu , Weitao Xiong , Jiacheng Deng , Peng Li , Tianyu Huang , Zhiyang Dou , Cheng Lin , Sai-Kit Yeung , Yuan Liu

We propose an efficient approach to exploiting motion information from consecutive frames of a video sequence to recover the 3D pose of people. Previous approaches typically compute candidate poses in individual frames and then link them in…

Computer Vision and Pattern Recognition · Computer Science 2016-09-05 Bugra Tekin , Artem Rozantsev , Vincent Lepetit , Pascal Fua

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein

The raise of collaborative robotics has led to wide range of sensor technologies to detect human-machine interactions: at short distances, proximity sensors detect nontactile gestures virtually occlusion-free, while at medium distances,…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Christoph Heindl , Markus Ikeda , Gernot Stübl , Andreas Pichler , Josef Scharinger

Accurate 6-DoF object pose estimation and tracking are critical for reliable robotic manipulation. However, zero-shot methods often fail under viewpoint-induced ambiguities and fixed-camera setups struggle when objects move or become…

Robotics · Computer Science 2026-03-10 Sheng Liu , Zhe Li , Weiheng Wang , Han Sun , Heng Zhang , Hongpeng Chen , Yusen Qin , Arash Ajoudani , Yizhao Wang

It is an exciting task to recover the scene's 3d-structure and camera pose from the video sequence. Most of the current solutions divide it into two parts, monocular depth recovery and camera pose estimation. The monocular depth recovery is…

Computer Vision and Pattern Recognition · Computer Science 2018-05-24 YanTong Wu , Yang Liu

Estimating the 6-DoF pose of a camera from a single image relative to a pre-computed 3D point-set is an important task for many computer vision applications. Perspective-n-Point (PnP) solvers are routinely used for camera pose estimation,…

Computer Vision and Pattern Recognition · Computer Science 2017-09-28 Dylan Campbell , Lars Petersson , Laurent Kneip , Hongdong Li

Most current action recognition methods heavily rely on appearance information by taking an RGB sequence of entire image regions as input. While being effective in exploiting contextual information around humans, e.g., human appearance and…

Computer Vision and Pattern Recognition · Computer Science 2021-04-16 Gyeongsik Moon , Heeseung Kwon , Kyoung Mu Lee , Minsu Cho

Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plausible locations and scales without interpenetration, aligning…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Qingyang Liu , Bingjie Gao , Weiheng Huang , Jun Zhang , Zhongqian Sun , Yang Wei , Fengrui Liu , Zelin Peng , Qianli Ma , Shuai Yang , Zhaohe Liao , Haonan Zhao , Li Niu

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Guangyuan Li , Bo Li , Jinwei Chen , Xiaobin Hu , Lei Zhao , Peng-Tao Jiang

We focus on the task of estimating a physically plausible articulated human motion from monocular video. Existing approaches that do not consider physics often produce temporally inconsistent output with motion artifacts, while…

Computer Vision and Pattern Recognition · Computer Science 2022-05-26 Erik Gärtner , Mykhaylo Andriluka , Hongyi Xu , Cristian Sminchisescu

Current approaches to pose generation rely heavily on intermediate representations, either through two-stage pipelines with quantization or autoregressive models that accumulate errors during inference. This fundamental limitation leads to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Yayuan Li , Filippos Bellos , Jason Corso

This paper proposes AutoScape, a long-horizon driving scene generation framework. At its core is a novel RGB-D diffusion model that iteratively generates sparse, geometrically consistent keyframes, serving as reliable anchors for the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jiacheng Chen , Ziyu Jiang , Mingfu Liang , Bingbing Zhuang , Jong-Chyi Su , Sparsh Garg , Ying Wu , Manmohan Chandraker

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

We consider the problem of segmenting dynamic regions in CrowdCam images, where a dynamic region is the projection of a moving 3D object on the image plane. Quite often, these regions are the most interesting parts of an image. CrowdCam…

Computer Vision and Pattern Recognition · Computer Science 2019-06-25 Nir Zarrabi , Shai Avidan , Yael Moses

Object pose estimation is an integral part of robot vision and AR. Previous 6D pose retrieval pipelines treat the problem either as a regression task or discretize the pose space to classify. We change this paradigm and reformulate the…

Computer Vision and Pattern Recognition · Computer Science 2020-12-02 Benjamin Busam , Hyun Jun Jung , Nassir Navab

Recent approaches have successfully focused on the segmentation of static reconstructions, thereby equipping downstream applications with semantic 3D understanding. However, the world in which we live is dynamic, characterized by numerous…

Robotics · Computer Science 2025-03-12 Tjark Behrens , René Zurbrügg , Marc Pollefeys , Zuria Bauer , Hermann Blum

Marker-less 3D human motion capture from a single colour camera has seen significant progress. However, it is a very challenging and severely ill-posed problem. In consequence, even the most accurate state-of-the-art approaches have…

Computer Vision and Pattern Recognition · Computer Science 2020-12-10 Soshi Shimada , Vladislav Golyanik , Weipeng Xu , Christian Theobalt