English
Related papers

Related papers: POMATO: Marrying Pointmap Matching with Temporal M…

200 papers

We propose a method for 3D object reconstruction and 6D-pose estimation from 2D images that uses knowledge about object shape as the primary key. In the proposed pipeline, recognition and labeling of objects in 2D images deliver 2D segment…

Computer Vision and Pattern Recognition · Computer Science 2022-03-03 Marcell Wolnitza , Osman Kaya , Tomas Kulvicius , Florentin Wörgötter , Babette Dellen

Directly learning multiple 3D objects motion from sequential images is difficult, while the geometric bundle adjustment lacks the ability to localize the invisible object centroid. To benefit from both the powerful object understanding…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Peiliang Li , Jieqi Shi , Shaojie Shen

We propose a lifelong 3D mapping framework that is modular, cloud-native by design and more importantly, works for both hand-held and robot-mounted 3D LiDAR mapping systems. Our proposed framework comprises of dynamic point removal,…

Robotics · Computer Science 2025-01-31 Liudi Yang , Sai Manoj Prakhya , Senhua Zhu , Ziyuan Liu

In this paper we introduce Co-Fusion, a dense SLAM system that takes a live stream of RGB-D images as input and segments the scene into different objects (using either motion or semantic cues) while simultaneously tracking and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-06 Martin Rünz , Lourdes Agapito

3D reconstruction is a fundamental issue in many applications and the feature point matching problem is a key step while reconstructing target objects. Conventional algorithms can only find a small number of feature points from two images…

Computer Vision and Pattern Recognition · Computer Science 2019-11-26 Zhihao Fang , He Ma , Xuemin Zhu , Xutao Guo , Ruixin Zhou

Make-up temporal video grounding (MTVG) aims to localize the target video segment which is semantically related to a sentence describing a make-up activity, given a long video. Compared with the general video grounding task, MTVG focuses on…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Jiaxiu Li , Kun Li , Jia Li , Guoliang Chen , Dan Guo , Meng Wang

We propose a real-time 3D human pose estimation and motion analysis method termed RePose for rehabilitation training. It is capable of real-time monitoring and evaluation of patients'motion during rehabilitation, providing immediate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-05 Junxiao Xue , Pavel Smirnov , Ziao Li , Yunyun Shi , Shi Chen , Xinyi Yin , Xiaohan Yue , Lei Wang , Yiduo Wang , Feng Lin , Yijia Chen , Xiao Ma , Xiaoran Yan , Qing Zhang , Fengjian Xue , Xuecheng Wu

Matching local geometric features on real-world depth images is a challenging task due to the noisy, low-resolution, and incomplete nature of 3D scan data. These difficulties limit the performance of current state-of-art methods, which are…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Andy Zeng , Shuran Song , Matthias Nießner , Matthew Fisher , Jianxiong Xiao , Thomas Funkhouser

Convolutional neural networks have enabled accurate image super-resolution in real-time. However, recent attempts to benefit from temporal correlations in video super-resolution have been limited to naive or inefficient architectures. In…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Jose Caballero , Christian Ledig , Andrew Aitken , Alejandro Acosta , Johannes Totz , Zehan Wang , Wenzhe Shi

Motion prediction is a classic problem in computer vision, which aims at forecasting future motion given the observed pose sequence. Various deep learning models have been proposed, achieving state-of-the-art performance on motion…

Computer Vision and Pattern Recognition · Computer Science 2022-01-10 Pengxiang Su , Zhenguang Liu , Shuang Wu , Lei Zhu , Yifang Yin , Xuanjing Shen

Reconstructing dense, volumetric models of real-world 3D scenes is important for many tasks, but capturing large scenes can take significant time, and the risk of transient changes to the scene goes up as the capture time increases. These…

Computer Vision and Pattern Recognition · Computer Science 2019-07-03 Stuart Golodetz , Tommaso Cavallari , Nicholas A Lord , Victor A Prisacariu , David W Murray , Philip H S Torr

We present a new combined approach for monocular model-based 3D tracking. A preliminary object pose is estimated by using a keypoint-based technique. The pose is then refined by optimizing the contour energy function. The energy determines…

Computer Vision and Pattern Recognition · Computer Science 2020-02-05 Bogdan Bugaev , Anton Kryshchenko , Roman Belov

This work proposes an end-to-end multi-camera 3D multi-object tracking (MOT) framework. It emphasizes spatio-temporal continuity and integrates both past and future reasoning for tracked objects. Thus, we name it "Past-and-Future reasoning…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Ziqi Pang , Jie Li , Pavel Tokmakov , Dian Chen , Sergey Zagoruyko , Yu-Xiong Wang

Accurate and temporally consistent modeling of human bodies is essential for a wide range of applications, including character animation, understanding human social behavior and AR/VR interfaces. Capturing human motion accurately from a…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Alexandra Zimmer , Anna Hilsmann , Wieland Morgenstern , Peter Eisert

Stereo matching provides depth estimation from binocular images for downstream applications. These applications mostly take video streams as input and require temporally consistent depth maps. However, existing methods mainly focus on the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Jiaxi Zeng , Chengtang Yao , Yuwei Wu , Yunde Jia

Due to the intrinsic complexity of high-dimensional (HD) data, dimensionality reduction (DR) techniques cannot preserve all the structural characteristics of the original data. Therefore, DR techniques focus on preserving either local…

Machine Learning · Computer Science 2025-11-18 Hyeon Jeon , Kwon Ko , Soohyun Lee , Jake Hyun , Taehyun Yang , Gyehun Go , Jaemin Jo , Jinwook Seo

Traditional novel view synthesis methods heavily rely on external camera pose estimation tools such as COLMAP, which often introduce computational bottlenecks and propagate errors. To address these challenges, we propose a unified framework…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Xianben Yang , Yuxuan Li , Tao Wang , Tao Wang , Yi Jin , Yidong Li , Haibin Ling

Recent advancements in multimodal large language models (MLLMs) have opened new avenues for video understanding. However, achieving high fidelity in zero-shot video tasks remains challenging. Traditional video processing methods rely…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yiming Zhang , Zhuokai Zhao , Zhaorun Chen , Zenghui Ding , Xianjun Yang , Yining Sun

Task and motion planning are long-standing challenges in robotics, especially when robots have to deal with dynamic environments exhibiting long-term dynamics, such as households or warehouses. In these environments, long-term dynamics…

Robotics · Computer Science 2025-09-23 Francesco Argenziano , Miguel Saavedra-Ruiz , Sacha Morin , Daniele Nardi , Liam Paull

We present a pose adaptive few-shot learning procedure and a two-stage data interpolation regularization, termed Pose Adaptive Dual Mixup (PADMix), for single-image 3D reconstruction. While augmentations via interpolating feature-label…

Computer Vision and Pattern Recognition · Computer Science 2021-12-24 Ta-Ying Cheng , Hsuan-Ru Yang , Niki Trigoni , Hwann-Tzong Chen , Tyng-Luh Liu