English
Related papers

Related papers: TAPVid-3D: A Benchmark for Tracking Any Point in 3…

200 papers

We tackle the problem of Persistent Independent Particles (PIPs), also called Tracking Any Point (TAP), in videos, which specifically aims at estimating persistent long-term trajectories of query points in videos. Previous methods attempted…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Weikang Bian , Zhaoyang Huang , Xiaoyu Shi , Yitong Dong , Yijin Li , Hongsheng Li

Animal Pose Estimation and Tracking (APT) is a critical task in detecting and monitoring the keypoints of animals across a series of video frames, which is essential for understanding animal behavior. Past works relating to animals have…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Yuxiang Yang , Yingqi Deng , Yufei Xu , Jing Zhang

Planar object tracking is an actively studied problem in vision-based robotic applications. While several benchmarks have been constructed for evaluating state-of-the-art algorithms, there is a lack of video sequences captured in the wild…

Computer Vision and Pattern Recognition · Computer Science 2018-05-23 Pengpeng Liang , Yifan Wu , Hu Lu , Liming Wang , Chunyuan Liao , Haibin Ling

We present a novel video generation framework that integrates 3-dimensional geometry and dynamic awareness. To achieve this, we augment 2D videos with 3D point trajectories and align them in pixel space. The resulting 3D-aware video…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Yunuo Chen , Junli Cao , Vidit Goel , Sergei Korolev , Chenfanfu Jiang , Jian Ren , Sergey Tulyakov , Anil Kag

Video Anomaly Detection (VAD), which aims to detect anomalies that deviate from expectation, has attracted increasing attention in recent years. Existing advancements in VAD primarily focus on model architectures and training strategies,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Zihao Liu , Xiaoyu Wu , Wenna Li , Linlin Yang , Shengjin Wang

Point tracking models often struggle to generalize to real-world videos because large-scale training data is predominantly synthetic$\unicode{x2014}$the only source currently feasible to produce at scale. Collecting real-world annotations,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Inès Hyeonsu Kim , Seokju Cho , Jahyeok Koo , Junghyun Park , Jiahui Huang , Honglak Lee , Joon-Young Lee , Seungryong Kim

3D single object tracking (SOT) is an important and challenging task for the autonomous driving and mobile robotics. Most existing methods perform tracking between two consecutive frames while ignoring the motion patterns of the target over…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Yu Lin , Zhiheng Li , Yubo Cui , Zheng Fang

Monocular 3D object tracking aims to estimate temporally consistent 3D object poses across video frames, enabling autonomous agents to reason about scene dynamics. However, existing state-of-the-art approaches are fully supervised and rely…

Robotics · Computer Science 2026-03-20 Nikhil Gosala , B. Ravi Kiran , Senthil Yogamani , Abhinav Valada

Spatio-temporal action recognition has been a challenging task that involves detecting where and when actions occur. Current state-of-the-art action detectors are mostly anchor-based, requiring sensitive anchor designs and huge computations…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Shentong Mo , Jingfei Xia , Xiaoqing Tan , Bhiksha Raj

We introduce ITTO, a challenging new benchmark suite for evaluating and diagnosing the capabilities and limitations of point tracking methods. Our videos are sourced from existing datasets and egocentric real-world recordings, with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Ilona Demler , Saumya Chauhan , Georgia Gkioxari

Recent studies on motion estimation have advocated an optimized motion representation that is globally consistent across the entire video, preferably for every pixel. This is challenging as a uniform representation may not account for the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Rui Li , Dong Liu

Tracking Any Point (TAP) has emerged as a fundamental tool for video understanding. Current approaches adapt Vision Foundation Models (VFMs) like DINOv2 via offline finetuning or test-time optimization. However, these VFMs rely on static…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Qiangqiang Wu , Tianyu Yang , Bo Fang , Jia Wan , Matias Di Martino , Guillermo Sapiro , Antoni B. Chan

Monocular 3D tracking aims to capture the long-term motion of pixels in 3D space from a single monocular video and has witnessed rapid progress in recent years. However, we argue that the existing monocular 3D tracking methods still fall…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Jiahao Lu , Weitao Xiong , Jiacheng Deng , Peng Li , Tianyu Huang , Zhiyang Dou , Cheng Lin , Sai-Kit Yeung , Yuan Liu

Traditional temporal action detection (TAD) usually handles untrimmed videos with small number of action instances from a single label (e.g., ActivityNet, THUMOS). However, this setting might be unrealistic as different classes of actions…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Jing Tan , Xiaotong Zhao , Xintian Shi , Bin Kang , Limin Wang

We propose a new long video dataset (called Track Long and Prosper - TLP) and benchmark for single object tracking. The dataset consists of 50 HD videos from real world scenarios, encompassing a duration of over 400 minutes (676K frames),…

Computer Vision and Pattern Recognition · Computer Science 2019-01-03 Abhinav Moudgil , Vineet Gandhi

We hypothesize that an agent that can look around in static scenes can learn rich visual representations applicable to 3D object tracking in complex dynamic scenes. We are motivated in this pursuit by the fact that the physical world itself…

Computer Vision and Pattern Recognition · Computer Science 2020-08-05 Adam W. Harley , Shrinidhi K. Lakshmikanth , Paul Schydlo , Katerina Fragkiadaki

In this paper, built upon TAPTRv2, we present TAPTRv3. TAPTRv2 is a simple yet effective DETR-like point tracking framework that works fine in regular videos but tends to fail in long videos. TAPTRv3 improves TAPTRv2 by addressing its…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Jinyuan Qu , Hongyang Li , Shilong Liu , Tianhe Ren , Zhaoyang Zeng , Lei Zhang

In recent years, vision-centric perception has flourished in various autonomous driving tasks, including 3D detection, semantic map construction, motion forecasting, and depth estimation. Nevertheless, the latency of vision-centric…

Computer Vision and Pattern Recognition · Computer Science 2022-12-20 Xiaofeng Wang , Zheng Zhu , Yunpeng Zhang , Guan Huang , Yun Ye , Wenbo Xu , Ziwei Chen , Xingang Wang

3D multi-object tracking (MOT) is essential to applications such as autonomous driving. Recent work focuses on developing accurate systems giving less attention to computational cost and system complexity. In contrast, this work proposes a…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Xinshuo Weng , Jianren Wang , David Held , Kris Kitani

Quantifying motion in 3D is important for studying the behavior of humans and other animals, but manual pose annotations are expensive and time-consuming to obtain. Self-supervised keypoint discovery is a promising strategy for estimating…