English
Related papers

Related papers: QueST: Persistent Queries as Semantic Monitors for…

200 papers

Accurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Rulin Zhou , Wenlong He , An Wang , Jianhang Zhang , Xuanhui Zeng , Xi Zhang , Chaowei Zhu , Haijun Hu , Hongliang Ren

In modern multimedia systems, efficient video processing is critical, especially in resource-constrained environments such as IoT-based camera networks, autonomous platforms, and wireless sensor multimedia systems. A key bottleneck in video…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Kakia Panagidi , Stathes Hadjieftymiadis

Data drift is the change in model input data that is one of the key factors leading to machine learning models performance degradation over time. Monitoring drift helps detecting these issues and preventing their harmful consequences.…

Computation and Language · Computer Science 2023-05-30 Ella Rabinovich , Matan Vetzler , Samuel Ackerman , Ateret Anaby-Tavor

Dense self-supervised learning has shown great promise for learning pixel- and patch-level representations, but extending it to videos remains challenging due to the complexity of motion dynamics. Existing approaches struggle as they rely…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Mohammadreza Salehi , Shashanka Venkataramanan , Ioana Simion , Efstratios Gavves , Cees G. M. Snoek , Yuki M Asano

Most scanning LiDAR sensors generate a sequence of point clouds in real-time. While conventional 3D object detectors use a set of unordered LiDAR points acquired over a fixed time interval, recent studies have revealed that substantial…

Computer Vision and Pattern Recognition · Computer Science 2022-12-22 Junho Koh , Junhyung Lee , Youngwoo Lee , Jaekyum Kim , Jun Won Choi

Structure-from-motion (SfM) largely relies on feature tracking. In image sequences, if disjointed tracks caused by objects moving in and out of the field of view, occasional occlusion, or image noise, are not handled well, corresponding SfM…

Computer Vision and Pattern Recognition · Computer Science 2016-10-13 Guofeng Zhang , Haomin Liu , Zilong Dong , Jiaya Jia , Tien-Tsin Wong , Hujun Bao

Streaming video recognition reasons about objects and their actions in every frame of a video. A good streaming recognition model captures both long-term dynamics and short-term changes of video. Unfortunately, in most existing methods, the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-20 Yue Zhao , Philipp Krähenbühl

Multi-Object Tracking (MOT) is evolving from geometric localization to Semantic MOT (SMOT) to answer complex relational queries, yet progress is hindered by semantic data scarcity and a structural disconnect between tracking architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Pan Liao , Feng Yang , Di Wu , Jinwen Yu , Yuhua Zhu , Wenhui Zhao , Dingwen Zhang

Video segmentation is a popular task, but applying image segmentation models frame-by-frame to videos does not preserve temporal consistency. In this paper, we propose a method to extend a query-based image segmentation model to video using…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Tsubasa Mizuno , Toru Tamaki

Online Multiple Target Tracking (MTT) is often addressed within the tracking-by-detection paradigm. Detections are previously extracted independently in each frame and then objects trajectories are built by maximizing specifically designed…

Computer Vision and Pattern Recognition · Computer Science 2015-09-15 Francesco Solera , Simone Calderara , Rita Cucchiara

Identifying mobility behaviors in rich trajectory data is of great economic and social interest to various applications including urban planning, marketing and intelligence. Existing work on trajectory clustering often relies on similarity…

Machine Learning · Computer Science 2020-03-04 Mingxuan Yue , Yaguang Li , Haoze Yang , Ritesh Ahuja , Yao-Yi Chiang , Cyrus Shahabi

Temporal sentence grounding aims to localize a target segment in an untrimmed video semantically according to a given sentence query. Most previous works focus on learning frame-level features of each whole frame in the entire video, and…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Daizong Liu , Xiang Fang , Wei Hu , Pan Zhou

Traditional video captioning requests a holistic description of the video, yet the detailed descriptions of the specific objects may not be available. Without associating the moving trajectories, these image-based data-driven methods cannot…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Fangyi Zhu , Jenq-Neng Hwang , Zhanyu Ma , Guang Chen , Jun Guo

In this paper, we propose a long-sequence modeling framework, named StreamPETR, for multi-view 3D object detection. Built upon the sparse query design in the PETR series, we systematically develop an object-centric temporal mechanism. The…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Shihao Wang , Yingfei Liu , Tiancai Wang , Ying Li , Xiangyu Zhang

This report introduces an improved method for the Tracking Any Point~(TAP), focusing on monitoring physical surfaces in video footage. Despite their success with short-sequence scenarios, TAP methods still face performance degradation and…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Yuxuan Zhang , Pengsong Niu , Kun Yu , Qingguo Chen , Yang Yang

Most existing real-time deep models trained with each frame independently may produce inconsistent results across the temporal axis when tested on a video sequence. A few methods take the correlations in the video sequence into…

Computer Vision and Pattern Recognition · Computer Science 2022-02-28 Yifan Liu , Chunhua Shen , Changqian Yu , Jingdong Wang

Modern video segmentation methods adopt object queries to perform inter-frame association and demonstrate satisfactory performance in tracking continuously appearing objects despite large-scale motion and transient occlusion. However, they…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Yikang Zhou , Tao Zhang , Shunping Ji , Shuicheng Yan , Xiangtai Li

Tracking dense 3D motion from monocular videos remains challenging, particularly when aiming for pixel-level precision over long sequences. We introduce DELTA, a novel method that efficiently tracks every pixel in 3D space, enabling…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Tuan Duc Ngo , Peiye Zhuang , Chuang Gan , Evangelos Kalogerakis , Sergey Tulyakov , Hsin-Ying Lee , Chaoyang Wang

Video instance segmentation aims at predicting object segmentation masks for each frame, as well as associating the instances across multiple frames. Recent end-to-end video instance segmentation methods are capable of performing object…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Quanzeng You , Jiang Wang , Peng Chu , Andre Abrantes , Zicheng Liu

Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Qing Liu , Vignesh Ramanathan , Dhruv Mahajan , Alan Yuille , Zhenheng Yang