English
Related papers

Related papers: Hopper: Multi-hop Transformer for Spatiotemporal R…

200 papers

Video object segmentation targets at segmenting a specific object throughout a video sequence, given only an annotated first frame. Recent deep learning based approaches find it effective by fine-tuning a general-purpose segmentation model…

Computer Vision and Pattern Recognition · Computer Science 2018-02-06 Linjie Yang , Yanran Wang , Xuehan Xiong , Jianchao Yang , Aggelos K. Katsaggelos

Spatio-temporal scene-graph approaches to video-based reasoning tasks, such as video question-answering (QA), typically construct such graphs for every video frame. These approaches often ignore the fact that videos are essentially…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Anoop Cherian , Chiori Hori , Tim K. Marks , Jonathan Le Roux

Modern video editing techniques have achieved high visual fidelity when inserting video objects. However, they focus on optimizing visual fidelity rather than physical causality, leading to edits that are physically inconsistent with their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Bohai Gu , Taiyi Wu , Dazhao Du , Jian Liu , Shuai Yang , Xiaotong Zhao , Alan Zhao , Song Guo

We propose a novel method for spatiotemporal multi-camera calibration using freely moving people in multiview videos. Since calibrating multiple cameras and finding matches across their views are inherently interdependent, performing both…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Sang-Eun Lee , Ko Nishino , Shohei Nobuhara

Predicting 3D human pose from a single monoscopic video can be highly challenging due to factors such as low resolution, motion blur and occlusion, in addition to the fundamental ambiguity in estimating 3D from 2D. Approaches that directly…

Computer Vision and Pattern Recognition · Computer Science 2021-04-26 Tao Jiang , Necati Cihan Camgoz , Richard Bowden

We address multimodal deepfake detection requiring both robustness and interpretability by proposing FakeHunter, a unified framework that combines memory guided retrieval, a structured Observation-Thought-Action reasoning loop, and adaptive…

Multimedia · Computer Science 2025-09-11 Chen Chen , Runze Li , Zejun Zhang , Pukun Zhao , Fanqing Zhou , Longxiang Wang , Haojian Huang

Transformer architectures have become the model of choice in natural language processing and are now being introduced into computer vision tasks such as image classification, object detection, and semantic segmentation. However, in the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Ce Zheng , Sijie Zhu , Matias Mendieta , Taojiannan Yang , Chen Chen , Zhengming Ding

In this report, we introduce a video hashing method for scalable video segment copy detection. The objective of video segment copy detection is to find the video (s) present in a large database, one of whose segments (cropped in time) is a…

Machine Learning · Computer Science 2019-11-22 Arjun Krishna , A S Akil Arif Ibrahim

Most existing transformer based video instance segmentation methods extract per frame features independently, hence it is challenging to solve the appearance deformation problem. In this paper, we observe the temporal information is…

Computer Vision and Pattern Recognition · Computer Science 2023-01-24 Zhenghao Zhang , Fangtao Shao , Zuozhuo Dai , Siyu Zhu

While human infants exhibit knowledge about object permanence from two months of age onwards, deep-learning approaches still largely fail to recognize objects' continued existence. We introduce a slot-based autoregressive deep learning…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Manuel Traub , Frederic Becker , Sebastian Otte , Martin V. Butz

In many computer vision tasks, the relevant information to solve the problem at hand is mixed to irrelevant, distracting information. This has motivated researchers to design attentional models that can dynamically focus on parts of images…

Computer Vision and Pattern Recognition · Computer Science 2017-02-14 Loris Bazzani , Hugo Larochelle , Lorenzo Torresani

Despite significant progress in video question answering (VideoQA), existing methods fall short of questions that require causal/temporal reasoning across frames. This can be attributed to imprecise motion representations. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Junwen Chen , Jie Zhu , Yu Kong

We present a large-scale study on unsupervised spatiotemporal representation learning from videos. With a unified perspective on four recent image-based frameworks, we study a simple objective that can easily generalize all these methods to…

Computer Vision and Pattern Recognition · Computer Science 2021-04-30 Christoph Feichtenhofer , Haoqi Fan , Bo Xiong , Ross Girshick , Kaiming He

Humans excel in grasping and manipulating objects because of their life-long experience and knowledge about the 3D shape and weight distribution of objects. However, the lack of such intuition in robots makes robotic grasping an…

Computer Vision and Pattern Recognition · Computer Science 2018-11-05 Ghazal Ghazaei , Iro Laina , Christian Rupprecht , Federico Tombari , Nassir Navab , Kianoush Nazarpour

Reading Comprehension has received significant attention in recent years as high quality Question Answering (QA) datasets have become available. Despite state-of-the-art methods achieving strong overall accuracy, Multi-Hop (MH) reasoning…

Computation and Language · Computer Science 2019-05-24 Alex Long , Joel Mason , Alan Blair , Wei Wang

With the increasing demand for video understanding, video moment and highlight detection (MHD) has emerged as a critical research topic. MHD aims to localize all moments and predict clip-wise saliency scores simultaneously. Despite progress…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Yifang Xu , Yunzhuo Sun , Yang Li , Yilei Shi , Xiaoxiang Zhu , Sidan Du

Accurately distinguishing each object is a fundamental goal of Multi-object tracking (MOT) algorithms. However, achieving this goal still remains challenging, primarily due to: (i) For crowded scenes with occluded objects, the high overlap…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Jiapeng Wu , Yichen Liu

In this contribution, a novel spatio-temporal prediction algorithm for video coding is introduced. This algorithm exploits temporal as well as spatial redundancies for effectively predicting the signal to be encoded. To achieve this, the…

Image and Video Processing · Electrical Eng. & Systems 2022-07-05 Jürgen Seiler , André Kaup

Although video summarization has achieved tremendous success benefiting from Recurrent Neural Networks (RNN), RNN-based methods neglect the global dependencies and multi-hop relationships among video frames, which limits the performance.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-23 Bin Zhao , Maoguo Gong , Xuelong Li

Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoning-the ability to comprehend object locations, orientations,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Haoran Tang , Meng Cao , Ruyang Liu , Xiaoxi Liang , Linglong Li , Ge Li , Xiaodan Liang