English
Related papers

Related papers: Temporal vs. Spatial: Comparing DINOv3 and V-JEPA2…

200 papers

Learning-based visual odometry (VO) algorithms achieve remarkable performance on common static scenes, benefiting from high-capacity models and massive annotated data, but tend to fail in dynamic, populated environments. Semantic…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Shihao Shen , Yilin Cai , Wenshan Wang , Sebastian Scherer

This paper proposes a novel pretext task to address the self-supervised video representation learning problem. Specifically, given an unlabeled video clip, we compute a series of spatio-temporal statistical summaries, such as the spatial…

Computer Vision and Pattern Recognition · Computer Science 2021-02-01 Jiangliu Wang , Jianbo Jiao , Linchao Bao , Shengfeng He , Wei Liu , Yun-hui Liu

Video analysis is a major computer vision task that has received a lot of attention in recent years. The current state-of-the-art performance for video analysis is achieved with Deep Neural Networks (DNNs) that have high computational costs…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Mireille El-Assal , Pierre Tirilly , Ioan Marius Bilasco

Self-attention learns pairwise interactions to model long-range dependencies, yielding great improvements for video action recognition. In this paper, we seek a deeper understanding of self-attention for temporal modeling in videos. We…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Bo He , Xitong Yang , Zuxuan Wu , Hao Chen , Ser-Nam Lim , Abhinav Shrivastava

Understanding the content of videos is one of the core techniques for developing various helpful applications in the real world, such as recognizing various human actions for surveillance systems or customer behavior analysis in an…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Chiwan Song , Woobin Im , Sung-eui Yoon

The present work proposes a Deep Learning architecture for the prediction of various consumer choice behaviors from time series of raw gaze or eye fixations on images of the decision environment, for which currently no foundational models…

Neural and Evolutionary Computing · Computer Science 2025-08-07 Moshe Unger , Alexander Tuzhilin , Michel Wedel

In the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes serve as the critical bridge connecting low-level geometric sensing with high-level semantic understanding. We present DINO\_4D, introducing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yiru Yang , Zhuojie Wu , Quentin Marguet , Nishant Kumar Singh , Max Schulthess

We propose a Variational Time Series Feature Extractor (VTSFE), inspired by the VAE-DMP model of Chen et al., to be used for action recognition and prediction. Our method is based on variational autoencoders. It improves VAE-DMP in that it…

Machine Learning · Computer Science 2018-09-27 Maxime Chaveroche , Adrien Malaisé , Francis Colas , François Charpillet , Serena Ivaldi

Since first proposed, Video Instance Segmentation(VIS) task has attracted vast researchers' focus on architecture modeling to boost performance. Though great advances achieved in online and offline paradigms, there are still insufficient…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Wenhe Jia , Lu Yang , Zilong Jia , Wenyi Zhao , Yilin Zhou , Qing Song

Motivated by the success of data-driven convolutional neural networks (CNNs) in object recognition on static images, researchers are working hard towards developing CNN equivalents for learning video features. However, learning video…

Computer Vision and Pattern Recognition · Computer Science 2015-05-19 Zhenzhong Lan , Dezhong Yao , Ming Lin , Shoou-I Yu , Alexander Hauptmann

The ability to detect similar actions across videos can be very useful for real-world applications in many fields. However, this task is still challenging for existing systems, since videos that present the same action, can be taken from…

Computer Vision and Pattern Recognition · Computer Science 2016-12-16 Michal Yarom , Michal Irani

Video understanding tasks have traditionally been modeled by two separate architectures, specially tailored for two distinct tasks. Sequence-based video tasks, such as action recognition, use a video backbone to directly extract…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Yucheng Zhao , Chong Luo , Chuanxin Tang , Dongdong Chen , Noel Codella , Zheng-Jun Zha

Recently, with the availability of cost-effective depth cameras coupled with real-time skeleton estimation, the interest in skeleton-based human action recognition is renewed. Most of the existing skeletal representation approaches use…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Zhize Wu , Thomas Weise , Le Zou , Fei Sun , Ming Tan

Masked autoencoders (MAEs) have emerged recently as art self-supervised spatiotemporal representation learners. Inheriting from the image counterparts, however, existing video MAEs still focus largely on static appearance learning whilst…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Haosen Yang , Deng Huang , Bin Wen , Jiannan Wu , Hongxun Yao , Yi Jiang , Xiatian Zhu , Zehuan Yuan

Temporal modelling is the key for efficient video action recognition. While understanding temporal information can improve recognition accuracy for dynamic actions, removing temporal redundancy and reusing past features can significantly…

Computer Vision and Pattern Recognition · Computer Science 2021-02-12 Yue Meng , Rameswar Panda , Chung-Ching Lin , Prasanna Sattigeri , Leonid Karlinsky , Kate Saenko , Aude Oliva , Rogerio Feris

This paper provides a review on representation learning for videos. We classify recent spatiotemporal feature learning methods for sequential visual data and compare their pros and cons for general video analysis. Building effective…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Elham Ravanbakhsh , Yongqing Liang , J. Ramanujam , Xin Li

Dynamic texture and scene classification are two fundamental problems in understanding natural video content. Extracting robust and effective features is a crucial step towards solving these problems. However the existing approaches suffer…

Computer Vision and Pattern Recognition · Computer Science 2015-02-03 Xianbiao Qi , Chun-Guang Li , Guoying Zhao , Xiaopeng Hong , Matti Pietikäinen

Human action Recognition for unknown views is a challenging task. We propose a view-invariant deep human action recognition framework, which is a novel integration of two important action cues: motion and shape temporal dynamics (STD). The…

Computer Vision and Pattern Recognition · Computer Science 2020-01-22 Chhavi Dhiman , Dinesh Kumar Vishwakarma

Self-supervised video representation methods typically focus on the representation of temporal attributes in videos. However, the role of stationary versus non-stationary attributes is less explored: Stationary features, which remain…

Computer Vision and Pattern Recognition · Computer Science 2021-09-27 Nadine Behrmann , Mohsen Fayyaz , Juergen Gall , Mehdi Noroozi

We introduce LLaVA-OneVision-2 (LLaVA-OV-2), the most capable vision-language model in the LLaVA-OneVision series to date, achieving superior performance across a broad range of multimodal benchmarks. The model builds on a native…

‹ Prev 1 8 9 10 Next ›