English
Related papers

Related papers: Orthogonal Temporal Interpolation for Zero-Shot Vi…

200 papers

In this paper, we explore the space-time video super-resolution task, which aims to generate a high-resolution (HR) slow-motion video from a low frame rate (LFR), low-resolution (LR) video. A simple solution is to split it into two…

Computer Vision and Pattern Recognition · Computer Science 2020-02-27 Xiaoyu Xiang , Yapeng Tian , Yulun Zhang , Yun Fu , Jan P. Allebach , Chenliang Xu

Tactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach…

Robotics · Computer Science 2024-09-17 Shiori Ueda , Atsushi Hashimoto , Masashi Hamaya , Kazutoshi Tanaka , Hideo Saito

Few-shot Video Object Detection (FSVOD) addresses the challenge of detecting novel objects in videos with limited labeled examples, overcoming the constraints of traditional detection methods that require extensive training data. This task…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Yogesh Kumar , Anand Mishra

Recent years have witnessed outstanding advances of large vision-language models (LVLMs). In order to tackle video understanding, most of them depend upon their implicit temporal understanding capacity. As such, they have not deciphered…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Thong Nguyen , Zhiyuan Hu , Xu Lin , Cong-Duy Nguyen , See-Kiong Ng , Luu Anh Tuan

Zero-Shot Temporal Action Localization (ZS-TAL) seeks to identify and locate actions in untrimmed videos unseen during training. Existing ZS-TAL methods involve fine-tuning a model on a large amount of annotated training data. While…

Computer Vision and Pattern Recognition · Computer Science 2024-04-12 Benedetta Liberatori , Alessandro Conti , Paolo Rota , Yiming Wang , Elisa Ricci

While recent large-scale video-language pre-training made great progress in video question answering, the design of spatial modeling of video-language models is less fine-grained than that of image-language models; existing practices of…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Hsin-Ying Lee , Hung-Ting Su , Bing-Chen Tsai , Tsung-Han Wu , Jia-Fong Yeh , Winston H. Hsu

This thesis explores the central question of how to leverage temporal relations among video elements to advance video understanding. Addressing the limitations of existing methods, the work presents a five-fold contribution: (1) an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Thong Thanh Nguyen

Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution(HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR by directly combining two…

Computer Vision and Pattern Recognition · Computer Science 2022-05-12 Mengshun Hu , Kui Jiang , Liang Liao , Jing Xiao , Junjun Jiang , Zheng Wang

Audio-visual generalised zero-shot learning for video classification requires understanding the relations between the audio and visual information in order to be able to recognise samples from novel, previously unseen classes at test time.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-21 Otniel-Bogdan Mercea , Thomas Hummel , A. Sophia Koepke , Zeynep Akata

Zero-shot Long Video Moment Retrieval (ZLVMR) is the task of identifying temporal segments in hour-long videos using a natural language query without task-specific training. The core technical challenge of LVMR stems from the computational…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Mingyu Jeon , Jisoo Yang , Sungjin Han , Jinkwon Hwang , Sunjae Yoon , Jonghee Kim , Junyeoung Kim

Few-shot video classification aims to learn new video categories with only a few labeled examples, alleviating the burden of costly annotation in real-world applications. However, it is particularly challenging to learn a class-invariant…

Computer Vision and Pattern Recognition · Computer Science 2021-05-12 Songyang Zhang , Jiale Zhou , Xuming He

Spatio-temporal feature learning is of central importance for action recognition in videos. Existing deep neural network models either learn spatial and temporal features independently (C2D) or jointly with unconstrained parameters (C3D).…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Chao Li , Qiaoyong Zhong , Di Xie , Shiliang Pu

Multimodal Large Language Models (MLLMs) have shown strong performance on Video Temporal Grounding (VTG). However, their coarse recognition capabilities are insufficient for fine-grained temporal understanding, making task-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jiwook Han , Geo Ahn , Youngrae Kim , Jinwoo Choi

Few-shot video object segmentation (FS-VOS) aims at segmenting video frames using a few labelled examples of classes not seen during initial training. In this paper, we present a simple but effective temporal transductive inference (TTI)…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Mennatullah Siam , Konstantinos G. Derpanis , Richard P. Wildes

Vision-language models (VLMs) have demonstrated remarkable performance across various visual tasks, leveraging joint learning of visual and textual representations. While these models excel in zero-shot image tasks, their application to…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Massimo Bosetti , Shibingfeng Zhang , Benedetta Liberatori , Giacomo Zara , Elisa Ricci , Paolo Rota

The number of categories for action recognition is growing rapidly and it has become increasingly hard to label sufficient training data for learning conventional models for all categories. Instead of collecting ever more data and labelling…

Computer Vision and Pattern Recognition · Computer Science 2016-12-05 Xun Xu , Timothy Hospedales , Shaogang Gong

Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions using knowledge learned from seen attribute-object compositions in the training set. Previous works mainly project an image and a composition into a common…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Tian Zhang , Kongming Liang , Ruoyi Du , Xian Sun , Zhanyu Ma , Jun Guo

This paper presents VideoLoom, a unified Video Large Language Model (Video LLM) for joint spatial-temporal understanding. To facilitate the development of fine-grained spatial and temporal localization capabilities, we curate LoomData-8.7k,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiapeng Shi , Junke Wang , Zuyao You , Bo He , Zuxuan Wu

Deep learning models have enjoyed great success for image related computer vision tasks like image classification and object detection. For video related tasks like human action recognition, however, the advancements are not as significant…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Xiaolin Song , Cuiling Lan , Wenjun Zeng , Junliang Xing , Jingyu Yang , Xiaoyan Sun

Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Peter Robicheaux , Matvei Popov , Anish Madan , Isaac Robinson , Joseph Nelson , Deva Ramanan , Neehar Peri