English
Related papers

Related papers: Beyond Play and Pause: Turning GPT-4o Spatial Weak…

200 papers

Extracting informative representations from videos is fundamental for effectively learning various downstream tasks. We present a novel approach for unsupervised learning of meaningful representations from videos, leveraging the concept of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Ali Younes , Simone Schaub-Meyer , Georgia Chalvatzaki

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Jun Zheng , Zhengze Xu , Mengting Chen , Jing Wang , Jinsong Lan , Xiaoyong Zhu , Kaifu Zhang , Bo Zheng , Xiaodan Liang

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

Human-Computer Interaction · Computer Science 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture,…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Dongho Lee , Jongseo Lee , Jinwoo Choi

This paper presents UniVST, a unified framework for localized video style transfer based on diffusion models. It operates without the need for training, offering a distinct advantage over existing diffusion methods that transfer style…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Quanjian Song , Mingbao Lin , Wengyi Zhan , Shuicheng Yan , Liujuan Cao , Rongrong Ji

In this paper, we introduce Attention Prompt Tuning (APT) - a computationally efficient variant of prompt tuning for video-based applications such as action recognition. Prompt tuning approaches involve injecting a set of learnable prompts…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Wele Gedara Chaminda Bandara , Vishal M. Patel

Recently unsupervised learning of depth from videos has made remarkable progress and the results are comparable to fully supervised methods in outdoor scenes like KITTI. However, there still exist great challenges when directly applying…

Computer Vision and Pattern Recognition · Computer Science 2019-10-22 Junsheng Zhou , Yuwang Wang , Kaihuai Qin , Wenjun Zeng

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Jongseo Lee , Joohyun Chang , Dongho Lee , Jinwoo Choi

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Zhanyu Wang , Longyue Wang , Zhen Zhao , Minghao Wu , Chenyang Lyu , Huayang Li , Deng Cai , Luping Zhou , Shuming Shi , Zhaopeng Tu

Previous deep learning-based video stabilizers require a large scale of paired unstable and stable videos for training, which are difficult to collect. Traditional trajectory-based stabilizers, on the other hand, divide the task into…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Yufei Xu , Jing Zhang , Stephen J. Maybank , Dacheng Tao

Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating…

Graphics · Computer Science 2025-10-07 Nilay Kumar , Priyansh Bhandari , G. Maragatham

The exponential growth of AI education has brought millions of learners to online platforms, yet this massive scale has simultaneously exposed critical pedagogical shortcomings. Traditional video-based instruction, while cost-effective and…

Human-Computer Interaction · Computer Science 2026-04-20 Mohammed Abraar , Raj Abhijit Dandekar , Rajat Dandekar , Sreedath Panat

Unsupervised multi-object scene decomposition is a fast-emerging problem in representation learning. Despite significant progress in static scenes, such models are unable to leverage important dynamic cues present in video. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Polina Zablotskaia , Edoardo A. Dominici , Leonid Sigal , Andreas M. Lehrmann

Traditional computer vision models often necessitate extensive data acquisition, annotation, and validation. These models frequently struggle in real-world applications, resulting in high false positive and negative rates, and exhibit poor…

Among various region embedding methods, graph-based region relation learning models stand out, owing to their strong structure representation ability for encoding spatial correlations with graph neural networks. Despite their effectiveness,…

Machine Learning · Computer Science 2023-05-09 Qianru Zhang , Chao Huang , Lianghao Xia , Zheng Wang , Zhonghang Li , Siuming Yiu

Miscommunication and communication challenges between instructors and students represents one of the primary barriers to post-secondary learning. Students often avoid or miss opportunities to ask questions during office hours due to…

Computers and Society · Computer Science 2024-02-02 Ramteja Sajja , Yusuf Sermet , David Cwiertny , Ibrahim Demir

Static appearance of video may impede the ability of a deep neural network to learn motion-relevant features in video action recognition. In this paper, we introduce a new concept, Dynamic Appearance (DA), summarizing the appearance…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Guoxi Huang , Adrian G. Bors

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation,…

Computation and Language · Computer Science 2023-05-29 Yixuan Su , Tian Lan , Huayang Li , Jialu Xu , Yan Wang , Deng Cai

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

Physics Education · Physics 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

Humans can collaborate and complete tasks based on visual signals and instruction from the environment. Training such a robot is difficult especially due to the understanding of the instruction and the complicated environment. Previous…

Artificial Intelligence · Computer Science 2023-05-12 Kairui Zhou