中文
相关论文

相关论文: Beyond Play and Pause: Turning GPT-4o Spatial Weak…

200 篇论文

Extracting informative representations from videos is fundamental for effectively learning various downstream tasks. We present a novel approach for unsupervised learning of meaningful representations from videos, leveraging the concept of…

计算机视觉与模式识别 · 计算机科学 2023-06-07 Ali Younes , Simone Schaub-Meyer , Georgia Chalvatzaki

Video Virtual Try-On (VVT) aims to seamlessly replace a garment on a person in a video with a new one. While existing methods have made significant strides in maintaining temporal consistency, they are predominantly confined to…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Jun Zheng , Zhengze Xu , Mengting Chen , Jing Wang , Jinsong Lan , Xiaoyong Zhu , Kaifu Zhang , Bo Zheng , Xiaodan Liang

Recent advancements in large multimodal models have provided blind or visually impaired (BVI) individuals with new capabilities to interpret and engage with the real world through interactive systems that utilize live video feeds. However,…

人机交互 · 计算机科学 2025-08-06 Ruei-Che Chang , Rosiana Natalie , Wenqian Xu , Jovan Zheng Feng Yap , Anhong Guo

Recognizing human actions in videos requires spatial and temporal understanding. Most existing action recognition models lack a balanced spatio-temporal understanding of videos. In this work, we propose a novel two-stream architecture,…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Dongho Lee , Jongseo Lee , Jinwoo Choi

This paper presents UniVST, a unified framework for localized video style transfer based on diffusion models. It operates without the need for training, offering a distinct advantage over existing diffusion methods that transfer style…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Quanjian Song , Mingbao Lin , Wengyi Zhan , Shuicheng Yan , Liujuan Cao , Rongrong Ji

In this paper, we introduce Attention Prompt Tuning (APT) - a computationally efficient variant of prompt tuning for video-based applications such as action recognition. Prompt tuning approaches involve injecting a set of learnable prompts…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Wele Gedara Chaminda Bandara , Vishal M. Patel

Recently unsupervised learning of depth from videos has made remarkable progress and the results are comparable to fully supervised methods in outdoor scenes like KITTI. However, there still exist great challenges when directly applying…

计算机视觉与模式识别 · 计算机科学 2019-10-22 Junsheng Zhou , Yuwang Wang , Kaihuai Qin , Wenjun Zeng

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Jongseo Lee , Joohyun Chang , Dongho Lee , Jinwoo Choi

While the recent advances in Multimodal Large Language Models (MLLMs) constitute a significant leap forward in the field, these models are predominantly confined to the realm of input-side multimodal comprehension, lacking the capacity for…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Zhanyu Wang , Longyue Wang , Zhen Zhao , Minghao Wu , Chenyang Lyu , Huayang Li , Deng Cai , Luping Zhou , Shuming Shi , Zhaopeng Tu

Previous deep learning-based video stabilizers require a large scale of paired unstable and stable videos for training, which are difficult to collect. Traditional trajectory-based stabilizers, on the other hand, divide the task into…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Yufei Xu , Jing Zhang , Stephen J. Maybank , Dacheng Tao

Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating…

图形学 · 计算机科学 2025-10-07 Nilay Kumar , Priyansh Bhandari , G. Maragatham

The exponential growth of AI education has brought millions of learners to online platforms, yet this massive scale has simultaneously exposed critical pedagogical shortcomings. Traditional video-based instruction, while cost-effective and…

人机交互 · 计算机科学 2026-04-20 Mohammed Abraar , Raj Abhijit Dandekar , Rajat Dandekar , Sreedath Panat

Unsupervised multi-object scene decomposition is a fast-emerging problem in representation learning. Despite significant progress in static scenes, such models are unable to leverage important dynamic cues present in video. We propose a…

计算机视觉与模式识别 · 计算机科学 2020-06-29 Polina Zablotskaia , Edoardo A. Dominici , Leonid Sigal , Andreas M. Lehrmann

Traditional computer vision models often necessitate extensive data acquisition, annotation, and validation. These models frequently struggle in real-world applications, resulting in high false positive and negative rates, and exhibit poor…

Among various region embedding methods, graph-based region relation learning models stand out, owing to their strong structure representation ability for encoding spatial correlations with graph neural networks. Despite their effectiveness,…

机器学习 · 计算机科学 2023-05-09 Qianru Zhang , Chao Huang , Lianghao Xia , Zheng Wang , Zhonghang Li , Siuming Yiu

Miscommunication and communication challenges between instructors and students represents one of the primary barriers to post-secondary learning. Students often avoid or miss opportunities to ask questions during office hours due to…

计算机与社会 · 计算机科学 2024-02-02 Ramteja Sajja , Yusuf Sermet , David Cwiertny , Ibrahim Demir

Static appearance of video may impede the ability of a deep neural network to learn motion-relevant features in video action recognition. In this paper, we introduce a new concept, Dynamic Appearance (DA), summarizing the appearance…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Guoxi Huang , Adrian G. Bors

We present PandaGPT, an approach to emPower large lANguage moDels with visual and Auditory instruction-following capabilities. Our pilot experiments show that PandaGPT can perform complex tasks such as detailed image description generation,…

计算与语言 · 计算机科学 2023-05-29 Yixuan Su , Tian Lan , Huayang Li , Jialu Xu , Yan Wang , Deng Cai

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

物理教育 · 物理学 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic

Humans can collaborate and complete tasks based on visual signals and instruction from the environment. Training such a robot is difficult especially due to the understanding of the instruction and the complicated environment. Previous…

人工智能 · 计算机科学 2023-05-12 Kairui Zhou