English
Related papers

Related papers: CamReasoner: Reinforcing Camera Movement Understan…

200 papers

Cause-and-effect reasoning in video is a significant challenge for Vision-Language Models (VLMs), as it requires going beyond surface-level perception to a deeper understanding of causal mechanisms. However, existing benchmarks rarely…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Mingfang Zhang , Jingjing Pan , Ashutosh Kumar , Rajat Saini , Mustafa Erdogan , Hsuan-Kung Yang , Caixin Kang , Yifei Huang , Yoichi Sato , Quan Kong

Grounding large language models (LLMs) in domain-specific tasks like post-hoc dash-cam driving video analysis is challenging due to their general-purpose training and lack of structured inductive biases. As vision is often the sole modality…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Manyi Yao , Bingbing Zhuang , Sparsh Garg , Amit Roy-Chowdhury , Christian Shelton , Manmohan Chandraker , Abhishek Aich

Spatial cognition is essential for human intelligence, enabling problem-solving through visual simulations rather than solely relying on verbal reasoning. However, existing AI benchmarks primarily assess verbal reasoning, neglecting the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Linjie Li , Mahtab Bigverdi , Jiawei Gu , Zixian Ma , Yinuo Yang , Ziang Li , Yejin Choi , Ranjay Krishna

Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, recent multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Yian Li , Yang Jiao , Bin Zhu , Tianwen Qian , Shaoxiang Chen , Jingjing Chen , Yu-Gang Jiang

Spatial reasoning is central to navigation and robotics, yet measuring model capabilities on these tasks remains difficult. Existing benchmarks evaluate models in a one-shot setting, requiring full solution generation in a single response,…

Artificial Intelligence · Computer Science 2026-04-13 Lars Benedikt Kaesberg , Tianyu Yang , Niklas Bauer , Terry Ruas , Jan Philip Wahle , Bela Gipp

Large Video-Language Models (Video-LMs) have achieved impressive progress in multimodal understanding, yet their reasoning remains weakly grounded in space and time. We present Know-Show, a new benchmark designed to evaluate spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Chinthani Sugandhika , Chen Li , Deepu Rajan , Basura Fernando

Understanding human motion from video is essential for a range of applications, including pose estimation, mesh recovery and action recognition. While state-of-the-art methods predominantly rely on transformer-based architectures, these…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Arnab Kumar Mondal , Stefano Alletto , Denis Tome

Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks. Several recent works, however, showed…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Navid Rajabi , Jana Kosecka

Interactive autonomous applications require robustness of the perception engine to artifacts in unconstrained videos. In this paper, we examine the effect of camera motion on the task of action detection. We develop a novel ranking method…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Burhan A. Mudassar , Sho Ko , Maojingjing Li , Priyabrata Saha , Saibal Mukhopadhyay

Unlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on explicit 3D inputs or specialized model architectures, we ask:…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Fangrui Zhu , Hanhui Wang , Yiming Xie , Jing Gu , Tianye Ding , Jianwei Yang , Huaizu Jiang

Cinematography understanding refers to the ability to recognize not only the visual content of a scene but also the cinematic techniques that shape narrative meaning. This capability is attracting increasing attention, as it enhances…

Artificial Intelligence · Computer Science 2025-10-06 Hang Wu , Yujun Cai , Haonan Ge , Hongkai Chen , Ming-Hsuan Yang , Yiwei Wang

Video Large Language Models (Video LLMs) have recently exhibited remarkable capabilities in general video understanding. However, they mainly focus on holistic comprehension and struggle with capturing fine-grained spatial and temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yuqian Yuan , Hang Zhang , Wentong Li , Zesen Cheng , Boqiang Zhang , Long Li , Xin Li , Deli Zhao , Wenqiao Zhang , Yueting Zhuang , Jianke Zhu , Lidong Bing

When humans face problems beyond their immediate capabilities, they rely on tools, providing a promising paradigm for improving visual reasoning in multimodal large language models (MLLMs). Effective reasoning, therefore, hinges on knowing…

Artificial Intelligence · Computer Science 2026-01-29 Mingyang Song , Haoyu Sun , Jiawei Gu , Linjie Li , Luxin Xu , Ranjay Krishna , Yu Cheng

Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Zhixuan Wu , Quanxing Zha , Teng Wang , Genbao Xu , Wenyuan Gu , Wei Rao , Nan Ma , Bo Cheng , Soujanya Poria

This work tackles the problem of geo-localization with a new paradigm using a large vision-language model (LVLM) augmented with human inference knowledge. A primary challenge here is the scarcity of data for training the LVLM - existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Ling Li , Yu Ye , Yao Zhou , Bingchuan Jiang , Wei Zeng

Existing benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understanding. However, how well do the models truly perform visual…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Ziyao Shangguan , Chuhan Li , Yuxuan Ding , Yanan Zheng , Yilun Zhao , Tesca Fitzgerald , Arman Cohan

While Multimodal Large Language Models (MLLMs) excel in semantic tasks, they frequently lack the "spatial sense" essential for sophisticated geometric reasoning. Current models typically suffer from exorbitant modality-alignment costs and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yi Zhang , Youya Xia , Yong Wang , Meng Song , Xin Wu , Wenjun Wan , Bingbing Liu , AiXue Ye , Hongbo Zhang , Feng Wen

Video Question Answering (VideoQA) demands models that jointly reason over spatial, temporal, and linguistic cues. However, the task's inherent complexity often requires multi-step reasoning that current large multimodal models (LMMs)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jason Nguyen , Ameet Rao , Alexander Chang , Ishaan Kumar , Erin Tan

Video Models have achieved remarkable success in high-fidelity video generation with coherent motion dynamics. Analogous to the development from text generation to text-based reasoning in language modeling, the development of video models…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Cheng Yang , Haiyuan Wan , Yiran Peng , Xin Cheng , Zhaoyang Yu , Jiayi Zhang , Junchi Yu , Xinlei Yu , Xiawu Zheng , Dongzhan Zhou , Chenglin Wu

Vision-Language Models (VLMs) have recently emerged as powerful tools, excelling in tasks that integrate visual and textual comprehension, such as image captioning, visual question answering, and image-text retrieval. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Ilias Stogiannidis , Steven McDonagh , Sotirios A. Tsaftaris