English
Related papers

Related papers: SPARROW: Learning Spatial Precision and Temporal R…

200 papers

Spatial understanding is essential for Multimodal Large Language Models (MLLMs) to support perception, reasoning, and planning in embodied environments. Despite recent progress, existing studies reveal that MLLMs still struggle with spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Wanyue Zhang , Yibin Huang , Yangbin Xu , JingJing Huang , Helu Zhi , Shuo Ren , Wang Xu , Jiajun Zhang

Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand, and reason about dynamic real-world scenarios. However, existing Multimodal Large…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Sunqi Fan , Jiashuo Cui , Meng-Hao Guo , Shuojin Yang

Recent video multimodal large language models achieve impressive results across various benchmarks. However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jiahao Meng , Tan Yue , Qi Xu , Haochen Wang , Zhongwei Ren , Weisong Liu , Yuhao Wang , Renrui Zhang , Yunhai Tong , Haodong Duan

Pre-trained on tremendous image-text pairs, vision-language models like CLIP have demonstrated promising zero-shot generalization across numerous image-based tasks. However, extending these capabilities to video tasks remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Zichen Liu , Kunlun Xu , Bing Su , Xu Zou , Yuxin Peng , Jiahuan Zhou

Fueled by the Large Language Models (LLMs) wave, Large Visual-Language Models (LVLMs) have emerged as a pivotal advancement, bridging the gap between image and text. However, video making it challenging for LVLMs to perform adequately due…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Yang Liu , Pengxiang Ding , Siteng Huang , Min Zhang , Han Zhao , Donglin Wang

This paper does not introduce a novel method but instead establishes a straightforward, incremental, yet essential baseline for video temporal grounding (VTG), a core capability in video understanding. While multimodal large language models…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jun Zhang , Teng Wang , Yuying Ge , Yixiao Ge , Xinhao Li , Ying Shan , Limin Wang

Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video frames and their associated information. Efficiently…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Xuyi Yang , Wenhao Zhang , Hongbo Jin , Lin Liu , Hongbo Xu , Yongwei Nie , Fei Yu , Fei Ma

Open-world referring segmentation requires grounding unconstrained language expressions to precise pixel-level regions. Existing multimodal large language models (MLLMs) exhibit strong open-world visual grounding, but their outputs remain…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yuan Yao , Qiushi Yang , Humen Zhong , Jiangning Wei , Yifang Men , Shuai Bai , Miaomiao Cui , Zhibo Yang

In this paper, we introduce Semantic Layering in Room Segmentation via LLMs (SeLRoS), an advanced method for semantic room segmentation by integrating Large Language Models (LLMs) with traditional 2D map-based segmentation. Unlike previous…

Robotics · Computer Science 2024-03-20 Taehyeon Kim , Byung-Cheol Min

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained vision language models (VLMs), recent approaches are still not qualified to accurately…

While Multimodal Large Language Models (MLLMs) have achieved impressive performance on semantic tasks, their spatial intelligence--crucial for robust and grounded AI systems--remains underdeveloped. Existing benchmarks fall short of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Mingrui Wu , Zhaozhi Wang , Fangjinhua Wang , Jiaolong Yang , Marc Pollefeys , Tong Zhang

While pre-training large-scale video-language models (VLMs) has shown remarkable potential for various downstream video-language tasks, existing VLMs can still suffer from certain commonly seen limitations, e.g., coarse-grained cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Hao Fei , Shengqiong Wu , Meishan Zhang , Min Zhang , Tat-Seng Chua , Shuicheng Yan

Vision-Language Models (VLMs) have been increasingly integrated into object navigation tasks for their rich prior knowledge and strong reasoning abilities. However, applying VLMs to navigation poses two key challenges: effectively…

Robotics · Computer Science 2025-09-17 Haokun Zhu , Zongtai Li , Zhixuan Liu , Wenshan Wang , Ji Zhang , Jonathan Francis , Jean Oh

Gloss-free Sign Language Translation (SLT) converts sign videos directly into spoken language sentences without relying on glosses. Recently, Large Language Models (LLMs) have shown remarkable translation performance in gloss-free methods…

Computation and Language · Computer Science 2025-02-25 Eui Jun Hwang , Sukmin Cho , Junmyeong Lee , Jong C. Park

We study the problem of segmenting moving objects in unconstrained videos. Given a video, the task is to segment all the objects that exhibit independent motion in at least one frame. We formulate this as a learning problem and design our…

Computer Vision and Pattern Recognition · Computer Science 2017-12-05 Pavel Tokmakov , Cordelia Schmid , Karteek Alahari

Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process a finite number of frames in a single inference,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yucheng Suo , Fan Ma , Linchao Zhu , Tianyi Wang , Fengyun Rao , Yi Yang

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Lennart Maack , Alexander Schlaefer

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive…

Long-horizon egocentric video presents significant challenges for visual navigation due to viewpoint drift and the absence of persistent geometric context. Although recent vision-language models perform well on image and short-video…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 James Tribble , Hao Wang , Si-En Hong , Chaoyi Zhou , Ashish Bastola , Siyu Huang , Abolfazl Razi

Radar sensors provide reliable perception across adverse weather, lighting, and long-range conditions, yet existing machine learning approaches remain fragmented and task-specific, with each downstream task employing distinct architectures…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Pushkal Mishra , Kshitiz Bansal , Dinesh Bharadia