English
Related papers

Related papers: CycliST: A Video Language Model Benchmark for Reas…

200 papers

Vision Language Models (VLMs) demonstrate remarkable proficiency in addressing a wide array of visual questions, which requires strong perception and reasoning faculties. Assessing these two competencies independently is crucial for model…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Yuxuan Qiao , Haodong Duan , Xinyu Fang , Junming Yang , Lin Chen , Songyang Zhang , Jiaqi Wang , Dahua Lin , Kai Chen

Cross-modal entity linking refers to the ability to align entities and their attributes across different modalities. While cross-modal entity linking is a fundamental skill needed for real-world applications such as multimodal code…

Computation and Language · Computer Science 2025-06-02 Iñigo Alonso , Gorka Azkune , Ander Salaberria , Jeremy Barnes , Oier Lopez de Lacalle

Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise intermediate visual states (IVS), much like a human would, which…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Xuecheng Wu , Jiaxing Liu , Danlei Huang , Yifan Wang , Yunyun Shi , Kedi Chen , Junxiao Xue , Yang Liu , Chunlin Chen , Hairong Dong , Dingkang Yang

In this work, we introduce SPLICE, a human-curated benchmark derived from the COIN instructional video dataset, designed to probe event-based reasoning across multiple dimensions: temporal, causal, spatial, contextual, and general…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Mohamad Ballout , Okajevo Wilfred , Seyedalireza Yaghoubi , Nohayr Muhammad Abdelmoneim , Julius Mayer , Elia Bruni

The advancement of Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs) and large vision-language models (LVLMs). However, a rigorous evaluation framework for video CoT reasoning…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Yukun Qi , Yiming Zhao , Yu Zeng , Xikun Bao , Wenxuan Huang , Lin Chen , Zehui Chen , Jie Zhao , Zhongang Qi , Feng Zhao

Understanding physical transformations is fundamental for reasoning in dynamic environments. While Vision Language Models (VLMs) show promise in embodied applications, whether they genuinely understand physical transformations remains…

Artificial Intelligence · Computer Science 2026-03-10 Dezhi Luo , Yijiang Li , Maijunxian Wang , Tianwei Zhao , Bingyang Wang , Siheng Wang , Pinyuan Feng , Pooyan Rahmanzadehgervi , Ziqiao Ma , Hokin Deng

We introduce OpenVLThinker, one of the first open-source large vision-language models (LVLMs) to exhibit sophisticated chain-of-thought reasoning, achieving notable performance gains on challenging visual reasoning tasks. While text-based…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Yihe Deng , Hritik Bansal , Fan Yin , Nanyun Peng , Wei Wang , Kai-Wei Chang

Multimodal Large Language Models (MLLMs) have achieved significant advancements in tasks like Visual Question Answering (VQA) by leveraging foundational Large Language Models (LLMs). However, their abilities in specific areas such as visual…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Mohamed Fazli Imam , Chenyang Lyu , Alham Fikri Aji

End-to-end text-image machine translation (TIMT), which directly translates textual content in images across languages, is crucial for real-world multilingual scene understanding. Despite advances in vision-language large models (VLLMs),…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Gengluo Li , Chengquan Zhang , Yupu Liang , Huawen Shen , Yaping Zhang , Pengyuan Lyu , Weinong Wang , Xingyu Wan , Gangyan Zeng , Han Hu , Can Ma , Yu Zhou

Recent research has increasingly focused on the reasoning capabilities of Large Language Models (LLMs) in multi-turn interactions, as these scenarios more closely mirror real-world problem-solving. However, analyzing the intricate reasoning…

Computation and Language · Computer Science 2025-11-14 Yiran Zhang , Mingyang Lin , Mark Dras , Usman Naseem

The development of Large Vision-Language Models (LVLMs) is striving to catch up with the success of Large Language Models (LLMs), yet it faces more challenges to be resolved. Very recent works enable LVLMs to localize object-level visual…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Zhipeng Huang , Zhizheng Zhang , Zheng-Jun Zha , Yan Lu , Baining Guo

Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zihui Cheng , Qiguang Chen , Jin Zhang , Hao Fei , Xiaocheng Feng , Wanxiang Che , Min Li , Libo Qin

Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance.…

Artificial Intelligence · Computer Science 2025-08-08 Xinrun Xu , Pi Bu , Ye Wang , Börje F. Karlsson , Ziming Wang , Tengtao Song , Qi Zhu , Jun Song , Zhiming Ding , Bo Zheng

Large vision-language models (LVLMs) have recently witnessed rapid advancements, exhibiting a remarkable capacity for perceiving, understanding, and processing visual information by connecting visual receptor with large language models…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Shuai Bai , Shusheng Yang , Jinze Bai , Peng Wang , Xingxuan Zhang , Junyang Lin , Xinggang Wang , Chang Zhou , Jingren Zhou

Vision-language models (VLMs) are widely assumed to exhibit in-context learning (ICL), a property similar to that of their language-only counterparts. While recent work suggests VLMs can perform multimodal ICL (MM-ICL), studies show they…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Chengyue Huang , Yuchen Zhu , Sichen Zhu , Jingyun Xiao , Moises Andrade , Shivang Chopra , Zsolt Kira

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Nilay Yilmaz , Maitreya Patel , Yiran Lawrence Luo , Tejas Gokhale , Chitta Baral , Suren Jayasuriya , Yezhou Yang

In recent years, video question answering based on multimodal large language models (MLLM) has garnered considerable attention, due to the benefits from the substantial advancements in LLMs. However, these models have a notable deficiency…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Jinglei Zhang , Yuanfan Guo , Rolandos Alexandros Potamias , Jiankang Deng , Hang Xu , Chao Ma

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Yuandong Wang , Yao Cui , Yuxin Zhao , Zhen Yang , Yangfu Zhu , Zhenzhou Shao

Human intelligence requires correctness and robustness, with the former being foundational for the latter. In video understanding, correctness ensures the accurate interpretation of visual content, and robustness maintains consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Yuanhan Zhang , Yunice Chew , Yuhao Dong , Aria Leo , Bo Hu , Ziwei Liu

Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications such as disaster management, traffic planning, embodied navigation, world modeling, and geography…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Azmine Toushik Wasi , Shahriyar Zaman Ridoy , Koushik Ahamed Tonmoy , Kinga Tshering , S. M. Muhtasimul Hasan , Wahid Faisal , Tasnim Mohiuddin , Md Rizwan Parvez
‹ Prev 1 3 4 5 6 7 10 Next ›