English
Related papers

Related papers: SUREON: A Benchmark and Vision-Language-Model for …

200 papers

Chain-of-thought (CoT) reasoning has emerged as a powerful tool for multimodal large language models on video understanding tasks. However, its necessity and advantages over direct answering remain underexplored. In this paper, we first…

Recent studies have demonstrated the effectiveness of Large Language Models (LLMs) as reasoning modules that can deconstruct complex tasks into more manageable sub-tasks, particularly when applied to visual reasoning tasks for images. In…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ahmad Mahmood , Ashmal Vayani , Muzammal Naseer , Salman Khan , Fahad Shahbaz Khan

Large Language Models (LLMs) have been shown to encode clinical knowledge. Many evaluations, however, rely on structured question-answer benchmarks, overlooking critical challenges of interpreting and reasoning about unstructured clinical…

Computation and Language · Computer Science 2026-04-01 Meghal Dani , Muthu Jeyanthi Prakash , Filip Rosa , Zeynep Akata , Stefanie Liebe

Surgery is a highly complex process, and artificial intelligence has emerged as a transformative force in supporting surgical guidance and decision-making. However, the unimodal nature of most current AI systems limits their ability to…

Computer Vision and Pattern Recognition · Computer Science 2026-01-26 Nakul Poudel , Richard Simon , Cristian A. Linte

The operating room (OR) is a dynamic and complex environment consisting of a multidisciplinary team working together in a high take environment to provide safe and efficient patient care. Additionally, surgeons are frequently exposed to…

Multimodal emotion analysis is shifting from static classification to generative reasoning. Beyond simple label prediction, robust affective reasoning must synthesize fine-grained signals such as facial micro-expressions and prosodic which…

Multimedia · Computer Science 2026-02-05 Zhixian Zhao , Wenjie Tian , Lei Xie

Surgery is a high-stakes domain where surgeons must navigate critical anatomical structures and actively avoid potential complications while achieving the main task at hand. Such surgical activity has been shown to affect long-term patient…

Despite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions. First, they overwhelmingly rely on static images, failing to capture the temporal complexity of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Zikui Cai , Andrew Wang , Anirudh Satheesh , Ankit Nakhawa , Hyunwoo Jae , Keenan Powell , Minghui Liu , Neel Jay , Sungbin Oh , Xiyao Wang , Yongyuan Liang , Tom Goldstein , Furong Huang

Clinically reliable perception of surgical scenes is essential for advancing intelligent, context-aware intraoperative assistance such as instrument handoff guidance, collision avoidance, and workflow-aware robotic support. Existing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Tajamul Ashraf , Abrar Ul Riyaz , Wasif Tak , Tavaheed Tariq , Sonia Yadav , Moloud Abdar , Janibul Bashir

As humans, we are natural any-horizon reasoners, i.e., we can decide whether to iteratively skim long videos or watch short ones in full when necessary for a given task. With this in mind, one would expect video reasoning models to reason…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Jitesh Jain , Jialuo Li , Zixian Ma , Jieyu Zhang , Chris Dongjoo Kim , Sangho Lee , Rohun Tripathi , Tanmay Gupta , Christopher Clark , Humphrey Shi

Human processes video reasoning in a sequential spatio-temporal reasoning logic, we first identify the relevant frames ("when") and then analyse the spatial relationships ("where") between key objects, and finally leverage these…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Zixu Cheng , Jian Hu , Ziquan Liu , Chenyang Si , Wei Li , Shaogang Gong

In surgical training for medical students, proficiency development relies on expert-led skill assessment, which is costly, time-limited, difficult to scale, and its expertise remains confined to institutions with available specialists.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Le Ma , Thiago Freitas dos Santos , Nadia Magnenat-Thalmann , Katarzyna Wac

Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). However, current VQA benchmarks overlook scenarios with visually…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jianing An , Luyang Jiang , Jie Luo , Wenjun Wu , Lei Huang

Surgery requires comprehensive medical knowledge, visual assessment skills, and procedural expertise. While recent surgical AI models have focused on solving task-specific problems, there is a need for general-purpose systems that can…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Samuel Schmidgall , Joseph Cho , Cyril Zakka , William Hiesinger

Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Hanoona Rasheed , Abdelrahman Shaker , Anqi Tang , Muhammad Maaz , Ming-Hsuan Yang , Salman Khan , Fahad Shahbaz Khan

We present VideoOrion, a Video Large Language Model (Video-LLM) that explicitly captures the key semantic information in videos - the spatial-temporal dynamics of objects throughout the videos. VideoOrion employs expert vision models to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yicheng Feng , Yijiang Li , Wanpeng Zhang , Hao Luo , Zihao Yue , Sipeng Zheng , Zongqing Lu

Vision-Language Models (VLMs) have enabled interpretable medical diagnosis by integrating visual perception with linguistic reasoning. Yet, existing medical chain-of-thought (CoT) models lack explicit mechanisms to represent and enforce…

Artificial Intelligence · Computer Science 2026-05-29 Jianxin Lin , Chunzheng Zhu , Peter J. Kneuertz , Yunfei Bai , Yuan Xue

Large Multimodal Models (LMMs) have achieved remarkable progress across various capabilities; however, complex video reasoning in the scientific domain remains a significant and challenging frontier. Current video benchmarks predominantly…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Andong Deng , Taojiannan Yang , Shoubin Yu , Lincoln Spencer , Mohit Bansal , Chen Chen , Serena Yeung-Levy , Xiaohan Wang

Medical Visual Language Models have shown great potential in various healthcare applications, including medical image captioning and diagnostic assistance. However, most existing models rely on text-based instructions, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Tan-Hanh Pham , Chris Ngo , Trong-Duong Bui , Minh Luu Quang , Tan-Huong Pham , Truong-Son Hy

State-of-the-art vision-language models (VLMs) score impressively on video benchmarks yet stumble on basic visual reasoning tasks involving spatial relations, navigation, and object selection that a preschooler solves easily. We hypothesize…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Bishoy Galoaa , Xiangyu Bai , Sarah Ostadabbas
‹ Prev 1 8 9 10 Next ›