English
Related papers

Related papers: Evaluating GPT-5 as a Multimodal Clinical Reasoner…

200 papers

We study GPT-3, a recent large language model, using tools from cognitive psychology. More specifically, we assess GPT-3's decision-making, information search, deliberation, and causal reasoning abilities on a battery of canonical…

Computation and Language · Computer Science 2023-02-22 Marcel Binz , Eric Schulz

Recent research has offered insights into the extraordinary capabilities of Large Multimodal Models (LMMs) in various general vision and language tasks. There is growing interest in how LMMs perform in more specialized domains. Social media…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Hanjia Lyu , Jinfa Huang , Daoan Zhang , Yongsheng Yu , Xinyi Mou , Jinsheng Pan , Zhengyuan Yang , Zhongyu Wei , Jiebo Luo

Recent advances in general medical AI have made significant strides, but existing models often lack the reasoning capabilities needed for complex medical decision-making. This paper presents GMAI-VL-R1, a multimodal medical reasoning model…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Yanzhou Su , Tianbin Li , Jiyao Liu , Chenglong Ma , Junzhi Ning , Cheng Tang , Sibo Ju , Jin Ye , Pengcheng Chen , Ming Hu , Shixiang Tang , Lihao Liu , Bin Fu , Wenqi Shao , Xiaowei Hu , Xiangwen Liao , Yuanfeng Ji , Junjun He

With the significant expansion of the context window in Large Language Models (LLMs), these models are theoretically capable of processing millions of tokens in a single pass. However, research indicates a significant gap between this…

Computation and Language · Computer Science 2026-02-25 Nima Esmi , Maryam Nezhad-Moghaddam , Fatemeh Borhani , Asadollah Shahbahrami , Amin Daemdoost , Georgi Gaydadjiev

The recent breakthroughs in OpenAI's GPT4o model have demonstrated surprisingly good capabilities in image generation and editing, resulting in significant excitement in the community. This technical report presents the first-look…

Computer Vision and Pattern Recognition · Computer Science 2025-05-05 Zhiyuan Yan , Junyan Ye , Weijia Li , Zilong Huang , Shenghai Yuan , Xiangyang He , Kaiqing Lin , Jun He , Conghui He , Li Yuan

Multimodal models are expected to be a critical component to future advances in artificial intelligence. This field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural…

Computation and Language · Computer Science 2024-06-11 Sai Munikoti , Ian Stewart , Sameera Horawalavithana , Henry Kvinge , Tegan Emerson , Sandra E Thompson , Karl Pazdernik

The ability to perform Chain-of-Thought (CoT) reasoning marks a major milestone for multimodal models (MMs), enabling them to solve complex visual reasoning problems. Yet a critical question remains: is such reasoning genuinely grounded in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Jusheng Zhang , Kaitong Cai , Xiaoyang Guo , Sidi Liu , Qinhan Lv , Ruiqi Chen , Jing Yang , Yijia Fan , Xiaofei Sun , Jian Wang , Ziliang Chen , Liang Lin , Keze Wang

OpenAI's latest large vision-language model (LVLM), GPT-4V(ision), has piqued considerable interest for its potential in medical applications. Despite its promise, recent studies and internal reviews highlight its underperformance in…

Computation and Language · Computer Science 2023-12-13 Pengcheng Chen , Ziyan Huang , Zhongying Deng , Tianbin Li , Yanzhou Su , Haoyu Wang , Jin Ye , Yu Qiao , Junjun He

Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remains: top-performing systems are either closed-source or computationally prohibitive,…

Computation and Language · Computer Science 2026-03-03 Kai Zhang , Zhengqing Yuan , Cheng Peng , Songlin Zhao , Mengxian Lyu , Ziyi Chen , Yanfang Ye , Wei Liu , Ying Zhang , Kaleb E Smith , Lifang He , Lichao Sun , Yonghui Wu

Machine reasoning has made great progress in recent years owing to large language models (LLMs). In the clinical domain, however, most NLP-driven projects mainly focus on clinical classification or reading comprehension, and under-explore…

Computation and Language · Computer Science 2024-05-13 Taeyoon Kwon , Kai Tzu-iunn Ong , Dongjin Kang , Seungjun Moon , Jeong Ryong Lee , Dosik Hwang , Yongsik Sim , Beomseok Sohn , Dongha Lee , Jinyoung Yeo

Multi-modality foundation models, as represented by GPT-4V, have brought a new paradigm for low-level visual perception and understanding tasks, that can respond to a broad range of natural human instructions in a model. While existing…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Haoning Wu , Zicheng Zhang , Erli Zhang , Chaofeng Chen , Liang Liao , Annan Wang , Kaixin Xu , Chunyi Li , Jingwen Hou , Guangtao Zhai , Geng Xue , Wenxiu Sun , Qiong Yan , Weisi Lin

Demographic inference plays a crucial role in understanding the representativeness and equity of social media-based research. However, existing methods typically rely on a single modality, such as text, image, or network, and are limited to…

Social and Information Networks · Computer Science 2025-12-04 Hao Yang , Angela Yao , Eric Chang , Hexiang Wang

We propose a new benchmark evaluating the performance of multimodal large language models on rebus puzzles. The dataset covers 333 original examples of image-based wordplay, cluing 13 categories such as movies, composers, major cities, and…

Recent work has shown promising performance of frontier large language models (LLMs) and their multimodal counterparts in medical quizzes and diagnostic tasks, highlighting their potential for broad clinical utility given their accessible,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Advait Gosai , Arun Kavishwar , Stephanie L. McNamara , Soujanya Samineni , Renato Umeton , Alexander Chowdhury , William Lotter

Large language models (LLMs) are entering clinician workflows, yet evaluations rarely measure how clinician reasoning shapes model behavior during clinical interactions. We combined 61 New England Journal of Medicine Case Records with 92…

Social problems stemming from the shortage of radiologists are intensifying, and artificial intelligence is being highlighted as a potential solution. Recently emerging large-scale generative AI has expanded from large language models…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Inwoo Seo , Eunkyoung Bae , Joo-Young Jeon , Young-Sang Yoon , Jiho Cha

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Deyao Zhu , Jun Chen , Xiaoqian Shen , Xiang Li , Mohamed Elhoseiny

Large Language Models such as GPTs (Generative Pre-trained Transformers) exhibit remarkable capabilities across a broad spectrum of applications. Nevertheless, due to their intrinsic complexity, these models present substantial challenges…

Machine Learning · Computer Science 2024-10-17 Ashkan Golgoon , Khashayar Filom , Arjun Ravi Kannan

In Multi-Agent Systems (MAS), agents are designed with social capabilities, allowing them to understand and reason about social concepts such as norms when interacting with others (e.g., inter-robot interactions). In Normative MAS (NorMAS),…

Multiagent Systems · Computer Science 2026-03-05 Oishik Chowdhury , Anushka Debnath , Bastin Tony Roy Savarimuthu

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their…

Computation and Language · Computer Science 2024-08-27 Florian Schneider , Sunayana Sitaram