中文
相关论文

相关论文: GROK: From Quantitative Biomarkers to Qualitative …

200 篇论文

Multimodal large language models (MLLMs) have achieved significant success in the general field of image processing. Their emerging task generalization and freeform conversational capabilities can greatly facilitate medical diagnostic…

图像与视频处理 · 电气工程与系统科学 2024-09-17 Youzhu Jin , Yichen Zhang

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Despite significant progress in Multi-modal Large Language Models (MLLMs), their clinical reasoning capacity for multi-modal diagnosis remains largely unexamined. Current benchmarks, mostly single-modality data, can't evaluate progressive…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Gui Wang , Zehao Zhong , YongSong Zhou , Yudong Li , Ende Wu , Wooi Ping Cheah , Rong Qu , Jianfeng Ren , Linlin Shen

In recent years, large language models (LLMs) have demonstrated remarkable potential across various medical applications. Building on this foundation, multimodal large language models (MLLMs) integrate LLMs with visual models to process…

计算与语言 · 计算机科学 2025-03-11 Xiaoyi Liang , Mouxiao Bian , Moxin Chen , Lihao Liu , Junjun He , Jie Xu , Lin Li

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…

Optical character recognition (OCR) and multilingual text understanding remain major failure modes of multimodal large language models (MLLMs), particularly in real-world images containing cluttered layouts, small fonts, blur, occlusion,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Qinwu Xu , Yifan Jiang , Haoyu Ren

Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastrointestinal endoscopy is currently hindered by two critical limitations: the misalignment between…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Huan Zheng , Yucheng Zhou , Tianyi Yan , Dubing Chen , Hongbo Lu , Wenlong Liao , Tao He , Pai Peng , Jianbing Shen

Most multimodal large language models (MLLMs) learn language-to-object grounding through causal language modeling where grounded objects are captured by bounding boxes as sequences of location tokens. This paradigm lacks pixel-level…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Yichi Zhang , Ziqiao Ma , Xiaofeng Gao , Suhaila Shakiah , Qiaozi Gao , Joyce Chai

Recently, multimodal deep learning, which integrates histopathology slides and molecular biomarkers, has achieved a promising performance in glioma grading. Despite great progress, due to the intra-modality complexity and inter-modality…

计算机视觉与模式识别 · 计算机科学 2024-08-19 Li Pan , Yupei Zhang , Qiushi Yang , Tan Li , Xiaohan Xing , Maximus C. F. Yeung , Zhen Chen

In the field of multimodal chain-of-thought (CoT) reasoning, existing approaches predominantly rely on reasoning on pure language space, which inherently suffers from language bias and is largely confined to math or science domains. This…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Jiacong Wang , Zijian Kang , Haochen Wang , Haiyong Jiang , Jiawen Li , Bohong Wu , Ya Wang , Jiao Ran , Xiao Liang , Chao Feng , Jun Xiao

In recent years, Multi-modal Large Language Models (MLLMs) have achieved strong performance in OCR-centric Visual Question Answering (VQA) tasks, illustrating their capability to process heterogeneous data and exhibit adaptability across…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Chen Duan , Zhentao Guo , Pei Fu , Zining Wang , Kai Zhou , Pengfei Yan

Large Vision Language Models (LVLMs) have made remarkable progress, enabling sophisticated vision-language interaction and dialogue applications. However, existing benchmarks primarily focus on reasoning tasks, often neglecting fine-grained…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Cong Pang , Hongtao Yu , Zixuan Chen , Lewei Lu , Xin Lou

MLLMs MLLMs are beginning to appear in clinical workflows, but their ability to perform complex medical reasoning remains unclear. We present Med-CMR, a fine-grained Medical Complex Multimodal Reasoning benchmark. Med-CMR distinguishes from…

Recent advances of large multi-modality models (LMM) have greatly improved the ability of image quality assessment (IQA) method to evaluate and explain the quality of visual content. However, these advancements are mostly focused on overall…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Chaofeng Chen , Sensen Yang , Haoning Wu , Liang Liao , Zicheng Zhang , Annan Wang , Wenxiu Sun , Qiong Yan , Weisi Lin

Multimodal large language models (MLLMs) show promising performance on medical visual question answering (VQA) and report generation, but these generation and explanation abilities do not reliably transfer to disease-specific…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Li Wang , Xi Chen , XiangWen Deng , HuaHui Yi , ZeKun Jiang , Kang Li , Jian Li

Multimodal language models (MLMs) show promise for clinical decision support and diagnostic reasoning, raising the prospect of end-to-end automated medical image interpretation. However, clinicians are highly selective in adopting AI tools;…

Referring Expression Comprehension (REC) is a popular multimodal task that aims to accurately detect target objects within a single image based on a given textual expression. However, due to the limitations of earlier models, traditional…

机器学习 · 计算机科学 2025-08-21 Guanghao Jin , Jingpei Wu , Tianpei Guo , Yiyi Niu , Weidong Zhou , Guoyang Liu

In recent years, multimodal large language models (MLLMs) have made significant strides by training on vast high-quality image-text datasets, enabling them to generally understand images well. However, the inherent difficulty in explicitly…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Yuanze Lin , Yunsheng Li , Dongdong Chen , Weijian Xu , Ronald Clark , Philip Torr , Lu Yuan

Multimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs…

Radiology Report Generation (RRG) through Vision-Language Models (VLMs) promises to reduce documentation burden, improve reporting consistency, and accelerate clinical workflows. However, their clinical adoption remains limited by the lack…

计算机视觉与模式识别 · 计算机科学 2026-02-18 Marco Salmè , Federico Siciliano , Fabrizio Silvestri , Paolo Soda , Rosa Sicilia , Valerio Guarrasi
‹ 上一页 1 2 3 10 下一页 ›