中文
相关论文

相关论文: MIMO: A medical vision language model with visual …

200 篇论文

Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs represent a bridge between object detection and segmentation, and report understanding and…

计算机视觉与模式识别 · 计算机科学 2026-03-23 Andrew Seohwan Yu , Mohsen Hariri , Kunio Nakamura , Mingrui Yang , Xiaojuan Li , Vipin Chaudhary

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Hozaifa Kassab , Ahmed Mahmoud , Mohamed Bahaa , Ammar Mohamed , Ali Hamdi

Several medical Multimodal Large Languange Models (MLLMs) have been developed to address tasks involving visual images with textual instructions across various medical modalities, achieving impressive results. Most current medical…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Lehan Wang , Haonan Wang , Honglong Yang , Jiaji Mao , Zehong Yang , Jun Shen , Xiaomeng Li

In hospitals, data are siloed to specific information systems that make the same information available under different modalities such as the different medical imaging exams the patient undergoes (CT scans, MRI, PET, Ultrasound, etc.) and…

计算机视觉与模式识别 · 计算机科学 2021-02-03 Tristan Sylvain , Francis Dutil , Tess Berthier , Lisa Di Jorio , Margaux Luck , Devon Hjelm , Yoshua Bengio

Multimodal Large Language Models (MLLMs) have demonstrated significant advances across numerous vision-language tasks. MLLMs have shown promising capability in aligning visual and textual modalities, allowing them to process image-text…

计算与语言 · 计算机科学 2025-09-29 Xiaolong Wang , Zhaolu Kang , Wangyuxuan Zhai , Xinyue Lou , Yunghwei Lai , Ziyue Wang , Yawen Wang , Kaiyu Huang , Yile Wang , Peng Li , Yang Liu

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to…

人工智能 · 计算机科学 2023-11-23 Jiayang Wu , Wensheng Gan , Zefeng Chen , Shicheng Wan , Philip S. Yu

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of…

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Fan Bai , Yuxin Du , Tiejun Huang , Max Q. -H. Meng , Bo Zhao

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Zeyuan Yang , Xueyang Yu , Delin Chen , Maohao Shen , Chuang Gan

In different multimodal scenarios, it needs to integrate and utilize information across modalities in a specific way based on the demands of the task. Different integration ways between modalities are referred to as "multimodal…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yu Miao , Zequn Yang , Yake Wei , Ziheng Chen , Haotian Ni , Haodong Duan , Kai Chen , Di Hu

The inability to interpret the model prediction in semantically and visually meaningful ways is a well-known shortcoming of most existing computer-aided diagnosis methods. In this paper, we propose MDNet to establish a direct multimodal…

计算机视觉与模式识别 · 计算机科学 2017-07-11 Zizhao Zhang , Yuanpu Xie , Fuyong Xing , Mason McGough , Lin Yang

Currently, dialogue systems have achieved high performance in processing text-based communication. However, they have not yet effectively incorporated visual information, which poses a significant challenge. Furthermore, existing models…

计算与语言 · 计算机科学 2023-12-19 Viktor Moskvoretskii , Anton Frolov , Denis Kuznetsov

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on machine translation,…

机器学习 · 计算机科学 2026-01-22 Piyush Singh Pasi

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional use of templates or pre-set information from training sets…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Yunhao Li , Xiaoqiong Liu , Luke Liu , Heng Fan , Libo Zhang

There has been a recent spike in interest in multi-modal Language and Vision problems. On the language side, most of these models primarily focus on English since most multi-modal datasets are monolingual. We try to bridge this gap with a…

机器学习 · 计算机科学 2021-09-17 Pranav Aggarwal , Ritiz Tambi , Ajinkya Kale

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal…

计算机视觉与模式识别 · 计算机科学 2025-09-11 Wenjun Yu , Yinchen Zhou , Jia-Xuan Jiang , Shubin Zeng , Yuee Li , Zhong Wang

We present UniModel, a unified generative model that jointly supports visual understanding and visual generation within a single pixel-to-pixel diffusion framework. Our goal is to achieve unification along three axes: the model, the tasks,…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Chi Zhang , Jiepeng Wang , Youming Wang , Yuanzhi Liang , Xiaoyan Yang , Zuoxin Li , Haibin Huang , Xuelong Li

With the increasing application of large language models (LLMs) in the medical domain, evaluating these models' performance using benchmark datasets has become crucial. This paper presents a comprehensive survey of various benchmark…

Recent advances in large vision-language models have led to impressive performance in visual question answering and multimodal reasoning. However, it remains unclear whether these models genuinely perform grounded visual reasoning or rely…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Chengfei Wu , Ronald Seoh , Bingxuan Li , Liqiang Zhang , Fengrong Han , Dan Goldwasser

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…