中文
相关论文

相关论文: MedTrinity-25M: A Large-scale Multimodal Dataset w…

200 篇论文

360 video captures the complete surrounding scenes with the ultra-large field of view of 360X180. This makes 360 scene understanding tasks, eg, segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Weiming Zhang , Dingwen Xiao , Aobotao Dai , Yexin Liu , Tianbo Pan , Shiqi Wen , Lei Chen , Lin Wang

Multimodal pre-training demonstrates its potential in the medical domain, which learns medical visual representations from paired medical reports. However, many pre-training tasks require extra annotations from clinicians, and most of them…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Tongkun Su , Jun Li , Xi Zhang , Haibo Jin , Hao Chen , Qiong Wang , Faqin Lv , Baoliang Zhao , Yin Hu

Retinal diseases spanning a broad spectrum can be effectively identified and diagnosed using complementary signals from multimodal data. However, multimodal diagnosis in ophthalmic practice is typically challenged in terms of data…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Lu Zhang , Huizhen Yu , Zuowei Wang , Fu Gui , Yatu Guo , Wei Zhang , Mengyu Jia

We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks. It is adept at concurrently processing and analyzing modalities of image, video, and text over 4K frames or 1M…

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

Large Language Models (LLMs), enhanced through agent tuning, have demonstrated remarkable capabilities in Chain-of-Thought (CoT) and tool utilization, significantly surpassing the performance of standalone models. However, the multimodal…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Tianhong Gao , Yannian Fu , Weiqun Wu , Haixiao Yue , Shanshan Liu , Gang Zhang

Acquiring properly annotated data is expensive in the medical field as it requires experts, time-consuming protocols, and rigorous validation. Active learning attempts to minimize the need for large annotated samples by actively sampling…

图像与视频处理 · 电气工程与系统科学 2023-06-22 Bidur Khanal , Binod Bhattarai , Bishesh Khanal , Danail Stoyanov , Cristian A. Linte

Prediction tasks in digital pathology are challenging due to the massive size of whole-slide images (WSIs) and the weak nature of training signals. Advances in computing, data availability, and self-supervised learning (SSL) have paved the…

图像与视频处理 · 电气工程与系统科学 2026-02-02 Vishwesh Ramanathan , Tony Xu , Pushpak Pati , Faruk Ahmed , Maged Goubran , Anne L. Martel

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Haoyi Tao , Chaozheng Huang , Nan Wang , Han Lyu , Linfeng Zhang , Guolin Ke , Xi Fang

Understanding 3D medical image volumes is critical in the medical field, yet existing 3D medical convolution and transformer-based self-supervised learning (SSL) methods often lack deep semantic comprehension. Recent advancements in…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Qiuhui Chen , Xuancheng Yao , Huping Ye , Yi Hong

With the rise of multimodal applications, instruction data has become critical for training multimodal language models capable of understanding complex image-based queries. Existing practices rely on powerful but costly large language…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Jieyu Zhang , Le Xue , Linxin Song , Jun Wang , Weikai Huang , Manli Shu , An Yan , Zixian Ma , Juan Carlos Niebles , Silvio Savarese , Caiming Xiong , Zeyuan Chen , Ranjay Krishna , Ran Xu

Multimodal Large Language Models (MLLMs) have exhibited immense potential across numerous medical specialties; yet, dentistry remains underexplored, in part due to limited domain-specific data, scarce dental expert annotations, insufficient…

The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from various language communities. Nevertheless, the main challenge…

计算与语言 · 计算机科学 2024-02-20 Zhongwei Wan , Che Liu , Mi Zhang , Jie Fu , Benyou Wang , Sibo Cheng , Lei Ma , César Quilodrán-Casas , Rossella Arcucci

Multimodal Large Language Model (MLLM) has recently garnered attention as a prominent research focus. By harnessing powerful LLM, it facilitates a transition of conversational generative AI from unimodal text to performing multimodal tasks.…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Xuechen Guo , Wenhao Chai , Shi-Yan Li , Gaoang Wang

Multimodal large language models (MLLMs) have shown remarkable potential in various domains, yet their application in the medical field is hindered by several challenges. General-purpose MLLMs often lack the specialized knowledge required…

人工智能 · 计算机科学 2025-09-29 Guanghao Zhu , Zhitian Hou , Zeyu Liu , Zhijie Sang , Congkai Xie , Hongxia Yang

This study introduces an evaluation framework for multimodal models in medical imaging diagnostics. We developed a pipeline incorporating data preprocessing, model inference, and preference-based evaluation, expanding an initial set of 500…

图像与视频处理 · 电气工程与系统科学 2024-12-10 Cailian Ruan , Chengyue Huang , Yahe Yang

Multimodal large language models (MLLMs) hold promise for integrating diverse data modalities, but current medical adaptations such as LLaVA-Med often fail to fully exploit the synergy between color fundus photography (CFP) and optical…

The lack of sufficient annotated image data is a common issue in medical image segmentation. For some organs and densities, the annotation may be scarce, leading to poor model training convergence, while other organs have plenty of…

图像与视频处理 · 电气工程与系统科学 2021-09-22 Anastasia Makarevich , Azade Farshad , Vasileios Belagiannis , Nassir Navab

This paper aims to build a model that can Segment Anything in 3D medical images, driven by medical terminologies as Text prompts, termed as SAT. Our main contributions are three-fold: (i) We construct the first multimodal knowledge tree on…

图像与视频处理 · 电气工程与系统科学 2025-07-21 Ziheng Zhao , Yao Zhang , Chaoyi Wu , Xiaoman Zhang , Xiao Zhou , Ya Zhang , Yanfeng Wang , Weidi Xie

Recent advancements in multimodal foundation models have showcased impressive capabilities in understanding and reasoning with visual and textual information. Adapting these foundation models trained for general usage to specialized domains…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Hejie Cui , Lingjun Mao , Xin Liang , Jieyu Zhang , Hui Ren , Quanzheng Li , Xiang Li , Carl Yang