中文
相关论文

相关论文: MedTrinity-25M: A Large-scale Multimodal Dataset w…

200 篇论文

Synthetic Aperture Radar (SAR) and optical imagery provide complementary strengths that constitute the critical foundation for transcending single-modality constraints and facilitating cross-modal collaborative processing and intelligent…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Peihao Wu , Yongxiang Yao , Yi Wan , Wenfei Zhang , Ruipeng Zhao , Jiayuan Li , Yongjun Zhang

This paper introduces a new challenge and datasets to foster research toward designing systems that can understand medical videos and provide visual answers to natural language questions. We believe medical videos may provide the best…

计算机视觉与模式识别 · 计算机科学 2022-02-01 Deepak Gupta , Kush Attal , Dina Demner-Fushman

In biomedical vision-language modeling, datasets are typically mined from scientific literature, pairing compound figures with captions that are short, context-dependent, and oftern partially informative. Prior work on subfigure extraction…

计算机视觉与模式识别 · 计算机科学 2025-12-08 Negin Baghbanzadeh , Mohammed Saidul Islam , Sajad Ashkezari , Elham Dolatabadi , Arash Afkanpour

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

计算机视觉与模式识别 · 计算机科学 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

Deep learning in medical imaging often requires large-scale, high-quality data or initiation with suitably pre-trained weights. However, medical datasets are limited by data availability, domain-specific knowledge, and privacy concerns, and…

Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform diverse tasks are limited by the (usually rather small)…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Roman Bachmann , Oğuzhan Fatih Kar , David Mizrahi , Ali Garjani , Mingfei Gao , David Griffiths , Jiaming Hu , Afshin Dehghan , Amir Zamir

Vision-Language Models (VLMs) have demonstrated significant potential in medical image analysis, yet their application in intraoral photography remains largely underexplored due to the lack of fine-grained, annotated datasets and…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Meng-Xun Li , Wen-Hui Deng , Zhi-Xing Wu , Chun-Xiao Jin , Jia-Min Wu , Yue Han , James Kit Hon Tsoi , Gui-Song Xia , Cui Huang

Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Rongsheng Wang , Junying Chen , Ke Ji , Zhenyang Cai , Shunian Chen , Yunjin Yang , Benyou Wang

Medical Visual Question Answering (MedVQA), which offers language responses to image-based medical inquiries, represents a challenging task and significant advancement in healthcare. It assists medical experts to swiftly interpret medical…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Xiaotang Gai , Chenyi Zhou , Jiaxiang Liu , Yang Feng , Jian Wu , Zuozhu Liu

Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is crucial for monitoring disease progression and guiding clinical decisions. Many recent automated…

计算与语言 · 计算机科学 2026-01-26 Xinyi Wang , Grazziela Figueredo , Ruizhe Li , Xin Chen

The application of machine learning in medicine and healthcare has led to the creation of numerous diagnostic and prognostic models. However, despite their success, current approaches generally issue predictions using data from a single…

机器学习 · 计算机科学 2025-05-13 Fergus Imrie , Stefan Denner , Lucas S. Brunschwig , Klaus Maier-Hein , Mihaela van der Schaar

Modern medical records include a vast amount of multimodal free text clinical data and imaging data from radiology, cardiology, and digital pathology. Fully mining such big data requires multitasking; otherwise, occult but important aspects…

图像与视频处理 · 电气工程与系统科学 2024-04-25 Chuang Niu , Qing Lyu , Christopher D. Carothers , Parisa Kaviani , Josh Tan , Pingkun Yan , Mannudeep K. Kalra , Christopher T. Whitlow , Ge Wang

Vision-and-language multi-modal pretraining and fine-tuning have shown great success in visual question answering (VQA). Compared to general domain VQA, the performance of biomedical VQA suffers from limited data. In this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Zheng Yuan , Qiao Jin , Chuanqi Tan , Zhengyun Zhao , Hongyi Yuan , Fei Huang , Songfang Huang

Existing Medical Large Vision-Language Models (Med-LVLMs), encapsulating extensive medical knowledge, demonstrate excellent capabilities in understanding medical images. However, there remain challenges in visual localization in medical…

计算与语言 · 计算机科学 2025-06-03 Yucheng Zhou , Lingran Song , Jianbing Shen

Multimodal deep learning has shown strong potential in medical applications by integrating heterogeneous data sources such as medical images and structured clinical variables. However, most existing approaches implicitly assume complete…

机器学习 · 计算机科学 2026-05-13 Camillo Maria Caruso , Valerio Guarrasi , Paolo Soda

In this study, we aim to initiate the development of Radiology Foundation Model, termed as RadFM. We consider the construction of foundational models from three perspectives, namely, dataset construction, model design, and thorough…

计算机视觉与模式识别 · 计算机科学 2023-11-17 Chaoyi Wu , Xiaoman Zhang , Ya Zhang , Yanfeng Wang , Weidi Xie

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in various multimodal tasks. However, their potential in the medical domain remains largely unexplored. A significant challenge arises from the scarcity of…

图像与视频处理 · 电气工程与系统科学 2024-04-23 Yutao Hu , Tianbin Li , Quanfeng Lu , Wenqi Shao , Junjun He , Yu Qiao , Ping Luo

Self-supervised learning has greatly facilitated medical image analysis by suppressing the training data requirement for real-world applications. Current paradigms predominantly rely on self-supervision within uni-modal image data, thereby…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Shaohao Rui , Lingzhi Chen , Zhenyu Tang , Lilong Wang , Mianxin Liu , Shaoting Zhang , Xiaosong Wang

The development of multi-label deep learning models for retinal disease classification is often hindered by the scarcity of large, expertly annotated clinical datasets due to patient privacy concerns and high costs. The recent release of…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Jerry Cao-Xue , Tien Comlekoglu , Keyi Xue , Guanliang Wang , Jiang Li , Gordon Laurie

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch