English
Related papers

Related papers: Diagnosing Shoulder Disorders Using Multimodal Lar…

200 papers

The Large Visual Language Models (LVLMs) enhances user interaction and enriches user experience by integrating visual modality on the basis of the Large Language Models (LLMs). It has demonstrated their powerful information processing and…

Artificial Intelligence · Computer Science 2024-10-22 Wei Lan , Wenyi Chen , Qingfeng Chen , Shirui Pan , Huiyu Zhou , Yi Pan

Large Vision Language Models (LVLMs) are becoming increasingly important in the medical domain, yet Medical LVLMs (Med-LVLMs) frequently generate hallucinations due to limited expertise and the complexity of medical applications. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Aofei Chang , Le Huang , Parminder Bhatia , Taha Kass-Hout , Fenglong Ma , Cao Xiao

The rapid advancement of Multimodal Large Language Models (MLLMs) has significantly impacted various multimodal tasks. However, these models face challenges in tasks that require spatial understanding within 3D environments. Efforts to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Duo Zheng , Shijia Huang , Liwei Wang

Combining multiple modalities carrying complementary information through multimodal learning (MML) has shown considerable benefits for diagnosing multiple pathologies. However, the robustness of multimodal models to missing modalities is…

Machine Learning · Computer Science 2024-07-31 Hava Chaptoukaev , Vincenzo Marcianó , Francesco Galati , Maria A. Zuluaga

Multimodal Vision Language Models (VLMs) have emerged as a transformative topic at the intersection of computer vision and natural language processing, enabling machines to perceive and reason about the world through both visual and textual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Zongxia Li , Xiyang Wu , Hongyang Du , Fuxiao Liu , Huy Nghiem , Guangyao Shi

The rising global prevalence of skin conditions, some of which can escalate to life-threatening stages if not timely diagnosed and treated, presents a significant healthcare challenge. This issue is particularly acute in remote areas where…

Computer Vision and Pattern Recognition · Computer Science 2024-02-19 Mahapara Khurshid , Mayank Vatsa , Richa Singh

Brain disorders are a major challenge to global health, causing millions of deaths each year. Accurate diagnosis of these diseases relies heavily on advanced medical imaging techniques such as Magnetic Resonance Imaging (MRI) and Computed…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Xuran Zhu

Recently, Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in visual understanding and reasoning across various vision-language tasks. However, we found that MLLMs cannot process effectively from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Bangyan Li , Wenxuan Huang , Zhenkun Gao , Yeqiang Wang , Yunhang Shen , Jingzhong Lin , Ling You , Yuxiang Shen , Shaohui Lin , Wanli Ouyang , Yuling Sun

Chronic diseases, including diabetes, hypertension, asthma, HIV-AIDS, epilepsy, and tuberculosis, necessitate rigorous adherence to medication to avert disease progression, manage symptoms, and decrease mortality rates. Adherence is…

Computer Vision and Pattern Recognition · Computer Science 2025-11-20 Md Asaduzzaman Jabin , Hanqi Jiang , Yiwei Li , Patrick Kaggwa , Eugene Douglass , Juliet N. Sekandi , Tianming Liu

Large language models (LLMs) have rapidly evolved from general-purpose systems to multimodal models capable of processing text, images, and audio. As both general-purpose LLMs (GLLMs) and multimodal LLMs (MLLMs) gain widespread adoption,…

Software Engineering · Computer Science 2026-04-08 Yujian Liu , Xiao Yu , Jacky Keung , Xing Hu , Xin Xia , Xiaoxue Ma

Multimodal large language models (MLLMs) can simultaneously process visual, textual, and auditory data, capturing insights that complement human analysis. However, existing video question-answering (VidQA) benchmarks and datasets often…

Machine Learning · Computer Science 2024-12-23 Jean Park , Kuk Jin Jang , Basam Alasaly , Sriharsha Mopidevi , Andrew Zolensky , Eric Eaton , Insup Lee , Kevin Johnson

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

Artificial Intelligence · Computer Science 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

Dense prediction is a fundamental requirement for many medical vision tasks such as medical image restoration, registration, and segmentation. The most popular vision model, Convolutional Neural Networks (CNNs), has reached bottlenecks due…

Image and Video Processing · Electrical Eng. & Systems 2023-11-29 Mingyuan Meng , Yuxin Xue , Dagan Feng , Lei Bi , Jinman Kim

With the rapid progress of large language models (LLMs), advanced multimodal large language models (MLLMs) have demonstrated impressive zero-shot capabilities on vision-language tasks. In the biomedical domain, however, even…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Siyuan Dai , Lunxiao Li , Kun Zhao , Eardi Lila , Paul K. Crane , Heng Huang , Dongkuan Xu , Haoteng Tang , Liang Zhan

Multimodal large language models (MLLMs) hold significant potential in medical applications, including disease diagnosis and clinical decision-making. However, these tasks require highly accurate, context-sensitive, and professionally…

Computation and Language · Computer Science 2025-09-01 Meidan Ding , Jipeng Zhang , Wenxuan Wang , Cheng-Yi Li , Wei-Chieh Fang , Hsin-Yu Wu , Haiqin Zhong , Wenting Chen , Linlin Shen

The challenge of Multimodal Deformable Image Registration (MDIR) lies in the conversion and alignment of features between images of different modalities. Generative models (GMs) cannot retain the necessary information enough from the source…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Mingrui Ma , Weijie Wang , Jie Ning , Jianfeng He , Nicu Sebe , Bruno Lepri

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Kanchana Ranasinghe , Xiang Li , Kumara Kahatapitiya , Michael S. Ryoo

Major depressive disorder (MDD) is a complex psychiatric disorder that affects the lives of hundreds of millions of individuals around the globe. Even today, researchers debate if morphological alterations in the brain are linked to MDD,…

Quantitative Methods · Quantitative Biology 2025-01-27 Roberto Goya-Maldonado , Tracy Erwin-Grabner , Ling-Li Zeng , Christopher R. K. Ching , Andre Aleman , Alyssa R. Amod , Zeynep Basgoze , Francesco Benedetti , Bianca Besteher , Katharina Brosch , Robin Bülow , Romain Colle , Colm G. Connolly , Emmanuelle Corruble , Baptiste Couvy-Duchesne , Kathryn Cullen , Udo Dannlowski , Christopher G. Davey , Annemiek Dols , Jan Ernsting , Jennifer W. Evans , Lukas Fisch , Paola Fuentes-Claramonte , Ali Saffet Gonul , Ian H. Gotlib , Hans J. Grabe , Nynke A. Groenewold , Dominik Grotegerd , Tim Hahn , J. Paul Hamilton , Laura K. M. Han , Ben J. Harrison , Tiffany C. Ho , Neda Jahanshad , Alec J. Jamieson , Andriana Karuk , Tilo Kircher , Bonnie Klimes-Dougan , Sheri-Michelle Koopowitz , Thomas Lancaster , Ramona Leenings , Meng Li , David E. J. Linden , Frank P. MacMaster , David M. A. Mehler , Susanne Meinert , Elisa Melloni , Bryon A. Mueller , Benson Mwangi , Igor Nenadić , Amar Ojha , Yasumasa Okamoto , Mardien L. Oudega , Brenda W. J. H. Penninx , Sara Poletti , Edith Pomarol-Clotet , Maria J. Portella , Elena Pozzi , Joaquim Radua , Elena Rodríguez-Cano , Matthew D. Sacchet , Raymond Salvador , Anouk Schrantee , Kang Sim , Jair C. Soares , Aleix Solanes , Dan J. Stein , Frederike Stein , Aleks Stolicyn , Sophia I. Thomopoulos , Yara J. Toenders , Aslihan Uyar-Demir , Eduard Vieta , Yolanda Vives-Gilabert , Henry Völzke , Martin Walter , Heather C. Whalley , Sarah Whittle , Nils Winter , Katharina Wittfeld , Margaret J. Wright , Mon-Ju Wu , Tony T. Yang , Carlos Zarate , Dick J. Veltman , Lianne Schmaal , Paul M. Thompson

Deep learning (DL) networks have recently been shown to outperform other segmentation methods on various public, medical-image challenge datasets [3,11,16], especially for large pathologies. However, in the context of diseases such as…

Computer Vision and Pattern Recognition · Computer Science 2018-10-18 Tanya Nair , Doina Precup , Douglas L. Arnold , Tal Arbel

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales