中文
相关论文

相关论文: CephGPT-4: An Interactive Multimodal Cephalometric…

200 篇论文

Vision-Language Foundation Models (VLMs), trained on large-scale multimodal datasets, have driven significant advances in Artificial Intelligence (AI) by enabling rich cross-modal reasoning. Despite their success in general domains,…

The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in the quantity and…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Junying Chen , Chi Gui , Ruyi Ouyang , Anningzhe Gao , Shunian Chen , Guiming Hardy Chen , Xidong Wang , Ruifei Zhang , Zhenyang Cai , Ke Ji , Guangjun Yu , Xiang Wan , Benyou Wang

Though Large Vision-Language Models (LVLMs) are being actively explored in medicine, their ability to conduct complex real-world telemedicine consultations combining accurate diagnosis with professional dialogue remains underexplored. This…

Multimodal medical image fusion plays a crucial role in medical diagnosis by integrating complementary information from different modalities to enhance image readability and clinical applicability. However, existing methods mainly follow…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Haozhe Xiang , Han Zhang , Yu Cheng , Xiongwen Quan , Wanwan Huang

A significant application of Large Language Models (LLMs), like ChatGPT, is their deployment as chat agents, which respond to human inquiries across a variety of domains. While current LLMs proficiently answer general questions, they often…

计算与语言 · 计算机科学 2024-04-16 Lang Cao

Large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation across various domains, including medicine. We present a comprehensive evaluation of GPT-4, a state-of-the-art LLM, on…

计算与语言 · 计算机科学 2023-04-13 Harsha Nori , Nicholas King , Scott Mayer McKinney , Dean Carignan , Eric Horvitz

Recent advances in medical multi-modal models focus on specialized image analysis like dermatology, pathology, or radiology. However, they do not fully capture the complexity of real-world clinical diagnostics, which involve heterogeneous…

计算机视觉与模式识别 · 计算机科学 2026-01-13 Jiao Xu , Junwei Liu , Jiangwei Lao , Qi Zhu , Yunpeng Zhao , Congyun Jin , Shinan Liu , Zhihong Lu , Lihe Zhang , Xin Chen , Jian Wang , Ping Wang

Multilingual capability is an essential aspect for large multimodal models, since they are usually deployed across various countries and languages. However, most existing benchmarks for multilingual multimodal reasoning struggle to…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Hongyu Wang , Jiayu Xu , Senwei Xie , Ruiping Wang , Jialin Li , Zhaojie Xie , Bin Zhang , Chuyan Xiong , Xilin Chen

We present MindGPT-4ov, a multimodal large language model (MLLM) that introduces a general post-training paradigm spanning data production, model training, and efficient deployment. It achieves state-of-the-art performance across multiple…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Wei Chen , Chaoqun Du , Feng Gu , Wei He , Qizhen Li , Zide Liu , Xuhao Pan , Chang Ren , Xudong Rao , Chenfeng Wang , Tao Wei , Chengjun Yu , Pengfei Yu , Yufei Zheng , Chunpeng Zhou , Pan Zhou , Xuhan Zhu

We introduce AnyGPT, an any-to-any multimodal language model that utilizes discrete representations for the unified processing of various modalities, including speech, text, images, and music. AnyGPT can be trained stably without any…

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them…

人工智能 · 计算机科学 2025-09-11 Ruiqi Wu , Yuang Yao , Tengfei Ma , Chenran Zhang , Na Su , Tao Zhou , Geng Chen , Wen Fan , Yi Zhou

Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are predominantly trained on professional literature, limiting their ability to communicate…

计算与语言 · 计算机科学 2026-04-08 Han Jang , Junhyeok Lee , Heeseong Eum , Kyu Sung Choi

Dementia, a progressive neurodegenerative disorder, affects memory, reasoning, and daily functioning, creating challenges for individuals and healthcare systems. Early detection is crucial for timely interventions that may slow disease…

神经元与认知 · 定量生物学 2025-03-04 Sahar Sinene Mehdoui , Abdelhamid Bouzid , Daniel Sierra-Sosa , Adel Elmaghraby

This paper presents a Multilingual Vision Large Language Model, named M-MiniGPT4. Our model exhibits strong vision-language understanding (VLU) capabilities across 11 languages. We utilize a mixture of native multilingual and translated…

计算与语言 · 计算机科学 2026-04-01 Seung Hun Han , Youssef Mohamed , Mohamed Elhoseiny

Medical image analysis is essential to clinical diagnosis and treatment, which is increasingly supported by multi-modal large language models (MLLMs). However, previous research has primarily focused on 2D medical images, leaving 3D images…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Fan Bai , Yuxin Du , Tiejun Huang , Max Q. -H. Meng , Bo Zhao

The growing convergence between Large Language Models (LLMs) and electroencephalography (EEG) research is enabling new directions in neural decoding, brain-computer interfaces (BCIs), and affective computing. This survey offers a systematic…

信号处理 · 电气工程与系统科学 2025-06-11 Naseem Babu , Jimson Mathew , A. P. Vinod

Large Language Models (LLMs), enhanced through agent tuning, have demonstrated remarkable capabilities in Chain-of-Thought (CoT) and tool utilization, significantly surpassing the performance of standalone models. However, the multimodal…

计算机视觉与模式识别 · 计算机科学 2025-07-30 Tianhong Gao , Yannian Fu , Weiqun Wu , Haixiao Yue , Shanshan Liu , Gang Zhang

The conventional pretraining-and-finetuning paradigm, while effective for common diseases with ample data, faces challenges in diagnosing data-scarce occupational diseases like pneumoconiosis. Recently, large language models (LLMs) have…

Large Language Models (LLMs) such as GPT developed by OpenAI, have already shown astonishing results, introducing quick changes in our society. This has been intensified by the release of ChatGPT which allows anyone to interact in a simple…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Ivan DeAndres-Tame , Ruben Tolosana , Ruben Vera-Rodriguez , Aythami Morales , Julian Fierrez , Javier Ortega-Garcia

Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cost. This limits their widespread deployment for dental screening outside specialist…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Kai Bian , Xucheng Guo , Bin Chen , Lingyan Ruan , Yiran Shen , Ting Dang , Hong Jia