English
Related papers

Related papers: MGA: Medical generalist agent through text-guided …

200 papers

Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in understanding common visual elements, largely due to their large-scale datasets and advanced training strategies. However, their effectiveness in medical…

Medical consultations are intrinsically speech-centric. However, most prior works focus on long-text-based interactions, which are cumbersome and patient-unfriendly. Recent advances in speech language models (SpeechLMs) have enabled more…

Computation and Language · Computer Science 2026-04-21 Sirry Chen , Jieyi Wang , Wei Chen , Zhongyu Wei

Memory-Augmented Generation (MAG) extends Large Language Models with external memory to support long-context reasoning, but existing approaches largely rely on semantic similarity over monolithic memory stores, entangling temporal, causal,…

Artificial Intelligence · Computer Science 2026-04-17 Dongming Jiang , Yi Li , Guanpeng Li , Bingzhe Li

Visual Question-Answering (VQA) is a challenging multimodal task that requires integrating visual and textual information to generate accurate responses. While multimodal Retrieval-Augmented Generation (mRAG) has shown promise in enhancing…

Computation and Language · Computer Science 2026-01-29 Zhuo Chen , Xinyu Geng , Xinyu Wang , Yong Jiang , Zhen Zhang , Pengjun Xie , Kewei Tu

The integration of deep learning-based glaucoma detection with large language models (LLMs) presents an automated strategy to mitigate ophthalmologist shortages and improve clinical reporting efficiency. However, applying general LLMs to…

Multiagent Systems · Computer Science 2025-12-18 Philip R. Liu , Sparsh Bansal , Jimmy Dinh , Aditya Pawar , Ramani Satishkumar , Shail Desai , Neeraj Gupta , Xin Wang , Shu Hu

Medical image-language pre-training aims to align medical images with clinically relevant text to improve model performance on various downstream tasks. However, existing models often struggle with the variability and ambiguity inherent in…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Shreyank N Gowda , Ruichi Zhang , Xiao Gu , Ying Weng , Lu Yang

Combining pre-trained expert models offers substantial potential for scalable multimodal reasoning, but building a unified framework remains challenging due to the increasing diversity of input modalities and task complexity. For instance,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Shoubin Yu , Yue Zhang , Ziyang Wang , Jaehong Yoon , Mohit Bansal

In the field of multimodal medical data analysis, leveraging diverse types of data and understanding their hidden relationships continues to be a research focus. The main challenges lie in effectively modeling the complex interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Xuhao Shan , Ruiquan Ge , Jikui Liu , Linglong Wu , Chi Zhang , Siqi Liu , Wenjian Qin , Wenwen Min , Ahmed Elazab , Changmiao Wang

Images in the medical domain are fundamentally different from the general domain images. Consequently, it is infeasible to directly employ general domain Visual Question Answering (VQA) models for the medical domain. Additionally, medical…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Yash Khare , Viraj Bagal , Minesh Mathew , Adithi Devi , U Deva Priyakumar , CV Jawahar

Decades' advances in digital health technologies, such as electronic health records, have largely streamlined routine clinical processes. Yet, most these systems are still hard to learn and use: Clinicians often face the burden of managing…

Artificial Intelligence · Computer Science 2025-09-16 Jared Zhu , Junde Wu

Artificial General Intelligence falls short when communicating role specific nuances to other systems. This is more pronounced when building autonomous LLM agents capable and designed to communicate with each other for real world problem…

Machine Learning · Computer Science 2024-03-19 Rabimba Karanjai , Weidong Shi

In healthcare intelligence, the ability to fuse heterogeneous, multi-intent information from diverse clinical sources is fundamental to building reliable decision-making systems. Large Language Model (LLM)-driven information interaction…

Computation and Language · Computer Science 2025-07-04 Dingkang Yang , Jinjie Wei , Mingcheng Li , Jiyao Liu , Lihao Liu , Ming Hu , Junjun He , Yakun Ju , Wei Zhou , Yang Liu , Lihua Zhang

We introduce a novel graph-based Retrieval-Augmented Generation (RAG) framework specifically designed for the medical domain, called \textbf{MedGraphRAG}, aimed at enhancing Large Language Model (LLM) capabilities for generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Junde Wu , Jiayuan Zhu , Yunli Qi , Jingkun Chen , Min Xu , Filippo Menolascina , Vicente Grau

While large-scale pretrained language models have significantly improved writing assistance functionalities such as autocomplete, more complex and controllable writing assistants have yet to be explored. We leverage advances in language…

Computation and Language · Computer Science 2021-09-21 Simeng Sun , Wenlong Zhao , Varun Manjunatha , Rajiv Jain , Vlad Morariu , Franck Dernoncourt , Balaji Vasan Srinivasan , Mohit Iyyer

Paucity of medical data severely limits the generalizability of diagnostic ML models, as the full spectrum of disease variability can not be represented by a small clinical dataset. To address this, diffusion models (DMs) have been…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Janet Wang , Yunbei Zhang , Zhengming Ding , Jihun Hamm

Large Language Models (LLMs) are becoming essential tools for various natural language processing tasks but often suffer from generating outdated or incorrect information. Retrieval-Augmented Generation (RAG) addresses this issue by…

Mobile agents show immense potential, yet current state-of-the-art (SoTA) agents exhibit inadequate success rates on real-world, long-horizon, cross-application tasks. We attribute this bottleneck to the agents' excessive reliance on…

Artificial Intelligence · Computer Science 2026-03-13 Yuxiang Zhou , Jichang Li , Yanhao Zhang , Haonan Lu , Guanbin Li

Recently, Multimodal Large Language Models (MLLMs) have been used as agents to control keyboard and mouse inputs by directly perceiving the Graphical User Interface (GUI) and generating corresponding commands. However, current agents…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Dongping Chen , Yue Huang , Siyuan Wu , Jingyu Tang , Liuyi Chen , Yilin Bai , Zhigang He , Chenlong Wang , Huichi Zhou , Yiqiang Li , Tianshuo Zhou , Yue Yu , Chujie Gao , Qihui Zhang , Yi Gui , Zhen Li , Yao Wan , Pan Zhou , Jianfeng Gao , Lichao Sun

The goal of automatic report generation is to generate a clinically accurate and coherent phrase from a single given X-ray image, which could alleviate the workload of traditional radiology reporting. However, in a real-world scenario,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Tiancheng Gu , Dongnan Liu , Zhiyuan Li , Weidong Cai

X-ray image-based medical report generation (MRG) is a pivotal area in artificial intelligence that can significantly reduce diagnostic burdens for clinicians and patient wait times. Existing MRG models predominantly rely on Large Language…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Mingzheng Zhang , Jinfeng Gao , Dan Xu , Jiangrui Yu , Yuhan Qiao , Lan Chen , Jin Tang , Xiao Wang