English
Related papers

Related papers: RadAgents: Multimodal Agentic Reasoning for Chest …

200 papers

The rapid advancements in large language models (LLMs) have unlocked their potential for multimodal tasks, where text and visual data are processed jointly. However, applying LLMs to medical imaging, particularly for chest X-rays (CXR),…

Image and Video Processing · Electrical Eng. & Systems 2025-02-11 Nicholas Evans , Stephen Baker , Miles Reed

Purpose: Echocardiographic interpretation requires video-level reasoning and guideline-based measurement analysis, which current deep learning models for cardiac ultrasound do not support. We present EchoAgent, a framework that enables…

We introduce GenAgent, unifying visual understanding and generation through an agentic multimodal model. Unlike unified models that face expensive training costs and understanding-generation trade-offs, GenAgent decouples these capabilities…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Kaixun Jiang , Yuzheng Wang , Junjie Zhou , Pandeng Li , Zhihang Liu , Chen-Wei Xie , Zhaoyu Chen , Yun Zheng , Wenqiang Zhang

Orthopantomograms (OPGs) are the standard panoramic radiograph in dentistry, used for full-arch screening across multiple diagnostic tasks. While Vision Language Models (VLMs) now allow multi-task OPG analysis through natural language, they…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhaolin Yu , Litao Yang , Ben Babicka , Ming Hu , Jing Hao , Anthony Huang , James Huang , Yueming Jin , Jiasong Wu , Zongyuan Ge

Identifying associations between imaging phenotypes, disease risk factors, and clinical outcomes is essential for understanding disease mechanisms. However, traditional approaches rely on human-driven hypothesis testing and selection of…

Artificial Intelligence · Computer Science 2025-09-09 Weitong Zhang , Mengyun Qiao , Chengqi Zang , Steven Niederer , Paul M Matthews , Wenjia Bai , Bernhard Kainz

Chest X-rays (CXRs) are the most frequently performed imaging examinations in clinical settings. Recent advancements in Large Multimodal Models (LMMs) have enabled automated CXR interpretation, enhancing diagnostic accuracy and efficiency.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Qingqiu Li , Zihang Cui , Seongsu Bae , Jilan Xu , Runtian Yuan , Yuejie Zhang , Rui Feng , Quanli Shen , Xiaobo Zhang , Junjun He , Shujun Wang

Automated chest radiographs interpretation requires both accurate disease classification and detailed radiology report generation, presenting a significant challenge in the clinical workflow. Current approaches either focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Difei Gu , Yunhe Gao , Yang Zhou , Mu Zhou , Dimitris Metaxas

Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yize Liu , Siyuan Yan , Ming Hu , Lie Ju , Xieji Li , Feilong Tang , Wei Feng , Zongyuan Ge

Continuous physiological monitoring is central to emergency care, yet deploying trustworthy AI is challenging. While LLMs can translate complex physiological signals into clinical narratives, it is unclear how agentic systems perform…

Machine Learning · Computer Science 2026-03-05 Davide Gabrielli , Paola Velardi , Stefano Faralli , Bardh Prenkaj

We introduce DriveAgent, a novel multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion to enhance situational understanding and decision-making. DriveAgent…

Robotics · Computer Science 2025-05-06 Xinmeng Hou , Wuqi Wang , Long Yang , Hao Lin , Jinglun Feng , Haigen Min , Xiangmo Zhao

Chest X-rays (CXRs) are among the most frequently performed imaging examinations worldwide, yet rising imaging volumes increase radiologist workload and the risk of diagnostic errors. Although artificial intelligence (AI) systems have shown…

In the realm of chest X-ray (CXR) image analysis, radiologists meticulously examine various regions, documenting their observations in reports. The prevalence of errors in CXR diagnoses, particularly among inexperienced radiologists and…

Image and Video Processing · Electrical Eng. & Systems 2024-05-01 Akash Awasthi , Safwan Ahmad , Bryant Le , Hien Van Nguyen

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Noriyuki Kugo , Xiang Li , Zixin Li , Ashish Gupta , Arpandeep Khatua , Nidhish Jain , Chaitanya Patel , Yuta Kyuragi , Yasunori Ishii , Masamoto Tanabiki , Kazuki Kozuka , Ehsan Adeli

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization but limited…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Jing Zhang , Wentao Jiang , Tao Huang , Zhiwei Wang , Jianxin Liu , Jian Chen , Ping Ye , Gang Wang , Zengmao Wang , Bo Du , Dacheng Tao

Radiology is essential to modern healthcare, yet rising demand and staffing shortages continue to pose major challenges. Recent advances in artificial intelligence have the potential to support radiologists and help address these…

Image and Video Processing · Electrical Eng. & Systems 2025-11-14 Phillip Sloan , Edwin Simpson , Majid Mirmehdi

Precision therapeutics require multimodal adaptive models that generate personalized treatment recommendations. We introduce TxAgent, an AI agent that leverages multi-step reasoning and real-time biomedical knowledge retrieval across a…

Artificial Intelligence · Computer Science 2025-03-17 Shanghua Gao , Richard Zhu , Zhenglun Kong , Ayush Noori , Xiaorui Su , Curtis Ginder , Theodoros Tsiligkaridis , Marinka Zitnik

While Large Language Models (LLMs) have demonstrated potential in healthcare, they often struggle with the complex, non-linear reasoning required for accurate clinical diagnosis. Existing methods typically rely on static, linear mappings…

Computation and Language · Computer Science 2026-05-28 Zhuohan Ge , Haoyang Li , Yubo Wang , Nicole Hu , Chen Jason Zhang , Qing Li

Radiological imaging is central to diagnosis, treatment planning, and clinical decision-making. Vision-language foundation models have spurred interest in automated radiology report generation (RRG), but safe deployment requires reliable…

Recent advances in Large Vision-Language Models (LVLMs) have shown strong potential for multi-modal radiological reasoning, particularly in tasks like diagnostic visual question answering (VQA) and radiology report generation. However, most…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Yannian Gu , Xizhuo Zhang , Linjie Mu , Yongrui Yu , Zhongzhen Huang , Shaoting Zhang , Xiaofan Zhang

Recent pathological foundation models have substantially advanced visual representation learning and multimodal interaction. However, most models still rely on a static inference paradigm in which whole-slide images are processed once to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Shengyi Hua , Jianfeng Wu , Tianle Shen , Kangzhe Hu , Zhongzhen Huang , Shujuan Ni , Zhihong Zhang , Yuan Li , Zhe Wang , Xiaofan Zhang