English
Related papers

Related papers: RadAgents: Multimodal Agentic Reasoning for Chest …

200 papers

Visual Language Models (VLMs) achieve promising results in medical reasoning but struggle with hallucinations, vague descriptions, inconsistent logic and poor localization. To address this, we propose a agent framework named Medical Visual…

Artificial Intelligence · Computer Science 2025-10-22 Guangfu Guo , Xiaoqian Lu , Yue Feng

Healthcare robotics requires robust multimodal perception and reasoning to ensure safety in dynamic clinical environments. Current Vision-Language Models (VLMs) demonstrate strong general-purpose capabilities but remain limited in temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Saurav Jha , Stefan K. Ehrlich

Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by…

Computation and Language · Computer Science 2024-10-08 Binxu Li , Tiankai Yan , Yuanting Pan , Jie Luo , Ruiyang Ji , Jiayuan Ding , Zhe Xu , Shilong Liu , Haoyu Dong , Zihao Lin , Yixin Wang

Heart diseases remain a leading cause of morbidity and mortality worldwide, necessitating accurate and trustworthy differential diagnosis. However, existing artificial intelligence-based diagnostic methods are often limited by insufficient…

Artificial intelligence is reshaping scientific exploration, but most methods automate procedural tasks without engaging in scientific reasoning, limiting autonomy in discovery. We introduce Materials Agents for Simulation and Theory in…

Understanding how two radiology image sets differ is critical for generating clinical insights and for interpreting medical AI systems. We introduce RadDiff, a multimodal agentic system that performs radiologist-style comparative reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Xiaoxian Shen , Yuhui Zhang , Sahithi Ankireddy , Xiaohan Wang , Maya Varma , Henry Guo , Curtis Langlotz , Serena Yeung-Levy

In clinical practice, radiology reporting is an essential yet complex, time-intensive, and error-prone task, particularly for 3D medical images. Existing automated approaches based on medical vision-language models primarily focus on…

Artificial Intelligence · Computer Science 2026-03-02 Yongrui Yu , Zhongzhen Huang , Linjie Mu , Shaoting Zhang , Xiaofan Zhang

Large scale vision language models have shown promise in automating chest Xray interpretation, yet their clinical utility remains limited by a gap between model outputs and radiologist reasoning. Most systems optimize for semantic…

Artificial Intelligence · Computer Science 2026-04-17 Kinhei Lee , Peiyuan Jing , Zhenxuan Zhang , Yue Yang , Tao Wang , Dominic C Marshall , Yingying Fang , Guang Yang

Current clinical agent built on small LLMs, such as TxAgent suffer from a \textit{Context Utilization Failure}, where models successfully retrieve biomedical evidence due to supervised finetuning but fail to ground their diagnosis in that…

Artificial Intelligence · Computer Science 2025-12-08 Ting-Ting Xie , Yixin Zhang

Cross-domain multimodal time series forecasting is a challenging task, requiring models to integrate precise numerical comprehension, cross-domain semantic understanding, and effective multimodal fusion. Existing approaches either build…

Artificial Intelligence · Computer Science 2026-05-29 Kun Feng , Ziwei Shan , Yuchen Fang , Yiyang Tan , Sihan Lu , Shuqi Gu , Lintao Ma , Xingyu Lu , Kan Ren

Retrieval-Augmented Generation (RAG) lifts the factuality of Large Language Models (LLMs) by injecting external knowledge, yet it falls short on problems that demand multi-step inference; conversely, purely reasoning-oriented approaches…

We introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents. Agentic Reasoning dynamically leverages web search, code execution, and structured memory to address…

Artificial Intelligence · Computer Science 2025-07-16 Junde Wu , Jiayuan Zhu , Yuyuan Liu , Min Xu , Yueming Jin

Chest imaging plays an essential role in diagnosing and predicting patients with COVID-19 with evidence of worsening respiratory status. Many deep learning-based approaches for pneumonia recognition have been developed to enable…

Image and Video Processing · Electrical Eng. & Systems 2024-01-17 Shengchao Chen , Sufen Ren , Guanjun Wang , Mengxing Huang , Chenyang Xue

We introduce PhysicalAgent, an agentic framework for robotic manipulation that integrates iterative reasoning, diffusion-based video generation, and closed-loop execution. Given a textual instruction, our method generates short video…

Traditional AI-based healthcare systems often rely on single-modal data, limiting diagnostic accuracy due to incomplete information. However, recent advancements in foundation models show promising potential for enhancing diagnosis…

Artificial Intelligence · Computer Science 2025-03-24 Sihan Wang , Suiyang Jiang , Yibo Gao , Boming Wang , Shangqi Gao , Xiahai Zhuang

Complex medical decision-making involves cooperative workflows operated by different clinicians. Designing AI multi-agent systems can expedite and augment human-level clinical decision-making. Existing multi-agent researches primarily focus…

Artificial Intelligence · Computer Science 2025-10-14 Kaitao Chen , Mianxin Liu , Daoming Zong , Chaoyue Ding , Shaohao Rui , Yankai Jiang , Mu Zhou , Xiaosong Wang

Medical imaging research is increasingly shifting from controlled benchmark evaluation toward real-world clinical deployment. In such settings, applying analytical methods extends beyond model design to require dataset-aware workflow…

Multiple myeloma is managed through sequential lines of therapy over years to decades, with each decision depending on cumulative disease history distributed across dozens to hundreds of heterogeneous clinical documents. Whether LLM-based…

Recent advances in large language models have enabled AI systems to achieve expert-level performance on domain-specific scientific tasks, yet these systems remain narrow and handcrafted. We introduce SciAgent, a unified multi-agent system…

Automated fetal ultrasound interpretation requires a workflow from visual perception, including plane recognition and anatomical segmentation, to clinical understanding, including biometric measurement and diagnostic reporting. However, the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xiaotian Hu , Mingxuan Liu , Junwei Huang , Kasidit Anmahapong , Yifei Chen , Yiming Huang , Xuguang Bai , Zihan Li , Hongjia Yang , Yingqi Hao , Hong Xu , Yu Jiang , Tian Tian , Yi Liao , Haibo Qu , Qiyuan Tian