English
Related papers

Related papers: CT-Agent: A Multimodal-LLM Agent for 3D CT Radiolo…

200 papers

Visual Question Answering (VQA) attracts much attention from both industry and academia. As a multi-modality task, it is challenging since it requires not only visual and textual understanding, but also the ability to align cross-modality…

Computer Vision and Pattern Recognition · Computer Science 2022-01-27 Peixi Xiong , Quanzeng You , Pei Yu , Zicheng Liu , Ying Wu

Large language models show potential for scalable mental-health support by simulating Cognitive Behavioral Therapy (CBT) counselors. However, existing methods often rely on static cognitive profiles and omniscient single-agent simulation,…

Computation and Language · Computer Science 2026-04-09 Chang Liu , Changsheng Ma , Yongfeng Tao , Bin Hu , Minqiang Yang

Visual Question Answering (VQA) presents a unique challenge as it requires the ability to understand and encode the multi-modal inputs - in terms of image processing and natural language processing. The algorithm further needs to learn how…

Computer Vision and Pattern Recognition · Computer Science 2017-09-26 Supriya Pandhre , Shagun Sodhani

Automated interpretation of CT images-particularly localizing and describing abnormal findings across multi-plane and whole-body scans-remains a significant challenge in clinical radiology. This work aims to address this challenge through…

Image and Video Processing · Electrical Eng. & Systems 2025-11-18 Ziheng Zhao , Lisong Dai , Ya Zhang , Yanfeng Wang , Weidi Xie

Complex table question answering (TQA) aims to answer questions that require complex reasoning, such as multi-step or multi-category reasoning, over data represented in tabular form. Previous approaches demonstrated notable performance by…

Computation and Language · Computer Science 2025-02-11 Wei Zhou , Mohsen Mesgar , Annemarie Friedrich , Heike Adel

Dermatological diagnosis requires integrating fine-grained visual perception with expert clinical knowledge. Although Multimodal Large Language Models (MLLMs) facilitate interactive medical image analysis, their application in dermatology…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yize Liu , Siyuan Yan , Ming Hu , Lie Ju , Xieji Li , Feilong Tang , Wei Feng , Zongyuan Ge

Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent self-attention based methods struggle to effectively handle cross-modal semantic alignments between vision and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Qiangguo Jin , Xianyao Zheng , Hui Cui , Changming Sun , Yuqi Fang , Cong Cong , Ran Su , Leyi Wei , Ping Xuan , Junbo Wang

Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent methods fail to fully leverage domain-specific medical knowledge, making it difficult to accurately associate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Xianyao Zheng , Hong Yu , Hui Cui , Changming Sun , Xiangyu Li , Ran Su , Leyi Wei , Jia Zhou , Junbo Wang , Qiangguo Jin

Chest X-ray plays a central role in thoracic diagnosis, and its interpretation inherently requires multi-step, evidence-grounded reasoning. However, large vision-language models (LVLMs) often generate plausible responses that are not…

Artificial Intelligence · Computer Science 2026-03-25 Hyungyung Lee , Hangyul Yoon , Edward Choi

Radiology report generation (RRG) aims to automatically produce diagnostic reports from medical images, with the potential to enhance clinical workflows and reduce radiologists' workload. While recent approaches leveraging multimodal large…

Artificial Intelligence · Computer Science 2025-05-16 Ziruo Yi , Ting Xiao , Mark V. Albert

This paper proposes CQ-VQA, a novel 2-level hierarchical but end-to-end model to solve the task of visual question answering (VQA). The first level of CQ-VQA, referred to as question categorizer (QC), classifies questions to reduce the…

Computer Vision and Pattern Recognition · Computer Science 2020-02-18 Aakansha Mishra , Ashish Anand , Prithwijit Guha

Cardiovascular diseases (CVDs) remain the foremost cause of mortality worldwide, a burden worsened by a severe deficit of healthcare workers. Artificial intelligence (AI) agents have shown potential to alleviate this gap through automated…

AI agents with tool-use capabilities show promise for integrating the domain expertise of various tools. In the medical field, however, tools are usually AI models that are inherently error-prone and can produce contradictory responses.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Zheang Huai , Honglong Yang , Xiaomeng Li

Visual Question Answering (VQA) is challenging due to the complex cross-modal relations. It has received extensive attention from the research community. From the human perspective, to answer a visual question, one needs to read the…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Hantao Huang , Tao Han , Wei Han , Deep Yap , Cheng-Ming Chiang

Recent advances in 3D medical vision-language models have enabled joint reasoning over volumetric images and text, showing strong performance in medical visual question-answering (VQA) and report generation. Despite this progress, it…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Mashrafi Monon , Umaima Rahman , Asif Hanif , Numan Saeed , Mohammad Yaqub

Visual Question Answering (VQA) requires integration of feature maps with drastically different structures and focus of the correct regions. Image descriptors have structures at multiple spatial scales, while lexical inputs inherently…

Computer Vision and Pattern Recognition · Computer Science 2018-07-20 Yang Shi , Tommaso Furlanello , Sheng Zha , Animashree Anandkumar

Recent advancements in Large Language Models (LLMs) have expanded their capabilities to multimodal contexts, including comprehensive video understanding. However, processing extensive videos such as 24-hour CCTV footage or full-length films…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Lu Zhang , Tiancheng Zhao , Heting Ying , Yibo Ma , Kyusong Lee

Recent advancements in Video Question Answering (VideoQA) have introduced LLM-based agents, modular frameworks, and procedural solutions, yielding promising results. These systems use dynamic agents and memory-based mechanisms to break down…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Tony Montes , Fernando Lozano

Automatic Chart Question Answering (ChartQA) is challenging due to the complex distribution of chart elements with patterns of the underlying data not explicitly displayed in charts. To address this challenge, we design a joint multimodal…

Computation and Language · Computer Science 2024-08-12 Yue Dai , Soyeon Caren Han , Wei Liu

We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yue Fan , Xiaojian Ma , Rujie Wu , Yuntao Du , Jiaqi Li , Zhi Gao , Qing Li