English
Related papers

Related papers: Eye-gaze Guided Multi-modal Alignment for Medical …

200 papers

Leveraging pre-trained visual language models has become a widely adopted approach for improving performance in downstream visual question answering (VQA) applications. However, in the specialized field of medical VQA, the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Gang Liu , Hongyang Li , Zerui He , Shenjun Zhong

Different medical imaging modalities capture diagnostic information at varying spatial resolutions, from coarse global patterns to fine-grained localized structures. However, most existing vision-language frameworks in the medical domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Shivang Chopra , Gabriela Sanchez-Rodriguez , Lingchao Mao , Andrew J Feola , Jing Li , Zsolt Kira

Human trajectory forecasting requires capturing the multimodal nature of pedestrian behavior. However, existing approaches suffer from prior misalignment. Their learned or fixed priors often fail to capture the full distribution of…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Chao Li , Rui Zhang , Siyuan Huang , Xian Zhong , Hongbo Jiang

Healthcare data now span EHRs, medical imaging, genomics, and wearable sensors, but most diagnostic models still process these modalities in isolation. This limits their ability to capture early, cross-modal disease signatures. This paper…

Machine Learning · Computer Science 2025-12-18 Md Talha Mohsin , Ismail Abdulrashid

World-wide-web, with the website and webpage as the main interface, facilitates the dissemination of important information. Hence it is crucial to optimize them for better user interaction, which is primarily done by analyzing users'…

Computer Vision and Pattern Recognition · Computer Science 2023-01-09 Ciheng Zhang , Decky Aspandi , Steffen Staab

Transformer-based deep learning models have demonstrated exceptional performance in medical imaging by leveraging attention mechanisms for feature representation and interpretability. However, these models are prone to learning spurious…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Shelley Zixin Shu , Haozhe Luo , Alexander Poellinger , Mauricio Reyes

Electronic health record (EHR) data is sparse and irregular as it is recorded at irregular time intervals, and different clinical variables are measured at each observation point. In this work, we propose a multi-view features integration…

Machine Learning · Computer Science 2021-01-27 Yurim Lee , Eunji Jun , Heung-Il Suk

Medical image reporting (MIR) aims to generate structured clinical descriptions from radiological images. Existing methods struggle with fine-grained feature extraction, multimodal alignment, and generalization across diverse imaging types,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-30 Amaan Izhar , Nurul Japar , Norisma Idris , Ting Dang

The rapid development of diagnostic technologies in healthcare is leading to higher requirements for physicians to handle and integrate the heterogeneous, yet complementary data that are produced during routine practice. For instance, the…

Machine Learning · Computer Science 2023-01-30 Can Cui , Haichun Yang , Yaohong Wang , Shilin Zhao , Zuhayr Asad , Lori A. Coburn , Keith T. Wilson , Bennett A. Landman , Yuankai Huo

Multimodal tabular-image fusion is an emerging task that has received increasing attention in various domains. However, existing methods may be hindered by gradient conflicts between modalities, misleading the optimization of the unimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Longfei Huang , Yang Yang

Understanding how novices acquire and hone visual search skills is crucial for developing and optimizing training methods across domains. Network analysis methods can be used to analyze graph representations of visual expertise. This study…

Human-Computer Interaction · Computer Science 2025-07-28 Pingjing Yang , Jennifer Cromley , Jana Diesner

Medicine is inherently a multimodal discipline. Medical images can reflect the pathological changes of cancer and tumors, while the expression of specific genes can influence their morphological characteristics. However, most deep learning…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Jiaying Zhou , Mingzhou Jiang , Junde Wu , Jiayuan Zhu , Ziyue Wang , Yueming Jin

The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Xun Zhu , Fanbin Mo , Zheng Zhang , Jiaxi Wang , Yiming Shi , Ming Wu , Chuang Zhang , Miao Li , Ji Wu

Recent advancements in medical Large Language Models (LLMs) have showcased their powerful reasoning and diagnostic capabilities. Despite their success, current unified multimodal medical LLMs face limitations in knowledge update costs,…

Computation and Language · Computer Science 2025-06-25 Yucheng Zhou , Lingran Song , Jianbing Shen

Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed…

Machine Learning · Computer Science 2022-05-03 Ahmed Abdou , Ekta Sood , Philipp Müller , Andreas Bulling

Medical image analysis often faces significant challenges due to limited expert-annotated data, hindering both model generalization and clinical adoption. We propose an expert-guided explainable few-shot learning framework that integrates…

Image and Video Processing · Electrical Eng. & Systems 2025-09-12 Ifrat Ikhtear Uddin , Longwei Wang , KC Santosh

Medical AI assistants support doctors in disease diagnosis, medical image analysis, and report generation. However, they still face significant challenges in clinical use, including limited accuracy with multimodal content and insufficient…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Haonan Wang , Jiaji Mao , Lehan Wang , Qixiang Zhang , Marawan Elbatel , Yi Qin , Huijun Hu , Baoxun Li , Wenhui Deng , Weifeng Qin , Hongrui Li , Jialin Liang , Jun Shen , Xiaomeng Li

Human eye gaze plays a significant role in many virtual and augmented reality (VR/AR) applications, such as gaze-contingent rendering, gaze-based interaction, or eye-based activity recognition. However, prior works on gaze analysis and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Zhiming Hu , Jiahui Xu , Syn Schmitt , Andreas Bulling

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models…

Computation and Language · Computer Science 2025-09-17 Ali Salamatian , Amirhossein Abaskohi , Wan-Cyuan Fan , Mir Rayat Imtiaz Hossain , Leonid Sigal , Giuseppe Carenini

Automated interpretation of electrocardiograms (ECG) has garnered significant attention with the advancements in machine learning methodologies. Despite the growing interest, most current studies focus solely on classification or regression…

Signal Processing · Electrical Eng. & Systems 2023-11-07 Jielin Qiu , Jiacheng Zhu , Shiqi Liu , William Han , Jingqi Zhang , Chaojing Duan , Michael Rosenberg , Emerson Liu , Douglas Weber , Ding Zhao
‹ Prev 1 8 9 10 Next ›