中文
相关论文

相关论文: GazeVaLM: A Multi-Observer Eye-Tracking Benchmark …

200 篇论文

Appearance-based gaze estimation aims to predict the 3D eye gaze direction from a single image. While recent deep learning-based approaches have demonstrated excellent performance, they usually assume one calibrated face in each input image…

计算机视觉与模式识别 · 计算机科学 2022-04-21 Mingfang Zhang , Yunfei Liu , Feng Lu

This paper proposes a novel multimodal DL architecture incorporating medical images and eye-tracking data for abnormality detection in chest x-rays. Our results show that applying eye gaze data directly into DL architectures does not show…

计算机视觉与模式识别 · 计算机科学 2023-02-07 André Luís , Chihcheng Hsieh , Isabel Blanco Nobre , Sandra Costa Sousa , Anderson Maciel , Catarina Moreira , Joaquim Jorge

While numerous recent benchmarks focus on evaluating generic Vision-Language Models (VLMs), they do not effectively address the specific challenges of geospatial applications. Generic VLM benchmarks are not designed to handle the…

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and…

Translating natural language to visualization (NL2VIS) has shown great promise for visual data analysis, but it remains a challenging task that requires multiple low-level implementations, such as natural language processing and…

人机交互 · 计算机科学 2024-08-08 Nan Chen , Yuge Zhang , Jiahang Xu , Kan Ren , Yuqing Yang

The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop…

计算机视觉与模式识别 · 计算机科学 2025-09-05 Yuanfeng Ji , Dan Lin , Xiyue Wang , Lu Zhang , Wenhui Zhou , Chongjian Ge , Ruihang Chu , Xiaoli Yang , Junhan Zhao , Junsong Chen , Xiangde Luo , Sen Yang , Jin Fang , Ping Luo , Ruijiang Li

Progress in remote PhotoPlethysmoGraphy (rPPG) is limited by the critical issues of existing publicly available datasets: small size, privacy concerns with facial videos, and lack of diversity in conditions. The paper introduces a novel…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Konstantin Egorov , Stepan Botman , Pavel Blinov , Galina Zubkova , Anton Ivaschenko , Alexander Kolsanov , Andrey Savchenko

Anatomical structure labels for chest radiographs are essential for medical image segmentation and a broad range of downstream diagnostic tasks. However, annotating anatomy directly on 2D chest radiographs is labor-intensive and…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Shuchang Ye , Mingyuan Meng , Hao Wang , Usman Naseem , Jinman Kim

Vision-language models (VLMs) have shown strong promise for medical image analysis, but most remain opaque, offering predictions without the transparent, stepwise reasoning clinicians rely on. We present a framework that brings…

Comprehensively interpreting human behavior is a core challenge in human-aware artificial intelligence. However, prior works typically focused on body behavior, neglecting the crucial role of eye gaze and its synergy with body motion. We…

人机交互 · 计算机科学 2025-11-21 Qing Chang , Zhiming Hu

Multimodal large language models (MLLMs) have shown remarkable performance in vision-language tasks. However, existing MLLMs are primarily trained on generic datasets, limiting their ability to reason on domain-specific visual cues such as…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Hatef Otroshi Shahreza , Sébastien Marcel

Echocardiography is the most widely used imaging modality in cardiology, yet its interpretation remains labor-intensive and inherently multimodal, requiring view recognition, quantitative measurements, qualitative assessments, and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Yuheng Li , Yue Zhang , Abdoul Aziz Amadou , Yuxiang Lai , Jike Zhong , Tiziano Passerini , Dorin Comaniciu , Puneet Sharma

Eye gaze, encompassing fixations and saccades, provides critical insights into human intentions and future actions. This study introduces a gaze-regularized framework that enhances Vision Language Models (VLMs) for egocentric behavior…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Anupam Pani , Yanchao Yang

Mammography screening is an essential tool for early detection of breast cancer. The speed and accuracy of mammography interpretation have the potential to be improved with deep learning methods. However, the development of a foundation…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Yuexi Du , Lihui Chen , Nicha C. Dvornek

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through…

人机交互 · 计算机科学 2025-04-09 Jun Rekimoto

Visual Language Models (VLMs) are now sufficiently advanced to support a broad range of applications, including answering complex visual questions, and are increasingly expected to interact with images in varied ways. To evaluate them,…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Ludovic Arnould , Salim Khazem , Hugues Ali Mehenni

AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no standardized benchmark for objectively evaluating their performance. To address this, we present…

计算机视觉与模式识别 · 计算机科学 2024-11-25 Xiaoman Zhang , Hong-Yu Zhou , Xiaoli Yang , Oishi Banerjee , Julián N. Acosta , Josh Miller , Ouwen Huang , Pranav Rajpurkar

Despite recent advances in medical vision-language pretraining, existing models still struggle to capture the diagnostic workflow: radiographs are typically treated as context-agnostic images, while radiologists' gaze -- a crucial cue for…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Kang Liu , Zhuoqi Ma , Siyu Liang , Yunan Li , Xiyue Gao , Chao Liang , Kun Xie , Qiguang Miao

Chest X-ray plays a central role in thoracic diagnosis, and its interpretation inherently requires multi-step, evidence-grounded reasoning. However, large vision-language models (LVLMs) often generate plausible responses that are not…

人工智能 · 计算机科学 2026-03-25 Hyungyung Lee , Hangyul Yoon , Edward Choi

Following the gaze of other people and analyzing the target they are looking at can help us understand what they are thinking, and doing, and predict the actions that may follow. Existing methods for gaze following struggle to perform well…

计算机视觉与模式识别 · 计算机科学 2025-02-05 Feiyang Liu , Dan Guo , Jingyuan Xu , Zihao He , Shengeng Tang , Kun Li , Meng Wang