中文
相关论文

相关论文: GazeVaLM: A Multi-Observer Eye-Tracking Benchmark …

200 篇论文

Human gaze provides essential cues for interpreting attention, intention, and social interaction in visual scenes, yet gaze understanding remains largely unexplored in current vision-language models (VLMs). While recent VLMs achieve strong…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Shijing Wang , Chaoqun Cui , Yaping Huang , Hyung Jin Chang , Yihua Cheng

Developing generalist foundation model has recently attracted tremendous attention among researchers in the field of AI for Medicine (AI4Medicine). A pivotal insight in developing these models is their reliance on dataset scaling, which…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Xiaoman Zhang , Chaoyi Wu , Ziheng Zhao , Jiayu Lei , Ya Zhang , Yanfeng Wang , Weidi Xie

Current vision-language models (VLMs) in medicine are primarily designed for categorical question answering (e.g., "Is this normal or abnormal?") or qualitative descriptive tasks. However, clinical decision-making often relies on…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongcheng Yao , Yongshuo Zong , Raman Dutt , Yongxin Yang , Sotirios A Tsaftaris , Timothy Hospedales

Eye-tracking analysis plays a vital role in medical imaging, providing key insights into how radiologists visually interpret and diagnose clinical cases. In this work, we first analyze radiologists' attention and agreement by measuring the…

AI alignment refers to models acting towards human-intended goals, preferences, or ethical principles. Given that most large-scale deep learning models act as black boxes and cannot be manually controlled, analyzing the similarity between…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Jiyoung Lee , Seungho Kim , Seunghyun Won , Joonseok Lee , Marzyeh Ghassemi , James Thorne , Jaeseok Choi , O-Kil Kwon , Edward Choi

In the field of chest X-ray (CXR) diagnosis, existing works often focus solely on determining where a radiologist looks, typically through tasks such as detection, segmentation, or classification. However, these approaches are often…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Trong Thang Pham , Jacob Brecheisen , Anh Nguyen , Hien Nguyen , Ngan Le

With the emergence of Virtual and Mixed Reality (XR) devices, eye tracking has received significant attention in the computer vision community. Eye gaze estimation is a crucial component in XR -- enabling energy efficient rendering,…

计算机视觉与模式识别 · 计算机科学 2020-03-20 Zhengyang Wu , Srivignesh Rajendran , Tarrence van As , Joelle Zimmermann , Vijay Badrinarayanan , Andrew Rabinovich

We present ReXGradient-160K, representing the largest publicly available chest X-ray dataset to date in terms of the number of patients. This dataset contains 160,000 chest X-ray studies with paired radiological reports from 109,487 unique…

图像与视频处理 · 电气工程与系统科学 2025-05-13 Xiaoman Zhang , Julián N. Acosta , Josh Miller , Ouwen Huang , Pranav Rajpurkar

Eye gaze offers valuable cues about attention, short-term intent, and future actions, making it a powerful signal for modeling egocentric behavior. In this work, we propose a gaze-regularized framework that enhances VLMs for two key…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Anupam Pani , Yanchao Yang

Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems.…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Trong Thang Pham , Ngoc-Vuong Ho , Nhat-Tan Bui , Thinh Phan , Patel Brijesh , Donald Adjeroh , Gianfranco Doretto , Anh Nguyen , Carol C. Wu , Hien Nguyen , Ngan Le

Understanding where people are looking is an informative social cue. In this work, we present Gaze360, a large-scale gaze-tracking dataset and method for robust 3D gaze estimation in unconstrained images. Our dataset consists of 238…

计算机视觉与模式识别 · 计算机科学 2019-10-23 Petr Kellnhofer , Adria Recasens , Simon Stent , Wojciech Matusik , Antonio Torralba

We introduce MultiMedEval, an open-source toolkit for fair and reproducible evaluation of large, medical vision-language models (VLM). MultiMedEval comprehensively assesses the models' performance on a broad array of six multi-modal tasks,…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Corentin Royer , Bjoern Menze , Anjany Sekuboyina

Medical Visual Question Answering (Med-VQA) combines computer vision and natural language processing to automatically answer clinical inquiries about medical images. However, current Med-VQA datasets exhibit two significant limitations: (1)…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Bo Liu , Ke Zou , Liming Zhan , Zexin Lu , Xiaoyu Dong , Yidi Chen , Chengqiang Xie , Jiannong Cao , Xiao-Ming Wu , Huazhu Fu

Gaze estimation is instrumental in modern virtual reality (VR) systems. Despite significant progress in remote-camera gaze estimation, VR gaze research remains constrained by data scarcity, particularly the lack of large-scale, accurately…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Gil Shapira , Ishay Goldin , Evgeny Artyomov , Donghoon Kim , Yosi Keller , Niv Zehngut

Recent advances in AI combine large language models (LLMs) with vision encoders that bring forward unprecedented technical capabilities to leverage for a wide range of healthcare applications. Focusing on the domain of radiology,…

Smart glasses with AI assistants are increasingly used in daily life. However, current systems lack awareness of the user's internal cognitive state, leaving them unable to proactively anticipate users' needs without access to cognitive…

Recent vision-language models (VLMs) have shown strong generalization and multimodal reasoning abilities in natural domains. However, their application to medical diagnosis remains limited by the lack of comprehensive and structured…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Sheng Lu , Hao Chen , Rui Yin , Juyan Ba , Yu Zhang , Yuanzhe Li

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Zhuangzhuang Dai , Vincent Gbouna Zakka , Luis J. Manso , Chen Li

Vietnamese medical research has become an increasingly vital domain, particularly with the rise of intelligent technologies aimed at reducing time and resource burdens in clinical diagnosis. Recent advances in vision-language models (VLMs),…

Large, labeled datasets have driven deep learning methods to achieve expert-level performance on a variety of medical imaging tasks. We present CheXpert, a large dataset that contains 224,316 chest radiographs of 65,240 patients. We design…