中文
相关论文

相关论文: GazeVaLM: A Multi-Observer Eye-Tracking Benchmark …

200 篇论文

The ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized hand avatars from images remains challenging due to complex…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Zicong Fan , Edoardo Remelli , David Dimond , Fadime Sener , Liuhao Ge , Bugra Tekin , Cem Keskin , Shreyas Hampali

Current LLM assistants are powerful at answering questions, but they have limited access to the behavioral context that reveals when and where a user is struggling. We present a gaze-grounded multimodal LLM assistant that uses egocentric…

人机交互 · 计算机科学 2026-04-10 Valdemar Danry , Javier Hernandez , Andrew Wilson , Pattie Maes , Judith Amores

Ultrasound imaging has become the preferred imaging modality for early cancer screening due to its advantages of non-ionizing radiation, low cost, and real-time imaging capabilities. However, conventional ultrasound diagnosis heavily relies…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Chaoyin She , Ruifang Lu , Lida Chen , Wei Wang , Qinghua Huang

Vision-Language Models (VLMs) are trained on vast amounts of data captured by humans emulating our understanding of the world. However, known as visual illusions, human's perception of reality isn't always faithful to the physical world.…

人工智能 · 计算机科学 2023-11-02 Yichi Zhang , Jiayi Pan , Yuchen Zhou , Rui Pan , Joyce Chai

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA), the task remains challenging, particularly when models…

Developing artificial intelligence (AI) for clinical research requires a comprehensive data foundation that supports model training and rigorous evaluation. Here, we introduce TrialPanorama, a large-scale structured resource that aggregates…

人工智能 · 计算机科学 2025-12-17 Zifeng Wang , Jiacheng Lin , Qiao Jin , Junyi Gao , Jathurshan Pradeepkumar , Pengcheng Jiang , Zhiyong Lu , Jimeng Sun

We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT images, and visual field…

图像与视频处理 · 电气工程与系统科学 2026-05-12 Jiao Wang , Chi Liu , Yiying Zhang , Hongchen Luo , Zhifen Guo , Ying Hu , Ke Xu , Jing Zhou , Hongyan Xu , Ruiting Zhou , Man Tang

Artificial intelligence (AI) is poised to transform healthcare by enabling personalized and efficient care through data-driven insights. Although radiology is at the forefront of AI adoption, in practice, the potential of AI models is often…

计算机视觉与模式识别 · 计算机科学 2025-02-17 Benjamin D. Killeen , Bohua Wan , Aditya V. Kulkarni , Nathan Drenkow , Michael Oberst , Paul H. Yi , Mathias Unberath

Medical Large Vision-Language Models (Med-LVLMs) demonstrate significant potential in healthcare, but their reliance on general medical data and coarse-grained global visual understanding limits them in intelligent ophthalmic diagnosis.…

计算机视觉与模式识别 · 计算机科学 2025-04-21 Sijing Li , Tianwei Lin , Lingshuai Lin , Wenqiao Zhang , Jiang Liu , Xiaoda Yang , Juncheng Li , Yucheng He , Xiaohui Song , Jun Xiao , Yueting Zhuang , Beng Chin Ooi

In this era of pandemic, the future of healthcare industry has never been more exciting. Artificial intelligence and machine learning (AI & ML) present opportunities to develop solutions that cater for very specific needs within the…

图像与视频处理 · 电气工程与系统科学 2022-11-29 Aravind Sasidharan Pillai

Rationales in the form of manually annotated input spans usually serve as ground truth when evaluating explainability methods in NLP. They are, however, time-consuming and often biased by the annotation process. In this paper, we debate…

计算与语言 · 计算机科学 2024-03-01 Stephanie Brandl , Oliver Eberle , Tiago Ribeiro , Anders Søgaard , Nora Hollenstein

Recent studies on appearance based gaze estimation indicate the ability of Neural Networks to decode gaze information from facial images encompassing pose information. In this paper, we propose Gaze-Net: A capsule network capable of…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Bhanuka Mahanama , Yasith Jayawardana , Sampath Jayarathna

With the rapid advancement of Artificial Intelligence Generated Content (AIGC) technologies, synthetic images have become increasingly prevalent in everyday life, posing new challenges for authenticity assessment and detection. Despite the…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Siwei Wen , Junyan Ye , Peilin Feng , Hengrui Kang , Zichen Wen , Yize Chen , Jiang Wu , Wenjun Wu , Conghui He , Weijia Li

When deep neural network (DNN) was first introduced to the medical image analysis community, researchers were impressed by its performance. However, it is evident now that a large number of manually labeled data is often a must to train a…

图像与视频处理 · 电气工程与系统科学 2022-05-16 Sheng Wang , Xi Ouyang , Tianming Liu , Qian Wang , Dinggang Shen

Integrating real-time artificial intelligence (AI) systems in clinical practices faces challenges such as scalability and acceptance. These challenges include data availability, biased outcomes, data quality, lack of transparency, and…

The recent development of data-driven AI promises to automate medical diagnosis; however, most AI functions as 'black boxes' to physicians with limited computational knowledge. Using medical imaging as a point of departure, we conducted…

人机交互 · 计算机科学 2020-01-22 Yao Xie , Melody Chen , David Kao , Ge Gao , Xiang 'Anthony' Chen

AI-driven models have shown great promise in detecting errors in radiology reports, yet the field lacks a unified benchmark for rigorous evaluation of error detection and further correction. To address this gap, we introduce CorBenchX, a…

人工智能 · 计算机科学 2025-05-20 Jing Zou , Qingqiu Li , Chenyu Lian , Lihao Liu , Xiaohan Yan , Shujun Wang , Jing Qin

Despite the progress in automatic detection of radiologic findings from chest X-ray (CXR) images in recent years, a quantitative evaluation of the explainability of these models is hampered by the lack of locally labeled datasets for…

Medical diagnosis is not a single prediction from a fully specified vignette. It is a sequential workup: clinicians decide what evidence to obtain, revise a differential diagnosis, and stop when the diagnosis is sufficiently supported. Most…

Conducting collaborative tasks, e.g., multi-user game, in virtual reality (VR) could enable us to explore more immersive and effective experience. However, for current VR systems, users cannot communicate properly with each other via their…

人机交互 · 计算机科学 2023-03-21 Song Zhao , Shiwei Cheng , Chenshuang Zhu