English
Related papers

Related papers: LG-Gaze: Learning Geometry-aware Continuous Prompt…

200 papers

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao

Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease…

Tissues and Organs · Quantitative Biology 2026-04-13 Aishik Konwer , Moinak Bhattacharya , Prateek Prasanna

Multiple datasets have been created for training and testing appearance-based gaze estimators. Intuitively, more data should lead to better performance. However, combining datasets to train a single esti-mator rarely improves gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Liang Wu , Bertram E. Shi

Visualization literacy assessments typically rely on correctness to classify performance, providing little evidence about how readers arrive at their answers. We argue that gaze can address this gap as an implicit process signal that…

Human-Computer Interaction · Computer Science 2026-03-25 Kathrin Schnizer

Medical vision-language models enable co-learning and integrating features from medical imaging and clinical text. However, these models are not easy to train and the latent representation space can be complex. Here we propose a novel way…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Che Liu , Sibo Cheng , Chen Chen , Mengyun Qiao , Weitong Zhang , Anand Shah , Wenjia Bai , Rossella Arcucci

We address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person's gaze target requires reasoning both about the person's appearance and the contents of the scene. Prior works…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Fiona Ryan , Ajay Bati , Sangmin Lee , Daniel Bolya , Judy Hoffman , James M. Rehg

Gaze-annotated facial data is crucial for training deep neural networks (DNNs) for gaze estimation. However, obtaining these data is labor-intensive and requires specialized equipment due to the challenge of accurately annotating the gaze…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Nerea Aranjuelo , Siyu Huang , Ignacio Arganda-Carreras , Luis Unzueta , Oihana Otaegui , Hanspeter Pfister , Donglai Wei

Spatial perception aims to estimate camera motion and scene structure from visual observations, a problem traditionally addressed through geometric modeling and physical consistency constraints. Recent learning-based methods have…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Haichao Zhu , Zhaorui Yang , Qian Zhang

The vision-language pre-training has enabled deep models to make a huge step forward in generalizing across unseen domains. The recent learning method based on the vision-language pre-training model is a great tool for domain generalization…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Liyuan Wang , Yan Jin , Zhen Chen , Jinlin Wu , Mengke Li , Yang Lu , Hanzi Wang

Fine-tuning large pretrained vision-language models (VLMs) has emerged as a prevalent paradigm for downstream adaptation, yet it faces a critical trade-off between domain specificity and domain generalization (DG) ability. Current methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Xinyao Li , Yinjie Min , Hongbo Chen , Zhekai Du , Fengling Li , Jingjing Li

Learning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Shiwei Jin , Zhen Wang , Lei Wang , Ning Bi , Truong Nguyen

Accurate 3D gaze estimation in unconstrained real-world environments remains a significant challenge due to variations in appearance, head pose, occlusion, and the limited availability of in-the-wild 3D gaze datasets. To address these…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Pierre Vuillecard , Jean-Marc Odobez

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large…

Robotics · Computer Science 2025-02-10 Houjian Yu , Mingen Li , Alireza Rezazadeh , Yang Yang , Changhyun Choi

Language learners should regularly engage in reading challenging materials as part of their study routine. Nevertheless, constantly referring to dictionaries is time-consuming and distracting. This paper presents a novel gaze-driven…

Computation and Language · Computer Science 2023-10-03 Taichi Higasa , Keitaro Tanaka , Qi Feng , Shigeo Morishima

In the field of medical image segmentation, the scarcity of labeled data poses a major challenge for existing models to accurately perceive target regions. Compared with manual annotation, gaze data is easier and cheaper to obtain. As a…

Image and Video Processing · Electrical Eng. & Systems 2026-04-14 Rongjun Ge , Chong Wang , Yuxin Liu , Chunqiang Lu , Cong Xia , Yehui Jiang , Fangyi Xu , Yinsu Zhu , Daoqiang Zhang , Chengyu Liu , Yang Chen , Shuo Li , Yuting He

Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons.…

Computer Vision and Pattern Recognition · Computer Science 2021-08-02 Yunfei Liu , Ruicong Liu , Haofei Wang , Feng Lu

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Eunji Chong , Nataniel Ruiz , Yongxin Wang , Yun Zhang , Agata Rozga , James Rehg

Automatic eye gaze estimation is an important problem in vision based assistive technology with use cases in different emerging topics such as augmented reality, virtual reality and human-computer interaction. Over the past few years, there…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Neeru Dubey , Shreya Ghosh , Abhinav Dhall

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios.…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Ming Dai , Wenxuan Cheng , Jiang-Jiang Liu , Lingfeng Yang , Zhenhua Feng , Wankou Yang , Jingdong Wang

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared…