中文
相关论文

相关论文: LG-Gaze: Learning Geometry-aware Continuous Prompt…

200 篇论文

Transformer-based general visual geometry frameworks have shown promising performance in camera pose estimation and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yangfan Xu , Lilian Zhang , Xiaofeng He , Pengdong Wu , Wenqi Wu , Jun Mao

Existing deep learning methods for radiology report generation enhance diagnostic efficiency but often overlook physician-informed medical priors. This leads to a suboptimal alignment between the structured explanations and disease…

组织与器官 · 定量生物学 2026-04-13 Aishik Konwer , Moinak Bhattacharya , Prateek Prasanna

Multiple datasets have been created for training and testing appearance-based gaze estimators. Intuitively, more data should lead to better performance. However, combining datasets to train a single esti-mator rarely improves gaze…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Liang Wu , Bertram E. Shi

Visualization literacy assessments typically rely on correctness to classify performance, providing little evidence about how readers arrive at their answers. We argue that gaze can address this gap as an implicit process signal that…

人机交互 · 计算机科学 2026-03-25 Kathrin Schnizer

Medical vision-language models enable co-learning and integrating features from medical imaging and clinical text. However, these models are not easy to train and the latent representation space can be complex. Here we propose a novel way…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Che Liu , Sibo Cheng , Chen Chen , Mengyun Qiao , Weitong Zhang , Anand Shah , Wenjia Bai , Rossella Arcucci

We address the problem of gaze target estimation, which aims to predict where a person is looking in a scene. Predicting a person's gaze target requires reasoning both about the person's appearance and the contents of the scene. Prior works…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Fiona Ryan , Ajay Bati , Sangmin Lee , Daniel Bolya , Judy Hoffman , James M. Rehg

Gaze-annotated facial data is crucial for training deep neural networks (DNNs) for gaze estimation. However, obtaining these data is labor-intensive and requires specialized equipment due to the challenge of accurately annotating the gaze…

计算机视觉与模式识别 · 计算机科学 2024-06-03 Nerea Aranjuelo , Siyu Huang , Ignacio Arganda-Carreras , Luis Unzueta , Oihana Otaegui , Hanspeter Pfister , Donglai Wei

Spatial perception aims to estimate camera motion and scene structure from visual observations, a problem traditionally addressed through geometric modeling and physical consistency constraints. Recent learning-based methods have…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Haichao Zhu , Zhaorui Yang , Qian Zhang

The vision-language pre-training has enabled deep models to make a huge step forward in generalizing across unseen domains. The recent learning method based on the vision-language pre-training model is a great tool for domain generalization…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Liyuan Wang , Yan Jin , Zhen Chen , Jinlin Wu , Mengke Li , Yang Lu , Hanzi Wang

Fine-tuning large pretrained vision-language models (VLMs) has emerged as a prevalent paradigm for downstream adaptation, yet it faces a critical trade-off between domain specificity and domain generalization (DG) ability. Current methods…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Xinyao Li , Yinjie Min , Hongbo Chen , Zhekai Du , Fengling Li , Jingjing Li

Learning-based gaze estimation methods require large amounts of training data with accurate gaze annotations. Facing such demanding requirements of gaze data collection and annotation, several image synthesis methods were proposed, which…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Shiwei Jin , Zhen Wang , Lei Wang , Ning Bi , Truong Nguyen

Accurate 3D gaze estimation in unconstrained real-world environments remains a significant challenge due to variations in appearance, head pose, occlusion, and the limited availability of in-the-wild 3D gaze datasets. To address these…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Pierre Vuillecard , Jean-Marc Odobez

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large…

机器人学 · 计算机科学 2025-02-10 Houjian Yu , Mingen Li , Alireza Rezazadeh , Yang Yang , Changhyun Choi

Language learners should regularly engage in reading challenging materials as part of their study routine. Nevertheless, constantly referring to dictionaries is time-consuming and distracting. This paper presents a novel gaze-driven…

计算与语言 · 计算机科学 2023-10-03 Taichi Higasa , Keitaro Tanaka , Qi Feng , Shigeo Morishima

In the field of medical image segmentation, the scarcity of labeled data poses a major challenge for existing models to accurately perceive target regions. Compared with manual annotation, gaze data is easier and cheaper to obtain. As a…

图像与视频处理 · 电气工程与系统科学 2026-04-14 Rongjun Ge , Chong Wang , Yuxin Liu , Chunqiang Lu , Cong Xia , Yehui Jiang , Fangyi Xu , Yinsu Zhu , Daoqiang Zhang , Chengyu Liu , Yang Chen , Shuo Li , Yuting He

Deep neural networks have significantly improved appearance-based gaze estimation accuracy. However, it still suffers from unsatisfactory performance when generalizing the trained model to new domains, e.g., unseen environments or persons.…

计算机视觉与模式识别 · 计算机科学 2021-08-02 Yunfei Liu , Ruicong Liu , Haofei Wang , Feng Lu

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…

计算机视觉与模式识别 · 计算机科学 2018-07-30 Eunji Chong , Nataniel Ruiz , Yongxin Wang , Yun Zhang , Agata Rozga , James Rehg

Automatic eye gaze estimation is an important problem in vision based assistive technology with use cases in different emerging topics such as augmented reality, virtual reality and human-computer interaction. Over the past few years, there…

计算机视觉与模式识别 · 计算机科学 2022-08-08 Neeru Dubey , Shreya Ghosh , Abhinav Dhall

Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios.…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Ming Dai , Wenxuan Cheng , Jiang-Jiang Liu , Lingfeng Yang , Zhenhua Feng , Wankou Yang , Jingdong Wang

Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models hinges on having a good connector that maps visual features generated by a vision encoder to a shared…