中文
相关论文

相关论文: LG-Gaze: Learning Geometry-aware Continuous Prompt…

200 篇论文

Deep learning-based appearance gaze estimation methods are gaining popularity due to their high accuracy and fewer constraints from the environment. However, existing high-precision models often rely on deeper networks, leading to problems…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zhang Cheng , Yanxia Wang

The success of vision-language models is primarily attributed to effective alignment across modalities such as vision and language. However, modality gaps persist in existing alignment algorithms and appear necessary for human perception as…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Hanqi Yan , Xiangxiang Cui , Lu Yin , Jindong Gu , Paul Pu Liang , Yulan He , Yifei Wang

Gaze estimation is pivotal in human scene comprehension tasks, particularly in medical diagnostic analysis. Eye-tracking technology facilitates the recording of physicians' ocular movements during image interpretation, thereby elucidating…

计算机视觉与模式识别 · 计算机科学 2024-08-13 Shaonan Liu , Wenting Chen , Jie Liu , Xiaoling Luo , Linlin Shen

Fine-tuning approaches for Vision-Language Models (VLMs) face a critical three-way trade-off between In-Distribution (ID) accuracy, Out-of-Distribution (OOD) generalization, and adversarial robustness. Existing robust fine-tuning strategies…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Shivang Chopra , Shaunak Halbe , Chengyue Huang , Brisa Maneechotesuwan , Zsolt Kira

Recent advancements in Vision-Language (VL) models have sparked interest in their deployment on edge devices, yet challenges in handling diverse visual modalities, manual annotation, and computational constraints remain. We introduce…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Kaiwen Cai , Zhekai Duan , Gaowen Liu , Charles Fleming , Chris Xiaoxuan Lu

In cognitive science and AI, a longstanding question is whether machines learn representations that align with those of the human mind. While current models show promise, it remains an open question whether this alignment is superficial or…

神经元与认知 · 定量生物学 2025-10-27 Craig Sanders , Billy Dickson , Sahaj Singh Maini , Robert Nosofsky , Zoran Tiganj

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Zero-shot image classification using auxiliary information, such as attributes describing discriminative object properties, requires time-consuming annotation by domain experts. We instead propose a method that relies on human gaze as…

计算机视觉与模式识别 · 计算机科学 2017-04-14 Nour Karessli , Zeynep Akata , Bernt Schiele , Andreas Bulling

Currently, inspired by the success of vision-language models (VLMs), an increasing number of researchers are focusing on improving VLMs and have achieved promising results. However, most existing methods concentrate on optimizing the…

计算机视觉与模式识别 · 计算机科学 2024-09-23 Dawei Yan , Pengcheng Li , Yang Li , Hao Chen , Qingguo Chen , Weihua Luo , Wei Dong , Qingsen Yan , Haokui Zhang , Chunhua Shen

Visual affordance learning is crucial for robots to understand and interact effectively with the physical world. Recent advances in this field attempt to leverage pre-trained knowledge of vision-language foundation models to learn…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Qian Zhang , Lin Zhang , Xing Fang , Mingxin Zhang , Zhiyuan Wei , Ran Song , Wei Zhang

In the medical multi-modal frameworks, the alignment of cross-modality features presents a significant challenge. However, existing works have learned features that are implicitly aligned from the data, without considering the explicit…

计算机视觉与模式识别 · 计算机科学 2024-06-17 Chong Ma , Hanqi Jiang , Wenting Chen , Yiwei Li , Zihao Wu , Xiaowei Yu , Zhengliang Liu , Lei Guo , Dajiang Zhu , Tuo Zhang , Dinggang Shen , Tianming Liu , Xiang Li

Training-free video anomaly detection (VAD) has recently emerged as a scalable alternative to supervised approaches, yet existing methods largely rely on static prompting and geometry-agnostic feature fusion. As a result, anomaly inference…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Ali Zia , Usman Ali , Muhammad Umer Ramzan , Hamza Abid , Abdul Rehman , Wei Xiang

Self-supervised depth estimation has evolved into an image reconstruction task that minimizes a photometric loss. While recent methods have made strides in indoor depth estimation, they often produce inconsistent depth estimation in…

计算机视觉与模式识别 · 计算机科学 2024-02-23 Anqi Cheng , Zhiyuan Yang , Haiyue Zhu , Kezhi Mao

Parameter-efficient prompt learning has become the de facto standard for adapting Vision-Language Models (VLMs) to downstream tasks. Existing approaches predominantly focus on aligning text prompts with first-order visual features (i.e.,…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Minglei Chen , Weilong Wang , Jiang Duan , Ye Deng

We address the challenge of unsupervised mistake detection in egocentric video of skilled human activities through the analysis of gaze signals. While traditional methods rely on manually labeled mistakes, our approach does not require…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Michele Mazzamuto , Antonino Furnari , Yoichi Sato , Giovanni Maria Farinella

Video diffusion models lack explicit geometric supervision during training, leading to inconsistency artifacts such as object deformation, spatial drift, and depth violations in generated videos. To address this limitation, we propose a…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Tengjiao Yin , Jinglei Shi , Heng Guo , Xi Wang

Previous studies have illustrated the potential of analysing gaze behaviours in collaborative learning to provide educationally meaningful information for students to reflect on their learning. Over the past decades, machine learning…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Junyuan Liang , Qi Zhou , Sahan Bulathwela , Mutlu Cukurova

Large language models have shown impressive results for multi-hop mathematical reasoning when the input question is only textual. Many mathematical reasoning problems, however, contain both text and image. With the ever-increasing adoption…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Mehran Kazemi , Hamidreza Alvari , Ankit Anand , Jialin Wu , Xi Chen , Radu Soricut

Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed…

机器学习 · 计算机科学 2022-05-03 Ahmed Abdou , Ekta Sood , Philipp Müller , Andreas Bulling

Geometry problem solving (GPS) is a challenging mathematical reasoning task requiring multi-modal understanding, fusion, and reasoning. Existing neural solvers take GPS as a vision-language task but are short in the representation of…

计算机视觉与模式识别 · 计算机科学 2024-02-21 Zhong-Zhi Li , Ming-Liang Zhang , Fei Yin , Cheng-Lin Liu