中文
相关论文

相关论文: Vision-language models for decoding provider atten…

200 篇论文

Recently, Transformer has made significant progress in various vision tasks. To balance computation and efficiency in video tasks, recent works heavily rely on factorized or window-based self-attention. However, these approaches split…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Bohao Xing , Deng Li , Rong Gao , Xin Liu , Heikki Kälviäinen

It is well known that human gaze carries significant information about visual attention. However, there are three main difficulties in incorporating the gaze data in an attention mechanism of deep neural networks: 1) the gaze fixation…

计算机视觉与模式识别 · 计算机科学 2020-11-10 Kyle Min , Jason J. Corso

This paper introduces a novel neural network-based reinforcement learning approach for robot gaze control. Our approach enables a robot to learn and to adapt its gaze control strategy for human-robot interaction neither with the use of…

机器人学 · 计算机科学 2019-02-18 Stéphane Lathuilière , Benoit Massé , Pablo Mesejo , Radu Horaud

Current multimodal large language models (MLLMs) cannot effectively utilize eye-gaze information for video understanding, even when gaze cues are supplied via visual overlays or text descriptions. We introduce GazeQwen, a parameter…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Trong Thang Pham , Hien Nguyen , Ngan Le

Medical ultrasound has been widely used to examine vascular structure in modern clinical practice. However, traditional ultrasound examination often faces challenges related to inter- and intra-operator variation. The robotic ultrasound…

机器人学 · 计算机科学 2025-02-10 Yuan Bi , Yang Su , Nassir Navab , Zhongliang Jiang

The spatial attention is a straightforward approach to enhance the performance for remote sensing image captioning. However, conventional spatial attention approaches consider only the attention distribution on one fixed coarse grid,…

计算机视觉与模式识别 · 计算机科学 2021-05-12 Chengze Wang , Zhiyu Jiang , Yuan Yuan

Appearance-based supervised methods with full-face image input have made tremendous advances in recent gaze estimation tasks. However, intensive human annotation requirement inhibits current methods from achieving industrial level accuracy…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Yangzhou Jiang , Yinxin Lin , Yaoming Wang , Teng Li , Bilian Ke , Bingbing Ni

In recent years we have witnessed an increasing number of interactive systems on handheld mobile devices which utilise gaze as a single or complementary interaction modality. This trend is driven by the enhanced computational power of these…

人机交互 · 计算机科学 2023-07-04 Yaxiong Lei , Shijing He , Mohamed Khamis , Juan Ye

The internal workings of modern deep learning models stay often unclear to an external observer, although spatial attention mechanisms are involved. The idea of this work is to translate these spatial attentions into natural language to…

计算机视觉与模式识别 · 计算机科学 2020-10-23 Philipp Sadler

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the…

计算机视觉与模式识别 · 计算机科学 2025-07-02 Zhuangzhuang Dai , Vincent Gbouna Zakka , Luis J. Manso , Chen Li

In the past few years the transformer model has been utilized for a variety of tasks such as image captioning, image classification natural language generation, and natural language understanding. As a key component of the transformer…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zanyar Zohourianshahzadi , Terrance E. Boult , Jugal K. Kalita

Where someone looks is a nonverbal communication cue that children and adults readily use. How well can Vision-Language Models (VLMs) infer gaze targets? To construct evaluation stimuli, we captured 1,360 real-world photos of scenes in…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Zory Zhang , Pinyuan Feng , Bingyang Wang , Tianwei Zhao , Suyang Yu , Qingying Gao , Hokin Deng , Ziqiao Ma , Yijiang Li , Dezhi Luo

During multi-party interactions, gaze direction is a key indicator of interest and intent, making it essential for social robots to direct their attention appropriately. Understanding the social context is crucial for robots to engage…

机器人学 · 计算机科学 2026-02-12 Ramtin Tabatabaei , Alireza Taheri

We present a real-time gaze tracking system that directly acquires task-relevant latent features using a fully passive optical encoder. Instead of forming and processing full-resolution images, our approach leverages a microlens array with…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Yidan Zheng , Matheus Souza , Kaizhang Kang , Qiang Fu , Hadi Amata , Wolfgang Heidrich

Learning harmful shortcuts such as spurious correlations and biases prevents deep neural networks from learning the meaningful and useful representations, thus jeopardizing the generalizability and interpretability of the learned…

We propose a deep visuo-tactile model for realtime estimation of the liquid inside a deformable container in a proprioceptive way.We fuse two sensory modalities, i.e., the raw visual inputs from the RGB camera and the tactile cues from our…

机器人学 · 计算机科学 2022-08-17 Fan Zhu , Ruixing Jia , Lei Yang , Youcan Yan , Zheng Wang , Jia Pan , Wenping Wang

While exploring visual scenes, humans' scanpaths are driven by their underlying attention processes. Understanding visual scanpaths is essential for various applications. Traditional scanpath models predict the where and when of gaze shifts…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Xianyu Chen , Ming Jiang , Qi Zhao

Representations learned by convolutional neural networks (CNNs) exhibit a remarkable resemblance to information processing patterns observed in the primate visual system on large neuroimaging datasets collected under diverse, naturalistic…

神经元与认知 · 定量生物学 2026-03-16 Dora Gozukara , Nasir Ahmad , Katja Seeliger , Djamari Oetringer , Linda Geerligs

This paper is interested in investigating whether human gaze signals can be leveraged to improve state-of-the-art search engine performance and how to incorporate this new input signal marked by human attention into existing neural…

信息检索 · 计算机科学 2022-07-06 Sibo Dong , Justin Goldstein , Grace Hui Yang

Gaze redirection aims at manipulating the gaze of a given face image with respect to a desired direction (i.e., a reference angle) and it can be applied to many real life scenarios, such as video-conferencing or taking group photos.…

计算机视觉与模式识别 · 计算机科学 2020-11-30 Jingjing Chen , Jichao Zhang , Enver Sangineto , Jiayuan Fan , Tao Chen , Nicu Sebe