English
Related papers

Related papers: ViTGaze: Gaze Following with Interaction Features …

200 papers

We propose a novel 3D gaze estimation approach that learns spatial relationships between the subject and objects in the scene, and outputs 3D gaze direction. Our method targets unconstrained settings, including cases where close-up views of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-19 Yuki Kawana , Shintaro Shiba , Quan Kong , Norimasa Kobori

Shortcut learning is common but harmful to deep learning models, leading to degenerated feature representations and consequently jeopardizing the model's generalizability and interpretability. However, shortcut learning in the widely used…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Chong Ma , Lin Zhao , Yuzhong Chen , David Weizhong Liu , Xi Jiang , Tuo Zhang , Xintao Hu , Dinggang Shen , Dajiang Zhu , Tianming Liu

Visual attention mechanisms play a crucial role in human perception and aesthetic evaluation. Recent advances in Vision Transformers (ViTs) have demonstrated remarkable capabilities in computer vision tasks, yet their alignment with human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Miguel Carrasco , César González-Martín , José Aranda , Luis Oliveros

We present VQA-MHUG - a novel 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker. We use our dataset to analyze the similarity between…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Ekta Sood , Fabian Kögel , Florian Strohm , Prajit Dhar , Andreas Bulling

We present GazeGen, a user interaction system that generates visual content (images and videos) for locations indicated by the user's eye gaze. GazeGen allows intuitive manipulation of visual content by targeting regions of interest with…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 He-Yen Hsieh , Ziyun Li , Sai Qian Zhang , Wei-Te Mark Ting , Kao-Den Chang , Barbara De Salvo , Chiao Liu , H. T. Kung

Machine Interpreting systems are currently implemented as unimodal, real-time speech-to-speech architectures, processing translation exclusively on the basis of the linguistic signal. Such reliance on a single modality, however, constrains…

Computation and Language · Computer Science 2025-09-30 Claudio Fantinuoli

Eye-tracking research has proven valuable in understanding numerous cognitive functions. Recently, Frey et al. provided an exciting deep learning method for learning eye movements from fMRI data. However, it needed to co-register fMRI into…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Xiuwen Wu , Rongjie Hu , Jie Liang , Yanming Wang , Bensheng Qiu , Xiaoxiao Wang

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

Recently, Vision Transformer and its variants have shown great promise on various computer vision tasks. The ability of capturing short- and long-range visual dependencies through self-attention is arguably the main source for the success.…

Computer Vision and Pattern Recognition · Computer Science 2021-07-02 Jianwei Yang , Chunyuan Li , Pengchuan Zhang , Xiyang Dai , Bin Xiao , Lu Yuan , Jianfeng Gao

Vision Transformers (ViTs) have become prominent models for solving various vision tasks. However, the interpretability of ViTs has not kept pace with their promising performance. While there has been a surge of interest in developing {\it…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Yao Qiang , Chengyin Li , Prashant Khanduri , Dongxiao Zhu

Vision Transformer (ViT) has shown high potential in video recognition, owing to its flexible design, adaptable self-attention mechanisms, and the efficacy of masked pre-training. Yet, it remains unclear how to adapt these pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Min Yang , Huan Gao , Ping Guo , Limin Wang

Multi-modal large language models (MLLMs) have advanced general-purpose video understanding but struggle with long, high-resolution videos -- they process every pixel equally in their vision transformers (ViTs) or LLMs despite significant…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Baifeng Shi , Stephanie Fu , Long Lian , Hanrong Ye , David Eigen , Aaron Reite , Boyi Li , Jan Kautz , Song Han , David M. Chan , Pavlo Molchanov , Trevor Darrell , Hongxu Yin

Eye gaze can provide rich information on human psychological activities, and has garnered significant attention in the field of Human-Robot Interaction (HRI). However, existing gaze estimation methods merely predict either the gaze…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Haoming Huang , Musen Zhang , Jianxin Yang , Zhen Li , Jinkai Li , Yao Guo

Human eye contact is a form of non-verbal communication and can have a great influence on social behavior. Since the location and size of the eye contact targets vary across different videos, learning a generic video-independent eye contact…

Computer Vision and Pattern Recognition · Computer Science 2022-10-06 Tianyi Wu , Yusuke Sugano

Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Manuel Benavent-Lledo , Konstantinos Bacharidis , Victoria Manousaki , Konstantinos Papoutsakis , Antonis Argyros , Jose Garcia-Rodriguez

We present CasualGaze, a novel eye-gaze-based target selection technique to support natural and casual eye-gaze input. Unlike existing solutions that require users to keep the eye-gaze center on the target actively, CasualGaze allows users…

Human-Computer Interaction · Computer Science 2024-08-26 Yingtian Shi , Yukang Yan , Zisu Li , Chen Liang , Yuntao Wang , Chun Yu , Yuanchun Shi

The goal of visual analytics is to create a symbiosis between human and computer by leveraging their unique strengths. While this model has demonstrated immense success, we are yet to realize the full potential of such a human-computer…

Human-Computer Interaction · Computer Science 2018-09-27 Ran Wan , Roman Garnett , Alvitta Ottley

For computer systems to effectively interact with humans using spoken language, they need to understand how the words being generated affect the users' moment-by-moment attention. Our study focuses on the incremental prediction of attention…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Sounak Mondal , Seoyoung Ahn , Zhibo Yang , Niranjan Balasubramanian , Dimitris Samaras , Gregory Zelinsky , Minh Hoai

Over the past few years, there has been an increasing interest to interpret gaze direction in an unconstrained environment with limited supervision. Owing to data curation and annotation issues, replicating gaze estimation method to other…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Shreya Ghosh , Abhinav Dhall , Jarrod Knibbe , Munawar Hayat

Vision transformer (ViT) expands the success of transformer models from sequential data to images. The model decomposes an image into many smaller patches and arranges them into a sequence. Multi-head self-attentions are then applied to the…

Machine Learning · Computer Science 2023-03-27 Yiran Li , Junpeng Wang , Xin Dai , Liang Wang , Chin-Chia Michael Yeh , Yan Zheng , Wei Zhang , Kwan-Liu Ma
‹ Prev 1 8 9 10 Next ›