English
Related papers

Related papers: Restoring Eye Contact to the Virtual Classroom wit…

200 papers

This paper explores the estimation of user attention in the setting of a cooperative handheld robot: a robot designed to behave as a handheld tool but that has levels of task knowledge. We use a tool-mounted gaze tracking system, which,…

Robotics · Computer Science 2018-10-16 Janis Stolzenwald , Walterio W. Mayol-Cuevas

In automotive domain, operation of secondary tasks like accessing infotainment system, adjusting air conditioning vents, and side mirrors distract drivers from driving. Though existing modalities like gesture and speech recognition systems…

Human-Computer Interaction · Computer Science 2021-06-11 Gowdham Prabhakar , Priyam Rajkhowa , Pradipta Biswas

This study presents a novel framework for 3D gaze tracking tailored for mixed-reality settings, aimed at enhancing joint attention and collaborative efforts in team-based scenarios. Conventional gaze tracking, often limited by monocular…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Eduardo Davalos , Yike Zhang , Ashwin T. S. , Joyce H. Fonteles , Umesh Timalsina , Guatam Biswas

We present a novel, automatic eye gaze tracking scheme inspired by smooth pursuit eye motion while playing mobile games or watching virtual reality contents. Our algorithm continuously calibrates an eye tracking system for a head mounted…

Computer Vision and Pattern Recognition · Computer Science 2016-12-22 Subarna Tripathi , Brian Guenter

Although pre-training on a large amount of data is beneficial for robot learning, current paradigms only perform large-scale pretraining for visual representations, whereas representations for other modalities are trained from scratch. In…

Robotics · Computer Science 2024-05-15 Jared Mejia , Victoria Dean , Tess Hellebrekers , Abhinav Gupta

Video-text retrieval, the task of retrieving videos based on a textual query or vice versa, is of paramount importance for video understanding and multimodal information retrieval. Recent methods in this area rely primarily on visual and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Boseung Jeong , Jicheol Park , Sungyeon Kim , Suha Kwak

Driver visual attention prediction is a critical task in autonomous driving and human-computer interaction (HCI) research. Most prior studies focus on estimating attention allocation at a single moment in time, typically using static RGB…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Kaiser Hamid , Khandakar Ashrafi Akbar , Nade Liang

Educational videos are a cornerstone of remote and blended learning. However, learners' fluctuating attention remains a significant barrier to effective information retention. Prior research has attempted to mitigate this by detecting and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Gabriel Becquet , Sébastien Lallé , Vanda Luengo , Ali Abou-Hassan

Video conferences play a vital role in our daily lives. However, many nonverbal cues are missing, including gaze and spatial information. We introduce LookAtChat, a web-based video conferencing system, which empowers remote users to…

Human-Computer Interaction · Computer Science 2021-08-06 Zhenyi He , Ruofei Du , Ken Perlin

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Bolin Lai , Fiona Ryan , Wenqi Jia , Miao Liu , James M. Rehg

State-of-the-art appearance-based gaze estimation methods, usually based on deep learning techniques, mainly rely on static features. However, temporal trace of eye gaze contains useful information for estimating a given gaze point. For…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Cristina Palmero , Oleg V. Komogortsev , Sachin S. Talathi

Emotional expressions are inherently multimodal -- integrating facial behavior, speech, and gaze -- but their automatic recognition is often limited to a single modality, e.g. speech during a phone call. While previous work proposed…

Machine Learning · Computer Science 2022-05-03 Ahmed Abdou , Ekta Sood , Philipp Müller , Andreas Bulling

Compositional Zero-Shot Learning (CZSL) aims to recognize novel attribute-object compositions by leveraging knowledge from seen compositions. Current methods align textual prototypes with visual features via Vision-Language Models (VLMs),…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Shiyu Zhang , Cheng Yan , Yang Liu , Chenchen Jing , Lei Zhou , Wenjun Wang

We present GazeDirector, a new approach for eye gaze redirection that uses model-fitting. Our method first tracks the eyes by fitting a multi-part eye region model to video frames using analysis-by-synthesis, thereby recovering eye region…

Computer Vision and Pattern Recognition · Computer Science 2017-05-01 Erroll Wood , Tadas Baltrusaitis , Louis-Philippe Morency , Peter Robinson , Andreas Bulling

Gaze estimation involves predicting where the person is looking at within an image or video. Technically, the gaze information can be inferred from two different magnification levels: face orientation and eye orientation. The inference is…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Ashesh , Chu-Song Chen , Hsuan-Tien Lin

Traditional sign language teaching methods face challenges such as limited feedback and diverse learning scenarios. Although 2D resources lack real-time feedback, classroom teaching is constrained by a scarcity of teacher. Methods based on…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Hongli Wen , Yang Xu , Lin Li , Xudong Ru , Xingce Wang , Zhongke Wu

Learning to classify video data from classes not included in the training data, i.e. video-based zero-shot learning, is challenging. We conjecture that the natural alignment between the audio and visual modalities in video data provides a…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Otniel-Bogdan Mercea , Lukas Riesch , A. Sophia Koepke , Zeynep Akata

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human activities through…

Human-Computer Interaction · Computer Science 2025-04-09 Jun Rekimoto

Sophisticated user interaction in the automotive industry is a fast emerging topic. Mid-air gestures and speech already have numerous applications for driver-car interaction. Additionally, multimodal approaches are being developed to…

Human-Computer Interaction · Computer Science 2020-12-29 Abdul Rafey Aftab , Michael von der Beeck , Michael Feld

This work proposes a biologically inspired approach that focuses on attention systems that are able to inhibit or constrain what is relevant at any one moment. We propose a radically new approach to making progress in human-robot joint…

Robotics · Computer Science 2016-06-09 Nick DePalma , Cynthia Breazeal
‹ Prev 1 4 5 6 7 8 10 Next ›