English
Related papers

Related papers: Learning from Observer Gaze:Zero-Shot Attention Pr…

200 papers

The interaction decoder utilized in prevalent Transformer-based HOI detectors typically accepts pre-composed human-object pairs as inputs. Though achieving remarkable performance, such paradigm lacks feasibility and cannot explore novel…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Liulei Li , Jianan Wei , Wenguan Wang , Yi Yang

Open Vocabulary Human-Object Interaction (HOI) detection aims to detect interactions between humans and objects while generalizing to novel interaction classes beyond the training set. Current methods often rely on Vision and Language…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Ting Lei , Shaofeng Yin , Qingchao Chen , Yuxin Peng , Yang Liu

This study presents a framework for capturing human attention in the spatio-temporal domain using eye-tracking glasses. Attention mapping is a key technology for human perceptual activity analysis or Human-Robot Interaction (HRI) to support…

Robotics · Computer Science 2021-07-09 Shuji Oishi , Kenji Koide , Masashi Yokozuka , Atsuhiko Banno

Reasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring…

Computer Vision and Pattern Recognition · Computer Science 2019-09-19 Bo Wan , Desen Zhou , Yongfei Liu , Rongjie Li , Xuming He

We present HOIGaze - a novel learning-based approach for gaze estimation during hand-object interactions (HOI) in extended reality (XR). HOIGaze addresses the challenging HOI setting by building on one key insight: The eye, hand, and head…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Zhiming Hu , Daniel Haeufle , Syn Schmitt , Andreas Bulling

Since detecting and recognizing individual human or object are not adequate to understand the visual world, learning how humans interact with surrounding objects becomes a core technology. However, convolution operations are weak in…

Computer Vision and Pattern Recognition · Computer Science 2020-03-12 Dongming Yang , YueXian Zou , Jian Zhang , Ge Li

The understanding of where humans look in a scene is a problem of great interest in visual perception and computer vision. When eye-tracking devices are not a viable option, models of human attention can be used to predict fixations. In…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Dario Zanca , Marco Gori

Cameras are essential vision instruments to capture images for pattern detection and measurement. Human-object interaction (HOI) detection is one of the most popular pattern detection approaches for captured human-centric visual scenes.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-28 Manli Zhu , Edmond S. L. Ho , Shuang Chen , Longzhi Yang , Hubert P. H. Shum

One of the most fundamental and information-laden actions humans do is to look at objects. However, a survey of current works reveals that existing gaze-related datasets annotate only the pixel being looked at, and not the boundaries of a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Henri Tomas , Marcus Reyes , Raimarc Dionido , Mark Ty , Jonric Mirando , Joel Casimiro , Rowel Atienza , Richard Guinto

This work proposes a biologically inspired approach that focuses on attention systems that are able to inhibit or constrain what is relevant at any one moment. We propose a radically new approach to making progress in human-robot joint…

Robotics · Computer Science 2016-06-09 Nick DePalma , Cynthia Breazeal

In human-object interactions (HOI) recognition, conventional methods consider the human body as a whole and pay a uniform attention to the entire body region. They ignore the fact that normally, human interacts with an object by using some…

Computer Vision and Pattern Recognition · Computer Science 2018-07-31 Hao-Shu Fang , Jinkun Cao , Yu-Wing Tai , Cewu Lu

Predicting driver attention is a critical problem for developing explainable autonomous driving systems and understanding driver behavior in mixed human-autonomous vehicle traffic scenarios. Although significant progress has been made…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Shreedhar Govil , Didier Stricker , Jason Rambach

Vehicles with a safety function for anticipating crashes in advance can enhance drivers' ability to avoid crashes. As dashboard cameras have become a low-cost sensor device accessible to almost every vehicle, deep neural networks for crash…

Human-Computer Interaction · Computer Science 2022-11-11 Yu Li , Muhammad Monjurul Karim , Ruwen Qin

Attention mechanisms have been widely applied in the Visual Question Answering (VQA) task, as they help to focus on the area-of-interest of both visual and textual information. To answer the questions correctly, the model needs to…

Computer Vision and Pattern Recognition · Computer Science 2017-09-20 Tingting Qiao , Jianfeng Dong , Duanqing Xu

Human-object contact (HOT) is designed to accurately identify the areas where humans and objects come into contact. Current methods frequently fail to account for scenarios where objects are frequently blocking the view, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yuxiao Wang , Wenpeng Neng , Zhenao Wei , Yu Lei , Weiying Xue , Nan Zhuang , Yanwu Xu , Xinyu Jiang , Qi Liu

Trajectory prediction for multi-agent interaction scenarios is a crucial challenge. Most advanced methods model agent interactions by efficiently factorized attention based on the temporal and agent axes. However, this static and foward…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Mingjin Zeng , Nan Ouyang , Wenkang Wan , Lei Ao , Qing Cai , Kai Sheng

Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is tied to physical priors such as force closure or generic…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Guangyi Han , Wei Zhai , Yuhang Yang , Yang Cao , Zheng-Jun Zha

Visual Question Answering (VQA) models employ attention mechanisms to discover image locations that are most relevant for answering a specific question. For this purpose, several multimodal fusion strategies have been proposed, ranging from…

Computer Vision and Pattern Recognition · Computer Science 2021-08-26 Moshiur R Farazi , Salman H Khan , Nick Barnes

Scene graph generation (SGG) and human-object interaction (HOI) detection are two important visual tasks aiming at localising and recognising relationships between objects, and interactions between humans and objects, respectively.…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Tao He , Lianli Gao , Jingkuan Song , Yuan-Fang Li

A visual hard attention model actively selects and observes a sequence of subregions in an image to make a prediction. The majority of hard attention models determine the attention-worthy regions by first analyzing a complete image.…

Computer Vision and Pattern Recognition · Computer Science 2021-11-16 Samrudhdhi B. Rangrej , James J. Clark