English
Related papers

Related papers: Mind the GAP: Glimpse-based Active Perception impr…

200 papers

Existing Temporal Action Detection (TAD) methods typically take a pre-processing step in converting an input varying-length video into a fixed-length snippet representation sequence, before temporal boundary estimation and action…

Computer Vision and Pattern Recognition · Computer Science 2023-03-06 Sauradip Nag , Xiatian Zhu , Yi-Zhe Song , Tao Xiang

Retailers have long been searching for ways to effectively understand their customers' behaviour in order to provide a smooth and pleasant shopping experience that attracts more customers everyday and maximises their revenue, consequently.…

Computer Vision and Pattern Recognition · Computer Science 2019-06-27 Mohammad Mahdi Kazemi Moghaddam , Ehsan Abbasnejad , Javen Shi

Eye movement is closely related to limb actions, so it can be used to infer movement intentions. More importantly, in some cases, eye movement is the only way for paralyzed and impaired patients with severe movement disorders to communicate…

Robotics · Computer Science 2020-12-17 Bo Yang , Jian Huang , Xiaolong Li , Xinxing Chen , Caihua Xiong , Yasuhisa Hasegawa

When working around other agents such as humans, it is important to model their perception capabilities to predict and make sense of their behavior. In this work, we consider agents whose perception capabilities are determined by their…

Robotics · Computer Science 2025-08-12 Maulik Bhatt , HongHao Zhen , Monroe Kennedy , Negar Mehr

This paper proposes an interactive system for mobile devices controlled by hand gestures aimed at helping people with visual impairments. This system allows the user to interact with the device by making simple static and dynamic hand…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Samer Alashhab , Antonio Javier Gallego , Miguel Ángel Lozano

Object-aware reasoning in vision-language tasks poses significant challenges for current models, particularly in handling unseen objects, reducing hallucinations, and capturing fine-grained relationships in complex visual scenes. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Antonio Carlos Rivera , Anthony Moore , Steven Robinson

Humans, this species expert in grasp detection, can grasp objects by taking into account hand-object positioning information. This work proposes a method to enable a robot manipulator to learn the same, grasping objects in the most optimal…

Robotic manipulation of highly deformable cloth presents a promising opportunity to assist people with several daily tasks, such as washing dishes; folding laundry; or dressing, bathing, and hygiene assistance for individuals with severe…

Robotics · Computer Science 2022-08-12 Yufei Wang , David Held , Zackory Erickson

To have a robot actively supporting a human during a collaborative task, it is crucial that robots are able to identify the current action in order to predict the next one. Common approaches make use of high-level knowledge, such as object…

Robotics · Computer Science 2017-03-08 Markus Eich , Sareh Shirazi , Gordon Wyeth

This paper addresses the challenging problem of estimating the general visual attention of people in images. Our proposed method is designed to work across multiple naturalistic social scenarios and provides a full picture of the subject's…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Eunji Chong , Nataniel Ruiz , Yongxin Wang , Yun Zhang , Agata Rozga , James Rehg

Humans inevitably develop a sense of the relationships between objects, some of which are based on their appearance. Some pairs of objects might be seen as being alternatives to each other (such as two pairs of jeans), while others may be…

Computer Vision and Pattern Recognition · Computer Science 2015-06-17 Julian McAuley , Christopher Targett , Qinfeng Shi , Anton van den Hengel

CLIP has emerged as a powerful multimodal model capable of connecting images and text through joint embeddings, but to what extent does it 'see' the same way humans do - especially when interpreting artworks? In this paper, we investigate…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Andrea Asperti , Leonardo Dessì , Maria Chiara Tonetti , Nico Wu

We cast visual imitation as a visual correspondence problem. Our robotic agent is rewarded when its actions result in better matching of relative spatial configurations for corresponding visual entities detected in its workspace and…

Robotics · Computer Science 2020-03-06 Maximilian Sieb , Zhou Xian , Audrey Huang , Oliver Kroemer , Katerina Fragkiadaki

Robots which interact with the physical world will benefit from a fine-grained tactile understanding of objects and surfaces. Additionally, for certain tasks, robots may need to know the haptic properties of an object before touching it. To…

Robotics · Computer Science 2016-04-13 Yang Gao , Lisa Anne Hendricks , Katherine J. Kuchenbecker , Trevor Darrell

Robotic grasping, the ability of robots to reliably secure and manipulate objects of varying shapes, sizes and orientations, is a complex task that requires precise perception and control. Deep neural networks have shown remarkable success…

Region-based artificial attention constitutes a framework for bio-inspired attentional processes on an intermediate abstraction level for the use in computer vision and mobile robotics. Segmentation algorithms produce regions of coherently…

Computer Vision and Pattern Recognition · Computer Science 2013-07-23 Jan Tünnermann , Dieter Enns , Bärbel Mertsching

Global Average Pooling (GAP) [4] has been used previously to generate class activation for image classification tasks. The motivation behind SIMILARnet comes from the fact that the convolutional filters possess position information of the…

Computer Vision and Pattern Recognition · Computer Science 2017-11-09 Arna Ghosh , Biswarup Bhattacharya , Somnath Basu Roy Chowdhury

Eye movements are crucial in understanding complex scenes. By predicting where humans look in natural scenes, we can understand how they percieve scenes and priotriaze information for further high-level processing. Here, we study the effect…

Computer Vision and Pattern Recognition · Computer Science 2015-05-15 Ali Borji , Mengyang Feng , Huchuan Lu

Joint attention is a core, early-developing form of social interaction. It is based on our ability to discriminate the third party objects that other people are looking at. While it has been shown that people can accurately determine…

Artificial Intelligence · Computer Science 2014-12-09 Tao Gao , Daniel Harari , Joshua Tenenbaum , Shimon Ullman

We introduce Gaussian masking for Language-Image Pre-Training (GLIP) a novel, straightforward, and effective technique for masking image patches during pre-training of a vision-language model. GLIP builds on Fast Language-Image Pre-Training…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Mingliang Liang , Martha Larson