English
Related papers

Related papers: Saccadic Predictive Vision Model with a Fovea

200 papers

Neuromorphic vision sensors or event cameras have made the visual perception of extremely low reaction time possible, opening new avenues for high-dynamic robotics applications. These event cameras' output is dependent on both motion and…

Structured prediction tasks pose a fundamental trade-off between the need for model complexity to increase predictive power and the limited computational resources for inference in the exponentially-sized output spaces such models require.…

Machine Learning · Statistics 2012-08-17 David Weiss , Benjamin Sapp , Ben Taskar

Parsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consider the geometry…

Computer Vision and Pattern Recognition · Computer Science 2017-08-09 Xin Li , Zequn Jie , Wei Wang , Changsong Liu , Jimei Yang , Xiaohui Shen , Zhe Lin , Qiang Chen , Shuicheng Yan , Jiashi Feng

Eye movements hold information about human perception, intention, and cognitive state. We propose a novel eye movement simulator that i) probabilistically simulates saccade movements as gamma distributions considering different peak…

Human-Computer Interaction · Computer Science 2018-09-11 Wolfgang Fuhl , Enkelejda Kasneci

In-flight objects capture is extremely challenging. The robot is required to complete trajectory prediction, interception position calculation and motion planning in sequence within tens of milliseconds. As in-flight uneven objects are…

Robotics · Computer Science 2021-03-16 Hongxiang Yu , Dashun Guo , Huan Yin , Anzhe Chen , Kechun Xu , Yue Wang , Rong Xiong

From falcons spotting preys to humans recognizing faces, rapid visual abilities depend on a foveated retinal organization which delivers high-acuity central vision while preserving low-resolution periphery. This organization is conserved…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jean-Nicolas Jérémie , Emmanuel Daucé , Laurent U Perrinet

The process of planning views to observe a scene is known as the Next Best View (NBV) problem. Approaches often aim to obtain high-quality scene observations while reducing the number of views, travel distance and computational cost.…

Robotics · Computer Science 2021-02-16 Rowan Border , Jonathan D. Gammell

Multimodal large language models (MLLMs) achieve ever-stronger performance on visual-language tasks. Even as traditional visual question answering (VQA) benchmarks approach saturation, reliable deployment requires satisfying low error…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Hector G. Rodriguez , Marcus Rohrbach

Convolutional layers are a major driving force behind the successes of deep learning. Pointwise convolution (PWC) is a 1x1 convolutional filter that is primarily used for parameter reduction. However, the PWC ignores the spatial information…

Computer Vision and Pattern Recognition · Computer Science 2020-02-07 Pratik Mazumder , Pravendra Singh , Vinay Namboodiri

Monocular visual odometry (VO) suffers severely from error accumulation during frame-to-frame pose estimation. In this paper, we present a self-supervised learning method for VO with special consideration for consistency over longer…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Yuliang Zou , Pan Ji , Quoc-Huy Tran , Jia-Bin Huang , Manmohan Chandraker

Recent advancements in sequence prediction have significantly improved the accuracy of video data interpretation; however, existing models often overlook the potential of attention-based mechanisms for next-frame prediction. This study…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Yiqiao Yin

We propose a novel attention model that can accurately attends to target objects of various scales and shapes in images. The model is trained to gradually suppress irrelevant regions in an input image via a progressive attentive process…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Paul Hongsuck Seo , Zhe Lin , Scott Cohen , Xiaohui Shen , Bohyung Han

Deep robot vision models are widely used for recognizing objects from camera images, but shows poor performance when detecting objects at untrained positions. Although such problem can be alleviated by training with large datasets, the…

Robotics · Computer Science 2022-10-26 Hyogo Hiruma , Hiroki Mori , Hiroshi Ito , Tetsuya Ogata

A custom head-mounted system to track smooth eye movements for control of a mouse cursor is implemented and evaluated. The system comprises a head-mounted infrared camera, an infrared light source, and a computer. Software-based image…

Human-Computer Interaction · Computer Science 2020-12-01 Adam Pantanowitz , Kimoon Kim , Chelsey Chewins , Isabel N. K. Tollman , David M. Rubin

Gaze object prediction (GOP) aims to predict the category and location of the object that a human is looking at. Previous methods utilized box-level supervision to identify the object that a person is looking at, but struggled with semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Yang Jin , Lei Zhang , Shi Yan , Bin Fan , Binglu Wang

Achieving human-like reasoning in deep learning models for complex tasks in unknown environments remains a critical challenge in embodied intelligence. While advanced vision-language models (VLMs) excel in static scene understanding, their…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jinzhou Tang , Jusheng zhang , Sidi Liu , Waikit Xiu , Qinhan Lv , Xiying Li

Systems involving human-robot collaboration necessarily require that steps be taken to ensure safety of the participating human. This is usually achievable if accurate, reliable estimates of the human's pose are available. In this paper, we…

Robotics · Computer Science 2023-10-30 Michael Zechmair , Alban Bornet , Yannick Morel

Low-cost autonomous agents including autonomous driving vehicles chiefly adopt monocular 3D object detection to perceive surrounding environment. This paper studies 3D intermediate representation methods which generate intermediate 3D…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Qian Ye , Ling Jiang , Wang Zhen , Yuyang Du

Vision models are often vulnerable to out-of-distribution (OOD) samples without adapting. While visual prompts offer a lightweight method of input-space adaptation for large-scale vision models, they rely on a high-dimensional additive…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Yun-Yun Tsai , Chengzhi Mao , Junfeng Yang

Semantic segmentation is an important branch of image processing and computer vision. With the popularity of deep learning, various convolutional neural networks have been proposed for pixel-level classification and segmentation tasks. In…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Xinyu Xu , Huazhen Liu , Tao Zhang , Huilin Xiong , Wenxian Yu
‹ Prev 1 3 4 5 6 7 10 Next ›