English
Related papers

Related papers: Saccadic Predictive Vision Model with a Fovea

200 papers

Modeling perception is critical for many applications and developments in computer graphics to optimize and evaluate content generation techniques. Most of the work to date has focused on central (foveal) vision. However, this is…

Graphics · Computer Science 2022-09-20 Cara Tursun , Piotr Didyk

Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this tension via foveation: a coarse view guides "where to look", while selectively acquired…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Juhong Min , Lazar Valkov , Vitali Petsiuk , Hossein Souri , Deen Dayal Mohan

Accurate object detection and prediction are critical to ensure the safety and efficiency of self-driving architectures. Predicting object trajectories and occupancy enables autonomous vehicles to anticipate movements and make decisions…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Miguel Antunes-García , Luis M. Bergasa , Santiago Montiel-Marín , Rafael Barea , Fabio Sánchez-García , Ángel Llamazares

The objects we perceive guide our eye movements when observing real-world dynamic scenes. Yet, gaze shifts and selective attention are critical for perceiving details and refining object boundaries. Object segmentation and gaze behavior…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Vito Mengers , Nicolas Roth , Oliver Brock , Klaus Obermayer , Martin Rolfs

The posterior parietal cortex is believed to direct eye movements, especially in regards to target tracking tasks, and a number of debates exist over the precise nature of the computations performed by the parietal cortex, with each side…

Neural and Evolutionary Computing · Computer Science 2012-10-16 Sacha Sokoloski

In this work we aim to mimic the human ability to acquire the intuition to estimate the performance of a design from visual inspection and experience alone. We study the ability of convolutional neural networks to predict static and dynamic…

Image and Video Processing · Electrical Eng. & Systems 2021-11-29 Philippe M. Wyder , Hod Lipson

Visual perception in modern Vision-Language Models (VLMs) is constrained by a perceptual bandwidth bottleneck: a broad field of view preserves global context but sacrifices the fine-grained details required for complex reasoning. We argue…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Anjie Liu , Ziqin Gong , Yan Song , Yuxiang Chen , Xiaolong Liu , Hengtong Lu , Kaike Zhang , Chen Wei , Jun Wang

Vision, as an inexpensive yet information rich sensor, is commonly used for perception on autonomous mobile robots. Unfortunately, accurate vision-based perception requires a number of assumptions about the environment to hold -- some…

Robotics · Computer Science 2019-08-01 Sadegh Rabiee , Joydeep Biswas

We present an approach to learn an object-centric forward model, and show that this allows us to plan for sequences of actions to achieve distant desired goals. We propose to model a scene as a collection of objects, each with an explicit…

Computer Vision and Pattern Recognition · Computer Science 2019-10-09 Yufei Ye , Dhiraj Gandhi , Abhinav Gupta , Shubham Tulsiani

The cerebellum plays a distinctive role within our motor control system to achieve fine and coordinated motions. While cerebellar lesions do not lead to a complete loss of motor functions, both action and perception are severally impacted.…

Robotics · Computer Science 2020-11-04 Omar Zahra , David Navarro-Alarcon , Silvia Tolu

Human eye-hand coordination relies on internal forward models that predict future states and compensate for sensory delays. During line tracing, the gaze typically leads the hand through predictive saccades, yet the extent to which this…

Neurons and Cognition · Quantitative Biology 2026-04-08 Emiko Shishido

Despite remarkable progress in recent years, Vision Language Models (VLMs) remain prone to overconfidence and hallucinations on tasks such as Visual Question Answering (VQA) and Visual Reasoning. Bayesian methods can potentially improve…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Tobias Jan Wieczorek , Nathalie Daun , Mohammad Emtiyaz Khan , Marcus Rohrbach

Field-of-View (FoV) adaptive streaming significantly reduces bandwidth requirement of immersive point cloud video (PCV) by only transmitting visible points in a viewer's FoV. The traditional approaches often focus on trajectory-based 6…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Chen Li , Tongyu Zong , Yueyu Hu , Yao Wang , Yong Liu

Recent self-supervised learning models simulate the development of semantic object representations by training on visual experience similar to that of toddlers. However, these models ignore the foveated nature of human vision with high/low…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Zhengyang Yu , Arthur Aubret , Chen Yu , Jochen Triesch

In modern video coding standards, block-based inter prediction is widely adopted, which brings high compression efficiency. However, in natural videos, there are usually multiple moving objects of arbitrary shapes, resulting in complex…

Image and Video Processing · Electrical Eng. & Systems 2024-09-13 Zhuoyuan Li , Zikun Yuan , Li Li , Dong Liu , Xiaohu Tang , Feng Wu

Humans combine prediction and perception to observe the world. When faced with rapidly moving birds or insects, we can only perceive them clearly by predicting their next position and focusing our gaze there. Inspired by this, this paper…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Song Zhang , Haoyu Chen , Ruibo Wang

This paper quantifies an error source that limits the accuracy of lidar scan matching, particularly for voxel-based methods. Lidar scan matching, which is used in dead reckoning (also known as lidar odometry) and mapping, computes the…

Robotics · Computer Science 2024-01-25 Jason Rife , Matthew McDermott

Inspired by foveal vision, hard attention models promise interpretability and parameter economy. However, existing models like the Recurrent Model of Visual Attention (RAM) and Deep Recurrent Attention Model (DRAM) failed to model the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Pengcheng Pan , Yonekura Shogo , Yasuo Kuniyoshi

Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While deep learning-based methods have recently shown superior accuracy compared to traditional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Dominik Kuczkowski , Laura Ruotsalainen

Objects are composed of a set of geometrically organized parts. We introduce an unsupervised capsule autoencoder (SCAE), which explicitly uses geometric relationships between parts to reason about objects. Since these relationships do not…

Machine Learning · Statistics 2019-12-03 Adam R. Kosiorek , Sara Sabour , Yee Whye Teh , Geoffrey E. Hinton