English
Related papers

Related papers: Saccadic Predictive Vision Model with a Fovea

200 papers

The ability to make educated predictions about their surroundings, and associate them with certain confidence, is important for intelligent systems, like autonomous vehicles and robots. It allows them to plan early and decide accordingly.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Liqian Ma , Stamatios Georgoulis , Xu Jia , Luc Van Gool

Fast reactions to changes in the surrounding visual environment require efficient attention mechanisms to reallocate computational resources to most relevant locations in the visual field. While current computational models keep improving…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Lapo Faggi , Alessandro Betti , Dario Zanca , Stefano Melacci , Marco Gori

Current video generation models produce physically inconsistent motion that violates real-world dynamics. We propose TrajVLM-Gen, a two-stage framework for physics-aware image-to-video generation. First, we employ a Vision Language Model to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Fan Yang , Zhiyang Chen , Yousong Zhu , Xin Li , Jinqiao Wang

In this paper, we introduce a rotational primitive prediction based 6D object pose estimation using a single image as an input. We solve for the 6D object pose of a known object relative to the camera using a single image with occlusion.…

Computer Vision and Pattern Recognition · Computer Science 2020-07-06 Myung-Hwan Jeon , Ayoung Kim

Accurate sarcopenia diagnosis via ultrasound remains challenging due to subtle imaging cues, limited labeled data, and the absence of clinical context in most models. We propose MedVQA-TREE, a multimodal framework that integrates a…

Image and Video Processing · Electrical Eng. & Systems 2025-08-28 Pardis Moradbeiki , Nasser Ghadiri , Sayed Jalal Zahabi , Uffe Kock Wiil , Kristoffer Kittelmann Brockhattingen , Ali Ebrahimi

Source-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which…

Computer Vision and Pattern Recognition · Computer Science 2026-01-22 Huizai Yao , Sicheng Zhao , Pengteng Li , Yi Cui , Shuo Lu , Weiyu Guo , Yunfan Lu , Yijie Xu , Hui Xiong

To fully understand the 3D context of a single image, a visual system must be able to segment both the visible and occluded regions of objects, while discerning their occlusion order. Ideally, the system should be able to handle any object…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Jiayang Ao , Qiuhong Ke , Krista A. Ehinger

Recently, multiple formulations of vision problems as probabilistic inversions of generative models based on computer graphics have been proposed. However, applications to 3D perception from natural images have focused on low-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2014-07-08 Tejas D. Kulkarni , Vikash K. Mansinghka , Pushmeet Kohli , Joshua B. Tenenbaum

The understanding of where humans look in a scene is a problem of great interest in visual perception and computer vision. When eye-tracking devices are not a viable option, models of human attention can be used to predict fixations. In…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Dario Zanca , Marco Gori

We propose a computational model of visual search that incorporates Bayesian interpretations of the neural mechanisms that underlie categorical perception and saccade planning. To enable meaningful comparisons between simulated and human…

Computer Vision and Pattern Recognition · Computer Science 2020-06-08 Maell Cullen , Jonathan Monney , M. Berk Mirza , Rosalyn Moran

Tracking vehicles in LIDAR point clouds is a challenging task due to the sparsity of the data and the dense search space. The lack of structure in point clouds impedes the use of convolution filters usually employed in 2D object tracking.…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Jesus Zarzar , Silvio Giancola , Bernard Ghanem

Instruction-following Vision Large Language Models (VLLMs) have achieved significant progress recently on a variety of tasks. These approaches merge strong pre-trained vision models and large language models (LLMs). Since these components…

Machine Learning · Computer Science 2024-02-20 Yiyang Zhou , Chenhang Cui , Rafael Rafailov , Chelsea Finn , Huaxiu Yao

We present a method to populate an unknown environment with models of previously seen objects, placed in a Euclidean reference frame that is inferred causally and on-line using monocular video along with inertial sensors. The system we…

Computer Vision and Pattern Recognition · Computer Science 2018-10-25 Xiaohan Fei , Stefano Soatto

Traditional geometric registration based estimation methods only exploit the CAD model implicitly, which leads to their dependence on observation quality and deficiency to occlusion. To address the problem,the paper proposes a bidirectional…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Yuhao Yang , Jun Wu , Yue Wang , Guangjian Zhang , Rong Xiong

As big spatial data becomes increasingly prevalent, classical spatiotemporal (ST) methods often do not scale well. While methods have been developed to account for high-dimensional spatial objects, the setting where there are exceedingly…

Applications · Statistics 2019-08-27 Samuel I. Berchuck , Felipe A. Medeiros , Sayan Mukherjee

Model-based eye tracking has been a dominant approach for eye gaze tracking because of its ability to generalize to different subjects, without the need of any training data and eye gaze annotations. Model-based eye tracking, however, is…

Computer Vision and Pattern Recognition · Computer Science 2021-06-28 Qiang Ji , Kang Wang

Surveying 3D scenes is a common task in robotics. Systems can do so autonomously by iteratively obtaining measurements. This process of planning observations to improve the model of a scene is called Next Best View (NBV) planning. NBV…

Robotics · Computer Science 2018-11-16 Rowan Border , Jonathan D. Gammell , Paul Newman

Estimating a semantically segmented bird's-eye-view (BEV) map from a single image has become a popular technique for autonomous control and navigation. However, they show an increase in localization error with distance from the camera.…

Computer Vision and Pattern Recognition · Computer Science 2022-04-07 Avishkar Saha , Oscar Mendez , Chris Russell , Richard Bowden

The ability to selectively attend to relevant stimuli while filtering out distractions is essential for agents that process complex, high-dimensional sensory input. This paper introduces a model of covert and overt visual attention through…

Computer Vision and Pattern Recognition · Computer Science 2025-05-09 Tin Mišić , Karlo Koledić , Fabio Bonsignorio , Ivan Petrović , Ivan Marković

Foveal vision makes up less than 1% of the visual field. The other 99% is peripheral vision. Precisely what human beings see in the periphery is both obvious and mysterious in that we see it with our own eyes but can't visualize what we…

Neural and Evolutionary Computing · Computer Science 2017-10-24 Lex Fridman , Benedikt Jenik , Shaiyan Keshvari , Bryan Reimer , Christoph Zetzsche , Ruth Rosenholtz