English
Related papers

Related papers: FoveaTer: Foveated Transformer for Image Classific…

200 papers

The Human visual perception of the world is of a large fixed image that is highly detailed and sharp. However, receptor density in the retina is not uniform: a small central region called the fovea is very dense and exhibits high…

Computer Vision and Pattern Recognition · Computer Science 2017-01-26 Alon Hazan , Yuval Harel , Ron Meir

This work proposes a novel face-swapping framework FlowFace++, utilizing explicit semantic flow supervision and end-to-end architecture to facilitate shape-aware face-swapping. Specifically, our work pretrains a facial shape discriminator…

Computer Vision and Pattern Recognition · Computer Science 2023-06-27 Yu Zhang , Hao Zeng , Bowen Ma , Wei Zhang , Zhimeng Zhang , Yu Ding , Tangjie Lv , Changjie Fan

A significant amount of redundancy exists between consecutive frames of a video. Object detectors typically produce detections for one image at a time, without any capabilities for taking advantage of this redundancy. Meanwhile, many…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Hughes Perreault , Guillaume-Alexandre Bilodeau , Nicolas Saunier , Maguelonne Héritier

Various stuff and things in visual data possess specific traits, which can be learned by deep neural networks and are implicitly represented as the visual prior, e.g., object location and shape, in the model. Such prior potentially impacts…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Jinheng Xie , Kai Ye , Yudong Li , Yuexiang Li , Kevin Qinghong Lin , Yefeng Zheng , Linlin Shen , Mike Zheng Shou

This survey explores the adaptation of visual transformer models in Autonomous Driving, a transition inspired by their success in Natural Language Processing. Surpassing traditional Recurrent Neural Networks in tasks like sequential image…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Quoc-Vinh Lai-Dang

Virtual Reality is regaining attention due to recent advancements in hardware technology. Immersive images / videos are becoming widely adopted to carry omnidirectional visual information. However, due to the requirements for higher spatial…

Image and Video Processing · Electrical Eng. & Systems 2021-06-15 Yize Jin , Anjul Patney , Alan Bovik

Spatial understanding of the semantics of the surroundings is a key capability needed by autonomous cars to enable safe driving decisions. Recently, purely vision-based solutions have gained increasing research interest. In particular,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Christian Witte , Jens Behley , Cyrill Stachniss , Marvin Raaijmakers

Current deep learning-based solutions for image analysis tasks are commonly incapable of handling problems to which multiple different plausible solutions exist. In response, posterior-based methods such as conditional Diffusion Models and…

Learning fine-grained details is a key issue in image aesthetic assessment. Most of the previous methods extract the fine-grained details via random cropping strategy, which may undermine the integrity of semantic information. Extensive…

Computer Vision and Pattern Recognition · Computer Science 2019-06-27 Xiaodan Zhang , Xinbo Gao , Wen Lu , Lihuo He

Translational invariance induced by pooling operations is an inherent property of convolutional neural networks, which facilitates numerous computer vision tasks such as classification. Yet to leverage rotational invariant tasks,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-11 Quentin Paletta , Anthony Hu , Guillaume Arbod , Philippe Blanc , Joan Lasenby

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

Robotics · Computer Science 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Transformer-based models have achieved great success in various NLP, vision, and speech tasks. However, the core of Transformer, the self-attention mechanism, has a quadratic time and memory complexity with respect to the sequence length,…

Computation and Language · Computer Science 2023-05-23 Chao-Hong Tan , Qian Chen , Wen Wang , Qinglin Zhang , Siqi Zheng , Zhen-Hua Ling

Background subtraction is a fundamental low-level processing task in numerous computer vision applications. The vast majority of algorithms process images on a pixel-by-pixel basis, where an independent decision is made for each pixel. A…

Computer Vision and Pattern Recognition · Computer Science 2013-03-19 Vikas Reddy , Conrad Sanderson , Brian C. Lovell

Dictionary learning algorithms or supervised deep convolution networks have considerably improved the efficiency of predefined feature representations such as SIFT. We introduce a deep scattering convolution network, with predefined wavelet…

Computer Vision and Pattern Recognition · Computer Science 2015-06-02 Edouard Oyallon , Stéphane Mallat

In this work, we present RadioTransformer, a novel visual attention-driven transformer framework, that leverages radiologists' gaze patterns and models their visuo-cognitive behavior for disease diagnosis on chest radiographs. Domain…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Moinak Bhattacharya , Shubham Jain , Prateek Prasanna

We show that adversarial examples, i.e., the visually imperceptible perturbations that result in Convolutional Neural Networks (CNNs) fail, can be alleviated with a mechanism based on foveations---applying the CNN in different image…

Machine Learning · Computer Science 2016-01-20 Yan Luo , Xavier Boix , Gemma Roig , Tomaso Poggio , Qi Zhao

Fast reactions to changes in the surrounding visual environment require efficient attention mechanisms to reallocate computational resources to most relevant locations in the visual field. While current computational models keep improving…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Lapo Faggi , Alessandro Betti , Dario Zanca , Stefano Melacci , Marco Gori

Research on human face processing using eye movements has provided evidence that we recognize face images successfully focusing our visual attention on a few inner facial regions, mainly on the eyes, nose and mouth. To understand how we…

Computer Vision and Pattern Recognition · Computer Science 2017-09-06 Carlos E. Thomaz , Vagner Amaral , Gilson A. Giraldi , Duncan F. Gillies , Daniel Rueckert

Integrating LiDAR and camera information in the bird's eye view (BEV) representation has demonstrated its effectiveness in 3D object detection. However, because of the fundamental disparity in geometric accuracy between these sensors,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Guowen Zhang , Chenhang He , Liyi Chen , Lei Zhang

Learning visual representations from observing actions to benefit robot visuo-motor policy generation is a promising direction that closely resembles human cognitive function and perception. Motivated by this, and further inspired by…