中文
相关论文

相关论文: FoveaTer: Foveated Transformer for Image Classific…

200 篇论文

In this paper, we tackle the challenge of actively attending to visual scenes using a foveated sensor. We introduce an end-to-end differentiable foveated active vision architecture that leverages a graph convolutional network to process…

计算机视觉与模式识别 · 计算机科学 2023-12-05 George Killick , Paul Henderson , Paul Siebert , Gerardo Aragon-Camarasa

We present a foveated object detector (FOD) as a biologically-inspired alternative to the sliding window (SW) approach which is the dominant method of search in computer vision object detection. Similar to the human visual system, the FOD…

计算机视觉与模式识别 · 计算机科学 2017-11-07 Emre Akbas , Miguel P. Eckstein

The human visual system processes images with varied degrees of resolution, with the fovea, a small portion of the retina, capturing the highest acuity region, which gradually declines toward the field of view's periphery. However, the…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Beatriz Paula , Plinio Moreno

From falcons spotting preys to humans recognizing faces, rapid visual abilities depend on a foveated retinal organization which delivers high-acuity central vision while preserving low-resolution periphery. This organization is conserved…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Jean-Nicolas Jérémie , Emmanuel Daucé , Laurent U Perrinet

The goal of this work is to characterize the representational impact that foveation operations have for machine vision systems, inspired by the foveated human visual system, which has higher acuity at the center of gaze and texture-like…

计算机视觉与模式识别 · 计算机科学 2021-06-24 Arturo Deza , Talia Konkle

Vision-language models benefit from high-resolution images, but the increase in visual-token count incurs high compute overhead. Humans resolve this tension via foveation: a coarse view guides "where to look", while selectively acquired…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Juhong Min , Lazar Valkov , Vitali Petsiuk , Hossein Souri , Deen Dayal Mohan

Foveated imaging provides a better tradeoff between situational awareness (field of view) and resolution and is critical in long-wavelength infrared regimes because of the size, weight, power, and cost of thermal sensors. We demonstrate…

Recently, virtual reality (VR) technology has been widely used in medical, military, manufacturing, entertainment, and other fields. These applications must simulate different complex material surfaces, various dynamic objects, and complex…

图形学 · 计算机科学 2022-11-22 Lili Wang , Xuehuai Shi , Yi Liu

Human vision is foveated, with variable resolution peaking at the center of a large field of view; this reflects an efficient trade-off for active sensing, allowing eye-movements to bring different parts of the world into focus with other…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Nicholas M. Blauch , George A. Alvarez , Talia Konkle

Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically reducing visual processing. In contrast, robot learning systems typically rely on passive, uniform…

机器人学 · 计算机科学 2025-09-23 Ian Chuang , Jinyu Zou , Andrew Lee , Dechen Gao , Iman Soltani

The problem of $\textit{visual metamerism}$ is defined as finding a family of perceptually indistinguishable, yet physically different images. In this paper, we propose our NeuroFovea metamer model, a foveated generative model that is based…

计算机视觉与模式识别 · 计算机科学 2019-01-01 Arturo Deza , Aditya Jonnalagadda , Miguel Eckstein

Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Brian Chao , Lior Yariv , Howard Xiao , Gordon Wetzstein

Emergent in the field of head mounted display design is a desire to leverage the limitations of the human visual system to reduce the computation, communication, and display workload in power and form-factor constrained systems. Fundamental…

Active perception and foveal vision are the foundations of the human visual system. While foveal vision reduces the amount of information to process during a gaze fixation, active perception will change the gaze direction to the most…

计算机视觉与模式识别 · 计算机科学 2022-08-25 Alexandre M. F. Dias , Luís Simões , Plinio Moreno , Alexandre Bernardino

Explainability in artificial intelligence (XAI) remains a crucial aspect for fostering trust and understanding in machine learning models. Current visual explanation techniques, such as gradient-based or class-activation-based methods,…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Mahadev Prasad Panda , Matteo Tiezzi , Martina Vilas , Gemma Roig , Bjoern M. Eskofier , Dario Zanca

Inspired by human vision, we propose a new periphery-fovea multi-resolution driving model that predicts vehicle speed from dash camera videos. The peripheral vision module of the model processes the full video frames in low resolution. Its…

计算机视觉与模式识别 · 计算机科学 2019-03-26 Ye Xia , Jinkyu Kim , John Canny , Karl Zipser , David Whitney

With the recent interest in virtual reality and augmented reality, there is a newfound demand for displays that can provide high resolution with a wide field of view (FOV). However, such displays incur significantly higher costs for…

图形学 · 计算机科学 2022-11-18 Susmija Jabbireddy , Xuetong Sun , Xiaoxu Meng , Amitabh Varshney

Parsing urban scene images benefits many applications, especially self-driving. Most of the current solutions employ generic image parsing models that treat all scales and locations in the images equally and do not consider the geometry…

计算机视觉与模式识别 · 计算机科学 2017-08-09 Xin Li , Zequn Jie , Wei Wang , Changsong Liu , Jimei Yang , Xiaohui Shen , Zhe Lin , Qiang Chen , Shuicheng Yan , Jiashi Feng

The core for tackling the fine-grained visual categorization (FGVC) is to learn subtle yet discriminative features. Most previous works achieve this by explicitly selecting the discriminative parts or integrating the attention mechanism via…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Jun Wang , Xiaohan Yu , Yongsheng Gao

Efficient processing of high-res video streams is safety-critical for many robotics applications such as autonomous driving. To maintain real-time performance, many practical systems downsample the video stream. But this can hurt downstream…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Chittesh Thavamani , Mengtian Li , Nicolas Cebron , Deva Ramanan
‹ 上一页 1 2 3 10 下一页 ›