English
Related papers

Related papers: Learning to Adapt to Position Bias in Vision Trans…

200 papers

Visual localization, which estimates a camera's pose within a known scene, is a fundamental capability for autonomous systems. While absolute pose regression (APR) methods have shown promise for efficient inference, they often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Sihang Li , Siqi Tan , Bowen Chang , Jing Zhang , Chen Feng , Yiming Li

Conditional neural processes (CNPs) are a flexible and efficient family of models that learn to learn a stochastic process from data. They have seen particular application in contextual image completion - observing pixel values at some…

Machine Learning · Computer Science 2024-02-20 Victor Prokhorov , Ivan Titov , N. Siddharth

Vision Transformers (ViTs) have achieved comparable or superior performance than Convolutional Neural Networks (CNNs) in computer vision. This empirical breakthrough is even more remarkable since, in contrast to CNNs, ViTs do not embed any…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Samy Jelassi , Michael E. Sander , Yuanzhi Li

Recent visual autonomous perception systems achieve remarkable performances with deep representation learning. However, they fail in scenarios with challenging illumination.While event cameras can mitigate this problem, there is a lack of a…

Robotics · Computer Science 2026-03-18 Jinghang Li , Shichao Li , Qing Lian , Peiliang Li , Xiaozhi Chen , Yi Zhou

We propose a novel probabilistic model for visual question answering (Visual QA). The key idea is to infer two sets of embeddings: one for the image and the question jointly and the other for the answers. The learning objective is to learn…

Computer Vision and Pattern Recognition · Computer Science 2018-06-12 Hexiang Hu , Wei-Lun Chao , Fei Sha

Object pose estimation methods allow finding locations of objects in unstructured environments. This is a highly desired skill for autonomous robot manipulation as robots need to estimate the precise poses of the objects in order to…

Robotics · Computer Science 2022-03-22 Tarik Kelestemur , Robert Platt , Taskin Padir

Vision transformers have shown great potential in various computer vision tasks owing to their strong capability to model long-range dependency using the self-attention mechanism. Nevertheless, they treat an image as a 1D sequence of visual…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Qiming Zhang , Yufei Xu , Jing Zhang , Dacheng Tao

Classification is a major tool of statistics and machine learning. A classification method first processes a training set of objects with given classes (labels), with the goal of afterward assigning new objects to one of these classes. When…

Machine Learning · Statistics 2024-07-08 Jakob Raymaekers , Peter J. Rousseeuw , Mia Hubert

Depictions of similar human body configurations can vary with changing viewpoints. Using only 2D information, we would like to enable vision algorithms to recognize similarity in human body poses across multiple views. This ability is…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Jennifer J. Sun , Jiaping Zhao , Liang-Chieh Chen , Florian Schroff , Hartwig Adam , Ting Liu

Accurate localization ability is fundamental in autonomous driving. Traditional visual localization frameworks approach the semantic map-matching problem with geometric models, which rely on complex parameter tuning and thus hinder…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Zhihuang Zhang , Meng Xu , Wenqiang Zhou , Tao Peng , Liang Li , Stefan Poslad

Transformers with causal attention can solve tasks that require positional information without using positional encodings. In this work, we propose and investigate a new hypothesis about how positional information can be stored without…

Computation and Language · Computer Science 2025-01-03 Chunsheng Zuo , Pavel Guerzhoy , Michael Guerzhoy

An accurate understanding of a self-driving vehicle's surrounding environment is crucial for its navigation system. To enhance the effectiveness of existing algorithms and facilitate further research, it is essential to provide…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Abtin Mahyar , Hossein Motamednia , Dara Rahmati

Transformers have emerged as a universal backbone across 3D perception, video generation, and world models for autonomous driving and embodied AI, where understanding camera geometry is essential for grounding visual observations in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Cheng Zhang , Boying Li , Meng Wei , Yan-Pei Cao , Camilo Cruz Gambardella , Dinh Phung , Jianfei Cai

Intelligent behaviour in the real-world requires the ability to acquire new knowledge from an ongoing sequence of experiences while preserving and reusing past knowledge. We propose a novel algorithm for unsupervised representation learning…

Estimating the camera's pose given images from a single camera is a traditional task in mobile robots and autonomous vehicles. This problem is called monocular visual odometry and often relies on geometric approaches that require…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 André O. Françani , Marcos R. O. A. Maximo

We study view-invariant imitation learning by explicitly conditioning policies on camera extrinsics. Using Plucker embeddings of per-pixel rays, we show that conditioning on extrinsics significantly improves generalization across viewpoints…

Visual pose regression models estimate the camera pose from a query image with a single forward pass. Current models learn pose encoding from an image using deep convolutional networks which are trained per scene. The resulting encoding is…

Computer Vision and Pattern Recognition · Computer Science 2020-12-23 Yoli Shavit , Ron Ferens

Machine learning for image classification is an active and rapidly developing field. With the proliferation of classifiers of different sizes and different architectures, the problem of choosing the right model becomes more and more…

Computer Vision and Pattern Recognition · Computer Science 2025-08-01 David A. Kelly , Akchunya Chanchal , Nathan Blake

Body shape plays an important role in determining what garments will best suit a given person, yet today's clothing recommendation methods take a "one shape fits all" approach. These body-agnostic vision methods and datasets are a barrier…

Computer Vision and Pattern Recognition · Computer Science 2020-03-31 Wei-Lin Hsiao , Kristen Grauman

We add one more invariance - the state invariance - to the more commonly used other invariances for learning object representations for recognition and retrieval. By state invariance, we mean robust with respect to changes in the structural…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Rohan Sarkar , Avinash Kak