English
Related papers

Related papers: FishRoPE: Projective Rotary Position Embeddings fo…

200 papers

Localization is one of the core parts of modern robotics. Classic localization methods typically follow the retrieve-then-register paradigm, achieving remarkable success. Recently, the emergence of end-to-end localization approaches has…

Robotics · Computer Science 2025-03-17 Ziyue Wang , Chenghao Shi , Neng Wang , Qinghua Yu , Xieyuanli Chen , Huimin Lu

Fisheye cameras are widely employed in automatic parking, and the video stream object detection (VSOD) of the fisheye camera is a fundamental perception function to ensure the safe operation of vehicles. In past research work, the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-30 Yixiong Yan , Liangzhu Cheng , Yongxu Li , Xinjuan Tuo

We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each…

Rotary Positional Encodings (RoPE) have emerged as a highly effective technique for one-dimensional sequences in Natural Language Processing spurring recent progress towards generalizing RoPE to higher-dimensional data such as images and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Chase van de Geijn , Timo Lüddecke , Polina Turishcheva , Alexander S. Ecker

Monocular visual odometry (VO) is a fundamental computer vision problem with applications in autonomous navigation, augmented reality and more. While deep learning-based methods have recently shown superior accuracy compared to traditional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-27 Dominik Kuczkowski , Laura Ruotsalainen

Multi-view camera-only 3D object detection largely follows two primary paradigms: exploiting bird's-eye-view (BEV) representations or focusing on perspective-view (PV) features, each with distinct advantages. Although several recent…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Zhe Huang , Yizhe Zhao , Hao Xiao , Chenyan Wu , Lingting Ge

Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream tasks such as retrieval, classification, and segmentation. However, such representations tend to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Jona Ruthardt , Manu Gaur , Deva Ramanan , Makarand Tapaswi , Yuki M. Asano

Autonomous driving systems require extensive data collection schemes to cover the diverse scenarios needed for building a robust and safe system. The data volumes are in the order of Exabytes and have to be stored for a long period of time…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Madhumitha Sakthi , Louis Kerofsky , Varun Ravi Kumar , Senthil Yogamani

A self-driving perception model aims to extract 3D semantic representations from multiple cameras collectively into the bird's-eye-view (BEV) coordinate frame of the ego car in order to ground downstream planner. Existing perception methods…

Computer Vision and Pattern Recognition · Computer Science 2022-08-19 Jiachen Lu , Zheyuan Zhou , Xiatian Zhu , Hang Xu , Li Zhang

Infrastructure-based perception plays a crucial role in intelligent transportation systems, offering global situational awareness and enabling cooperative autonomy. However, existing camera-based detection models often underperform in such…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Yun Zhang , Zhaoliang Zheng , Johnson Liu , Zhiyu Huang , Zewei Zhou , Zonglin Meng , Tianhui Cai , Jiaqi Ma

Although recent learning-based calibration methods can predict extrinsic and intrinsic camera parameters from a single image, the accuracy of these methods is degraded in fisheye images. This degradation is caused by mismatching between the…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Nobuhiko Wakai , Satoshi Sato , Yasunori Ishii , Takayoshi Yamashita

Achieving robust and real-time 3D perception is fundamental for autonomous vehicles. While most existing 3D perception methods prioritize detection accuracy, they often overlook critical aspects such as computational efficiency, onboard…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Trung Pham , Mehran Maghoumi , Wanli Jiang , Bala Siva Sashank Jujjavarapu , Mehdi Sajjadi , Xin Liu , Hsuan-Chu Lin , Bor-Jeng Chen , Giang Truong , Chao Fang , Junghyun Kwon , Minwoo Park

View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most prominent VTM…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Zhiqi Li , Zhiding Yu , Wenhai Wang , Anima Anandkumar , Tong Lu , Jose M. Alvarez

Underwater visual localization remains challenging due to wavelength-dependent attenuation, poor texture, and non-Gaussian sensor noise. We introduce MARVO, a physics-aware, learning-integrated odometry framework that fuses underwater image…

Robotics · Computer Science 2025-12-01 Sacchin Sundar , Atman Kikani , Aaliya Alam , Sumukh Shrote , A. Nayeemulla Khan , A. Shahina

Goal-driven mobile robot navigation in map-less environments requires effective state representations for reliable decision-making. Inspired by the favorable properties of Bird's-Eye View (BEV) in point clouds for visual perception, this…

Robotics · Computer Science 2024-09-04 Jiahao Jiang , Yuxiang Yang , Yingqi Deng , Chenlong Ma , Jing Zhang

Significant advances have been made recently in Visual Place Recognition (VPR), feature correspondence, and localization due to the proliferation of deep-learning-based methods. However, existing approaches tend to address, partially or…

Computer Vision and Pattern Recognition · Computer Science 2020-12-22 Satyajit Tourani , Dhagash Desai , Udit Singh Parihar , Sourav Garg , Ravi Kiran Sarvadevabhatla , Michael Milford , K. Madhava Krishna

The task of Visual Place Recognition (VPR) is to predict the location of a query image from a database of geo-tagged images. Recent studies in VPR have highlighted the significant advantage of employing pre-trained foundation models like…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Issar Tzachor , Boaz Lerner , Matan Levy , Michael Green , Tal Berkovitz Shalev , Gavriel Habib , Dvir Samuel , Noam Korngut Zailer , Or Shimshi , Nir Darshan , Rami Ben-Ari

Accurate environment perception is essential for automated driving. When using monocular cameras, the distance estimation of elements in the environment poses a major challenge. Distances can be more easily estimated when the camera…

Computer Vision and Pattern Recognition · Computer Science 2020-05-11 Lennart Reiher , Bastian Lampe , Lutz Eckstein

Recent advancements in Bird's Eye View (BEV) fusion for map construction have demonstrated remarkable mapping of urban environments. However, their deep and bulky architecture incurs substantial amounts of backpropagation memory and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-23 Minsu Kim , Giseop Kim , Sunwook Choi

Transformer architectures rely on position encodings to model the spatial structure of input data. Rotary Position Encoding (RoPE) is a widely used method in language models that encodes relative positions through fixed, block-diagonal,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Sophie Ostmeier , Brian Axelrod , Maya Varma , Michael E. Moseley , Akshay Chaudhari , Curtis Langlotz