English
Related papers

Related papers: Co-visual pattern augmented generative transformer…

200 papers

The ability to transform location-centric geospatial data into meaningful computational representations has become fundamental to modern spatial analysis and decision-making. Geospatial Representation Learning (GRL), the process of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Xixuan Hao , Yutian Jiang , Xingchen Zou , Jiabo Liu , Yifang Yin , Song Gao , Flora Salim , Tianrui Li , Yuxuan Liang

Learning to generate natural scenes has always been a daunting task in computer vision. This is even more laborious when generating images with very different views. When the views are very different, the view fields have little overlap or…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Hao Ding , Songsong Wu , Hao Tang , Fei Wu , Guangwei Gao , Xiao-Yuan Jing

Similar to vision-and-language navigation (VLN) tasks that focus on bridging the gap between vision and language for embodied navigation, the new Rendezvous (RVS) task requires reasoning over allocentric spatial relationships (independent…

Computation and Language · Computer Science 2024-07-01 Tzuf Paz-Argaman , John Palowitch , Sayali Kulkarni , Reut Tsarfaty , Jason Baldridge

Feed-forward surround-view autonomous driving scene reconstruction offers fast, generalizable inference ability, which faces the core challenge of ensuring generalization while elevating novel view quality. Due to the surround-view with…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Junhong Lin , Kangli Wang , Shunzhou Wang , Songlin Fan , Ge Li , Wei Gao

Bird's-Eye View (BEV) maps provide a structured, top-down abstraction that is crucial for autonomous-driving perception. In this work, we employ Cross-View Transformers (CVT) for learning to map camera images to three BEV's channels - road,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Felipe Carlos dos Santos , Eric Aislan Antonelo , Gustavo Claudio Karl Couto

The goal of cross-view image based geo-localization is to determine the location of a given street view image by matching it against a collection of geo-tagged satellite images. This task is notoriously challenging due to the drastic…

Computer Vision and Pattern Recognition · Computer Science 2021-03-12 Aysim Toker , Qunjie Zhou , Maxim Maximov , Laura Leal-Taixé

A key goal for the advancement of AI is to develop technologies that serve the needs not just of one group but of all communities regardless of their geographical region. In fact, a significant proportion of knowledge is locally shared by…

Computer Vision and Pattern Recognition · Computer Science 2023-01-06 Da Yin , Feng Gao , Govind Thattai , Michael Johnston , Kai-Wei Chang

Drone-view Geo-Localization (DVGL) aims to achieve accurate localization of drones by retrieving the most relevant GPS-tagged satellite images. However, most existing methods heavily rely on strictly pre-paired drone-satellite images for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-27 Zhongwei Chen , Zhao-Xu Yang , Hai-Jun Rong , Guoqi Li

Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong zero-shot reasoning…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Siddhant Bharadwaj , Ashish Vashist , Fahimul Aleem , Shruti Vyas

Age estimation of face images is a crucial task with various practical applications in areas such as video surveillance and Internet access control. While deep learning-based age estimation frameworks, e.g., convolutional neural network…

Computer Vision and Pattern Recognition · Computer Science 2023-07-03 Yuntao Shou , Xiangyong Cao , Deyu Meng

Convolutional Neural Networks (CNNs) are models that are utilized extensively for the hierarchical extraction of features. Vision transformers (ViTs), through the use of a self-attention mechanism, have recently achieved superior modeling…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Ali Jamali , Swalpa Kumar Roy , Danfeng Hong , Peter M Atkinson , Pedram Ghamisi

Existing psychophysical studies have revealed that the cross-modal visual-tactile perception is common for humans performing daily activities. However, it is still challenging to build the algorithmic mapping from one modality space to…

Computer Vision and Pattern Recognition · Computer Science 2021-07-13 Shaoyu Cai , Kening Zhu , Yuki Ban , Takuji Narumi

Visual localization has traditionally been formulated as a pair-wise pose regression problem. Existing approaches mainly estimate relative poses between two images and employ a late-fusion strategy to obtain absolute pose estimates.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Tianchen Deng , Wenhua Wu , Kunzhen Wu , Guangming Wang , Siting Zhu , Shenghai Yuan , Xun Chen , Guole Shen , Zhe Liu , Hesheng Wang

Most visual scene understanding tasks in the field of computer vision involve identification of the objects present in the scene. Image regions like hideouts, turns, & other obscured regions of the scene also contain crucial information,…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Binoy Saha , Sukhendu Das

Incorporating prior structure information into the visual state estimation could generally improve the localization performance. In this letter, we aim to address the paradox between accuracy and efficiency in coupling visual factors with…

Robotics · Computer Science 2020-10-06 Huaiyang Huang , Haoyang Ye , Yuxiang Sun , Ming Liu

This paper develops a deep-learning framework to synthesize a ground-level view of a location given an overhead image. We propose a novel conditional generative adversarial network (cGAN) in which the trained generator generates realistic…

Computer Vision and Pattern Recognition · Computer Science 2019-02-21 Xueqing Deng , Yi Zhu , Shawn Newsam

Planet-scale photo geolocalization involves the intricate task of estimating the geographic location depicted in an image purely based on its visual features. While deep learning models, particularly convolutional neural networks (CNNs),…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 David Faget , José Luis Lisani , Miguel Colom

Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yedong Shen , Shiqi Zhang , Sha Zhang , Yifan Duan , Xinran Zhang , Wenhao Yu , Lu Zhang , Jiajun Deng , Yanyong Zhang

We propose an image-based cross-view geolocalization method that estimates the global pose of a UAV with the aid of georeferenced satellite imagery. Our method consists of two Siamese neural networks that extract relevant features despite…

Robotics · Computer Science 2018-09-18 Akshay Shetty , Grace Xingxin Gao

In the field of autonomous vehicles (AVs), accurately discerning commander intent and executing linguistic commands within a visual context presents a significant challenge. This paper introduces a sophisticated encoder-decoder framework,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-19 Haicheng Liao , Huanming Shen , Zhenning Li , Chengyue Wang , Guofa Li , Yiming Bie , Chengzhong Xu