English
Related papers

Related papers: Retouchdown: Adding Touchdown to StreetLearn as a …

200 papers

Spatial Description Resolution, as a language-guided localization task, is proposed for target location in a panoramic street view, given corresponding language descriptions. Explicitly characterizing an object-level relationship while…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Peiyao Wang , Weixin Luo , Yanyu Xu , Haojie Li , Shugong Xu , Jianyu Yang , Shenghua Gao

Street-view imagery provides us with novel experiences to explore different places remotely. Carefully calibrated street-view images (e.g. Google Street View) can be used for different downstream tasks, e.g. navigation, map features…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Wenmiao Hu , Yichen Zhang , Yuxuan Liang , Yifang Yin , Andrei Georgescu , An Tran , Hannes Kruppa , See-Kiong Ng , Roger Zimmermann

While Multimodal Large Language Models have achieved human-like performance on many visual and textual reasoning tasks, their proficiency in fine-grained spatial understanding, such as route tracing on maps remains limited. Unlike humans,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Artemis Panagopoulou , Aveek Purohit , Achin Kulshrestha , Soroosh Yazdani , Mohit Goyal

In vision-and-language navigation (VLN), an embodied agent is required to navigate in realistic 3D environments following natural language instructions. One major bottleneck for existing VLN approaches is the lack of sufficient training…

Computer Vision and Pattern Recognition · Computer Science 2022-08-26 Shizhe Chen , Pierre-Louis Guhur , Makarand Tapaswi , Cordelia Schmid , Ivan Laptev

Training a deep network to perform semantic segmentation requires large amounts of labeled data. To alleviate the manual effort of annotating real images, researchers have investigated the use of synthetic data, which can be labeled…

Computer Vision and Pattern Recognition · Computer Science 2018-07-18 Fatemeh Sadat Saleh , Mohammad Sadegh Aliakbarian , Mathieu Salzmann , Lars Petersson , Jose M. Alvarez

Uniform and variable environments still remain a challenge for stable visual localization and mapping in mobile robot navigation. One of the possible approaches suitable for such environments is appearance-based teach-and-repeat navigation,…

Robotics · Computer Science 2025-03-18 Václav Truhlařík , Tomáš Pivoňka , Michal Kasarda , Libor Přeučil

In this work, we propose a modular approach for the Vision-Language Navigation (VLN) task by decomposing the problem into four sub-modules that use state-of-the-art Large Language Models (LLMs) and Vision-Language Models (VLMs) in a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Navid Rajabi , Jana Kosecka

This paper presents OmniCity, a new dataset for omnipotent city understanding from multi-level and multi-view images. More precisely, the OmniCity contains multi-view satellite images as well as street-level panorama and mono-view images,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Weijia Li , Yawen Lai , Linning Xu , Yuanbo Xiangli , Jinhua Yu , Conghui He , Gui-Song Xia , Dahua Lin

Despite vision-language models' (VLMs) remarkable capabilities as versatile visual assistants, two substantial challenges persist within the existing VLM frameworks: (1) lacking task diversity in pretraining and visual instruction tuning,…

Computation and Language · Computer Science 2024-02-20 Zhiyang Xu , Chao Feng , Rulin Shao , Trevor Ashby , Ying Shen , Di Jin , Yu Cheng , Qifan Wang , Lifu Huang

Building upon recent Deep Neural Network architectures, current approaches lying in the intersection of computer vision and natural language processing have achieved unprecedented breakthroughs in tasks like automatic captioning or image…

Computer Vision and Pattern Recognition · Computer Science 2016-03-24 Arnau Ramisa , Fei Yan , Francesc Moreno-Noguer , Krystian Mikolajczyk

The goal of cross-view image based geo-localization is to determine the location of a given street view image by matching it against a collection of geo-tagged satellite images. This task is notoriously challenging due to the drastic…

Computer Vision and Pattern Recognition · Computer Science 2021-03-12 Aysim Toker , Qunjie Zhou , Maxim Maximov , Laura Leal-Taixé

This article presents UrbanTwin datasets, high-fidelity, realistic replicas of three public roadside lidar datasets: LUMPI, V2X-Real-IC, and TUMTraf-I. Each UrbanTwin dataset contains 10K annotated frames corresponding to one of the public…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Muhammad Shahbaz , Shaurya Agarwal

Curb ramps are critical for urban accessibility, but robustly detecting them in images remains an open problem due to the lack of large-scale, high-quality datasets. While prior work has attempted to improve data availability with…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 John S. O'Meara , Jared Hwang , Zeyu Wang , Michael Saugstad , Jon E. Froehlich

We address the problem of jointly learning vision and language to understand the object in a fine-grained manner. The key idea of our approach is the use of object descriptions to provide the detailed understanding of an object. Based on…

Computer Vision and Pattern Recognition · Computer Science 2018-03-19 Anh Nguyen , Thanh-Toan Do , Ian Reid , Darwin G. Caldwell , Nikos G. Tsagarakis

Computer-assisted surgery research requires large, deeply annotated video datasets that capture clinical and technical variability. Existing cataract surgery resources lack the diversity and annotation depth required to train generalizable…

During the last half decade, convolutional neural networks (CNNs) have triumphed over semantic segmentation, which is one of the core tasks in many applications such as autonomous driving. However, to train CNNs requires a considerable…

Computer Vision and Pattern Recognition · Computer Science 2018-11-15 Yang Zhang , Philip David , Boqing Gong

Text-based person retrieval aims to identify a target individual from an image gallery using a natural language description. Existing methods primarily focus on appearance-driven cross-modal retrieval, yet face significant challenges due to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Yingjia Xu , Jinlin Wu , Daming Gao , Zhen Chen , Yang Yang , Min Cao , Mang Ye , Zhen Lei

Lane-level scene annotations provide invaluable data in autonomous vehicles for trajectory planning in complex environments such as urban areas and cities. However, obtaining such data is time-consuming and expensive since lane annotations…

Computer Vision and Pattern Recognition · Computer Science 2021-05-04 Jannik Zürn , Johan Vertens , Wolfram Burgard

Semantic scene understanding is crucial for robotics and computer vision applications. In autonomous driving, 3D semantic segmentation plays an important role for enabling safe navigation. Despite significant advances in the field, the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Lucas Nunes , Rodrigo Marcuzzi , Jens Behley , Cyrill Stachniss

We introduce the task of 3D visual grounding in large-scale dynamic scenes based on natural linguistic descriptions and online captured multi-modal visual data, including 2D images and 3D LiDAR point clouds. We present a novel method,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Zhenxiang Lin , Xidong Peng , Peishan Cong , Ge Zheng , Yujin Sun , Yuenan Hou , Xinge Zhu , Sibei Yang , Yuexin Ma