中文
相关论文

相关论文: Retouchdown: Adding Touchdown to StreetLearn as a …

200 篇论文

Cross-view image retrieval, particularly street-to-satellite matching, is a critical task for applications such as autonomous navigation, urban planning, and localization in GPS-denied environments. However, existing approaches often…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Jeongho Min , Dongyoung Kim , Jaehyup Lee

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes.…

机器人学 · 计算机科学 2026-01-16 Yanghong Mei , Yirong Yang , Longteng Guo , Qunbo Wang , Ming-Ming Yu , Xingjian He , Wenjun Wu , Jing Liu

Promising performance has been achieved for visual perception on the point cloud. However, the current methods typically rely on labour-extensive annotations on the scene scans. In this paper, we explore how synthetic models alleviate the…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Runnan Chen , Xinge Zhu , Nenglun Chen , Dawei Wang , Wei Li , Yuexin Ma , Ruigang Yang , Wenping Wang

City-scale 3D point cloud is a promising way to express detailed and complicated outdoor structures. It encompasses both the appearance and geometry features of segmented city components, including cars, streets, and buildings, that can be…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Taiki Miyanishi , Fumiya Kitamori , Shuhei Kurita , Jungdae Lee , Motoaki Kawanabe , Nakamasa Inoue

Images represent a commonly used form of visual communication among people. Nevertheless, image classification may be a challenging task when dealing with unclear or non-common images needing more context to be correctly annotated. Metadata…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Tobia Tesan , Pasquale Coscia , Lamberto Ballan

High-level 3D scene understanding is essential in many applications. However, the challenges of generating accurate 3D annotations make development of deep learning models difficult. We turn to recent advancements in automatic retrieval of…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Yuchen Rao , Stefan Ainetter , Sinisa Stekovic , Vincent Lepetit , Friedrich Fraundorfer

Vision-and-Language Navigation (VLN) is a natural language grounding task where an agent learns to follow language instructions and navigate to specified destinations in real-world environments. A key challenge is to recognize and stop at…

计算机视觉与模式识别 · 计算机科学 2020-10-20 Jiannan Xiang , Xin Eric Wang , William Yang Wang

Visual Place Recognition (VPR) in indoor environments is beneficial to humans and robots for better localization and navigation. It is challenging due to appearance changes at various frequencies, and difficulties of obtaining ground truth…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Diwei Sheng , Anbang Yang , John-Ross Rizzo , Chen Feng

Location recognition is commonly treated as visual instance retrieval on "street view" imagery. The dataset items and queries are panoramic views, i.e. groups of images taken at a single location. This work introduces a novel…

计算机视觉与模式识别 · 计算机科学 2017-04-24 Ahmet Iscen , Giorgos Tolias , Yannis Avrithis , Teddy Furon , Ondrej Chum

Deep learning-based approaches to delineating 3D structure depend on accurate annotations to train the networks. Yet, in practice, people, no matter how conscientious, have trouble precisely delineating in 3D and on a large scale, in part…

计算机视觉与模式识别 · 计算机科学 2022-12-27 Doruk Oner , Leonardo Citraro , Mateusz Koziński , Pascal Fua

Geospatial predictions are crucial for diverse fields such as disaster management, urban planning, and public health. Traditional machine learning methods often face limitations when handling unstructured or multi-modal data like street…

计算与语言 · 计算机科学 2024-11-25 Zongrong Li , Junhao Xu , Siqin Wang , Yifan Wu , Haiyang Li

Natural Language (NL) descriptions can be one of the most convenient or the only way to interact with systems built to understand and detect city scale traffic patterns and vehicle-related events. In this paper, we extend the widely adopted…

计算机视觉与模式识别 · 计算机科学 2021-04-07 Qi Feng , Vitaly Ablavsky , Stan Sclaroff

We introduce a unified framework to jointly model images, text, and human attention traces. Our work is built on top of the recent Localized Narratives annotation framework [30], where each word of a given caption is paired with a mouse…

计算机视觉与模式识别 · 计算机科学 2021-05-14 Zihang Meng , Licheng Yu , Ning Zhang , Tamara Berg , Babak Damavandi , Vikas Singh , Amy Bearman

Despite the recent success of deep-learning based semantic segmentation, deploying a pre-trained road scene segmenter to a city whose images are not presented in the training set would not achieve satisfactory performance due to dataset…

计算机视觉与模式识别 · 计算机科学 2017-04-28 Yi-Hsin Chen , Wei-Yu Chen , Yu-Ting Chen , Bo-Cheng Tsai , Yu-Chiang Frank Wang , Min Sun

In a busy city street, a pedestrian surrounded by distractions can pick out a single sign if it is relevant to their route. Artificial agents in outdoor Vision-and-Language Navigation (VLN) are also confronted with detecting supervisory…

机器学习 · 计算机科学 2022-11-21 Jason Armitage , Leonardo Impett , Rico Sennrich

We introduce Synscapes -- a synthetic dataset for street scene parsing created using photorealistic rendering techniques, and show state-of-the-art results for training and validation as well as new types of analysis. We study the behavior…

计算机视觉与模式识别 · 计算机科学 2018-10-23 Magnus Wrenninge , Jonas Unger

We introduce DialNav, a novel collaborative embodied dialog task, where a navigation agent (Navigator) and a remote guide (Guide) engage in multi-turn dialog to reach a goal location. Unlike prior work, DialNav aims for holistic evaluation…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Leekyeung Han , Hyunji Min , Gyeom Hwangbo , Jonghyun Choi , Paul Hongsuck Seo

Urban waste management remains a critical challenge for the development of smart cities. Despite the growing number of litter detection datasets, the problem of monitoring overflowing waste containers, particularly from images captured by…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Diogo J. Paulo , João Martins , Hugo Proença , João C. Neves

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene under-…

计算与语言 · 计算机科学 2017-04-25 Spandana Gella , Frank Keller

Vision-and-Language Navigation (VLN) requires grounding instructions, such as "turn right and stop at the door", to routes in a visual environment. The actual grounding can connect language to the environment through multiple modalities,…

计算与语言 · 计算机科学 2019-06-11 Ronghang Hu , Daniel Fried , Anna Rohrbach , Dan Klein , Trevor Darrell , Kate Saenko