English
Related papers

Related papers: Systematically Evaluating Equivalent Purpose for D…

200 papers

Foreground map evaluation is crucial for gauging the progress of object segmentation algorithms, in particular in the filed of salient object detection where the purpose is to accurately detect and segment the most salient object in a…

Computer Vision and Pattern Recognition · Computer Science 2017-08-03 Deng-Ping Fan , Ming-Ming Cheng , Yun Liu , Tao Li , Ali Borji

Locational measures of accessibility are widely used in urban and transportation planning to understand the impact of the transportation system on influencing people's access to places. However, there is a considerable lack of measurement…

Econometrics · Economics 2024-04-09 Rajat Verma , Mithun Debnath , Shagun Mittal , Satish V. Ukkusuri

Large Vision-Language Models (LVLMs) demonstrate a promising direction for assisting individuals with blindness or low-vision (BLV). Yet, measuring their true utility in real-world scenarios is challenging because evaluating whether their…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Eunki Kim , Na Min An , Wan Ju Kang , Sangryul Kim , James Thorne , Hyunjung Shim

Multi-scale tile maps are essential for geographic information services, serving as fundamental outcomes of surveying and cartographic workflows. While existing image generation networks can produce map-like outputs from remote sensing…

Image and Video Processing · Electrical Eng. & Systems 2025-05-29 Chenxing Sun , Yongyang Xu , Xuwei Xu , Xixi Fan , Jing Bai , Xiechun Lu , Zhanlong Chen

In this paper, we introduce MapBench-the first dataset specifically designed for human-readable, pixel-based map-based outdoor navigation, curated from complex path finding scenarios. MapBench comprises over 1600 pixel space map path…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Shuo Xing , Zezhou Sun , Shuangyu Xie , Kaiyuan Chen , Yanjia Huang , Yuping Wang , Jiachen Li , Dezhen Song , Zhengzhong Tu

Multi-goal reaching is an important problem in reinforcement learning needed to achieve algorithmic generalization. Despite recent advances in this field, current algorithms suffer from three major challenges: high sample complexity,…

Map construction methods automatically produce and/or update road network datasets using vehicle tracking data. Enabled by the ubiquitous generation of georeferenced tracking data, there has been a recent surge in map construction…

Computational Geometry · Computer Science 2014-11-19 Mahmuda Ahmed , Sophia Karagiorgou , Dieter Pfoser , Carola Wenk

This study introduces a novel approach to online embedding of multi-scale CLIP (Contrastive Language-Image Pre-Training) features into 3D maps. By harnessing CLIP, this methodology surpasses the constraints of conventional…

Robotics · Computer Science 2024-03-28 Shun Taguchi , Hideki Deguchi

Grounding language to a navigating agent's observations can leverage pretrained multimodal foundation models to match perceptions to object or event descriptions. However, previous approaches remain disconnected from environment mapping,…

Robotics · Computer Science 2025-06-10 Chenguang Huang , Oier Mees , Andy Zeng , Wolfram Burgard

Semantic maps allow a robot to reason about its surroundings to fulfill tasks such as navigating known environments, finding specific objects, and exploring unmapped areas. Traditional mapping approaches provide accurate geometric…

Robotics · Computer Science 2026-02-03 Felix Igelbrink , Lennart Niecksch , Marian Renz , Martin Günther , Martin Atzmueller

Recent progress in large vision-language models has driven improvements in language-based semantic navigation, where an embodied agent must reach a target object described in natural language. Yet we still lack a clear, language-focused…

The Global South faces unique challenges in achieving digital inclusion due to a heavy reliance on mobile devices for internet access and the prevalence of slow or unreliable networks. While numerous studies have investigated web…

Software Engineering · Computer Science 2025-01-29 Masudul Hasan Masud Bhuiyan , Matteo Varvello , Cristian-Alexandru Staicu , Yasir Zaki

Recent advancements in foundation models have improved autonomous tool usage and reasoning, but their capabilities in map-based reasoning remain underexplored. To address this, we introduce MapEval, a benchmark designed to assess foundation…

This paper presents a cognition-inspired agnostic framework for building a map for Visual Place Recognition. This framework draws inspiration from human-memorability, utilizes the traditional image entropy concept and computes the static…

Computer Vision and Pattern Recognition · Computer Science 2019-03-22 Mubariz Zaffar , Shoaib Ehsan , Michael Milford , Klaus Mcdonald Maier

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Sara Ghazanfari , Siddharth Garg , Nicolas Flammarion , Prashanth Krishnamurthy , Farshad Khorrami , Francesco Croce

Although large language models (LLMs) have advanced rapidly, robust automation of complex software workflows remains an open problem. In long-horizon settings, agents frequently suffer from cascading errors and environmental stochasticity;…

Artificial Intelligence · Computer Science 2026-03-30 Yenchia Feng , Chirag Sharma , Karime Maamari

Providing an equitable and inclusive user experience (UX) for people with disabilities (PWD) is a central goal of accessible design. In the specific case of Deaf users, whose hearing impairments impact language development and…

Human-Computer Interaction · Computer Science 2025-08-12 Andrés Eduardo Fuentes-Cortázar , Alejandra Rivera-Hernández , José Rafael Rojano-Cáceres

Multimodal large language models (MLLMs), equipped with increasingly advanced planning and tool-use capabilities, are evolving into autonomous agents capable of performing multimodal web browsing and deep search in open-world environments.…

Geolocation, the task of identifying an image's location, requires complex reasoning and is crucial for navigation, monitoring, and cultural preservation. However, current methods often produce coarse, imprecise, and non-interpretable…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Zirui Song , Jingpu Yang , Yuan Huang , Jonathan Tonglet , Zeyu Zhang , Tao Cheng , Meng Fang , Iryna Gurevych , Xiuying Chen

Interactive streetscape mapping tools such as Google Street View (GSV) and Meta Mapillary enable users to virtually navigate and experience real-world environments via immersive 360{\deg} imagery but remain fundamentally inaccessible to…

Human-Computer Interaction · Computer Science 2025-09-29 Jon E. Froehlich , Alexander Fiannaca , Nimer Jaber , Victor Tsaran , Shaun Kane