中文
相关论文

相关论文: egenioussBench: A New Dataset for Geospatial Visua…

200 篇论文

Geoscience intelligence is expected to understand, reason about, and predict earth system changes to support human decision-making in critical domains such as disaster response, climate adaptation and environmental protection. Although…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yushuo Zheng , Zicheng Zhang , Huiyu Duan , Chunyi Li , Zijian Chen , Ziheng Jia , Yue Shi , Ke Gu , Xiongkuo Min , Guangtao Zhai

Estimating the location where an image was taken based solely on the contents of the image is a challenging task, even for humans, as properly labeling an image in such a fashion relies heavily on contextual information, and is not as…

计算机视觉与模式识别 · 计算机科学 2017-12-29 Jesse M. Johns , Jeremiah Rounds , Michael J. Henry

Egocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yifei Cao , Yu Liu , Guolong Wang , Zhu Liu , Kai Wang , Xianjie Zhang , Jizhe Yu , Xun Tu

Unmanned Aerial Vehicles (UAVs) rely on satellite systems for stable positioning. However, due to limited satellite coverage or communication disruptions, UAVs may lose signals from satellite-based positioning systems. In such situations,…

计算机视觉与模式识别 · 计算机科学 2023-08-14 Ming Dai , Enhui Zheng , Zhenhua Feng , Jiedong Zhuang , Wankou Yang

We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of \textbf{756 images} of pairwise non-intersecting Jordan curves across easy, polygonal, topographic-inspired, maze-like,…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Amirreza Mohseni , Mona Mohammadi , Morteza Saghafian , Naser Talebizadeh Sardari

In this paper, we address the problem of global-scale image geolocation, proposing a mixed classification-retrieval scheme. Unlike other methods that strictly tackle the problem as a classification or retrieval task, we combine the two…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Giorgos Kordopatis-Zilos , Panagiotis Galopoulos , Symeon Papadopoulos , Ioannis Kompatsiaris

Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and an overreliance on…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Fengxian Ji , Jingpu Yang , Zirui Song , Yuanxi Wang , Zhexuan Cui , Yuke Li , Qian Jiang , Xiuying Chen

Spatial intelligence is essential for multimodal large language models, yet current benchmarks largely assess it only from an understanding perspective. We ask whether modern generative or unified multimodal models also possess generative…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Muzhi Zhu , Shunyao Jiang , Huanyi Zheng , Zekai Luo , Hao Zhong , Anzhou Li , Kaijun Wang , Jintao Rong , Yang Liu , Hao Chen , Tao Lin , Chunhua Shen

In this paper, we present a generalizable model-free 6-DoF object pose estimator called Gen6D. Existing generalizable pose estimators either need high-quality object models or require additional depth maps or object masks in test time,…

计算机视觉与模式识别 · 计算机科学 2023-01-30 Yuan Liu , Yilin Wen , Sida Peng , Cheng Lin , Xiaoxiao Long , Taku Komura , Wenping Wang

Multimodal retrieval is becoming a crucial component of modern AI applications, yet its evaluation lags behind the demands of more realistic and challenging scenarios. Existing benchmarks primarily probe surface-level semantic…

Large-scale foundation models in Earth Observation can learn versatile, label-efficient representations by leveraging massive amounts of unlabeled data. However, existing public datasets are often limited in scale, geographic coverage, or…

Humans can imagine and manipulate visual images mentally, a capability known as spatial visualization. While many multi-modal benchmarks assess reasoning on visible visual information, the ability to infer unseen relationships through…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Siting Wang , Minnan Pei , Luoyang Sun , Cheng Deng , Yuchen Li , Kun Shao , Zheng Tian , Haifeng Zhang , Jun Wang

Geospatial raster data, such as that collected by satellite-based imaging systems at different times and spectral bands, hold immense potential for enabling a wide range of high-impact applications. This potential stems from the rich…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Haozhe Si , Yuxuan Wan , Minh Do , Deepak Vasisht , Han Zhao , Hendrik F. Hamann

Recent advances in MLLMs are reframing segmentation from fixed-category prediction to instruction-grounded localization. While reasoning based segmentation has progressed rapidly in natural scenes, remote sensing lacks a generalizable…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Lifan Jiang , Yuhang Pei , oxi Wu , Yan Zhao , Tianrun Wu , Shulong Yu , Lihui Zhang , Deng Cai

Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplete evidence use and imperfect evidence integration can…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Zihang Lin , Huaiyuan Qin , Muli Yang , Hongyuan Zhu

Recent progress in zero-shot 6D object pose estimation has been driven largely by large-scale models and cloud-based inference. However, these approaches often introduce high latency, elevated energy consumption, and deployment risks…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Javier Villena Toro , Mehdi Tarkian

The ability to locate an object in an image according to natural language instructions is crucial for many real-world applications. In this work we propose LocateBench, a high-quality benchmark dedicated to evaluating this ability. We…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Ting-Rui Chiang , Joshua Robinson , Xinyan Velocity Yu , Dani Yogatama

Generalizable dense feature matching in endoscopic images is crucial for robot-assisted tasks, including 3D reconstruction, navigation, and surgical scene understanding. Yet, it remains a challenge due to difficult visual conditions (e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Bingyu Yang , Qingyao Tian , Yimeng Geng , Huai Liao , Xinyan Huang , Jiebo Luo , Hongbin Liu

We present EgoHumans, a new multi-view multi-human video benchmark to advance the state-of-the-art of egocentric human 3D pose estimation and tracking. Existing egocentric benchmarks either capture single subject or indoor-only scenarios,…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Rawal Khirodkar , Aayush Bansal , Lingni Ma , Richard Newcombe , Minh Vo , Kris Kitani

Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in this field is the establishment of a comprehensive evaluation…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Yiwei Ma , Jiayi Ji , Ke Ye , Weihuang Lin , Zhibin Wang , Yonghan Zheng , Qiang Zhou , Xiaoshuai Sun , Rongrong Ji