中文
相关论文

相关论文: AerialVG: A Challenging Benchmark for Aerial Visua…

200 篇论文

3D visual grounding (3DVG) identifies objects in 3D scenes from language descriptions. Existing zero-shot approaches leverage 2D vision-language models (VLMs) by converting 3D spatial information (SI) into forms amenable to VLM processing,…

计算机视觉与模式识别 · 计算机科学 2025-12-11 Yuanyuan Liu , Haiyang Mei , Dongyang Zhan , Jiayue Zhao , Dongsheng Zhou , Bo Dong , Xin Yang

The ability to understand and reason about spatial relationships between objects in images is an important component of visual reasoning. This skill rests on the ability to recognize and localize objects of interest and determine their…

计算与语言 · 计算机科学 2024-10-14 Navid Rajabi , Jana Kosecka

Visual grounding in text-rich document images is a critical yet underexplored challenge for Document Intelligence and Visual Question Answering (VQA) systems. We present DRISHTIKON, a multi-granular and multi-block visual grounding…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Badri Vishal Kasuba , Parag Chaudhuri , Ganesh Ramakrishnan

Medical Visual Grounding (MVG) aims to identify diagnostically relevant phrases from free-text radiology reports and localize their corresponding regions in medical images, providing interpretable visual evidence to support clinical…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yifan Gao , Tao Zhou , Yi Zhou , Ke Zou , Yizhe Zhang , Huazhu Fu

Recent advancements in multimodal large language models have driven breakthroughs in visual question answering. Yet, a critical gap persists, `conceptualization'-the ability to recognize and reason about the same concept despite variations…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Zahra Babaiee , Peyman M. Kiasari , Daniela Rus , Radu Grosu

Learning to ground natural language queries to target objects or regions in 3D point clouds is quite essential for 3D scene understanding. Nevertheless, existing 3D visual grounding approaches require a substantial number of bounding box…

计算机视觉与模式识别 · 计算机科学 2025-08-27 Xiaoxu Xu , Yitian Yuan , Qiudan Zhang , Wenhui Wu , Zequn Jie , Lin Ma , Xu Wang

Aerial Vision-and-Language Navigation (VLN) aims to enable unmanned aerial vehicles (UAVs) to interpret natural language instructions and navigate complex urban environments using onboard visual observation. This task holds promise for…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Huilin Xu , Zhuoyang Liu , Yixiang Luomei , Feng Xu

Visual relationship detection, as a challenging task used to find and distinguish the interactions between object pairs in one image, has received much attention recently. In this work, we propose a novel visual relationship detection…

计算机视觉与模式识别 · 计算机科学 2019-11-05 Hao Zhou , Chongyang Zhang , Chuanping Hu

Aerial image categorization plays an indispensable role in remote sensing and artificial intelligence. In this paper, we propose a new aerial image categorization framework, focusing on organizing the local patches of each aerial image into…

计算机视觉与模式识别 · 计算机科学 2016-11-04 Yuxin Hu , Luming Zhang

Understanding relationships between objects is central to visual intelligence, with applications in embodied AI, assistive systems, and scene understanding. Yet, most visual relationship detection (VRD) models rely on a fixed predicate set,…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Shanmukha Vellamcheti , Sanjoy Kundu , Sathyanarayanan N. Aakur

Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual question answering.…

计算与语言 · 计算机科学 2025-02-10 Akshar Tumu , Parisa Kordjamshidi

Visual Geo-localization (VG) is the task of estimating the position where a given photo was taken by comparing it with a large database of images of known locations. To investigate how existing techniques would perform on a real-world…

计算机视觉与模式识别 · 计算机科学 2022-04-08 Gabriele Berton , Carlo Masone , Barbara Caputo

We propose a novel fine-grained cross-view localization method that estimates the 3 Degrees of Freedom pose of a ground-level image in an aerial image of the surroundings by matching fine-grained features between the two images. The pose is…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Zimin Xia , Alexandre Alahi

The existing works on object-level language grounding with 3D objects mostly focus on improving performance by utilizing the off-the-shelf pre-trained models to capture features, such as viewpoint selection or geometric priors. However,…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Penglei Sun , Yaoxian Song , Xinglin Pan , Peijie Dong , Xiaofei Yang , Qiang Wang , Zhixu Li , Tiefeng Li , Xiaowen Chu

Place recognition and visual localization are particularly challenging in wide baseline configurations. In this paper, we contribute with the \emph{Danish Airs and Grounds} (DAG) dataset, a large collection of street-level and aerial images…

计算机视觉与模式识别 · 计算机科学 2022-02-07 Andrea Vallone , Frederik Warburg , Hans Hansen , Søren Hauberg , Javier Civera

Cross-view localization aims to estimate the 3-DoF pose of a ground-view image by aligning it with aerial or satellite imagery. Existing methods typically address this task through direct regression or feature alignment in a shared…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Panwang Xia , Qiong Wu , Lei Yu , Yi Liu , Mingtao Xiong , Xudong Lu , Yi Liu , Haoyu Guo , Yongxiang Yao , Junjian Zhang , Xiangyuan Cai , Hongwei Hu , Zhi Zheng , Yongjun Zhang , Yi Wan

Tracking by natural language specification aims to locate the referred target in a sequence based on the natural language description. Existing algorithms solve this issue in two steps, visual grounding and tracking, and accordingly deploy…

计算机视觉与模式识别 · 计算机科学 2023-03-22 Li Zhou , Zikun Zhou , Kaige Mao , Zhenyu He

Visual grounding aims to align visual information of specific regions of images with corresponding natural language expressions. Current visual grounding methods leverage pre-trained visual and language backbones independently to obtain…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Jiaxi Wang , Wenhui Hu , Xueyang Liu , Beihu Wu , Yuting Qiu , YingYing Cai

Recently, researchers have attempted to investigate the capability of LLMs in handling videos and proposed several video LLM models. However, the ability of LLMs to handle video grounding (VG), which is an important time-related video task…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Wei Feng , Xin Wang , Hong Chen , Zeyang Zhang , Houlun Chen , Zihan Song , Yuwei Zhou , Yuekui Yang , Haiyang Wu , Wenwu Zhu

This paper presents Vision-Language Global Localization (VLG-Loc), a novel global localization method that uses human-readable labeled footprint maps containing only names and areas of distinctive visual landmarks in an environment. While…

机器人学 · 计算机科学 2025-12-19 Mizuho Aoki , Kohei Honda , Yasuhiro Yoshimura , Takeshi Ishita , Ryo Yonetani