中文
相关论文

相关论文: Iterative Robust Visual Grounding with Masked Refe…

200 篇论文

We propose a method to improve Visual Question Answering (VQA) with Retrieval-Augmented Generation (RAG) by introducing text-grounded object localization. Rather than retrieving information based on the entire image, our approach enables…

人工智能 · 计算机科学 2025-10-01 Xinxi Chen , Tianyang Chen , Lijia Hong

Visual grounding, a crucial vision-language task involving the understanding of the visual context based on the query expression, necessitates the model to capture the interactions between objects, as well as various spatial and attribute…

计算机视觉与模式识别 · 计算机科学 2023-12-27 Haozhan Shen , Tiancheng Zhao , Mingwei Zhu , Jianwei Yin

Existing Multimodal Large Language Models (MLLMs) for image forgery detection and localization predominantly operate under a text-centric Chain-of-Thought (CoT) paradigm. However, forcing these models to textually characterize imperceptible…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Youqi Wang , Shen Chen , Haowei Wang , Rongxuan Peng , Taiping Yao , Shunquan Tan , Changsheng Chen , Bin Li , Shouhong Ding

Large vision-and-language models (VLMs) trained to match images with text on large-scale datasets of image-text pairs have shown impressive generalization ability on several vision and language tasks. Several recent works, however, showed…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Navid Rajabi , Jana Kosecka

3D visual grounding aims to identify the target object within a 3D point cloud scene referred to by a natural language description. Previous works usually require significant data relating to point color and their descriptions to exploit…

计算机视觉与模式识别 · 计算机科学 2025-02-25 Tung-Yu Wu , Sheng-Yu Huang , Yu-Chiang Frank Wang

As the complexity of 3D digital content grows exponentially, understanding human visual attention is critical for optimizing rendering and processing resources. Therefore, reliable 3D mesh saliency ground truth (GT) is essential for…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Guoquan Zheng , Jie Hao , Huiyu Duan , Long Tang , Shuo Yang , Yucheng Zhu , Yongming Han , Liang Yuan , Patrick Le Callet , Guangtao Zhai

The advancement of Large Vision-Language Models (LVLMs) requires precise local region-based reasoning that faithfully grounds the model's logic in actual visual evidence. However, existing datasets face limitations in scalability due to…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Byeonggeuk Lim , Kyeonghyun Kim , JungMin Yun , YoungBin Kim

Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Feng Lu , Xiangyuan Lan , Lijun Zhang , Dongmei Jiang , Yaowei Wang , Chun Yuan

The traditional visual-inertial SLAM system often struggles with stability under low-light or motion-blur conditions, leading to potential lost of trajectory tracking. High accuracy and robustness are essential for the long-term and stable…

机器人学 · 计算机科学 2024-11-05 Hongkun Luo , Yang Liu , Chi Guo , Zengke Li , Weiwei Song

Multi-view clustering (MVC) has emerged as a powerful technique for extracting valuable insights from data characterized by multiple perspectives or modalities. Despite significant advancements, existing MVC methods struggle with…

人工智能 · 计算机科学 2024-12-24 Lijian Li

3D Visual Grounding (3DVG) involves localizing target objects in 3D point clouds based on natural language. While prior work has made strides using textual descriptions, leveraging spoken language-known as Audio-based 3D Visual…

In this paper, we explore a novel task named visual Relation Grounding in Videos (vRGV). The task aims at spatio-temporally localizing the given relations in the form of subject-predicate-object in the videos, so as to provide supportive…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Junbin Xiao , Xindi Shang , Xun Yang , Sheng Tang , Tat-Seng Chua

Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for existing large vision-language models (LVLMs) often limited by…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Jingwei Peng , Jiehao Chen , Mateo Alejandro Rojas , Meilin Zhang

Video temporal grounding (VTG) is typically tackled with dataset-specific models that transfer poorly across domains and query styles. Recent efforts to overcome this limitation have adapted large multimodal language models (MLLMs) to VTG,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Joungbin An , Agrim Jain , Kristen Grauman

Recent advances in large language models have significantly improved textual reasoning through the effective use of Chain-of-Thought (CoT) and reinforcement learning. However, extending these successes to vision-language tasks remains…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Minheng Ni , Zhengyuan Yang , Linjie Li , Chung-Ching Lin , Kevin Lin , Wangmeng Zuo , Lijuan Wang

In this paper, we propose a new open-source benchmarking framework for Visual Geo-localization (VG) that allows to build, train, and test a wide range of commonly used architectures, with the flexibility to change individual components of a…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Gabriele Berton , Riccardo Mereu , Gabriele Trivigno , Carlo Masone , Gabriela Csurka , Torsten Sattler , Barbara Caputo

Cross-view geo-localisation identifies coarse geographical position of an automated vehicle by matching a ground-level image to a geo-tagged satellite image from a database. Despite the advancements in Cross-view geo-localisation,…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Barkin Dagda , Muhammad Awais , Saber Fallah

A hierarchical cross-modal fusion model is proposed for vision-language question answering (VLQA) in industrial robotics, targeting the challenges of semantic ambiguity, complex environmental layouts, and domain-specific terminology common…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Ping Li , Bartlomiej Brzozka

Vision-based interception using multicopters equipped strapdown camera is challenging due to camera-motion coupling and evasive targets. This paper proposes a method integrating Image-Based Visual Servoing (IBVS) with proportional…

机器人学 · 计算机科学 2025-04-07 Hailong Yan , Kun Yang , Yixiao Cheng , Zihao Wang , Dawei Li

Cross-view geo-localization (CVGL) aims to accurately localize street-view images through retrieval of corresponding geo-tagged satellite images. While prior works have achieved nearly perfect performance on certain standard datasets, their…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Le Wu , Lv Bo , Songsong Ouyang , Yingying Zhu