中文
相关论文

相关论文: AMIGO: Agentic Multi-Image Grounding Oracle Benchm…

200 篇论文

Due to the advantages of fusing information from various modalities, multimodal learning is gaining increasing attention. Being a fundamental task of multimodal learning, Visual Grounding (VG), aims to locate objects in images through…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Zhiyuan Chang , Mingyang Li , Junjie Wang , Cheng Li , Boyu Wu , Fanjiang Xu , Qing Wang

Multi-Modal Knowledge Graphs (MMKGs) have proven valuable for various downstream tasks. However, scaling them up is challenging because building large-scale MMKGs often introduces mismatched images (i.e., noise). Most entities in KGs belong…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Yikai Zhang , Qianyu He , Xintao Wang , Siyu Yuan , Jiaqing Liang , Yanghua Xiao

Vision-Language Models (VLMs) are increasingly susceptible to sophisticated adversarial attacks, including adaptive strategies specifically designed to bypass existing defenses. To address this vulnerability, we propose MirrorCheck, a…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Samar Fares , Klea Ziu , Toluwani Aremu , Nikita Durasov , Martin Takáč , Pascal Fua , Ivan Laptev , Karthik Nandakumar

Referring segmentation aims to segment the target objects in images or videos based on the textual query. Despite remarkable progress over the past years, existing works always assume that the user-provided queries are already precise and…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Yuting Yang , Haichao Jiang , Tianming Liang , Quan Zhang , Jian-Fang Hu

Agentic methods have emerged as a powerful and autonomous paradigm that enhances reasoning, collaboration, and adaptive control, enabling systems to coordinate and independently solve complex tasks. We extend this paradigm to safety…

人工智能 · 计算机科学 2025-10-30 Juan Ren , Mark Dras , Usman Naseem

Multimodal Large Language Models (MLLMs) have made notable advances in visual understanding, yet their abilities to recognize objects modified by specific attributes remain an open question. To address this, we explore MLLMs' reasoning…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Jiaxuan Li , Junwen Mo , MinhDuc Vo , Akihiro Sugimoto , Hideki Nakayama

In multimedia application scenarios, images captured under low-illumination conditions often lead to lower accuracy in visual perception tasks compared to those taken in well-lit environments. To tackle this challenge, we propose AMIEOD, an…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Xiaochen Huang , Honggang Chen , Weicheng Zhang , Xiaobo Dai , Yongyi Li , Linbo Qing , Xiaohai He

The human brain can effortlessly recognize and localize objects, whereas current 3D object detection methods based on LiDAR point clouds still report inferior performance for detecting occluded and distant objects: the point cloud…

计算机视觉与模式识别 · 计算机科学 2022-08-25 Liang Du , Xiaoqing Ye , Xiao Tan , Edward Johns , Bo Chen , Errui Ding , Xiangyang Xue , Jianfeng Feng

Visual grounding (VG) typically focuses on locating regions of interest within an image using natural language, and most existing VG methods are limited to single-image interpretations. This limits their applicability in real-world…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Wenxuan Wang , Zijia Zhao , Yisi Zhang , Yepeng Tang , Erdong Hu , Xinlong Wang , Jing Liu

Enterprise agents increasingly operate inside scoped retrieval systems, delegated workflows, and policy-constrained evidence environments. In these settings, access control can be enforced correctly while the system still produces an answer…

人工智能 · 计算机科学 2026-05-08 Krti Tallam

Existing multi-agent perception systems assume that every agent utilizes the same model with identical parameters and architecture. The performance can be degraded with different perception models due to the mismatch in their confidence…

机器人学 · 计算机科学 2023-03-14 Runsheng Xu , Weizhe Chen , Hao Xiang , Lantao Liu , Jiaqi Ma

Image geo-localization aims to determine where a photograph was taken, a task that often requires more than recognizing visible landmarks. Human experts typically solve it through an iterative workflow: they inspect informative regions,…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Yong Li , Furong Jia , Dacheng Yin , Kang Rong , Fengyun Rao , Jing Lyu , Fan Zhang

Real-world image restoration (IR) is inherently complex and often requires combining multiple specialized models to address diverse degradations. Inspired by human problem-solving, we propose AgenticIR, an agentic system that mimics the…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Kaiwen Zhu , Jinjin Gu , Zhiyuan You , Yu Qiao , Chao Dong

Multi-view clustering has wide applications in many image processing scenarios. In these scenarios, original image data often contain missing instances and noises, which is ignored by most multi-view clustering methods. However, missing…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Xiang Fang , Yuchong Hu , Pan Zhou , Dapeng Oliver Wu

A serious issue that harms the performance of zero-shot visual recognition is named objective misalignment, i.e., the learning objective prioritizes improving the recognition accuracy of seen classes rather than unseen classes, while the…

计算机视觉与模式识别 · 计算机科学 2024-10-14 Jiannan Ge , Lingxi Xie , Hongtao Xie , Pandeng Li , Xiaopeng Zhang , Yongdong Zhang , Qi Tian

While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Tri Cao , Khoi Le , Thong Nguyen , Cong-Duy Nguyen , Quynh Vo , Anh Tuan Luu , Chunyan Miao , See-Kiong Ng , Shuicheng Yan , Bryan Hooi

Appreciating multi-figure paintings requires understanding how characters relate through subtle cues like gaze alignment, gesture, and spatial arrangement. We present MIRAGE, an evidence-centric framework designed to scaffold the…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jui-Cheng Chiu , Yu-Chao Wang , Shengyang Luo , Tongyan Wang , Qi Yang , Nabin Khanal , Yingjie Victor Chen

The detection and grounding of manipulated content in multimodal data has emerged as a critical challenge in media forensics. While existing benchmarks demonstrate technical progress, they suffer from misalignment artifacts that poorly…

计算机视觉与模式识别 · 计算机科学 2025-09-17 Jinjie Shen , Yaxiong Wang , Lechao Cheng , Nan Pu , Zhun Zhong

Adversarial Imitation Learning (AIL) allows the agent to reproduce expert behavior with low-dimensional states and actions. However, challenges arise in handling visual states due to their less distinguishable representation compared to…

机器学习 · 计算机科学 2024-01-23 Yunke Wang , Linwei Tao , Bo Du , Yutian Lin , Chang Xu

We propose Unified-IO, a model that performs a large variety of AI tasks spanning classical computer vision tasks, including pose estimation, object detection, depth estimation and image generation, vision-and-language tasks such as region…

计算机视觉与模式识别 · 计算机科学 2022-10-06 Jiasen Lu , Christopher Clark , Rowan Zellers , Roozbeh Mottaghi , Aniruddha Kembhavi