中文
相关论文

相关论文: Beyond Referring Expressions: Scenario Comprehensi…

200 篇论文

Semantic segmentation is an essential step for many vision applications in order to understand a scene and the objects within. Recent progress in hyperspectral imaging technology enables the application in driving scenarios and the hope is…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Nick Theisen , Robin Bartsch , Dietrich Paulus , Peer Neubert

Existing diffusion-based super-resolution approaches often exhibit semantic ambiguities due to inaccuracies and incompleteness in their text conditioning, coupled with the inherent tendency for cross-attention to divert towards irrelevant…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Chen Chen , Majid Abdolshah , Violetta Shevchenko , Hongdong Li , Chang Xu , Pulak Purkait

Learning commonsense reasoning from visual contexts and scenes in real-world is a crucial step toward advanced artificial intelligence. However, existing video reasoning benchmarks are still inadequate since they were mainly designed for…

计算机视觉与模式识别 · 计算机科学 2024-05-20 Andong Wang , Bo Wu , Sunli Chen , Zhenfang Chen , Haotian Guan , Wei-Ning Lee , Li Erran Li , Chuang Gan

Establishing stable mappings between natural language expressions and visual percepts is a foundational problem for both cognitive science and artificial intelligence. Humans routinely ground linguistic reference in noisy, ambiguous…

人工智能 · 计算机科学 2026-02-24 Joseph Bingham

We introduce \textsc{MathSticks}, a benchmark for Visual Symbolic Compositional Reasoning (VSCR), which unifies visual perception, symbolic manipulation, and arithmetic consistency. Each task presents an incorrect matchstick equation that…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yuheng Ji , Huajie Tan , Cheng Chi , Yijie Xu , Yuting Zhao , Enshen Zhou , Huaihai Lyu , Pengwei Wang , Zhongyuan Wang , Shanghang Zhang , Xiaolong Zheng

Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Zhicheng Wang , Zhiyu Pan , Zhan Peng , Jian Cheng , Liwen Xiao , Wei Jiang , Zhiguo Cao

Remote Sensing Visual Grounding (RSVG) aims to localize target objects in large-scale aerial imagery based on natural language descriptions. Owing to the vast spatial scale and high semantic ambiguity of remote sensing scenes, these…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Shiqi Huang , Shuting He , Bihan Wen

Visual grounding is a task to ground referring expressions in images, e.g., localize "the white truck in front of the yellow one". To resolve this task fundamentally, the model should first find out the contextual objects (e.g., the…

计算机视觉与模式识别 · 计算机科学 2020-04-13 Daqing Liu , Hanwang Zhang , Zheng-Jun Zha , Meng Wang , Qianru Sun

As a critical clue of video super-resolution (VSR), inter-frame alignment significantly impacts overall performance. However, accurate pixel-level alignment is a challenging task due to the intricate motion interweaving in the video. In…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Qi Tang , Yao Zhao , Meiqin Liu , Jian Jin , Chao Yao

Despite recent advancements in vision-language models, their performance remains suboptimal on images from non-western cultures due to underrepresentation in training datasets. Various benchmarks have been proposed to test models' cultural…

计算与语言 · 计算机科学 2024-07-02 Mehar Bhatia , Sahithya Ravi , Aditya Chinchure , Eunjeong Hwang , Vered Shwartz

Prior studies on 3D scene understanding have primarily developed specialized models for specific tasks or required task-specific fine-tuning. In this study, we propose Grounded 3D-LLM, which explores the potential of 3D large multi-modal…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Yilun Chen , Shuai Yang , Haifeng Huang , Tai Wang , Runsen Xu , Ruiyuan Lyu , Dahua Lin , Jiangmiao Pang

Video Referring Expression Comprehension (REC) aims to localize a target object in video frames referred by the natural language expression. Recently, the Transformerbased methods have greatly boosted the performance limit. However, we…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Ji Jiang , Meng Cao , Tengtao Song , Yuexian Zou

While reinforcement learning (RL) over chains of thought has significantly advanced language models in tasks such as mathematics and coding, visual reasoning introduces added complexity by requiring models to direct visual attention,…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Gabriel Sarch , Snigdha Saha , Naitik Khandelwal , Ayush Jain , Michael J. Tarr , Aviral Kumar , Katerina Fragkiadaki

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limits the deployment of…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Manjunath Prasad Holenarasipura Rajiv , B. M. Vidyavathi

Despite recent successes, test-time scaling - i.e., dynamically expanding the token budget during inference as needed - remains brittle for vision-language models (VLMs): unstructured chains-of-thought about images entangle perception and…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Niccolo Avogaro , Nayanika Debnath , Li Mi , Thomas Frick , Junling Wang , Zexue He , Hang Hua , Konrad Schindler , Mattia Rigotti

Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpretable scalar scores rather than characterizing physics-driven RS degradations, deviating…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Chen Zhong , Xiao An , Jiaxing Sun , Zihan Gui , Guangyi Yang , Wei He

We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision-language models for remote sensing images. Although several vision-language datasets in remote sensing have been proposed to pursue this…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Xiang Li , Jian Ding , Mohamed Elhoseiny

Scene graph generation (SGG) is a sophisticated task that suffers from both complex visual features and dataset long-tail problem. Recently, various unbiased strategies have been proposed by designing novel loss functions and data balancing…

计算机视觉与模式识别 · 计算机科学 2023-03-03 Xiaoguang Chang , Teng Wang , Shaowei Cai , Changyin Sun

The human language is one of the most natural interfaces for humans to interact with robots. This paper presents a robot system that retrieves everyday objects with unconstrained natural language descriptions. A core issue for the system is…

机器人学 · 计算机科学 2017-07-19 Mohit Shridhar , David Hsu

Given a question-image input, the Visual Commonsense Reasoning (VCR) model can predict an answer with the corresponding rationale, which requires inference ability from the real world. The VCR task, which calls for exploiting the…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Xuejiao Tang , Wenbin Zhang