中文
相关论文

相关论文: GroundingSuite: Measuring Complex Multi-Granular P…

200 篇论文

The emergence of Multimodal Large Language Models (MLLMs) has driven significant advances in Graphical User Interface (GUI) agent capabilities. Nevertheless, existing GUI agent training and inference techniques still suffer from a dilemma…

人工智能 · 计算机科学 2026-04-09 Shuquan Lian , Yuhang Wu , Jia Ma , Yifan Ding , Zihan Song , Bingqi Chen , Xiawu Zheng , Hui Li , Rongrong Ji

Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these models excel in various downstream tasks, including visual…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Madhukar Reddy Vongala , Saurabh Srivastava , Jana Košecká

Referring Expression Segmentation (RES) aims to generate a segmentation mask for the object described by a given language expression. Existing classic RES datasets and methods commonly support single-target expressions only, i.e., one…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Chang Liu , Henghui Ding , Xudong Jiang

Approaches to Grounded Language Learning typically focus on a single task-based final performance measure that may not depend on desirable properties of the learned hidden representations, such as their ability to predict salient attributes…

Unlike Object Detection, Visual Grounding task necessitates the detection of an object described by complex free-form language. To simultaneously model such complex semantic and visual representations, recent state-of-the-art studies adopt…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Weitai Kang , Luowei Zhou , Junyi Wu , Changchang Sun , Yan Yan

In the domain of the U.S. Army modeling and simulation, the availability of high quality annotated 3D data is pivotal to creating virtual environments for training and simulations. Traditional methodologies for 3D semantic and instance…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Jiuyi Xu , Meida Chen , Andrew Feng , Zifan Yu , Yangming Shi

Benefiting from strong generalization ability, pre-trained vision language models (VLMs), e.g., CLIP, have been widely utilized in zero-shot scene understanding. Unlike simple recognition tasks, grounded situation recognition (GSR) requires…

计算机视觉与模式识别 · 计算机科学 2024-04-25 Jiaming Lei , Lin Li , Chunping Wang , Jun Xiao , Long Chen

Land Use and Land Cover (LULC) mapping is a fundamental task in Earth Observation (EO). However, current LULC models are typically developed for a specific modality and a fixed class taxonomy, limiting their generability and broader…

计算机视觉与模式识别 · 计算机科学 2026-05-04 Chenying Liu , Wei Huang , Xiao Xiang Zhu

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Current research…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Keyan Chen , Jiafan Zhang , Chenyang Liu , Zhengxia Zou , Zhenwei Shi

In recent years, language-guided open-set aerial object detection has gained significant attention due to its better alignment with real-world application needs. However, due to limited datasets, most existing language-guided methods…

计算机视觉与模式识别 · 计算机科学 2025-09-29 Guoting Wei , Yu Liu , Xia Yuan , Xizhe Xue , Linlin Guo , Yifan Yang , Chunxia Zhao , Zongwen Bai , Haokui Zhang , Rong Xiao

General robotic grasping systems require accurate object affordance perception in diverse open-world scenarios following human instructions. However, current studies suffer from the problem of lacking reasoning-based large-scale affordance…

计算机视觉与模式识别 · 计算机科学 2025-08-01 Dongming Wu , Yanping Fu , Saike Huang , Yingfei Liu , Fan Jia , Nian Liu , Feng Dai , Tiancai Wang , Rao Muhammad Anwer , Fahad Shahbaz Khan , Jianbing Shen

Recently, large multimodal models have built a bridge from visual to textual information, but they tend to underperform in remote sensing scenarios. This underperformance is due to the complex distribution of objects and the significant…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Cong Yang , Zuchao Li , Lefei Zhang

Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding methods heavily rely on…

计算机视觉与模式识别 · 计算机科学 2019-07-02 Zhiyuan Fang , Shu Kong , Charless Fowlkes , Yezhou Yang

Annotating a large-scale in-the-wild person re-identification dataset especially of marathon runners is a challenging task. The variations in the scenarios such as camera viewpoints, resolution, occlusion, and illumination make the problem…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Pranjal Singh Rajput , Yeshwanth Napolean , Jan van Gemert

Evaluating the alignment between textual prompts and generated images is critical for ensuring the reliability and usability of text-to-image (T2I) models. However, most existing evaluation methods rely on coarse-grained metrics or static…

计算机视觉与模式识别 · 计算机科学 2026-02-24 Fulin Shi , Wenyi Xiao , Bin Chen , Liang Din , Leilei Gan

Geospatial pixel reasoning aims to generate segmentation masks in remote sensing imagery directly from natural-language instructions. Most existing approaches follow a paradigm that fine-tunes multimodal large language models under…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Chengjie Jiang , Yunqi Zhou , Jiafeng Yan , Jing Li , Jiayang Li , Yue Zhou , Hongjie He , Jonathan Li

Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Xing Zi , Jinghao Xiao , Yunxiao Shi , Xian Tao , Jun Li , Ali Braytee , Mukesh Prasad

Extending image-based Large Multimodal Models (LMMs) to videos is challenging due to the inherent complexity of video data. The recent approaches extending image-based LMMs to videos either lack the grounding capabilities (e.g., VideoChat,…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Shehan Munasinghe , Rusiru Thushara , Muhammad Maaz , Hanoona Abdul Rasheed , Salman Khan , Mubarak Shah , Fahad Khan

Current 3D visual grounding tasks only process sentence level detection or segmentation, which critically fails to leverage the rich compositional contextual reasonings within natural language expressions. To address this challenge, we…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Qi Chen , Changli Wu , Jiayi Ji , Yiwei Ma , Liujuan Cao

Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves as an ideal testing…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Junzhuo Liu , Xuzheng Yang , Weiwei Li , Peng Wang
‹ 上一页 1 8 9 10 下一页 ›