中文
相关论文

相关论文: GRASP: Geospatial pixel Reasoning viA Structured P…

200 篇论文

The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large…

机器人学 · 计算机科学 2025-02-10 Houjian Yu , Mingen Li , Alireza Rezazadeh , Yang Yang , Changhyun Choi

Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct supervision falls far…

With the flourishing prosperity of generative models, manipulated facial images have become increasingly accessible, raising concerns regarding privacy infringement and societal trust. In response, proactive defense strategies embed…

密码学与安全 · 计算机科学 2025-10-03 Yue Li , Linying Xue , Dongdong Lin , Qiushi Li , Hui Tian , Hongxia Wang

Representation alignment (REPA) guides generative training by distilling representations from a strong, pretrained vision encoder to intermediate diffusion features. We investigate a fundamental question: what aspect of the target…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Jaskirat Singh , Xingjian Leng , Zongze Wu , Liang Zheng , Richard Zhang , Eli Shechtman , Saining Xie

Reinforcement Learning (RL) can directly enhance the reasoning capabilities of large language models without extensive reliance on Supervised Fine-Tuning (SFT). In this work, we revisit the traditional Policy Gradient (PG) mechanism and…

机器学习 · 计算机科学 2026-02-04 Xiangxiang Chu , Hailang Huang , Xiao Zhang , Fei Wei , Yong Wang

Existing visual perception systems focus on region-level segmentation in single-turn dialogues, relying on complex and explicit query instructions. Such systems cannot reason at the pixel level and comprehend dynamic user intent that…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Dexian Cai , Xiaocui Yang , Yongkang Liu , Daling Wang , Shi Feng , Yifei Zhang , Soujanya Poria

Recent advancements in robotic grasping have led to its integration as a core module in many manipulation systems. For instance, language-driven semantic segmentation enables the grasping of any designated object or object part. However,…

机器人学 · 计算机科学 2025-07-09 Yun Du , Mengao Zhao , Tianwei Lin , Yiwei Jin , Chaodong Huang , Zhizhong Su

Enabling VLA models to predict environmental dynamics, known as world modeling, has been recognized as essential for improving robotic reasoning and generalization. However, current approaches face two main issues: 1. The training objective…

机器人学 · 计算机科学 2026-02-20 Han Zhao , Jingbo Wang , Wenxuan Song , Shuai Chen , Yang Liu , Yan Wang , Haoang Li , Donglin Wang

Segmentation is a fundamental task in computer vision, underpinning pixel-level scene understanding and serving as a cornerstone for applications ranging from autonomous perception to medical image analysis. For complex referring…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Mario Markov , Stefan Maria Ailuro , Mohammad Mahdi , Luc Van Gool , Danda Pani Paudel

Face segmentation is the task of densely labeling pixels on the face according to their semantics. While current methods place an emphasis on developing sophisticated architectures, use conditional random fields for smoothness, or rather…

计算机视觉与模式识别 · 计算机科学 2020-03-31 Iacopo Masi , Joe Mathai , Wael AbdAlmageed

Existing reinforcement learning approaches for Large Language Models typically perform policy optimization at the granularity of individual tokens or entire response sequences. However, such formulations often misalign with the natural…

人工智能 · 计算机科学 2026-05-08 Lei Gao , Zhuoming Li , Mengxi Jia , Jiakang Yuan , Hongbo Sun , Hao Sun , Xuelong Li

Self-attention is of vital importance in semantic segmentation as it enables modeling of long-range context, which translates into improved performance. We argue that it is equally important to model short-range context, especially to…

计算机视觉与模式识别 · 计算机科学 2022-12-29 Hasib Zunair , A. Ben Hamza

Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Tao Zhang , Jinyong Wen , Zhen Chen , Kun Ding , Shiming Xiang , Chunhong Pan

Open-vocabulary image segmentation has been advanced through the synergy between mask generators and vision-language models like Contrastive Language-Image Pre-training (CLIP). Previous approaches focus on generating masks while aligning…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Quan-Sheng Zeng , Yunheng Li , Daquan Zhou , Guanbin Li , Qibin Hou , Ming-Ming Cheng

Previous methods for image geo-localization have typically treated the task as either classification or retrieval, often relying on black-box decisions that lack interpretability. The rise of large vision-language models (LVLMs) has enabled…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Ling Li , Yao Zhou , Yuxuan Liang , Fugee Tsung , Jiaheng Wei

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image…

计算机视觉与模式识别 · 计算机科学 2025-01-24 Akashah Shabbir , Mohammed Zumri , Mohammed Bennamoun , Fahad S. Khan , Salman Khan

Despite impressive advancements in Visual-Language Models (VLMs) for multi-modal tasks, their reliance on RGB inputs limits precise spatial understanding. Existing methods for integrating spatial cues, such as point clouds or depth, either…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Yang Liu , Ming Ma , Xiaomin Yu , Pengxiang Ding , Han Zhao , Mingyang Sun , Siteng Huang , Donglin Wang

Large Language Models (LLMs) have demonstrated impressive capabilities in language processing, yet they often struggle with tasks requiring genuine visual spatial reasoning. In this paper, we introduce a novel two-stage training framework…

计算与语言 · 计算机科学 2025-02-26 Alan Dao , Dinh Bach Vu

Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting their value as synthetic augmentation for rare conditions. We trace this to low head-versus-tail…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Felix Nützel , Mischa Dombrowski , Bernhard Kainz

Instruction-driven segmentation in remote sensing generates masks from guidance, offering great potential for accessible and generalizable applications. However, existing methods suffer from fragmented task formulations and limited…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Shuo Ni , Di Wang , He Chen , Haonan Guo , Ning Zhang , Jing Zhang