English
Related papers

Related papers: UGround: Towards Unified Visual Grounding with Unr…

200 papers

We propose VISO-Grasp, a novel vision-language-informed system designed to systematically address visibility constraints for grasping in severely occluded environments. By leveraging Foundation Models (FMs) for spatial reasoning and active…

Robotics · Computer Science 2025-08-07 Yitian Shi , Di Wen , Guanqi Chen , Edgar Welte , Sheng Liu , Kunyu Peng , Rainer Stiefelhagen , Rania Rayyes

Foundation segmentation models such as Segment Anything Model (SAM) are now routinely used in iterative pipelines, where each predicted mask is fed back as the next prompt. This practice turns segmentation into a closed-loop dynamical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 H. M. Shadman Tabib , Md. Shamsuzzoha Bayzid , M Sohel Rahman

We present a fast, spatio-temporal scene understanding framework based on Visual Geometry Grounded Transformer (VGGT). The proposed pipeline is designed to enable efficient, close to real-time performance, supporting applications including…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Gergely Dinya , Péter Halász , András Lőrincz , Kristóf Karacs , Anna Gelencsér-Horváth

Robots navigating dynamic, cluttered, and semantically complex environments must integrate perception, symbolic reasoning, and spatial planning to generalize across diverse layouts and object categories. Existing methods often rely on…

Robotics · Computer Science 2025-10-14 Ahmed Alanazi , Duy Ho , Yugyung Lee

Fine-grained image-text alignment is a pivotal challenge in multimodal learning, underpinning key applications such as visual question answering, image captioning, and vision-language navigation. Unlike global alignment, fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jiale Liu , Haoming Zhou , Yishu Liu , Bingzhi Chen , Yuncheng Jiang

We focus on the confounding bias between language and location in the visual grounding pipeline, where we find that the bias is the major visual reasoning bottleneck. For example, the grounding process is usually a trivial language-location…

Computer Vision and Pattern Recognition · Computer Science 2022-01-03 Jianqiang Huang , Yu Qin , Jiaxin Qi , Qianru Sun , Hanwang Zhang

Recently, generative graph models have shown promising results in learning graph representations through self-supervised methods. However, most existing generative graph representation learning (GRL) approaches rely on random masking across…

Machine Learning · Computer Science 2026-05-08 Xinyue Hu , Zhibin Duan , Xinyang Liu , Yuxin Li , Bo Chen , Chaojie Wang , Yilin He , Hongwei Liu , Mingyuan Zhou

Unified segmentation of 3D point clouds is crucial for scene understanding, but is hindered by its sparse structure, limited annotations, and the challenge of distinguishing fine-grained object classes in complex environments. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Zongyan Han , Mohamed El Amine Boudjoghra , Jiahua Dong , Jinhong Wang , Rao Muhammad Anwer

Prompt tuning, like CoOp, has recently shown promising vision recognizing and transfer learning ability on various downstream tasks with the emergence of large pre-trained vision-language models like CLIP. However, we identify that existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yongzhu Miao , Shasha Li , Jintao Tang , Ting Wang

The rapid advancement of vision-language models has catalyzed the emergence of GUI agents, which hold immense potential for automating complex tasks, from online shopping to flight booking, thereby alleviating the burden of repetitive…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Zhongyin Zhao , Yuan Liu , Yikun Liu , Haicheng Wang , Le Tian , Xiao Zhou , Yangxiu You , Zilin Yu , Yang Yu , Jie Zhou

3D Visual Grounding (3DVG) seeks to locate target objects in 3D scenes using natural language descriptions, enabling downstream applications such as augmented reality and robotics. Existing approaches typically rely on labeled 3D data and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Rong Li , Shijie Li , Lingdong Kong , Xulei Yang , Junwei Liang

Inspection of confined infrastructure such as culverts often requires accessing hidden spaces whose entrances are reachable primarily from elevated viewpoints. Aerial-ground cooperation enables a UAV to deploy a compact UGV for interior…

Robotics · Computer Science 2026-03-17 Seoyoung Lee , Shaekh Mohammad Shithil , Durgakant Pushp , Lantao Liu , Zhangyang Wang

The growing availability of co-located geospatial data spanning aerial imagery, street-level views, elevation models, text, and geographic coordinates offers a unique opportunity for multimodal representation learning. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Guillaume Astruc , Eduard Trulls , Jan Hosang , Loic Landrieu , Paul-Edouard Sarlin

Dexterous grasping aims to produce diverse grasping postures with a high grasping success rate. Regression-based methods that directly predict grasping parameters given the object may achieve a high success rate but often lack diversity.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Jiaxin Lu , Hao Kang , Haoxiang Li , Bo Liu , Yiding Yang , Qixing Huang , Gang Hua

We tackle the challenge of open-vocabulary segmentation, where we need to identify objects from a wide range of categories in different environments, using text prompts as our input. To overcome this challenge, existing methods often use…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Yu-Jhe Li , Xinyang Zhang , Kun Wan , Lantao Yu , Ajinkya Kale , Xin Lu

Embodied vision-based real-world systems, such as mobile robots, require a careful balance between energy consumption, compute latency, and safety constraints to optimize operation across dynamic tasks and contexts. As local computation…

In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such requests require agents to identify task-relevant entities,…

Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning isolated modules per task leads to parameter explosion.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Xuezhi Cui , Dongbo Zhou , Wang Guo , Zeyuan Wang , Ziyu Li , Gaozhi Zhou , Xian Li , Ling Zhao , Wentao Yang , Chao Tao , Haifeng Li

We rethink the segment anything model (SAM) and propose a novel multiprompt network called COMPrompter for camouflaged object detection (COD). SAM has zero-shot generalization ability beyond other models and can provide an ideal framework…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Xiaoqin Zhang , Zhenni Yu , Li Zhao , Deng-Ping Fan , Guobao Xiao

This work addresses a fundamental challenge in applying deep learning to power systems: developing neural network models that transfer across significant system changes, including networks with entirely different topologies and…

Systems and Control · Electrical Eng. & Systems 2025-09-11 Tong Wu , Anna Scaglione , Sandy Miguel , Daniel Arnold