English
Related papers

Related papers: UGround: Towards Unified Visual Grounding with Unr…

200 papers

Understanding and localizing objects in complex 3D environments from natural language descriptions, known as 3D Visual Grounding (3DVG), is a foundational challenge in embodied AI, with broad implications for robotics, augmented reality,…

Robotics · Computer Science 2026-03-10 Jiaxi Zhang , Yunheng Wang , Wei Lu , Taowen Wang , Weisheng Xu , Shuning Zhang , Yixiao Feng , Yuetong Fang , Renjing Xu

Recent advancements in robotic manipulation have highlighted the potential of intermediate representations for improving policy generalization. In this work, we explore grounding masks as an effective intermediate representation, balancing…

Robotics · Computer Science 2025-05-01 Haifeng Huang , Xinyi Chen , Yilun Chen , Hao Li , Xiaoshen Han , Zehan Wang , Tai Wang , Jiangmiao Pang , Zhou Zhao

Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning…

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

Robotics · Computer Science 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Remote sensing (RS) visual grounding aims to use natural language expression to locate specific objects (in the form of the bounding box or segmentation mask) in RS images, enhancing human interaction with intelligent RS interpretation…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Yue Zhou , Mengcheng Lan , Xiang Li , Litong Feng , Yiping Ke , Xue Jiang , Qingyun Li , Xue Yang , Wayne Zhang

Graphical User Interface (GUI) grounding aims to translate natural language instructions into executable screen coordinates, enabling automated GUI interaction. Nevertheless, incorrect grounding can result in costly, hard-to-reverse actions…

Artificial Intelligence · Computer Science 2026-02-04 Qingni Wang , Yue Fan , Xin Eric Wang

UAV-ground visual tracking (UGVT) aims to simultaneously track the same object from both the UAV and the ground view. However, existing two-stream methods suffer from isolated feature extraction and rely heavily on implicit appearance…

Computer Vision and Pattern Recognition · Computer Science 2026-05-07 Boyue Xu , Ruichao Hou , Tongwei Ren , Gangshan Wu

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD).…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Minseok Kang , Minhyeok Lee , Minjung Kim , Donghyeong Kim , Sangyoun Lee

Inspired by the activity-silent and persistent activity mechanisms in human visual perception biology, we design a Unified Static and Dynamic Network (UniSDNet), to learn the semantic association between the video and text/audio queries in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jingjing Hu , Dan Guo , Kun Li , Zhan Si , Xun Yang , Xiaojun Chang , Meng Wang

Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We posit this limitation arises from the scarcity of high-quality, large-scale clinical…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Mengmeng Zhang , Xiaoping Wu , Hao Luo , Fan Wang , Yisheng Lv

Robust 3D representation learning forms the perceptual foundation of spatial intelligence, enabling downstream tasks in scene understanding and embodied AI. However, learning such representations directly from unposed multi-view images…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Bo Zhou , Qiuxia Lai , Zeren Sun , Xiangbo Shu , Yazhou Yao , Wenguan Wang

In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Weixin Ye , Wei Wang , Yahui Liu , Yue Song , Bin Ren , Wei Bi , Rita Cucchiara , Nicu Sebe

Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frameworks. Moment-level approaches directly predict the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Xing Cheng , Xiangyu Wu , Dong Shen , Hezheng Lin , Fan Yang

Constructing 4D language fields is crucial for embodied AI, augmented/virtual reality, and 4D scene understanding, as they provide enriched semantic representations of dynamic environments and enable open-vocabulary querying in complex…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Xianfeng Wu , Yajing Bai , Minghan Li , Xianzu Wu , Xueqi Zhao , Zhongyuan Lai , Wenyu Liu , Xinggang Wang

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored language-guided…

Robotics · Computer Science 2025-12-25 Zebin Jiang , Tianle Jin , Xiangtong Yao , Alois Knoll , Hu Cao

A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate grounding model. Incompatible decoders and box heads hinder…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yani Zhang , Dongming Wu , Hao Shi , Yingfei Liu , Tiancai Wang , Xingping Dong

Nonlinear finite element crash simulations are accurate but computationally expensive, limiting their use in iterative design optimisation. Machine-learning surrogate models based on graph neural networks (GNNs) offer a faster alternative.…

Machine Learning · Computer Science 2026-05-18 Haoran Li , Tobias Lehrer , Yingxue Zhao , Haosu Zhou , Philipp Stocker , Tobias Pfaff , Nan Li

Open-vocabulary 3D visual grounding and reasoning aim to localize objects in a scene based on implicit language descriptions, even when they are occluded. This ability is crucial for tasks such as vision-language navigation and autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Zhenyang Liu , Yikai Wang , Sixiao Zheng , Tongying Pan , Longfei Liang , Yanwei Fu , Xiangyang Xue

Open-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally struggle to segment unified objects consistently across diverse text reference expressions.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Zhihua Liu , Amrutha Saseendran , Lei Tong , Xilin He , Fariba Yousefi , Nikolay Burlutskiy , Dino Oglic , Tom Diethe , Philip Teare , Huiyu Zhou , Chen Jin

Given an untrimmed video, temporal sentence grounding (TSG) aims to locate a target moment semantically according to a sentence query. Although previous respectable works have made decent success, they only focus on high-level visual…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Xiang Fang , Daizong Liu , Pan Zhou , Guoshun Nan
‹ Prev 1 2 3 10 Next ›