English
Related papers

Related papers: TaskGround: Structured Executable Task Inference f…

200 papers

Grounding natural language in 3D environments is a critical step toward achieving robust 3D vision-language alignment. Current datasets and models for 3D visual grounding predominantly focus on identifying and localizing objects from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Zhuofan Zhang , Ziyu Zhu , Junhao Li , Pengxiang Li , Tianxu Wang , Tengyu Liu , Xiaojian Ma , Yixin Chen , Baoxiong Jia , Siyuan Huang , Qing Li

Real-world robots localize objects from natural-language instructions while scenes around them keep changing. Yet most of the existing 3D visual grounding (3DVG) method still assumes a reconstructed and up-to-date point cloud, an assumption…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Miao Hu , Zhiwei Huang , Tai Wang , Jiangmiao Pang , Dahua Lin , Nanning Zheng , Runsen Xu

Despite recent advances in text-to-image generation, models still struggle to accurately render prompt-specified text with correct spatial layout -- especially in multi-span, structured settings. This challenge is driven not only by the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Dongxing Mao , Yilin Wang , Linjie Li , Zhengyuan Yang , Alex Jinpeng Wang

Large multimodal models have demonstrated impressive problem-solving abilities in vision and language tasks, and have the potential to encode extensive world knowledge. However, it remains an open challenge for these models to perceive,…

Artificial Intelligence · Computer Science 2024-09-24 Yew Ken Chia , Qi Sun , Lidong Bing , Soujanya Poria

GUI grounding aims to align natural language instructions with precise regions in complex user interfaces. Advanced multimodal large language models show strong ability in visual GUI grounding but still struggle with small or visually…

Artificial Intelligence · Computer Science 2025-12-02 Aiden Yiliu Li , Bizhi Yu , Daoan Lei , Tianhe Ren , Shilong Liu

LLM agents achieve 85-96% success on tasks where instructions fully specify the action, but drop to 29-53% when action feasibility depends on environmental state that the instruction does not mention. We argue that this gap reflects a…

Computation and Language · Computer Science 2026-05-29 Zixuan Wang , Dingming Li , Hongxing Li , Yanrui Miao , Shuo Chen , Yuchen Yan , Wenqi Zhang , Yongliang Shen , Weiming Lu , Jun Xiao , Yueting Zhuang

We present TartanGround, a large-scale, multi-modal dataset to advance the perception and autonomy of ground robots operating in diverse environments. This dataset, collected in various photorealistic simulation environments includes…

Robotics · Computer Science 2025-07-31 Manthan Patel , Fan Yang , Yuheng Qiu , Cesar Cadena , Sebastian Scherer , Marco Hutter , Wenshan Wang

Vision-Language Models (VLMs) empower embodied agents to execute complex instructions, yet they remain vulnerable to contextual safety risks where benign commands become hazardous due to subtle environmental states. Existing safeguards…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xiaoya Lu , Yijin Zhou , Zeren Chen , Ruocheng Wang , Bingrui Sima , Enshen Zhou , Lu Sheng , Dongrui Liu , Jing Shao

Following the success of the in-context learning paradigm in large-scale language and computer vision models, the recently emerging field of in-context reinforcement learning is experiencing a rapid growth. However, its development has been…

Machine Learning · Computer Science 2025-03-04 Alexander Nikulin , Ilya Zisman , Alexey Zemtsov , Vladislav Kurenkov

Spatio-Temporal Video Grounding (STVG) aims to localize target objects in videos based on natural language descriptions. Despite recent advances in Multimodal Large Language Models, a significant gap remains between current models and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Hong Gao , Jingyu Wu , Xiangkai Xu , Kangni Xie , Yunchen Zhang , Bin Zhong , Xurui Gao , Min-Ling Zhang

We present UGround, a \textbf{U}nified visual \textbf{Ground}ing paradigm that dynamically selects intermediate layers across \textbf{U}nrolled transformers as ``mask as prompt,'' diverging from the prevailing pipeline that leverages the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Rui Qian , Xin Yin , Chuanhang Deng , Zhiyuan Peng , Jian Xiong , Wei Zhai , Dejing Dou

A core task in embodied intelligence is ego-centric 3D visual grounding. Existing methods typically adopt two-stage, heterogeneous pipelines that pair a detector with a separate grounding model. Incompatible decoders and box heads hinder…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Yani Zhang , Dongming Wu , Hao Shi , Yingfei Liu , Tiancai Wang , Xingping Dong

Accurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the…

Robotics · Computer Science 2024-05-13 René Zurbrügg , Yifan Liu , Francis Engelmann , Suryansh Kumar , Marco Hutter , Vaishakh Patil , Fisher Yu

Global perception is essential for embodied agents in 360{\deg} spaces, yet current affordance grounding remains largely object-centric and restricted to perspective views. To bridge this gap, we introduce a novel task: Holistic Affordance…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Guoliang Zhu , Wanjun Jia , Caoyang Shao , Yuheng Zhang , Zhiyong Li , Kailun Yang

Complex reasoning problems often involve implicit spatial and geometric relationships that are not explicitly encoded in text. While recent reasoning models perform well across many domains, purely text-based reasoning struggles to capture…

Computation and Language · Computer Science 2026-01-07 Meiqi Chen , Fandong Meng , Jie Zhou

Task planning for robots in real-life settings presents significant challenges. These challenges stem from three primary issues: the difficulty in identifying grounded sequences of steps to achieve a goal; the lack of a standardized mapping…

For robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it…

Robotics · Computer Science 2016-02-03 Chengxi Ye , Yezhou Yang , Cornelia Fermuller , Yiannis Aloimonos

In this study, we explore the sophisticated domain of task planning for robust household embodied agents, with a particular emphasis on the intricate task of selecting substitute objects. We introduce the CommonSense Object Affordance Task…

Artificial Intelligence · Computer Science 2024-10-24 Ayush Agrawal , Raghav Prabhakar , Anirudh Goyal , Dianbo Liu

Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DINOv2, offer excellent features but are pretrained on object-centric data. We find that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Thinesh Thiyakesan Ponbagavathi , Chengzheng Yang , Alina Roitberg

Vickrey-Clarke-Groves (VCG) mechanisms are often used to allocate tasks to selfish and rational agents. VCG mechanisms are incentive compatible, direct mechanisms that are efficient (i.e., maximise social utility) and individually rational…