English
Related papers

Related papers: Dream2Real: Zero-Shot 3D Object Rearrangement with…

200 papers

Deploying reinforcement learning policies trained in simulation to real autonomous vehicles remains a fundamental challenge, particularly for VLM-guided RL frameworks whose policies are typically learned with simulator-native observations…

Robotics · Computer Science 2026-04-07 Zilin Huang , Zhengyang Wan , Zihao Sheng , Boyue Wang , Junwei You , Yue Leng , Sikai Chen

Vision-Language Models (VLMs) exhibit strong visual reasoning capabilities, yet they still struggle with 3D understanding. In particular, VLMs often fail to infer a text-consistent goal 6D pose of a target object in a 3D scene. However, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Sangwon Baik , Gunhee Kim , Mingi Choi , Hanbyul Joo

The choice of scene representation is crucial in both the shape inference algorithms it requires and the smart applications it enables. We present efficient and optimisable multi-class learned object descriptors together with a novel…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Edgar Sucar , Kentaro Wada , Andrew Davison

Interpreting object-referential language and grounding objects in 3D with spatial relations and attributes is essential for robots operating alongside humans. However, this task is often challenging due to the diversity of scenes, large…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Nader Zantout , Haochen Zhang , Pujith Kachana , Jinkai Qiu , Guofei Chen , Ji Zhang , Wenshan Wang

Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and embodied agents. We attribute this to a modality gap between…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Yifan Liu , Fangneng Zhan , Kaichen Zhou , Yilun Du , Paul Pu Liang , Hanspeter Pfister

The ability to understand and reason the 3D real world is a crucial milestone towards artificial general intelligence. The current common practice is to finetune Large Language Models (LLMs) with 3D data and texts to enable 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Sha Zhang , Di Huang , Jiajun Deng , Shixiang Tang , Wanli Ouyang , Tong He , Yanyong Zhang

Vision--language models reliably name objects in a scene, but do they represent the 3D layout those objects inhabit? We introduce a 3,034-sample human-curated benchmark targeting three components of spatial understanding: depth-ordered…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Animesh Maheshwari , Divyansh Sahu , Nishit Verma

We present an approach to improve 3D vehicle labeling in self-driving applications through zero-shot inference of vehicle information, leveraging Vehicle Make and Model Recognition (VMMR) methods. The proposed approach utilizes a Vision…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Steven Chen , Shivesh Khaitan , Nemanja Djuric

We consider the problem of finding spatial configurations of multiple objects in images, e.g., a mobile inspection robot is tasked to localize abandoned tools on the floor. We define the spatial configuration of objects by first-order logic…

Generalization is a pivotal challenge for agents following natural language instructions. To approach this goal, we leverage a vision-language model (VLM) for visual grounding and transfer its vision-language knowledge into reinforcement…

Artificial Intelligence · Computer Science 2024-08-06 Haobin Jiang , Zongqing Lu

Real-world applications, such as autonomous driving and humanoid robot manipulation, require precise spatial perception. However, it remains underexplored how Vision-Language Models (VLMs) recognize spatial relationships and perceive…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Fei Kong , Jinhao Duan , Kaidi Xu , Zhenhua Guo , Xiaofeng Zhu , Xiaoshuang Shi

The strong performance of large vision-language models (VLMs) trained with reinforcement learning (RL) has motivated similar approaches for fine-tuning vision-language-action (VLA) models in robotics. Many recent works fine-tune VLAs…

Robotics · Computer Science 2026-03-31 Andrew Choi , Xinjie Wang , Zhizhong Su , Wei Xu

Vision-Language Pre-training (VLP) has advanced the performance of many vision-language tasks, such as image-text retrieval, visual entailment, and visual reasoning. The pre-training mostly utilizes lexical databases and image queries in…

Computation and Language · Computer Science 2023-06-30 Yasmine Karoui , Rémi Lebret , Negar Foroutan , Karl Aberer

Language-guided embodied navigation requires an agent to interpret object-referential instructions, search across multiple rooms, localize the referenced target, and execute reliable motion toward it. Existing systems remain limited in real…

Robotics · Computer Science 2026-03-19 Zhongyuang Liu , Min He , Shaonan Yu , Xinhang Xu , Muqing Cao , Jianping Li , Jianfei Yang , Lihua Xie

The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also perceive and adapt…

Robotics · Computer Science 2025-08-19 Zhichen Lou , Kechun Xu , Zhongxiang Zhou , Rong Xiong

Commanding a robot to navigate with natural language instructions is a long-term goal for grounded language understanding and robotics. But the dominant language is English, according to previous studies on vision-language navigation (VLN).…

Computation and Language · Computer Science 2020-12-08 An Yan , Xin Eric Wang , Jiangtao Feng , Lei Li , William Yang Wang

Previous Vision-Language-Action models face critical limitations in navigation: scarce, diverse data from labor-intensive collection and static representations that fail to capture temporal dynamics and physical laws. We propose NavDreamer,…

Robotics · Computer Science 2026-02-11 Xijie Huang , Weiqi Gai , Tianyue Wu , Congyu Wang , Zhiyang Liu , Xin Zhou , Yuze Wu , Fei Gao

We present Real2Code, a novel approach to reconstructing articulated objects via code generation. Given visual observations of an object, we first reconstruct its part geometry using an image segmentation model and a shape completion model.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Zhao Mandi , Yijia Weng , Dominik Bauer , Shuran Song

Multimodal 3D grounding has garnered considerable interest in Vision-Language Models (VLMs) \cite{yin2025spatial} for advancing spatial reasoning in complex environments. However, these models suffer from a severe "2D semantic bias" that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Yutong Zhong

Abstract semantic 3D scene understanding is a problem of critical importance in robotics. As robots still lack the common-sense knowledge about household objects and locations of an average human, we investigate the use of pre-trained…

Robotics · Computer Science 2023-11-09 William Chen , Siyi Hu , Rajat Talak , Luca Carlone
‹ Prev 1 3 4 5 6 7 10 Next ›