English
Related papers

Related papers: Precise Mobile Manipulation of Small Everyday Obje…

200 papers

There is a growing interest in applying large language models (LLMs) in robotic tasks, due to their remarkable reasoning ability and extensive knowledge learned from vast training corpora. Grounding LLMs in the physical world remains an…

Robotics · Computer Science 2024-04-11 Wenqiang Lai , Yuan Gao , Tin Lun Lam

We present a method for visual object classification using only a single feature, transformed color SIFT with a variant of Spatial Pyramid Matching (SPM) that we called Sliding Spatial Pyramid Matching (SSPM), trained with an ensemble of…

Computer Vision and Pattern Recognition · Computer Science 2012-12-19 Hao Wooi Lim , Yong Haur Tay

Combining gradient-based trajectory optimization with differentiable physics simulation is an efficient technique for solving soft-body manipulation problems. Using a well-crafted optimization objective, the solver can quickly converge onto…

Machine Learning · Computer Science 2023-12-12 Zhiao Huang , Feng Chen , Yewen Pu , Chunru Lin , Hao Su , Chuang Gan

Recently, large language models (LLMs) and vision-language models (VLMs) have achieved significant success, demonstrating remarkable capabilities in understanding various images and videos, particularly in classification and detection…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Fei Wang , Chengcheng Chen , Hongyu Chen , Yugang Chang , Weiming Zeng

Transformers have revolutionized deep learning based computer vision with improved performance as well as robustness to natural corruptions and adversarial attacks. Transformers are used predominantly for 2D vision tasks, including image…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Hemang Chawla , Arnav Varma , Elahe Arani , Bahram Zonooz

Desktop cleaning demands open-vocabulary recognition and precise manipulation for heterogeneous debris. We propose a hierarchical framework integrating reflective Vision-Language Model (VLM) planning with dual-arm execution via structured…

Robotics · Computer Science 2025-06-24 Yufan Liu , Yi Wu , Gweneth Ge , Haoliang Cheng , Rui Liu

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text annotation of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Shijie Wang , Dahun Kim , Ali Taalimi , Chen Sun , Weicheng Kuo

Object search is a fundamental task for robots deployed in indoor building environments, yet challenges arise due to observation instability, especially for open-vocabulary models. While foundation models (LLMs/VLMs) enable reasoning about…

Robotics · Computer Science 2025-03-05 Qianwei Wang , Yifan Xu , Vineet Kamat , Carol Menassa

In this paper, we demonstrate that mobile manipulation policies utilizing a 3D latent map achieve stronger spatial and temporal reasoning than policies relying solely on images. We introduce Seeing the Bigger Picture (SBP), an end-to-end…

Robotics · Computer Science 2026-03-06 Sunghwan Kim , Woojeh Chung , Zhirui Dai , Dwait Bhatt , Arth Shukla , Hao Su , Yulun Tian , Nikolay Atanasov

Currently, manipulation tasks for deformable objects often focus on activities like folding clothes, handling ropes, and manipulating bags. However, research on contact-rich tasks involving deformable objects remains relatively…

Robotics · Computer Science 2026-03-06 Yuhang Zhang , Jinming Ma , Feng Wu

Following the recent popularity of Large Language Models (LLMs), several attempts have been made to extend them to the visual domain. From having a visual assistant that could guide us through unfamiliar environments to generative models…

There is a gap in the understanding of occluded objects in existing large-scale visual language multi-modal models. Current state-of-the-art multi-modal models fail to provide satisfactory results in describing occluded objects through…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Shuxin Yang , Xinhan Di

With the development of large language models, many remarkable linguistic systems like ChatGPT have thrived and achieved astonishing success on many tasks, showing the incredible power of foundation models. In the spirit of unleashing the…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Dingyuan Zhang , Dingkang Liang , Hongcheng Yang , Zhikang Zou , Xiaoqing Ye , Zhe Liu , Xiang Bai

Robotic manipulation in complex scenes demands precise perception of task-relevant details, yet fixed or suboptimal viewpoints often impair fine-grained perception and induce occlusions, constraining imitation-learned policies. We present…

Recent advances in robotics and autonomous systems have broadened the use of robots in laboratory settings, including automated synthesis, scalable reaction workflows, and collaborative tasks in self-driving laboratories (SDLs). This paper…

Robotics · Computer Science 2025-10-23 Shifa Sulaiman , Tobias Busk Jensen , Stefan Hein Bengtson , Simon Bøgh

Solving long-horizon tasks requires robots to integrate high-level semantic reasoning with low-level physical interaction. While vision-language models (VLMs) and video generation models can decompose tasks and imagine outcomes, they often…

Object Goal Navigation (ObjectNav) challenges robots to find objects in unseen environments, demanding sophisticated reasoning. While Vision-Language Models (VLMs) show potential, current ObjectNav methods often employ them superficially,…

Robotics · Computer Science 2025-06-23 Mobin Habibpour , Fatemeh Afghah

Robotic manipulation requires sophisticated commonsense reasoning, a capability naturally possessed by large-scale Vision-Language Models (VLMs). While VLMs show promise as zero-shot planners, their lack of grounded physical understanding…

Robotics · Computer Science 2026-03-18 Emily Yue-Ting Jia , Weiduo Yuan , Tianheng Shi , Vitor Guizilini , Jiageng Mao , Yue Wang

Video Understanding, Scene Interpretation and Commonsense Reasoning are highly challenging tasks enabling the interpretation of visual information, allowing agents to perceive, interact with and make rational decisions in its environment.…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Nicolas Schuler , Lea Dewald , Nick Baldig , Jürgen Graf

This paper presents an image-based visual servo control (IBVS) method for a first-person-view (FPV) quadrotor to conduct aggressive aerial tracking. There are three major challenges to maneuvering an underactuated vehicle using IBVS: (i)…

Robotics · Computer Science 2022-12-20 Chao Qin , Qiuyu Yu , Hugh H. T. Liu
‹ Prev 1 8 9 10 Next ›