中文
相关论文

相关论文: You Only Demonstrate Once: Category-Level Manipula…

200 篇论文

Large-scale robot learning has made progress on complex manipulation tasks, yet long horizon, contact rich problems, especially those involving deformable objects, remain challenging due to inconsistent demonstration quality. We propose a…

机器人学 · 计算机科学 2026-04-28 Qianzhong Chen , Justin Yu , Mac Schwager , Pieter Abbeel , Yide Shentu , Philipp Wu

Learning real-world robotic manipulation is challenging, particularly when limited demonstrations are available. Existing methods for few-shot manipulation often rely on simulation-augmented data or pre-built modules like grasping and pose…

The optimisation of crop harvesting processes for commonly cultivated crops is of great importance in the aim of agricultural industrialisation. Nowadays, the utilisation of machine vision has enabled the automated identification of crops,…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Hongyu Zhao , Zezhi Tang , Zhenhong Li , Yi Dong , Yuancheng Si , Mingyang Lu , George Panoutsos

Due to burdensome data requirements, learning from demonstration often falls short of its promise to allow users to quickly and naturally program robots. Demonstrations are inherently ambiguous and incomplete, making correct generalization…

机器学习 · 计算机科学 2019-04-29 Wonjoon Goo , Scott Niekum

This article presents a semantic tracker which simultaneously tracks a single target and recognises its category. In general, it is hard to design a tracking model suitable for all object categories, e.g., a rigid tracker for a car is not…

计算机视觉与模式识别 · 计算机科学 2016-11-22 Jingjing Xiao , Qiang Lan , Linbo Qiao , Ales Leonardis

We propose DINOBot, a novel imitation learning framework for robot manipulation, which leverages the image-level and pixel-level capabilities of features extracted from Vision Transformers trained with DINO. When interacting with a novel…

机器人学 · 计算机科学 2024-02-21 Norman Di Palo , Edward Johns

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

机器人学 · 计算机科学 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Given a demonstration of a complex manipulation task, such as pouring liquid from one container to another, we seek to generate a motion plan for a new task instance involving objects with different geometries. This is nontrivial since we…

机器人学 · 计算机科学 2026-02-03 Dibyendu Das , Aditya Patankar , Nilanjan Chakraborty , C. R. Ramakrishnan , I. V. Ramakrishnan

Image-point class incremental learning helps the 3D-points-vision robots continually learn category knowledge from 2D images, improving their perceptual capability in dynamic environments. However, some incremental learning methods address…

计算机视觉与模式识别 · 计算机科学 2025-04-14 Chao Qi , Jianqin Yin , Ren Zhang

Many robot manipulation tasks require the robot to make and break contact with objects and surfaces. The dynamics of such changing-contact robot manipulation tasks are discontinuous when contact is made or broken, and continuous elsewhere.…

机器人学 · 计算机科学 2021-06-22 Saif Sidhik , Mohan Sridharan , Dirk Ruiken

Garment manipulation (e.g., unfolding, folding and hanging clothes) is essential for future robots to accomplish home-assistant tasks, while highly challenging due to the diversity of garment configurations, geometries and deformations.…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Ruihai Wu , Haoran Lu , Yiyan Wang , Yubo Wang , Hao Dong

Video is a promising source of knowledge for embodied agents to learn models of the world's dynamics. Large deep networks have become increasingly effective at modeling complex video data in a self-supervised manner, as evaluated by metrics…

计算机视觉与模式识别 · 计算机科学 2023-04-27 Stephen Tian , Chelsea Finn , Jiajun Wu

Large-scale demonstration data has powered key breakthroughs in robot manipulation, but collecting that data remains costly and time-consuming. We present Constraint-Preserving Data Generation (CP-Gen), a method that uses a single expert…

机器人学 · 计算机科学 2025-08-07 Kevin Lin , Varun Ragunath , Andrew McAlinden , Aaditya Prasad , Jimmy Wu , Yuke Zhu , Jeannette Bohg

Category-level object pose estimation aims to predict the 6D pose and 3D size of objects within given categories. Existing approaches for this task rely solely on 6D poses as supervisory signals without explicitly capturing the intrinsic…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Zhujun Li , Shuo Zhang , Ioannis Stamos

Concept Bottleneck Models (CBMs) enable interpretable image classification by structuring predictions around human-understandable concepts, but extending this paradigm to video remains challenging due to the difficulty of extracting…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Patrick Knab , Sascha Marton , Philipp J. Schubert , Drago Guggiana , Christian Bartelt

Does multi-view demonstration truly improve robot manipulation, or merely enhance cross-view robustness? We present a systematic study quantifying the performance gains, scaling behavior, and underlying mechanisms of multi-view data for…

Fine-grained robot manipulation, such as lifting and rotating a bottle to display the label on the cap, requires robust reasoning about object parts and their relationships with intended tasks. Despite recent advances in training…

机器人学 · 计算机科学 2025-06-18 Yifan Yin , Zhengtao Han , Shivam Aarya , Jianxin Wang , Shuhang Xu , Jiawei Peng , Angtian Wang , Alan Yuille , Tianmin Shu

Visuomotor control (VMC) is an effective means of achieving basic manipulation tasks such as pushing or pick-and-place from raw images. Conditioning VMC on desired goal states is a promising way of achieving versatile skill primitives.…

机器人学 · 计算机科学 2021-09-27 Oliver Groth , Chia-Man Hung , Andrea Vedaldi , Ingmar Posner

The canonical approach to video action recognition dictates a neural model to do a classic and standard 1-of-N majority vote task. They are trained to predict a fixed set of predefined categories, limiting their transferable ability on new…

计算机视觉与模式识别 · 计算机科学 2021-09-20 Mengmeng Wang , Jiazheng Xing , Yong Liu

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Wenbin Tan , Jiawen Lin , Fangyong Wang , Yuan Xie , Yong Xie , Yachao Zhang , Yanyun Qu