English
Related papers

Related papers: RoVi-Aug: Robot and Viewpoint Augmentation for Cro…

200 papers

Semantic image segmentation aims to obtain object labels with precise boundaries, which usually suffers from overfitting. Recently, various data augmentation strategies like regional dropout and mix strategies have been proposed to address…

Computer Vision and Pattern Recognition · Computer Science 2021-04-22 Jiawei Zhang , Yanchun Zhang , Xiaowei Xu

If generalist robots are to operate in truly unstructured environments, they need to be able to recognize and reason about novel objects and scenarios. Such objects and scenarios might not be present in the robot's own training data. We…

Robotics · Computer Science 2023-10-17 Kevin Black , Mitsuhiko Nakamoto , Pranav Atreya , Homer Walke , Chelsea Finn , Aviral Kumar , Sergey Levine

Vision Transformers (ViT) have been shown to attain highly competitive performance for a wide range of vision applications, such as image classification, object detection and semantic image segmentation. In comparison to convolutional…

Computer Vision and Pattern Recognition · Computer Science 2022-06-24 Andreas Steiner , Alexander Kolesnikov , Xiaohua Zhai , Ross Wightman , Jakob Uszkoreit , Lucas Beyer

Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large…

To aid humans in everyday tasks, robots need to know which objects exist in the scene, where they are, and how to grasp and manipulate them in different situations. Therefore, object recognition and grasping are two key functionalities for…

Robotics · Computer Science 2022-12-07 Hamidreza Kasaei , Sha Luo , Remo Sasso , Mohammadreza Kasaei

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled…

Robotics · Computer Science 2026-05-06 Zhiyuan Li , Wenyan Yang , Wenshuai Zhao , Yue Ma , Yuanpeng Tu , Pekka Marttinen , Joni Pajarinen

The deployment of humanoid robots for dexterous manipulation in unstructured environments remains challenging due to perceptual limitations that constrain the effective workspace. In scenarios where physical constraints prevent the robot…

Robotics · Computer Science 2026-03-09 Pei Qu , Zheng Li , Yufei Jia , Ziyun Liu , Liang Zhu , Haoang Li , Jinni Zhou , Jun Ma

The rise of foundation models paves the way for generalist robot policies in the physical world. Existing methods relying on text-only instructions often struggle to generalize to unseen scenarios. We argue that interleaved image-text…

Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense…

Robotics · Computer Science 2026-03-03 Zilong Xie , Jingyu Gong , Xin Tan , Zhizhong Zhang , Yuan Xie

The rise of generalist robotic policies has created an exponential demand for large-scale training data. However, on-robot data collection is labor-intensive and often limited to specific environments. In contrast, open-world images capture…

Generic object detection algorithms have proven their excellent performance in recent years. However, object detection on underwater datasets is still less explored. In contrast to generic datasets, underwater images usually have color…

Computer Vision and Pattern Recognition · Computer Science 2020-03-25 Wei-Hong Lin , Jia-Xing Zhong , Shan Liu , Thomas Li , Ge Li

Active perception in vision-based robotic manipulation aims to move the camera toward more informative observation viewpoints, thereby providing high-quality perceptual inputs for downstream tasks. Most existing active perception methods…

Robotics · Computer Science 2026-01-21 Deyun Qin , Zezhi Liu , Hanqian Luo , Xiao Liang , Yongchun Fang

Robots can use Visual Imitation Learning (VIL) to learn manipulation tasks from video demonstrations. However, translating visual observations into actionable robot policies is challenging due to the high-dimensional nature of video data.…

Robotics · Computer Science 2025-01-22 Ananth Jonnavittula , Sagar Parekh , Dylan P. Losey

Vision-based imitation learning has shown promising capabilities of endowing robots with various motion skills given visual observation. However, current visuomotor policies fail to adapt to drastic changes in their visual observations. We…

Robotics · Computer Science 2025-01-03 Pingcheng Jian , Easop Lee , Zachary Bell , Michael M. Zavlanos , Boyuan Chen

Segmenting unseen objects is a crucial ability for the robot since it may encounter new environments during the operation. Recently, a popular solution is leveraging RGB-D features of large-scale synthetic data and directly applying the…

Computer Vision and Pattern Recognition · Computer Science 2023-02-23 Lu Zhang , Siqi Zhang , Xu Yang , Hong Qiao , Zhiyong Liu

In recent years, the demand for social robots has grown, requiring them to adapt their behaviors based on users' states. Accurately assessing user experience (UX) in human-robot interaction (HRI) is crucial for achieving this adaptability.…

Robotics · Computer Science 2025-08-01 Ryo Miyoshi , Yuki Okafuji , Takuya Iwamoto , Junya Nakanishi , Jun Baba

The rapidly growing number of product categories in large-scale e-commerce makes accurate object identification for automated packing in warehouses substantially more difficult. As the catalog grows, intra-class variability and a long tail…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Xingwu Zhang , Guanxuan Li , Zhuocheng Zhang , Zijun Long

Estimating robot pose from RGB images is a crucial problem in computer vision and robotics. While previous methods have achieved promising performance, most of them presume full knowledge of robot internal states, e.g. ground-truth robot…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Shikun Ban , Juling Fan , Xiaoxuan Ma , Wentao Zhu , Yu Qiao , Yizhou Wang

Autonomous robot manipulation is a complex and continuously evolving robotics field. This paper focuses on data augmentation methods in imitation learning. Imitation learning consists of three stages: data collection from experts, learning…

Robotics · Computer Science 2024-10-08 Masato Kobayashi , Thanpimon Buamanee , Yuki Uranishi

Recent advancements utilizing large-scale video data for learning video generation models demonstrate significant potential in understanding complex physical dynamics. It suggests the feasibility of leveraging diverse robot trajectory data…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Youpeng Wen , Junfan Lin , Yi Zhu , Jianhua Han , Hang Xu , Shen Zhao , Xiaodan Liang