中文
相关论文

相关论文: Act3D: 3D Feature Field Transformers for Multi-Tas…

200 篇论文

Recently, multi-view diffusion-based 3D generation methods have gained significant attention. However, these methods often suffer from shape and texture misalignment across generated multi-view images, leading to low-quality 3D generation…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Zhuojiang Cai , Yiheng Zhang , Meitong Guo , Mingdao Wang , Yuwang Wang

We study the task of embodied visual active learning, where an agent is set to explore a 3d environment with the goal to acquire visual scene understanding by actively selecting views for which to request annotation. While accurate on some…

计算机视觉与模式识别 · 计算机科学 2020-12-18 David Nilsson , Aleksis Pirinen , Erik Gärtner , Cristian Sminchisescu

The field of collaborative robotics and human-robot interaction often focuses on the prediction of human behaviour, while assuming the information about the robot setup and configuration being known. This is often the case with fixed…

机器人学 · 计算机科学 2019-02-18 Justinas Miseikis , Inka Brijacak , Saeed Yahyanejad , Kyrre Glette , Ole Jakob Elle , Jim Torresen

Robotic manipulation requires precise spatial understanding to interact with objects in the real world. Point-based methods suffer from sparse sampling, leading to the loss of fine-grained semantics. Image-based methods typically feed RGB…

机器人学 · 计算机科学 2026-01-14 Hao Shi , Bin Xie , Yingfei Liu , Yang Yue , Tiancai Wang , Haoqiang Fan , Xiangyu Zhang , Gao Huang

Functionality segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing methods rely on fragmented pipelines that suffer from visual…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Jiaying Lin , Dan Xu

Manipulator robots are increasingly being deployed in retail environments, yet contact rich edge cases still trigger costly human teleoperation. A prominent example is upright lying beverage bottles, where purely visual cues are often…

机器人学 · 计算机科学 2025-09-30 Ryo Watanabe , Maxime Alvarez , Pablo Ferreiro , Pavel Savkin , Genki Sano

Soft manipulators are known for their superiority in coping with high-safety-demanding interaction tasks, e.g., robot-assisted surgeries, elderly caring, etc. Yet the challenges residing in real-time contact feedback have hindered further…

机器人学 · 计算机科学 2024-09-13 Xiaohan Zhu , Ran Bu , Zhen Li , Fan Xu , Hesheng Wang

Human actions manipulating articulated objects, such as opening and closing a drawer, can be categorized into multiple modalities we define as interaction modes. Traditional robot learning approaches lack discrete representations of these…

机器人学 · 计算机科学 2024-10-29 Liquan Wang , Ankit Goyal , Haoping Xu , Animesh Garg

We would like robots to achieve purposeful manipulation by placing any instance from a category of objects into a desired set of goal states. Existing manipulation pipelines typically specify the desired configuration as a target 6-DOF pose…

机器人学 · 计算机科学 2019-10-30 Lucas Manuelli , Wei Gao , Peter Florence , Russ Tedrake

A promising direction for pre-training 3D point clouds is to leverage the massive amount of data in 2D, whereas the domain gap between 2D and 3D creates a fundamental challenge. This paper proposes a novel approach to point-cloud…

计算机视觉与模式识别 · 计算机科学 2024-04-30 Siming Yan , Chen Song , Youkang Kong , Qixing Huang

This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast Feature Field ($\text{F}^3$). We learn this representation by predicting future events from past…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Richeek Das , Kostas Daniilidis , Pratik Chaudhari

We explore how intermediate policy representations can facilitate generalization by providing guidance on how to perform manipulation tasks. Existing representations such as language, goal images, and trajectory sketches have been shown to…

机器人学 · 计算机科学 2024-11-06 Soroush Nasiriany , Sean Kirmani , Tianli Ding , Laura Smith , Yuke Zhu , Danny Driess , Dorsa Sadigh , Ted Xiao

Learning sensorimotor control policies from high-dimensional images crucially relies on the quality of the underlying visual representations. Prior works show that structured latent space such as visual keypoints often outperforms…

机器学习 · 计算机科学 2021-06-15 Boyuan Chen , Pieter Abbeel , Deepak Pathak

Active perception in vision-based robotic manipulation aims to move the camera toward more informative observation viewpoints, thereby providing high-quality perceptual inputs for downstream tasks. Most existing active perception methods…

机器人学 · 计算机科学 2026-01-21 Deyun Qin , Zezhi Liu , Hanqian Luo , Xiao Liang , Yongchun Fang

We study how visual representations pre-trained on diverse human video data can enable data-efficient learning of downstream robotic manipulation tasks. Concretely, we pre-train a visual representation using the Ego4D human video dataset…

机器人学 · 计算机科学 2022-11-21 Suraj Nair , Aravind Rajeswaran , Vikash Kumar , Chelsea Finn , Abhinav Gupta

Diffusion policies are powerful visuomotor models for robotic manipulation, yet they often fail to generalize to manipulators or end-effectors unseen during training and struggle to accommodate new task requirements at inference time.…

机器人学 · 计算机科学 2025-09-16 Xiangtong Yao , Yirui Zhou , Yuan Meng , Yanwen Liu , Liangyu Dong , Zitao Zhang , Zhenshan Bing , Kai Huang , Fuchun Sun , Alois Knoll

We present a novel Learning from Demonstration (LfD) method, Deformable Manipulation from Demonstrations (DMfD), to solve deformable manipulation tasks using states or images as inputs, given expert demonstrations. Our method uses…

机器人学 · 计算机科学 2022-07-22 Gautam Salhotra , I-Chun Arthur Liu , Marcus Dominguez-Kuhne , Gaurav S. Sukhatme

Despite great strides in language-guided manipulation, existing work has been constrained to table-top settings. Table-tops allow for perfect and consistent camera angles, properties are that do not hold in mobile manipulation. Task plans…

机器人学 · 计算机科学 2023-11-08 Priyam Parashar , Vidhi Jain , Xiaohan Zhang , Jay Vakil , Sam Powers , Yonatan Bisk , Chris Paxton

This paper tackles spatial perception and manipulation challenges in Vision-Language-Action (VLA) models. To address depth ambiguity from monocular input, we leverage a pre-trained multi-view diffusion model to synthesize latent novel views…

The concept of 3D scene graphs is increasingly recognized as a powerful semantic and hierarchical representation of the environment. Current approaches often address this at a coarse, object-level resolution. In contrast, our goal is to…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Dennis Rotondi , Fabio Scaparro , Hermann Blum , Kai O. Arras