中文
相关论文

相关论文: Actron3D: Learning Actionable Neural Functions fro…

200 篇论文

Learning policies for complex humanoid tasks remains both challenging and compelling. Inspired by how infants and athletes rely on external support--such as parental walkers or coach-applied guidance--to acquire skills like walking,…

机器人学 · 计算机科学 2025-07-01 Zhanxiang Cao , Yang Zhang , Buqing Nie , Huangxuan Lin , Haoyang Li , Yue Gao

Traditional autonomous vehicle pipelines that follow a modular approach have been very successful in the past both in academia and industry, which has led to autonomy deployed on road. Though this approach provides ease of interpretation,…

机器学习 · 计算机科学 2021-01-18 Tanmay Agarwal , Hitesh Arora , Jeff Schneider

Diffusion-based visuomotor policies effectively capture multimodal action distributions through iterative denoising, but their high inference latency limits real-time robotic control. Recent flow matching and consistency-based methods…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Chongyang Xu , Yixian Zou , Ziliang Feng , Fanman Meng , Shuaicheng Liu

A robot's ability to act is fundamentally constrained by what it can perceive. Many existing approaches to visual representation learning utilize general-purpose training criteria, e.g. image reconstruction, smoothness in latent space, or…

Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and how to grasp an object by taking its functionality into…

机器人学 · 计算机科学 2025-03-04 Yingbo Tang , Shuaike Zhang , Xiaoshuai Hao , Pengwei Wang , Jianlong Wu , Zhongyuan Wang , Shanghang Zhang

In this work, we aim to learn a unified vision-based policy for multi-fingered robot hands to manipulate a variety of objects in diverse poses. Though prior work has shown benefits of using human videos for policy learning, performance…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Zerui Chen , Shizhe Chen , Etienne Arlaud , Ivan Laptev , Cordelia Schmid

The success of deep neural networks generally requires a vast amount of training data to be labeled, which is expensive and unfeasible in scale, especially for video collections. To alleviate this problem, in this paper, we propose…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Longlong Jing , Xiaodong Yang , Jingen Liu , Yingli Tian

Understanding fine-grained object affordances is imperative for robots to manipulate objects in unstructured environments given open-ended task instructions. However, existing methods of visual affordance predictions often rely on manually…

机器人学 · 计算机科学 2025-08-27 Yihe Tang , Wenlong Huang , Yingke Wang , Chengshu Li , Roy Yuan , Ruohan Zhang , Jiajun Wu , Li Fei-Fei

Learning to understand and infer object functionalities is an important step towards robust visual intelligence. Significant research efforts have recently focused on segmenting the object parts that enable specific types of human-object…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Spyridon Thermos , Petros Daras , Gerasimos Potamianos

We present a novel paradigm of building an animatable 3D human representation from a monocular video input, such that it can be rendered in any unseen poses and views. Our method is based on a dynamic Neural Radiance Field (NeRF) rigged by…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Gusi Te , Xiu Li , Xiao Li , Jinglu Wang , Wei Hu , Yan Lu

Grasp detection in clutter requires the robot to reason about the 3D scene from incomplete and noisy perception. In this work, we draw insight that 3D reconstruction and grasp learning are two intimately connected tasks, both of which…

机器人学 · 计算机科学 2021-07-22 Zhenyu Jiang , Yifeng Zhu , Maxwell Svetlik , Kuan Fang , Yuke Zhu

Recent development in autonomous driving involves high-level computer vision and detailed road scene understanding. Today, most autonomous vehicles are using mediated perception approach for path planning and control, which highly rely on…

计算机视觉与模式识别 · 计算机科学 2019-03-22 Chen Sun , Jean M. Uwabeza Vianney , Dongpu Cao

Learning from videos offers a promising path toward generalist robots by providing rich visual and temporal priors beyond what real robot datasets contain. While existing video generative models produce impressive visual predictions, they…

We present a cross robot visuomotor learning framework that integrates diffusion policy based control with 3D semantic scene representations from D3Fields to enable category level generalization in manipulation. Its modular design supports…

Affordances are key attributes of what must be perceived by an autonomous robotic agent in order to effectively interact with novel objects. Historically, the concept derives from the literature in psychology and cognitive science, where…

机器人学 · 计算机科学 2020-04-17 Paola Ardón , Èric Pairet , Katrin S. Lohan , Subramanian Ramamoorthy , Ronald P. A. Petrick

We address a class of manipulation problems where the robot perceives the scene with a depth sensor and can move its end effector in a space with six degrees of freedom -- 3D position and orientation. Our approach is to formulate the…

机器人学 · 计算机科学 2018-09-28 Marcus Gualtieri , Robert Platt

Deep Reinforcement learning holds the guarantee of empowering self-ruling robots to master enormous collections of conduct abilities with negligible human mediation. The improvements brought by this technique enables robots to perform…

人工智能 · 计算机科学 2021-05-21 Maxence Mahe , Pierre Belamri , Jesus Bujalance Martin

Robot learning provides a number of ways to teach robots simple skills, such as grasping. However, these skills are usually trained in open, clutter-free environments, and therefore would likely cause undesirable collisions in more complex,…

机器人学 · 计算机科学 2022-12-13 Vitalis Vosylius , Edward Johns

Observing that the key for robotic action planning is to understand the target-object motion when its associated part is manipulated by the end effector, we propose to generate the 3D object-part scene flow and extract its transformations…

We introduce Knowledge Fusion Transformers for video action classification. We present a self-attention based feature enhancer to fuse action knowledge in 3D inception based spatio-temporal context of the video clip intended to be…

计算机视觉与模式识别 · 计算机科学 2020-10-01 Ganesh Samarth , Sheetal Ojha , Nikhil Pareek