中文
相关论文

相关论文: Training-Free Robot Pose Estimation using Off-the-…

200 篇论文

As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Hieu Bui , Nathaniel E. Chodosh , Arash Tavakoli

Autonomy in robot-assisted minimally invasive surgery has the potential to reduce surgeon cognitive and task load, thereby increasing procedural efficiency. However, implementing accurate autonomous control can be difficult due to poor…

机器人学 · 计算机科学 2026-03-18 Shuyuan Yang , Zonghe Chua

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models…

In this work, we tackle the problem of online camera-to-robot pose estimation from single-view successive frames of an image sequence, a crucial task for robots to interact with the world.

机器人学 · 计算机科学 2023-07-25 Yang Tian , Jiyao Zhang , Zekai Yin , Hao Dong

The widespread use of cameras in our society has created an overwhelming amount of video data, far exceeding the capacity for human monitoring. This presents a critical challenge for public safety and security, as the timely detection of…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Pascal Benschop , Cristian Meo , Justin Dauwels , Jelte P. Mense

Learning reward functions remains the bottleneck to equip a robot with a broad repertoire of skills. Large Language Models (LLM) contain valuable task-related knowledge that can potentially aid in the learning of reward functions. However,…

机器人学 · 计算机科学 2024-05-17 Yuwei Zeng , Yao Mu , Lin Shao

Vision-Language Models (VLMs), such as Flamingo and GPT-4V, have shown immense potential by integrating large language models with vision systems. Nevertheless, these models face challenges in the fundamental computer vision task of object…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Michael Dorkenwald , Nimrod Barazani , Cees G. M. Snoek , Yuki M. Asano

The prevalence of Vision-Language Models (VLMs) raises important questions about privacy in an era where visual information is increasingly available. While foundation VLMs demonstrate broad knowledge and learned capabilities, we…

计算机视觉与模式识别 · 计算机科学 2025-02-21 Neel Jay , Hieu Minh Nguyen , Trung Dung Hoang , Jacob Haimes

Large Language Models (LLMs) and strong vision models have enabled rapid research and development in the field of Vision-Language-Action models that enable robotic control. The main objective of these methods is to develop a generalist…

机器人学 · 计算机科学 2024-06-25 Omkar Joglekar , Tal Lancewicki , Shir Kozlovsky , Vladimir Tchuiev , Zohar Feldman , Dotan Di Castro

Recent progress in vision language foundation models has shown their ability to understand multimodal data and resolve complicated vision language tasks, including robotics manipulation. We seek a straightforward way of making use of…

机器人学 · 计算机科学 2024-02-06 Xinghang Li , Minghuan Liu , Hanbo Zhang , Cunjun Yu , Jie Xu , Hongtao Wu , Chilam Cheang , Ya Jing , Weinan Zhang , Huaping Liu , Hang Li , Tao Kong

We present a novel, reusable and task-agnostic primitive for assessing the outcome of a force-interaction robotic skill, useful e.g.\ for applications such as quality control in industrial manufacturing. The proposed method is easily…

机器人学 · 计算机科学 2018-05-14 Xiang Zhang , Athanasios S. Polydoros , Justus Piater

Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce…

机器人学 · 计算机科学 2025-09-29 Asher J. Hancock , Xindi Wu , Lihan Zha , Olga Russakovsky , Anirudha Majumdar

We present an open-source robotic framework that integrates computer vision and machine learning based inverse kinematics to enable low-cost laboratory automation tasks such as colony picking and liquid handling. The system uses a custom…

机器人学 · 计算机科学 2026-03-24 Zachary Logan , Andrew Dudash , Daniel Negrón

Video generative models (VGMs) pretrained on large-scale internet data can produce temporally coherent rollout videos that capture rich object dynamics, offering a compelling foundation for zero-shot robotic manipulation. However, VGMs…

机器人学 · 计算机科学 2026-03-09 Gehao Zhang , Zhenyang Ni , Payal Mohapatra , Han Liu , Ruohan Zhang , Qi Zhu

Robotic manipulation requires sophisticated commonsense reasoning, a capability naturally possessed by large-scale Vision-Language Models (VLMs). While VLMs show promise as zero-shot planners, their lack of grounded physical understanding…

机器人学 · 计算机科学 2026-03-18 Emily Yue-Ting Jia , Weiduo Yuan , Tianheng Shi , Vitor Guizilini , Jiageng Mao , Yue Wang

Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies…

机器人学 · 计算机科学 2025-06-23 Kaiyuan Chen , Shuangyu Xie , Zehan Ma , Pannag R Sanketi , Ken Goldberg

Using learned reward functions (LRFs) as a means to solve sparse-reward reinforcement learning (RL) tasks has yielded some steady progress in task-complexity through the years. In this work, we question whether today's LRFs are best-suited…

机器学习 · 计算机科学 2023-08-24 Ademi Adeniji , Amber Xie , Carmelo Sferrazza , Younggyo Seo , Stephen James , Pieter Abbeel

Traditional approaches to off-road autonomy rely on separate models for terrain classification, height estimation, and quantifying slip or slope conditions. Utilizing several models requires training each component separately, having task…

机器人学 · 计算机科学 2026-04-07 Abdelmoamen Nasser , Yousef Baba'a , Murad Mebrahtu , Nadya Abdel Madjid , Jorge Dias , Majid Khonji

When humans see a scene, they can roughly imagine the forces applied to objects based on their experience and use them to handle the objects properly. This paper considers transferring this "force-visualization" ability to robots. We…

机器人学 · 计算机科学 2023-04-13 Ryo Hanai , Yukiyasu Domae , Ixchel G. Ramirez-Alpizar , Bruno Leme , Tetsuya Ogata

Zero-shot object pose estimation enables the retrieval of object poses from images without necessitating object-specific training. In recent approaches this is facilitated by vision foundation models (VFM), which are pre-trained models that…

计算机视觉与模式识别 · 计算机科学 2024-11-26 Bernd Von Gimborn , Philipp Ausserlechner , Markus Vincze , Stefan Thalhammer
‹ 上一页 1 8 9 10 下一页 ›