中文
相关论文

相关论文: Kinova Gemini: Interactive Robot Grasping with Vis…

200 篇论文

The growing presence of service robots in human-centric environments, such as warehouses, demands seamless and intuitive human-robot collaboration. In this paper, we propose a collaborative shelf-picking framework that combines multimodal…

机器人学 · 计算机科学 2025-04-10 Abhinav Pathak , Kalaichelvi Venkatesan , Tarek Taha , Rajkumar Muthusamy

We introduce Language-Informed Latent Actions (LILA), a framework for learning natural language interfaces in the context of human-robot collaboration. LILA falls under the shared autonomy paradigm: in addition to providing discrete…

机器人学 · 计算机科学 2021-11-08 Siddharth Karamcheti , Megha Srivastava , Percy Liang , Dorsa Sadigh

Human-robot collaboration requires robots to quickly infer user intent, provide transparent reasoning, and assist users in achieving their goals. Our recent work introduced GUIDER, our framework for inferring navigation and manipulation…

机器人学 · 计算机科学 2025-08-18 Cesar Alan Contreras , Manolis Chiou , Alireza Rastegarpanah , Michal Szulik , Rustam Stolkin

A polynomial solution to the inverse kinematic problem of the Kinova Gen3 Lite robot is proposed in this paper. This serial robot is based on a 6R kinematic chain and is not wrist-partitioned. We first start from the forward kinematics…

机器人学 · 计算机科学 2021-02-03 Hamed Montazer Zohour , Bruno Belzile , David St-Onge

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes.…

This paper introduces CognitiveDog, a pioneering development of quadruped robot with Large Multi-modal Model (LMM) that is capable of not only communicating with humans verbally but also physically interacting with the environment through…

Gemini is a natural language understanding system developed for spoken language applications. The paper describes the architecture of Gemini, paying particular attention to resolving the tension between robustness and overgeneration. Gemini…

cmp-lg · 计算机科学 2008-02-03 John Dowding , Jean Mark Gawron , Doug Appelt , John Bear , Lynn Cherny , Robert Moore , Douglas Moran

We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a new set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes in which a single…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Gracjan Góral , Alicja Ziarko , Piotr Miłoś , Michał Nauman , Maciej Wołczyk , Michał Kosiński

Robotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of…

机器人学 · 计算机科学 2026-03-17 Shaojun Cai , Nuwan Janaka , Ashwin Ram , Janidu Shehan , Yingjia Wan , Kotaro Hara , David Hsu

We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the…

机器人学 · 计算机科学 2025-06-19 Haoyu Xiong , Xiaomeng Xu , Jimmy Wu , Yifan Hou , Jeannette Bohg , Shuran Song

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always on" mode,…

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in…

In this paper, we identify challenges in children's current information retrieval process, and propose conversational robots as an opportunity to ease this process in a responsible way. Tools children currently use in this process, such as…

信息检索 · 计算机科学 2021-06-16 T. Beelen , E. Velner , R. Ordelman , K. P. Truong , V. Evers , T. Huibers

Gestures are a natural form of communication between humans and can also be leveraged for human-robot interaction. This work presents a gesture-based user interface for object selection using pointing and click gestures. An experiment with…

机器人学 · 计算机科学 2026-04-08 Bijan Kavousian , Oliver Petrovic , Werner Herfs

The goal of the system presented in this paper is to develop a natural talking gesture generation behavior for a humanoid robot, by feeding a Generative Adversarial Network (GAN) with human talking gestures recorded by a Kinect. A direct…

机器人学 · 计算机科学 2019-09-05 Unai Zabala , Igor Rodriguez , José María Martínez-Otzeta , Elena Lazkano

Arguably, the visual perception of conversational agents to the physical world is a key way for them to exhibit the human-like intelligence. Image-grounded conversation is thus proposed to address this challenge. Existing works focus on…

计算与语言 · 计算机科学 2021-06-24 Zujie Liang , Huang Hu , Can Xu , Chongyang Tao , Xiubo Geng , Yining Chen , Fan Liang , Daxin Jiang

The recently released Google Gemini class of models are the first to comprehensively report results that rival the OpenAI GPT series across a wide variety of tasks. In this paper, we do an in-depth exploration of Gemini's language…

We envision robots that can collaborate and communicate seamlessly with humans. It is necessary for such robots to decide both what to say and how to act, while interacting with humans. To this end, we introduce a new task, dialogue object…

机器人学 · 计算机科学 2021-07-23 Monica Roy , Kaiyu Zheng , Jason Liu , Stefanie Tellex

Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions…

人机交互 · 计算机科学 2025-09-03 Xuetong Wang , Ching Christie Pang , Pan Hui

Autonomous robot navigation and manipulation in open environments require reasoning and replanning with closed-loop feedback. In this work, we present COME-robot, the first closed-loop robotic system utilizing the GPT-4V vision-language…

机器人学 · 计算机科学 2025-03-10 Peiyuan Zhi , Zhiyuan Zhang , Yu Zhao , Muzhi Han , Zeyu Zhang , Zhitian Li , Ziyuan Jiao , Baoxiong Jia , Siyuan Huang