English
Related papers

Related papers: Kinova Gemini: Interactive Robot Grasping with Vis…

200 papers

The growing presence of service robots in human-centric environments, such as warehouses, demands seamless and intuitive human-robot collaboration. In this paper, we propose a collaborative shelf-picking framework that combines multimodal…

Robotics · Computer Science 2025-04-10 Abhinav Pathak , Kalaichelvi Venkatesan , Tarek Taha , Rajkumar Muthusamy

We introduce Language-Informed Latent Actions (LILA), a framework for learning natural language interfaces in the context of human-robot collaboration. LILA falls under the shared autonomy paradigm: in addition to providing discrete…

Robotics · Computer Science 2021-11-08 Siddharth Karamcheti , Megha Srivastava , Percy Liang , Dorsa Sadigh

Human-robot collaboration requires robots to quickly infer user intent, provide transparent reasoning, and assist users in achieving their goals. Our recent work introduced GUIDER, our framework for inferring navigation and manipulation…

Robotics · Computer Science 2025-08-18 Cesar Alan Contreras , Manolis Chiou , Alireza Rastegarpanah , Michal Szulik , Rustam Stolkin

A polynomial solution to the inverse kinematic problem of the Kinova Gen3 Lite robot is proposed in this paper. This serial robot is based on a 6R kinematic chain and is not wrist-partitioned. We first start from the forward kinematics…

Robotics · Computer Science 2021-02-03 Hamed Montazer Zohour , Bruno Belzile , David St-Onge

This paper proposes to solve the problem of Vision-and-Language Navigation with legged robots, which not only provides a flexible way for humans to command but also allows the robot to navigate through more challenging and cluttered scenes.…

This paper introduces CognitiveDog, a pioneering development of quadruped robot with Large Multi-modal Model (LMM) that is capable of not only communicating with humans verbally but also physically interacting with the environment through…

Gemini is a natural language understanding system developed for spoken language applications. The paper describes the architecture of Gemini, paying particular attention to resolving the tension between robustness and overgeneration. Gemini…

cmp-lg · Computer Science 2008-02-03 John Dowding , Jean Mark Gawron , Doug Appelt , John Bear , Lynn Cherny , Robert Moore , Douglas Moran

We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a new set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes in which a single…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Gracjan Góral , Alicja Ziarko , Piotr Miłoś , Michał Nauman , Maciej Wołczyk , Michał Kosiński

Robotic guidance systems have shown promise in supporting blind and visually impaired (BVI) individuals with wayfinding and obstacle avoidance. However, most existing systems assume a clear path and do not support a critical aspect of…

Robotics · Computer Science 2026-03-17 Shaojun Cai , Nuwan Janaka , Ashwin Ram , Janidu Shehan , Yingjia Wan , Kotaro Hara , David Hsu

We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the…

Robotics · Computer Science 2025-06-19 Haoyu Xiong , Xiaomeng Xu , Jimmy Wu , Yifan Hou , Jeannette Bohg , Shuran Song

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always on" mode,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-31 Yifei Huang , Jilan Xu , Baoqi Pei , Yuping He , Guo Chen , Lijin Yang , Xinyuan Chen , Yaohui Wang , Zheng Nie , Jinyao Liu , Guoshun Fan , Dechen Lin , Fang Fang , Kunpeng Li , Chang Yuan , Yali Wang , Yu Qiao , Limin Wang

The surge of interest towards Multi-modal Large Language Models (MLLMs), e.g., GPT-4V(ision) from OpenAI, has marked a significant trend in both academia and industry. They endow Large Language Models (LLMs) with powerful capabilities in…

In this paper, we identify challenges in children's current information retrieval process, and propose conversational robots as an opportunity to ease this process in a responsible way. Tools children currently use in this process, such as…

Information Retrieval · Computer Science 2021-06-16 T. Beelen , E. Velner , R. Ordelman , K. P. Truong , V. Evers , T. Huibers

Gestures are a natural form of communication between humans and can also be leveraged for human-robot interaction. This work presents a gesture-based user interface for object selection using pointing and click gestures. An experiment with…

Robotics · Computer Science 2026-04-08 Bijan Kavousian , Oliver Petrovic , Werner Herfs

The goal of the system presented in this paper is to develop a natural talking gesture generation behavior for a humanoid robot, by feeding a Generative Adversarial Network (GAN) with human talking gestures recorded by a Kinect. A direct…

Robotics · Computer Science 2019-09-05 Unai Zabala , Igor Rodriguez , José María Martínez-Otzeta , Elena Lazkano

Arguably, the visual perception of conversational agents to the physical world is a key way for them to exhibit the human-like intelligence. Image-grounded conversation is thus proposed to address this challenge. Existing works focus on…

Computation and Language · Computer Science 2021-06-24 Zujie Liang , Huang Hu , Can Xu , Chongyang Tao , Xiubo Geng , Yining Chen , Fan Liang , Daxin Jiang

The recently released Google Gemini class of models are the first to comprehensively report results that rival the OpenAI GPT series across a wide variety of tasks. In this paper, we do an in-depth exploration of Gemini's language…

We envision robots that can collaborate and communicate seamlessly with humans. It is necessary for such robots to decide both what to say and how to act, while interacting with humans. To this end, we introduce a new task, dialogue object…

Robotics · Computer Science 2021-07-23 Monica Roy , Kaiyu Zheng , Jason Liu , Stefanie Tellex

Virtual assistants (VAs) have become ubiquitous in daily life, integrated into smartphones and smart devices, sparking interest in AI companions that enhance user experiences and foster emotional connections. However, existing companions…

Human-Computer Interaction · Computer Science 2025-09-03 Xuetong Wang , Ching Christie Pang , Pan Hui

Autonomous robot navigation and manipulation in open environments require reasoning and replanning with closed-loop feedback. In this work, we present COME-robot, the first closed-loop robotic system utilizing the GPT-4V vision-language…

Robotics · Computer Science 2025-03-10 Peiyuan Zhi , Zhiyuan Zhang , Yu Zhao , Muzhi Han , Zeyu Zhang , Zhitian Li , Ziyuan Jiao , Baoxiong Jia , Siyuan Huang