English
Related papers

Related papers: MindEye-OmniAssist: A Gaze-Driven LLM-Enhanced Ass…

200 papers

With the rise in popularity of smart glasses, users' attention has been integrated into Vision-Language Models (VLMs) to streamline multi-modal querying in daily scenarios. However, leveraging gaze data to model users' attention may…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Zeyu Wang , Baiyu Chen , Kun Yan , Hongjing Piao , Hao Xue , Flora D. Salim , Yuanchun Shi , Yuntao Wang

Human teams exhibit both implicit and explicit intention sharing. To further development of human-robot collaboration, intention recognition is crucial on both sides. Present approaches rely on a vast sensor suite on and around the robot to…

Robotics · Computer Science 2021-02-23 David Puljiz , Bowen Zhou , Ke Ma , Björn Hein

Within this work, we explore intention inference for user actions in the context of a handheld robot setup. Handheld robots share the shape and properties of handheld tools while being able to process task information and aid manipulation.…

Robotics · Computer Science 2018-10-16 Janis Stolzenwald , Walterio W. Mayol-Cuevas

People are proficient at communicating their intentions in order to avoid conflicts when navigating in narrow, crowded environments. In many situations mobile robots lack both the ability to interpret human intentions and the ability to…

Robotics · Computer Science 2019-11-07 Justin Hart , Reuth Mirsky , Stone Tejeda , Bonny Mahajan , Jamin Goo , Kathryn Baldauf , Sydney Owen , Peter Stone

Constraint-aware estimation of human intent is essential for robots to physically collaborate and interact with humans. Further, to achieve fluid collaboration in dynamic tasks intent estimation should be achieved in real-time. In this…

Robotics · Computer Science 2024-09-04 Yifei Simon Shao , Tianyu Li , Shafagh Keyvanian , Pratik Chaudhari , Vijay Kumar , Nadia Figueroa

Training a Multimodal Large Language Model (MLLM) from scratch, like GPT-4, is resource-intensive. Regarding Large Language Models (LLMs) as the core processor for multimodal information, our paper introduces LMEye, a human-like eye with a…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Yunxin Li , Baotian Hu , Xinyu Chen , Lin Ma , Yong Xu , Min Zhang

Enabling robots to understand human gaze target is a crucial step to allow capabilities in downstream tasks, for example, attention estimation and movement anticipation in real-world human-robot interactions. Prior works have addressed the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Zhuangzhuang Dai , Vincent Gbouna Zakka , Luis J. Manso , Chen Li

Large Vision-Language Models (VLMs) often exhibit text inertia, where attention drifts from visual evidence toward linguistic priors, resulting in object hallucinations. Existing decoding strategies intervene only at the output logits and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Weijue Bu , Guan Yuan , Guixian Zhang

Multimodal large language models (MLLMs) have shown remarkable capabilities in cross-modal understanding and reasoning, offering new opportunities for intelligent assistive systems, yet existing systems still struggle with risk-aware…

Robotics · Computer Science 2026-04-08 Renjun Gao

This paper introduces a prototype for a new approach to assistive robotics, integrating edge computing with Natural Language Processing (NLP) and computer vision to enhance the interaction between humans and robotic systems. Our proof of…

Robotics · Computer Science 2024-10-11 Pascal Sikorski , Kaleb Yu , Lucy Billadeau , Flavio Esposito , Hadi AliAkbarpour , Madi Babaiasl

Physically assistive robots present an opportunity to significantly increase the well-being and independence of individuals with motor impairments or other forms of disability who are unable to complete activities of daily living (ADLs).…

It is crucial that robots' performance can be improved after deployment, as they are inherently likely to encounter novel scenarios never seen before. This paper presents an innovative solution: an interactive learning-based robot system…

Human-Computer Interaction · Computer Science 2025-08-01 Kohou Wang , ZhaoXiang Liu , Lin Bai , Kun Fan , Xiang Liu , Huan Hu , Kai Wang , Shiguo Lian

One of the current trends in robotics is to employ large language models (LLMs) to provide non-predefined command execution and natural human-robot interaction. It is useful to have an environment map together with its language…

Robotics · Computer Science 2025-01-09 Evgenii Kruzhkov , Sven Behnke

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Hengfei Wang , Anshul Gupta , Pierre Vuillecard , Jean-Marc Odobez

When humans work together to complete a joint task, each person builds an internal model of the situation and how it will evolve. Efficient collaboration is dependent on how these individual models overlap to form a shared mental model…

Robotics · Computer Science 2022-08-26 Wesley P. Chan , Morgan Crouch , Khoa Hoang , Charlie Chen , Nicole Robinson , Elizabeth Croft

Task-oriented communications are an important element in future intelligent IoT systems. Existing IoT systems, however, are limited in their capacity to handle complex tasks, particularly in their interactions with humans to accomplish…

Information Theory · Computer Science 2024-08-12 Hongwei Cui , Yuyang Du , Qun Yang , Yulin Shao , Soung Chang Liew

This work presents a next-generation human-robot interface that can infer and realize the user's manipulation intention via sight only. Specifically, we develop a system that integrates near-eye-tracking and robotic manipulation to enable…

Robotics · Computer Science 2023-05-16 Shaochen Wang , Wei Zhang , Zhangli Zhou , Jiaxi Cao , Ziyang Chen , Kang Chen , Bin Li , Zhen Kan

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- inputs are…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Pengzhan Sun , Junbin Xiao , Tze Ho Elden Tse , Yicong Li , Arjun Akula , Angela Yao

The recent development of Agentic AI systems, empowered by autonomous large language models (LLMs) agents with planning and tool-usage capabilities, enables new possibilities for the evolution of industrial automation and reduces the…

Machine Learning · Computer Science 2026-01-07 Marcos Lima Romero , Ricardo Suyama

In high-conflict mixed-traffic scenarios involving human-driven and autonomous vehicles, most existing autonomous driving systems default to overly conservative behaviors, lack proactive interaction, and consequently suffer from limited…

Robotics · Computer Science 2026-04-28 Xinwei Dong , Jiyang Li , Jiabin Xie , Yang Yi , Tianshang Jia , Shiyu Fang , Ye Tian , Peng Hang
‹ Prev 1 3 4 5 6 7 10 Next ›