English
Related papers

Related papers: GestLLM: Advanced Hand Gesture Interpretation via …

200 papers

The advancement of embodied intelligence is accelerating the integration of robots into daily life as human assistants. This evolution requires robots to not only interpret high-level instructions and plan tasks but also perceive and adapt…

Robotics · Computer Science 2025-08-19 Zhichen Lou , Kechun Xu , Zhongxiang Zhou , Rong Xiong

Large language models (LLMs) can provide rich physical descriptions of most worldly objects, allowing robots to achieve more informed and capable grasping. We leverage LLMs' common sense physical reasoning and code-writing abilities to…

Robotics · Computer Science 2024-09-10 William Xie , Maria Valentini , Jensen Lavering , Nikolaus Correll

Large Language Models (LLMs) pre-trained on internet-scale datasets have shown impressive capabilities in code understanding, synthesis, and general purpose question-and-answering. Key to their performance is the substantial prior knowledge…

Robotics · Computer Science 2023-11-03 Andrea Tagliabue , Kota Kondo , Tong Zhao , Mason Peterson , Claudius T. Tewari , Jonathan P. How

Enabling home-assistant robots to perceive and manipulate a diverse range of 3D objects based on human language instructions is a pivotal challenge. Prior research has predominantly focused on simplistic and task-oriented instructions,…

Robotics · Computer Science 2024-03-14 Ran Xu , Yan Shen , Xiaoqi Li , Ruihai Wu , Hao Dong

Accurate prediction of human behavior is crucial for AI systems to effectively support real-world applications, such as autonomous robots anticipating and assisting with human tasks. Real-world scenarios frequently present challenges such…

Human-Computer Interaction · Computer Science 2025-07-21 Kojiro Takeyama , Yimeng Liu , Misha Sra

We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate…

Millimeter-wave radar offers a privacy-preserving and environment-robust alternative to vision-based sensing, enabling human motion analysis in challenging conditions such as low light, occlusions, rain, or smoke. However, its sparse point…

Machine Learning · Computer Science 2025-11-18 Zengyuan Lai , Jiarui Yang , Songpengcheng Xia , Lizhou Lin , Lan Sun , Renwen Wang , Jianran Liu , Qi Wu , Ling Pei

Collaborative robots became a popular tool for increasing productivity in partly automated manufacturing plants. Intuitive robot teaching methods are required to quickly and flexibly adapt the robot programs to new tasks. Gestures have an…

Robotics · Computer Science 2024-01-04 Petr Vanc , Jan Kristof Behrens , Karla Stepanova

Augmented reality (AR) offers immersive interaction but remains inaccessible for users with motor impairments or limited dexterity due to reliance on precise input methods. This study proposes a gesture-based interaction system for AR…

Human-Computer Interaction · Computer Science 2025-06-19 Yikan Wang

Lately, there has been an increasing interest in hand gesture analysis systems. Recent works have employed pattern recognition techniques and have focused on the development of systems with more natural user interfaces. These systems may…

Human-Computer Interaction · Computer Science 2013-12-18 Renata Cristina Barros Madeo , Priscilla Koch Wagner , Sarajane Marques Peres

Large Multimodal Models (LMMs) have shown strong potential for assisting users in tasks, such as programming, content creation, and information access, yet their interaction remains largely limited to traditional interfaces such as desktops…

Human-Computer Interaction · Computer Science 2026-02-12 Liuchuan Yu , Yongqi Zhang , Lap-Fai Yu

This paper presents an in-depth survey on the use of multimodal Generative Artificial Intelligence (GenAI) and autoregressive Large Language Models (LLMs) for human motion understanding and generation, offering insights into emerging…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Muhammad Islam , Tao Huang , Euijoon Ahn , Usman Naseem

Gesture recognition is mainly apprehensive on analyzing the functionality of human wits. The main goal of gesture recognition is to create a system which can recognize specific human gestures and use them to convey information or for device…

Artificial Intelligence · Computer Science 2010-12-02 Harshith C , Karthik R. Shastry , Manoj Ravindran , M. V. V. N. S. Srikanth , Naveen Lakshmikhanth

Large vision-language models (LVLMs), such as the Generative Pre-trained Transformer 4-omni (GPT-4o), are emerging multi-modal foundation models which have great potential as powerful artificial-intelligence (AI) assistance tools for a…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Keshav Bimbraw , Ye Wang , Jing Liu , Toshiaki Koike-Akino

The Large Language Models (LLM) are increasingly being deployed in robotics to generate robot control programs for specific user tasks, enabling embodied intelligence. Existing methods primarily focus on LLM training and prompt design that…

Robotics · Computer Science 2025-08-27 ZhenDong Chen , ZhanShang Nie , ShiXing Wan , JunYi Li , YongTian Cheng , Shuai Zhao

The ability to wield tools was once considered exclusive to human intelligence, but it's now known that many other animals, like crows, possess this capability. Yet, robotic systems still fall short of matching biological dexterity. In this…

Robotics · Computer Science 2026-01-12 Hoi-Yin Lee , Peng Zhou , Anqing Duan , Wanyu Ma , Chenguang Yang , David Navarro-Alarcon

Information workers' productivity is significantly influenced by their cognitive states and physiological responses. AI assistants such as ChatGPT, Copilot, and others have become integral components of knowledge-intensive workplaces. These…

Human-Computer Interaction · Computer Science 2026-05-12 Amog Rao , Utkarsh Agarwal , Amol Harsh , Siddharth Siddharth

Large Language Models (LLMs) and strong vision models have enabled rapid research and development in the field of Vision-Language-Action models that enable robotic control. The main objective of these methods is to develop a generalist…

As robots become increasingly prevalent in both homes and industrial settings, the demand for intuitive and efficient human-machine interaction continues to rise. Gesture recognition offers an intuitive control method that does not require…

Robotics · Computer Science 2025-09-16 Yuqing Song , Cesare Tonola , Stefano Savazzi , Sanaz Kianoush , Nicola Pedrocchi , Stephan Sigg

Recent advancements in multimodal Human-Robot Interaction (HRI) datasets have highlighted the fusion of speech and gesture, expanding robots' capabilities to absorb explicit and implicit HRI insights. However, existing speech-gesture HRI…

Robotics · Computer Science 2024-03-05 Snehesh Shrestha , Yantian Zha , Saketh Banagiri , Ge Gao , Yiannis Aloimonos , Cornelia Fermuller