中文
相关论文

相关论文: How do Foundation Models Compare to Skeleton-Based…

200 篇论文

Human gesture recognition has assumed a capital role in industrial applications, such as Human-Machine Interaction. We propose an approach for segmentation and classification of dynamic gestures based on a set of handcrafted features, which…

计算机视觉与模式识别 · 计算机科学 2020-08-27 André Brás , Miguel Simão , Pedro Neto

We investigate the ability of Vision Language Models (VLMs) to perform visual perspective taking using a new set of visual tasks inspired by established human tests. Our approach leverages carefully controlled scenes in which a single…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Gracjan Góral , Alicja Ziarko , Piotr Miłoś , Michał Nauman , Maciej Wołczyk , Michał Kosiński

Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole existing solution GestureGPT suffers from limited recognition…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Zhuoming Li , Aitong Liu , Mengxi Jia , Yubi Lu , Tengxiang Zhang , Changzhi Sun , Dell Zhang , Xuelong Li

In recent years robots have become an important part of our day-to-day lives with various applications. Human-robot interaction creates a positive impact in the field of robotics to interact and communicate with the robots. Gesture…

机器人学 · 计算机科学 2024-09-11 Sajjad Hussain , Khizer Saeed , Almas Baimagambetov , Shanay Rab , Md Saad

Intuitive user interfaces are indispensable to interact with the human centric smart environments. In this paper, we propose a unified framework that recognizes both static and dynamic gestures, using simple RGB vision (without depth…

计算机视觉与模式识别 · 计算机科学 2021-03-18 Osama Mazhar , Sofiane Ramdani , Andrea Cherubini

For embodied agents to effectively understand and interact within the world around them, they require a nuanced comprehension of human actions grounded in physical space. Current action recognition models, often relying on RGB video, learn…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Nicholas Babey , Tiffany Gu , Yiheng Li , Cristian Meo , Kevin Zhu

In autonomous driving, it is crucial to correctly interpret traffic gestures (TGs), such as those of an authority figure providing orders or instructions, or a pedestrian signaling the driver, to ensure a safe and pleasant traffic…

计算机视觉与模式识别 · 计算机科学 2025-04-16 Tonko E. W. Bossen , Andreas Møgelmose , Ross Greer

Human-Robot Interaction (HRI) has become increasingly important as robots are being integrated into various aspects of daily life. One key aspect of HRI is gesture recognition, which allows robots to interpret and respond to human gestures…

人机交互 · 计算机科学 2024-01-10 Sandeep Reddy Sabbella , Sara Kaszuba , Francesco Leotta , Pascal Serrarens , Daniele Nardi

Hand Gesture Recognition (HGR) enables intuitive human-computer interactions in various real-world contexts. However, existing frameworks often struggle to meet the real-time requirements essential for practical HGR applications. This study…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Oluwaleke Yusuf , Maki Habib , Mohamed Moustafa

As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Hieu Bui , Nathaniel E. Chodosh , Arash Tavakoli

Gesture recognition based on surface electromyographic signal (sEMG) is one of the most used methods. The traditional manual feature extraction can only extract some low-level signal features, this causes poor classifier performance and low…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Mingjin Zhang , Jiahao Wang , Jianming Wang , Qi Wang

We introduce MERGE, a system for situational grounding of actors, objects, and events in dynamic human-robot group interactions. Effective collaboration in such settings requires consistent situational awareness, built on persistent…

Online recognition of gestures is critical for intuitive human-robot interaction (HRI) and further push collaborative robotics into the market, making robots accessible to more people. The problem is that it is difficult to achieve accurate…

机器人学 · 计算机科学 2023-04-17 M. A. Simão , O. Gibaru , P. Neto

Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigorous biometric tasks remains unexplored. This work presents an exploratory study…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Marta Robledo-Moreno , Ruben Vera-Rodriguez , Ruben Tolosana , Javier Ortega-Garcia

While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for…

机器人学 · 计算机科学 2025-08-19 Deqing Song , Weimin Yang , Maryam Rezayati , Hans Wernher van de Venn

The expansion of instruction-tuning data has enabled foundation language models to exhibit improved instruction adherence and superior performance across diverse downstream tasks. Semantically-rich 3D human motion is being progressively…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Lei Hu , Yongjing Ye , Shihong Xia

In robot learning, it is common to either ignore the environment semantics, focusing on tasks like whole-body control which only require reasoning about robot-environment contacts, or conversely to ignore contact dynamics, focusing on…

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial…

机器人学 · 计算机科学 2026-05-22 Wenxuan Guo , Ziyuan Li , Meng Zhang , Yichen Liu , Yimeng Dong , Chuxi Xu , Yunfei Wei , Ze Chen , Erjin Zhou , Jianjiang Feng

Gesture recognition is mainly apprehensive on analyzing the functionality of human wits. The main goal of gesture recognition is to create a system which can recognize specific human gestures and use them to convey information or for device…

Current video foundation models, including the strongest self-supervised models such as V-JEPA2, fail to capture how humans organize social information in dynamic scenes. For example, across a range of diverse vision models tested, none…

神经元与认知 · 定量生物学 2026-05-14 Kathy Garcia , Leyla Isik
‹ 上一页 1 2 3 10 下一页 ›