中文
相关论文

相关论文: Grounding Foundational Vision Models with 3D Human…

200 篇论文

Human action recognition in video is an active yet challenging research topic due to high variation and complexity of data. In this paper, a novel video based action recognition framework utilizing complementary cues is proposed to handle…

计算机视觉与模式识别 · 计算机科学 2019-09-10 Muhammad Usman Khalid , Jie Yu

Human Action Recognition is an important task of Human Robot Interaction as cooperation between robots and humans requires that artificial agents recognise complex cues from the environment. A promising approach is using trained classifiers…

计算机视觉与模式识别 · 计算机科学 2019-08-26 Frederico Belmonte Klein , Angelo Cangelosi

Action recognition is a critical task for social robots to meaningfully engage with their environment. 3D human skeleton-based action recognition is an attractive research area in recent years. Although, the existing approaches are good at…

计算机视觉与模式识别 · 计算机科学 2021-01-01 Hui Feng , Shanshan Wang , Shuzhi Sam Ge

Human motion understanding has advanced rapidly through vision-based progress in recognition, tracking, and captioning. However, most existing methods overlook physical cues such as joint actuation forces that are fundamental in…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Anh Dao , Manh Tran , Yufei Zhang , Xiaoming Liu , Zijun Cui

Vision-Language-Action (VLA) models show promise for robotic control, yet performance in complex household environments remains sub-optimal. Mobile manipulation requires reasoning about global scene layout, fine-grained geometry, and…

机器人学 · 计算机科学 2026-03-25 Ruisen Tu , Arth Shukla , Sohyun Yoo , Xuanlin Li , Junxi Li , Jianwen Xie , Hao Su , Zhuowen Tu

Action recognition and anticipation are key to the success of many computer vision applications. Existing methods can roughly be grouped into those that extract global, context-aware representations of the entire image or sequence, and…

计算机视觉与模式识别 · 计算机科学 2016-11-21 Mohammad Sadegh Aliakbarian , Fatemehsadat Saleh , Basura Fernando , Mathieu Salzmann , Lars Petersson , Lars Andersson

Recognizing and categorizing human actions is an important task with applications in various fields such as human-robot interaction, video analysis, surveillance, video retrieval, health care system and entertainment industry. This thesis…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Zahra Gharaee

Active perception, the ability of a robot to proactively adjust its viewpoint to acquire task-relevant information, is essential for robust operation in unstructured real-world environments. While critical for downstream tasks such as…

机器人学 · 计算机科学 2026-03-03 Yongxi Huang , Zhuohang Wang , Wenjing Tang , Cewu Lu , Panpan Cai

Human-robot collaboration requires the establishment of methods to guarantee the safety of participating operators. A necessary part of this process is ensuring reliable human pose estimation. Established vision-based modalities encounter…

机器人学 · 计算机科学 2024-06-28 Michael Zechmair , Yannick Morel

Learning predictive world models from unlabelled video is a foundational challenge in artificial intelligence. While Joint Embedding Predictive Architectures (JEPA) have set new benchmarks in semantic classification, they often remain…

计算机视觉与模式识别 · 计算机科学 2026-05-18 Santosh Kumar Paidi

We propose embodied scene-aware human pose estimation where we estimate 3D poses based on a simulated agent's proprioception and scene awareness, along with external third-person observations. Unlike prior methods that often resort to…

计算机视觉与模式识别 · 计算机科学 2022-10-17 Zhengyi Luo , Shun Iwase , Ye Yuan , Kris Kitani

Multimodal-based action recognition methods have achieved high success using pose and RGB modality. However, skeletons sequences lack appearance depiction and RGB images suffer irrelevant noise due to modality limitations. To address this,…

计算机视觉与模式识别 · 计算机科学 2024-01-05 Jinfu Liu , Runwei Ding , Yuhang Wen , Nan Dai , Fanyang Meng , Shen Zhao , Mengyuan Liu

Human skeleton, as a compact representation of human action, has received increasing attention in recent years. Many skeleton-based action recognition methods adopt graph convolutional networks (GCN) to extract features on top of human…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Haodong Duan , Yue Zhao , Kai Chen , Dahua Lin , Bo Dai

Reliable perception is essential for robots that interact with the world. But sensors alone are often insufficient to provide this capability, and they are prone to errors due to various conditions in the environment. Furthermore, there is…

机器人学 · 计算机科学 2021-07-08 Ying Siu Liang , Dongkyu Choi , Kenneth Kwok

Pose detection is one of the fundamental steps for the recognition of human actions. In this paper we propose a novel trainable detector for recognizing human poses based on the analysis of the skeleton. The main idea is that a skeleton…

计算机视觉与模式识别 · 计算机科学 2017-08-03 Alessia Saggese , Nicola Strisciuglio , Mario Vento , Nicolai Petkov

Understanding human actions in visual data is tied to advances in complementary research areas including object recognition, human dynamics, domain adaptation and semantic segmentation. Over the last decade, human action analysis evolved…

计算机视觉与模式识别 · 计算机科学 2017-02-02 Samitha Herath , Mehrtash Harandi , Fatih Porikli

Group activity detection in multi-person scenes is challenging due to complex human interactions, occlusions, and variations in appearance over time. This work presents a computer vision based framework for group activity recognition and…

People infer rich social information from others' actions. These inferences are often constrained by the physical world: what agents can do, what obstacles permit, and how the physical actions of agents causally change an environment and…

神经元与认知 · 定量生物学 2026-03-31 Lance Ying , Aydan Y. Huang , Aviv Netanyahu , Andrei Barbu , Boris Katz , Joshua B. Tenenbaum , Tianmin Shu

Recognizing human actions is fundamentally a spatio-temporal reasoning problem, and should be, at least to some extent, invariant to the appearance of the human and the objects involved. Motivated by this hypothesis, in this work, we take…

计算机视觉与模式识别 · 计算机科学 2021-11-04 Gorjan Radevski , Marie-Francine Moens , Tinne Tuytelaars

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

机器人学 · 计算机科学 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi