English
Related papers

Related papers: PALUM: Part-based Attention Learning for Unified M…

200 papers

Recognition of human actions and associated interactions with objects and the environment is an important problem in computer vision due to its potential applications in a variety of domains. The most versatile methods can generalize to…

Computer Vision and Pattern Recognition · Computer Science 2019-12-10 Behnoosh Parsa , Athma Narayanan , Behzad Dariush

Existing multimodal-based human action recognition approaches are computationally intensive, limiting their deployment in real-time applications. In this work, we present a novel and efficient pose-driven attention-guided multimodal network…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Ahmed Abdelkawy , Asem Ali , Aly Farag

Human motion prediction aims to generate future motions based on the observed human motions. Witnessing the success of Recurrent Neural Networks (RNN) in modeling the sequential data, recent works utilize RNN to model human-skeleton motion…

Computer Vision and Pattern Recognition · Computer Science 2019-10-02 Xiangbo Shu , Liyan Zhang , Guo-Jun Qi , Wei Liu , Jinhui Tang

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Kunhang Li , Jason Naradowsky , Yansong Feng , Yusuke Miyao

Real-world objects are composed of distinctive, object-specific parts. Identifying these parts is key to performing fine-grained, compositional reasoning-yet, large multimodal models (LMMs) struggle to perform this seemingly straightforward…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Ansel Blume , Jeonghwan Kim , Hyeonjeong Ha , Elen Chatikyan , Xiaomeng Jin , Khanh Duy Nguyen , Nanyun Peng , Kai-Wei Chang , Derek Hoiem , Heng Ji

In recent years, remarkable results have been achieved in self-supervised action recognition using skeleton sequences with contrastive learning. It has been observed that the semantic distinction of human action features is often…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Yilei Hua , Wenhan Wu , Ce Zheng , Aidong Lu , Mengyuan Liu , Chen Chen , Shiqian Wu

This paper presents a hybrid real-time camera pose estimation framework with a novel partitioning scheme and introduces motion averaging to monocular Simultaneous Localization and Mapping (SLAM) systems. Breaking through the limitations of…

Computer Vision and Pattern Recognition · Computer Science 2020-11-04 Xinyi Li , Haibin Ling

Skeleton-based action segmentation requires recognizing composable actions in untrimmed videos. Current approaches decouple this problem by first extracting local visual features from skeleton sequences and then processing them by a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-22 Di Yang , Yaohui Wang , Antitza Dantcheva , Quan Kong , Lorenzo Garattoni , Gianpiero Francesca , Francois Bremond

Existing automatic approaches for 3D virtual character motion synthesis supporting scene interactions do not generalise well to new objects outside training distributions, even when trained on extensive motion capture datasets with diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Wanyue Zhang , Rishabh Dabral , Thomas Leimkühler , Vladislav Golyanik , Marc Habermann , Christian Theobalt

In this work, we propose TextIM, a novel framework for synthesizing TEXT-driven human Interactive Motions, with a focus on the precise alignment of part-level semantics. Existing methods often overlook the critical roles of interactive body…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Siyuan Fan , Bo Du , Xiantao Cai , Bo Peng , Longling Sun

We present an attention-based modular neural framework for computer vision. The framework uses a soft attention mechanism allowing models to be trained with gradient descent. It consists of three modules: a recurrent attention module…

Machine Learning · Computer Science 2016-04-29 Samira Ebrahimi Kahou , Vincent Michalski , Roland Memisevic

Recovering world-coordinate human motion from monocular videos with humanoid robot retargeting is significant for embodied intelligence and robotics. To avoid complex SLAM pipelines or heavy temporal models, we propose a lightweight,…

Robotics · Computer Science 2025-12-29 Zhangzheng Tu , Kailun Su , Shaolong Zhu , Yukun Zheng

Skeleton-based action recognition aims to project skeleton sequences to action categories, where skeleton sequences are derived from multiple forms of pre-detected points. Compared with earlier methods that focus on exploring single-form…

Computer Vision and Pattern Recognition · Computer Science 2022-07-01 Xuanhan Wang , Yan Dai , Lianli Gao , Jingkuan Song

Although large vision-language models (LVLMs) have demonstrated remarkable capabilities, they are prone to hallucinations in multi-image tasks. We attribute this issue to limitations in existing attention mechanisms and insufficient…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Xiaochen Yang , Hao Fang , Jiawei Kong , Yaoxin Mao , Bin Chen , Shu-Tao Xia

We propose an end-to-end trainable, cross-category method for reconstructing multiple man-made articulated objects from a single RGBD image, focusing on part-level shape reconstruction and pose and kinematics estimation. We depart from…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuki Kawana , Tatsuya Harada

Skeleton-based Temporal Action Segmentation involves the dense action classification of variable-length skeleton sequences. Current approaches primarily apply graph-based networks to extract framewise, whole-body-level motion…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Bowen Chen , Haoyu Ji , Zhiyong Wang , Benjamin Filtjens , Chunzhuo Wang , Weihong Ren , Bart Vanrumste , Honghai Liu

Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Bizhu Wu , Jinheng Xie , Keming Shen , Zhe Kong , Jianfeng Ren , Ruibin Bai , Rong Qu , Linlin Shen

Text-motion retrieval aims to learn a semantically aligned latent space between natural language descriptions and 3D human motion skeleton sequences, enabling bidirectional search across the two modalities. Most existing methods use a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Yao Zhang , Zhuchenyang Liu , Yanlan He , Thomas Ploetz , Yu Xiao

Leveraging wearable devices for motion reconstruction has emerged as an economical and viable technique. Certain methodologies employ sparse Inertial Measurement Units (IMUs) on the human body and harness data-driven strategies to model…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Xueyuan Yang , Chao Yao , Xiaojuan Ban

Recently, tactile sensing has attracted great interest in robotics, especially for facilitating exploration of unstructured environments and effective manipulation. A detailed understanding of the surface textures via tactile sensing is…

Robotics · Computer Science 2020-08-12 Guanqun Cao , Yi Zhou , Danushka Bollegala , Shan Luo