English
Related papers

Related papers: Motion-X++: A Large-Scale Multimodal 3D Whole-body…

200 papers

This study investigates the use of large language models (LLMs) for human behavior understanding by jointly leveraging motion and video data. We argue that integrating these complementary modalities is essential for capturing both…

Computer Vision and Pattern Recognition · Computer Science 2026-01-08 Rajan Das Gupta , Lei Wei , Md Yeasin Rahat , Nafiz Fahad , Abir Ahmed , Liew Tze Hui

Recent approaches in depth-based human activity analysis achieved outstanding performance and proved the effectiveness of 3D representation for classification of action classes. Currently available depth-based and RGB+D-based action…

Computer Vision and Pattern Recognition · Computer Science 2016-04-12 Amir Shahroudy , Jun Liu , Tian-Tsong Ng , Gang Wang

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Bharat Lal Bhatnagar , Xianghui Xie , Ilya A. Petrov , Cristian Sminchisescu , Christian Theobalt , Gerard Pons-Moll

To understand how people look, interact, or perform tasks, we need to quickly and accurately capture their 3D body, face, and hands together from an RGB image. Most existing methods focus only on parts of the body. A few recent approaches…

Computer Vision and Pattern Recognition · Computer Science 2020-08-21 Vasileios Choutas , Georgios Pavlakos , Timo Bolkart , Dimitrios Tzionas , Michael J. Black

Human pose estimation is a critical task in computer vision and sports biomechanics, with applications spanning sports science, rehabilitation, and biomechanical research. While significant progress has been made in monocular 3D pose…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Calvin Yeung , Tomohiro Suzuki , Ryota Tanaka , Zhuoer Yin , Keisuke Fujii

Sign language recognition is a challenging and often underestimated problem comprising multi-modal articulators (handshape, orientation, movement, upper body and face) that integrate asynchronously on multiple streams. Learning powerful…

Computer Vision and Pattern Recognition · Computer Science 2019-11-22 Hamid Reza Vaezi Joze , Oscar Koller

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Chuan Guo , Xinxin Zuo , Sen Wang , Li Cheng

The great success of wearables and smartphone apps for provision of extensive physical workout instructions boosts a whole industry dealing with consumer oriented sensors and sports equipment. But with these opportunities there are also new…

Computers and Society · Computer Science 2017-11-23 Andre Ebert , Marie Kiermeier , Chadly Marouane , Claudia Linnhoff-Popien

In this paper, we propose H-MoRe, a novel pipeline for learning precise human-centric motion representation. Our approach dynamically preserves relevant human motion while filtering out background movement. Notably, unlike previous methods…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Zhanbo Huang , Xiaoming Liu , Yu Kong

This paper extends the popular task of multi-object tracking to multi-object tracking and segmentation (MOTS). Towards this goal, we create dense pixel-level annotations for two existing tracking datasets using a semi-automatic annotation…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Paul Voigtlaender , Michael Krause , Aljosa Osep , Jonathon Luiten , Berin Balachandar Gnana Sekar , Andreas Geiger , Bastian Leibe

We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion…

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Understanding how humans interact with the world necessitates accurate 3D hand pose estimation, a task complicated by the hand's high degree of articulation, frequent occlusions, self-occlusions, and rapid motions. While most existing…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Enes Duran , Muhammed Kocabas , Vasileios Choutas , Zicong Fan , Michael J. Black

Hand pose estimation plays a vital role in capturing subtle nonverbal cues essential for understanding human affect. However, collecting diverse, expressive real-world data remains challenging due to labor-intensive manual annotation that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Masum Hasan , Cengiz Ozel , Nina Long , Alexander Martin , Samuel Potter , Tariq Adnan , Sangwu Lee , Ehsan Hoque

Understanding objects through multiple sensory modalities is fundamental to human perception, enabling cross-sensory integration and richer comprehension. For AI and robotic systems to replicate this ability, access to diverse, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Samuel Clarke , Suzannah Wistreich , Yanjie Ze , Jiajun Wu

We present SignAvatars, the first large-scale, multi-prompt 3D sign language (SL) motion dataset designed to bridge the communication gap for Deaf and hard-of-hearing individuals. While there has been an exponentially growing number of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Zhengdi Yu , Shaoli Huang , Yongkang Cheng , Tolga Birdal

Human Activity Recognition in RGB-D videos has been an active research topic during the last decade. However, no efforts have been found in the literature, for recognizing human activity in RGB-D videos where several performers are…

Computer Vision and Pattern Recognition · Computer Science 2018-07-10 Snehasis Mukherjee , Leburu Anvitha , T. Mohana Lahari

We present a real-time approach for multi-person 3D motion capture at over 30 fps using a single RGB camera. It operates successfully in generic scenes which may contain occlusions by objects and by other people. Our method operates in…

The ability of intelligent systems to predict human behaviors is crucial, particularly in fields such as autonomous vehicle navigation and social robotics. However, the complexity of human motion have prevented the development of a…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Yang Gao , Po-Chien Luan , Alexandre Alahi

Comprehensive capturing of human motions requires both accurate captures of complex poses and precise localization of the human within scenes. Most of the HPE datasets and methods primarily rely on RGB, LiDAR, or IMU data. However, solely…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Ming Yan , Yan Zhang , Shuqiang Cai , Shuqi Fan , Xincheng Lin , Yudi Dai , Siqi Shen , Chenglu Wen , Lan Xu , Yuexin Ma , Cheng Wang