English
Related papers

Related papers: PoseGPT: Quantization-based 3D Human Motion Genera…

200 papers

Human movement is goal-directed and influenced by the spatial layout of the objects in the scene. To plan future human motion, it is crucial to perceive the environment -- imagine how hard it is to navigate a new room with lights off.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-03 Zhe Cao , Hang Gao , Karttikeya Mangalam , Qi-Zhi Cai , Minh Vo , Jitendra Malik

Motion generation from discrete quantization offers many advantages over continuous regression, but at the cost of inevitable approximation errors. Previous methods usually quantize the entire body pose into one code, which not only faces…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Weihao Yuan , Weichao Shen , Yisheng He , Yuan Dong , Xiaodong Gu , Zilong Dong , Liefeng Bo , Qixing Huang

Motion prediction, recently popularized as world models, refers to the anticipation of future agent states or scene evolution, which is rooted in human cognition, bridging perception and decision-making. It enables intelligent systems, such…

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Yin Wang , Zhiying Leng , Frederick W. B. Li , Shun-Cheng Wu , Xiaohui Liang

Human pose forecasting is inherently multimodal since multiple futures exist for an observed pose sequence. However, evaluating multimodality is challenging since the task is ill-posed. Therefore, we first propose an alternative paradigm to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Reyhaneh Hosseininejad , Megh Shukla , Saeed Saadatnejad , Mathieu Salzmann , Alexandre Alahi

Accurate human motion prediction with well-calibrated uncertainty is critical for safe human-robot collaboration (HRC), where robots must anticipate and react to human movements in real time. We propose a structured multitask variational…

Robotics · Computer Science 2026-03-10 Jinger Chong , Xiaotong Zhang , Kamal Youcef-Toumi

A photorealistic and immersive human avatar experience demands capturing fine, person-specific details such as cloth and hair dynamics, subtle facial expressions, and characteristic motion patterns. Achieving this requires large,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Michael Steiner , Zhang Chen , Alexander Richard , Vasu Agrawal , Markus Steinberger , Michael Zollhöfer

We propose the use of a proportional-derivative (PD) control based policy learned via reinforcement learning (RL) to estimate and forecast 3D human pose from egocentric videos. The method learns directly from unsegmented egocentric videos…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Ye Yuan , Kris Kitani

The attention mechanism provides a sequential prediction framework for learning spatial models with enhanced implicit temporal consistency. In this work, we show a systematic design (from 2D to 3D) for how conventional networks and other…

Computer Vision and Pattern Recognition · Computer Science 2021-03-05 Ruixu Liu , Ju Shen , He Wang , Chen Chen , Sen-ching Cheung , Vijayan K. Asari

We present SketchGPT, a flexible framework that employs a sequence-to-sequence autoregressive model for sketch generation, and completion, and an interpretation case study for sketch recognition. By mapping complex sketches into simplified…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Adarsh Tiwari , Sanket Biswas , Josep Lladós

Recent feed-forward networks have achieved remarkable progress in sparse-view 3D reconstruction by predicting dense point maps directly from RGB images. However, they often suffer from geometric inconsistencies and limited fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Yutong Chen , Yiming Wang , Xucong Zhang , Sergey Prokudin , Siyu Tang

In this paper, we develop a neural network model to predict future human motion from an observed human motion history. We propose a non-autoregressive transformer architecture to leverage its parallel nature for easier training and fast,…

Robotics · Computer Science 2025-01-20 Mohammad Mahdavian , Payam Nikdel , Mahdi TaherAhmadi , Mo Chen

Estimation of 3D human pose from monocular image has gained considerable attention, as a key step to several human-centric applications. However, generalizability of human pose estimation models developed using supervision on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Jogendra Nath Kundu , Siddharth Seth , Rahul M , Mugalodi Rakesh , R. Venkatesh Babu , Anirban Chakraborty

Short-term human pose prediction plays a crucial role in interactive systems, assistive robots, and emotion-aware human-computer interaction[1-3]. While current trajectory prediction models primarily rely on geometric motion cues, they…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Jingni Huang , Peter Bloodsworth

Human pose estimation aims to accurately estimate a wide variety of human poses. However, existing datasets often follow a long-tailed distribution that unusual poses only occupy a small portion, which further leads to the lack of diversity…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Wentao Jiang , Sheng Jin , Wentao Liu , Chen Qian , Ping Luo , Si Liu

Sequential recommender systems aim to predict a user's future interests by extracting temporal patterns from their behavioral history. Existing approaches typically employ transformer-based architectures to process long sequences of user…

Information Retrieval · Computer Science 2026-02-24 Adamya Shyam , Venkateswara Rao Kagita , Bharti Rana , Vikas Kumar

Accepted in the ICIP 2025 We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Edward Effendy , Kuan-Wei Tseng , Rei Kawakami

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2023-12-04 Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Sen Wang , Li Cheng

We present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos. VideoGPT uses VQ-VAE that learns downsampled discrete latent representations of a raw video by employing 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Wilson Yan , Yunzhi Zhang , Pieter Abbeel , Aravind Srinivas

3D human pose estimation from 2D images is a challenging problem due to depth ambiguity and occlusion. Because of these challenges the task is underdetermined, where there exists multiple -- possibly infinite -- poses that are plausible…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Francis Snelgar , Ming Xu , Stephen Gould , Liang Zheng , Akshay Asthana