English
Related papers

Related papers: PoseGPT: Quantization-based 3D Human Motion Genera…

200 papers

We build a Generative Pre-trained Transformer (GPT) model from scratch to solve sequential decision making tasks arising in contexts of operations research and management science which we call OMGPT. We first propose a general sequence…

Machine Learning · Computer Science 2025-05-21 Hanzhao Wang , Guanting Chen , Kalyan Talluri , Xiaocheng Li

We consider the problem of vision-based pose estimation for autonomous systems. While deep neural networks have been successfully used for vision-based tasks, they inherently lack provable guarantees on the correctness of their output,…

Robotics · Computer Science 2026-01-27 Ulices Santa Cruz , Mahmoud Elfar , Yasser Shoukry

Natural and lifelike locomotion remains a fundamental challenge for humanoid robots to interact with human society. However, previous methods either neglect motion naturalness or rely on unstable and ambiguous style rewards. In this paper,…

Robotics · Computer Science 2025-03-13 Haodong Zhang , Liang Zhang , Zhenghan Chen , Lu Chen , Yue Wang , Rong Xiong

Generative adversarial networks achieve great performance in photorealistic image synthesis in various domains, including human images. However, they usually employ latent vectors that encode the sampled outputs globally. This does not…

Computer Vision and Pattern Recognition · Computer Science 2021-03-15 Kripasindhu Sarkar , Lingjie Liu , Vladislav Golyanik , Christian Theobalt

Generating holistic co-speech gestures that integrate full-body motion with facial expressions suffers from semantically incoherent coordination on body motion and spatially unstable meaningless movements due to existing part-decomposed or…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Xuanmeng Sha , Liyun Zhang , Tomohiro Mashita , Naoya Chiba , Yuki Uranishi

Human motion prediction is consisting in forecasting future body poses from historically observed sequences. It is a longstanding challenge due to motion's complex dynamics and uncertainty. Existing methods focus on building up complicated…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Zhihao Wang , Yulin Zhou , Ningyu Zhang , Xiaosong Yang , Jun Xiao , Zhao Wang

Human motion prediction aims to forecast future human poses given a historical motion. Whether based on recurrent or feed-forward neural networks, existing learning based methods fail to model the observation that human motion tends to…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Wei Mao , Miaomiao Liu , Mathieu Salzmann , Hongdong Li

We tackle the task of diverse 3D human motion prediction, that is, forecasting multiple plausible future 3D poses given a sequence of observed 3D poses. In this context, a popular approach consists of using a Conditional Variational…

Machine Learning · Computer Science 2020-12-08 Sadegh Aliakbarian , Fatemeh Sadat Saleh , Lars Petersson , Stephen Gould , Mathieu Salzmann

Robotics has long been a field riddled with complex systems architectures whose modules and connections, whether traditional or learning-based, require significant human expertise and prior knowledge. Inspired by large pre-trained language…

Robotics · Computer Science 2022-09-27 Rogerio Bonatti , Sai Vemprala , Shuang Ma , Felipe Frujeri , Shuhang Chen , Ashish Kapoor

Aligning multiple modalities in a latent space, such as images and texts, has shown to produce powerful semantic visual representations, fueling tasks like image captioning, text-to-image generation, or image grounding. In the context of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Francesc Moreno-Noguer , Grégory Rogez

We address the task of identifying distracted driving by analyzing in-car videos using efficient transformers. Although transformer models have achieved outstanding performance in human action recognition tasks, their high computational…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Ricardo Pizarro , Roberto Valle , Rafael Barea , Jose M. Buenaposada , Luis Baumela , Luis Miguel Bergasa

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

In this paper, we tackle the task of scene-aware 3D human motion forecasting, which consists of predicting future human poses given a 3D scene and a past human motion. A key challenge of this task is to ensure consistency between the human…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Wei Mao , Miaomiao Liu , Richard Hartley , Mathieu Salzmann

Generative Pre-trained Transformer (GPT) is a state-of-the-art machine learning model capable of generating human-like text through natural language processing (NLP). GPT is trained on massive amounts of text data and uses deep learning…

Human pose forecasting is a challenging problem involving complex human body motion and posture dynamics. In cases that there are multiple people in the environment, one's motion may also be influenced by the motion and dynamic movements of…

Computer Vision and Pattern Recognition · Computer Science 2022-08-31 Edward Vendrow , Satyajit Kumar , Ehsan Adeli , Hamid Rezatofighi

Recent advancements in trajectory-guided video generation have achieved notable progress. However, existing models still face challenges in generating object motions with potentially changing 6D poses under wide-range rotations, due to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Longbin Ji , Lei Zhong , Pengfei Wei , Changjian Li

In this paper we propose a convolutional autoencoder to address the problem of motion infilling for 3D human motion data. Given a start and end sequence, motion infilling aims to complete the missing gap in between, such that the filled in…

Computer Vision and Pattern Recognition · Computer Science 2021-11-17 Manuel Kaufmann , Emre Aksan , Jie Song , Fabrizio Pece , Remo Ziegler , Otmar Hilliges

Articulated objects are central to interactive 3D applications, including embodied AI, robotics, and VR/AR, where functional part decomposition and kinematic motion are essential. Yet producing high-fidelity articulated assets remains…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Qingming Liu , Xinyue Yao , Shuyuan Zhang , Yueci Deng , Guiliang Liu , Zhen Liu , Kui Jia

The ability of intelligent systems to predict human behaviors is crucial, particularly in fields such as autonomous vehicle navigation and social robotics. However, the complexity of human motion have prevented the development of a…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Yang Gao , Po-Chien Luan , Alexandre Alahi

This paper studies the task of full generative modelling of realistic images of humans, guided only by coarse sketch of the pose, while providing control over the specific instance or type of outfit worn by the user. This is a difficult…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Xu Chen , Jie Song , Otmar Hilliges