English
Related papers

Related papers: Mem-MLP: Real-Time 3D Human Motion Generation from…

200 papers

A large number of cameras embedded on smart-phones, drones or inside cars have a direct access to external motion sensing from gyroscopes and accelerometers. On these power-limited devices, video compression must be of low-complexity. For…

Image and Video Processing · Electrical Eng. & Systems 2020-02-03 Karim El Khoury , Pascal Pellegrin , Antonin Descampe , Sébastien Lugan , Benoit Macq

Recent Multi-Modal Large Language Models (MLLMs) have demonstrated strong capabilities in learning joint representations from text and images. However, their spatial reasoning remains limited. We introduce 3DFroMLLM, a novel framework that…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Noor Ahmed , Cameron Braunstein , Steffen Eger , Eddy Ilg

Recovering world-coordinate human motion from monocular videos with humanoid robot retargeting is significant for embodied intelligence and robotics. To avoid complex SLAM pipelines or heavy temporal models, we propose a lightweight,…

Robotics · Computer Science 2025-12-29 Zhangzheng Tu , Kailun Su , Shaolong Zhu , Yukun Zheng

Tracking the full body motions of users in XR (AR/VR) devices is a fundamental challenge to bring a sense of authentic social presence. Due to the absence of dedicated leg sensors, currently available body tracking methods adopt a synthesis…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Denys Rozumnyi , Nadine Bertsch , Othman Sbai , Filippo Arcadu , Yuhua Chen , Artsiom Sanakoyeu , Manoj Kumar , Catherine Herold , Robin Kips

Accurate human motion prediction (HMP) is critical for seamless human-robot collaboration, particularly in handover tasks that require real-time adaptability. Despite the high accuracy of state-of-the-art models, their computational…

Robotics · Computer Science 2025-03-04 Gerard Gómez-Izquierdo , Javier Laplaza , Alberto Sanfeliu , Anaís Garrell

Imitation learning is a promising approach for training humanoid robots to both walk and manipulate, but it requires a large number of demonstrations, which are time-intensive and difficult to collect via teleoperation. Existing…

We present the first marker-less approach for temporally coherent 3D performance capture of a human with general clothing from monocular video. Our approach reconstructs articulated human skeleton motion as well as medium-scale non-rigid…

Computer Vision and Pattern Recognition · Computer Science 2018-02-26 Weipeng Xu , Avishek Chatterjee , Michael Zollhöfer , Helge Rhodin , Dushyant Mehta , Hans-Peter Seidel , Christian Theobalt

Rehabilitation is important to improve quality of life for mobility-impaired patients. Smart walkers are a commonly used solution that should embed automatic and objective tools for data-driven human-in-the-loop control and monitoring.…

Computer Vision and Pattern Recognition · Computer Science 2021-07-06 Manuel Palermo , Sara Moccia , Lucia Migliorelli , Emanuele Frontoni , Cristina P. Santos

With the advent of Transformer-based one-stream trackers that possess strong capability in inter-frame relation modeling, recent research has increasingly focused on how to introduce spatio-temporal context. However, most existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Wenrui Cai , Zhenyi Lu , Yuzhe Li , Yongchao Feng , Jinqing Zhang , Qingjie Liu , Yunhong Wang

The cognitive system for human action and behavior has evolved into a deep learning regime, and especially the advent of Graph Convolution Networks has transformed the field in recent years. However, previous works have mainly focused on…

Computer Vision and Pattern Recognition · Computer Science 2021-07-16 Feng Shi , Chonghan Lee , Liang Qiu , Yizhou Zhao , Tianyi Shen , Shivran Muralidhar , Tian Han , Song-Chun Zhu , Vijaykrishnan Narayanan

The ability to predict motion in real time is fundamental to many maneuvering activities in animals, particularly those critical for survival, such as attack and escape responses. Given its significance, it is no surprise that motion…

Image and Video Processing · Electrical Eng. & Systems 2025-09-16 Subhradip Chakraborty , Shay Snyder , Md Abdullah-Al Kaiser , Maryam Parsa , Gregory Schwartz , Akhilesh R. Jaiswal

Monocular video human mesh recovery is essential for digital humans, avatar animation, and embodied simulation, where both temporal stability and expressive whole-body motion are required. Existing video HMR methods produce coherent body…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Wenhao Shen , Ming Zhou , Hengyuan Zhang , Siyuan Bian , Youjiang Xu , Xi Lin

We present an approach that uses a deep learning model, in particular, a MultiLayer Perceptron (MLP), for estimating the missing values of a variable in multivariate time series data. We focus on filling a long continuous gap (e.g.,…

The attention mechanism has become a go-to technique for natural language processing and computer vision tasks. Recently, the MLP-Mixer and other MLP-based architectures, based simply on multi-layer perceptrons (MLPs), are also powerful…

Computer Vision and Pattern Recognition · Computer Science 2022-05-31 Tian Lv , Chongyang Bai , Chaojie Wang

Despite recent advancements in the Large Reconstruction Model (LRM) demonstrating impressive results, when extending its input from single image to multiple images, it exhibits inefficiencies, subpar geometric and texture quality, as well…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Mengfei Li , Xiaoxiao Long , Yixun Liang , Weiyu Li , Yuan Liu , Peng Li , Wenhan Luo , Wenping Wang , Yike Guo

In multi-view 3D human pose estimation, models typically rely on images captured simultaneously from different camera views to predict a pose at a specific moment. While providing accurate spatial information, this traditional approach…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Ling Li , Changjie Chen , Yuyan Wang , Jiaqing Lyu , Kenglun Chang , Yiyun Chen , Zhidong Deng

Estimating 3D full-body avatars from AR/VR devices is essential for creating immersive experiences in AR/VR applications. This task is challenging due to the limited input from Head Mounted Devices, which capture only sparse observations…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Han Feng , Wenchao Ma , Quankai Gao , Xianwei Zheng , Nan Xue , Huijuan Xu

Structural coloration is commonly modeled using wave optics for reliable and photorealistic rendering of natural, quasi-periodic and complex nanostructures. Such models often rely on dense, preliminary or preprocessed data to accurately…

Graphics · Computer Science 2025-07-03 Narayan Kandel , Daljit Singh J. S. Dhillon

We introduce a dynamic sparse training algorithm based on linearized Bregman iterations / mirror descent that exploits the naturally incurred sparsity by alternating between periods of static and dynamic sparsity pattern updates. The key…

Machine Learning · Computer Science 2026-05-19 Yannick Lunk , Sebastian J. Scott , Leon Bungert

Neural fields, a category of neural networks trained to represent high-frequency signals, have gained significant attention in recent years due to their impressive performance in modeling complex 3D data, such as signed distance (SDFs) or…

Computer Vision and Pattern Recognition · Computer Science 2024-02-13 Marko Mihajlovic , Sergey Prokudin , Marc Pollefeys , Siyu Tang