English
Related papers

Related papers: ViPS: Video-informed Pose Spaces for Auto-Rigged M…

200 papers

Today's Mixed Reality head-mounted displays track the user's head pose in world space as well as the user's hands for interaction in both Augmented Reality and Virtual Reality scenarios. While this is adequate to support user input, it…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Jiaxi Jiang , Paul Streli , Huajian Qiu , Andreas Fender , Larissa Laich , Patrick Snape , Christian Holz

We propose a novel efficient and lightweight model for human pose estimation from a single image. Our model is designed to achieve competitive results at a fraction of the number of parameters and computational cost of various…

Computer Vision and Pattern Recognition · Computer Science 2020-02-11 Hossam Isack , Christian Haene , Cem Keskin , Sofien Bouaziz , Yuri Boykov , Shahram Izadi , Sameh Khamis

Animation of humanoid characters is essential in various graphics applications, but requires significant time and cost to create realistic animations. We propose an approach to synthesize 4D animated sequences of input static 3D humanoid…

Graphics · Computer Science 2025-03-21 Marc Benedí San Millán , Angela Dai , Matthias Nießner

3D human pose estimation is a key enabling technology for applications such as healthcare monitoring, human-robot collaboration, and immersive gaming, but real-world deployment remains challenged by viewpoint variations. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yejia Liu , Hengle Jiang , Haoxian Liu , Runxi Huang , Xiaomin Ouyang

Pretrained vision-language models (VLMs), such as CLIP, have shown remarkable potential in few-shot image classification and led to numerous effective transfer learning strategies. These methods leverage the pretrained knowledge of VLMs to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Dexia Chen , Qianjie Zhu , Weibing Li , Yue Yu , Tong Zhang , Ruixuan Wang

Video Frame Interpolation (VFI) remains a cornerstone in video enhancement, enabling temporal upscaling for tasks like slow-motion rendering, frame rate conversion, and video restoration. While classical methods rely on optical flow and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Priyansh Srivastava , Romit Chatterjee , Abir Sen , Aradhana Behura , Ratnakar Dash

Several video-based 3D pose and shape estimation algorithms have been proposed to resolve the temporal inconsistency of single-image-based methods. However it still remains challenging to have stable and accurate reconstruction. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-10-25 Ziwen Li , Bo Xu , Han Huang , Cheng Lu , Yandong Guo

Predicting future video frames is a challenging task with many downstream applications. Previous work has shown that procedural knowledge enables deep models for complex dynamical settings, however their model ViPro assumed a given ground…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Patrick Takenaka , Johannes Maucher , Marco F. Huber

Monocular 3D reconstruction of articulated object categories is challenging due to the lack of training data and the inherent ill-posedness of the problem. In this work we use video self-supervision, forcing the consistency of consecutive…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Filippos Kokkinos , Iasonas Kokkinos

Motion capture from a monocular video is fundamental and crucial for us humans to naturally experience and interact with each other in Virtual Reality (VR) and Augmented Reality (AR). However, existing methods still struggle with…

Computer Vision and Pattern Recognition · Computer Science 2022-10-31 Xin Chen , Zhuo Su , Lingbo Yang , Pei Cheng , Lan Xu , Bin Fu , Gang Yu

Articulated 3D object generation is fundamental for creating realistic, functional, and interactable virtual assets which are not simply static. We introduce MeshArt, a hierarchical transformer-based approach to generate articulated 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Daoyi Gao , Yawar Siddiqui , Lei Li , Angela Dai

Pose Estimation techniques rely on visual cues available through observations represented in the form of pixels. But the performance is bounded by the frame rate of the video and struggles from motion blur, occlusions, and temporal…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Snehesh Shrestha , Cornelia Fermüller , Tianyu Huang , Pyone Thant Win , Adam Zukerman , Chethan M. Parameshwara , Yiannis Aloimonos

Gaussian Splatting has achieved remarkable progress in multi-view surface reconstruction, yet it exhibits notable degradation when only few views are available. Although recent efforts alleviate this issue by enhancing multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Jimin Tang , Wenyuan Zhang , Junsheng Zhou , Zian Huang , Kanle Shi , Shenkun Xu , Yu-Shen Liu , Zhizhong Han

The design of functional seating furniture is a complicated process which often requires extensive manual design effort and empirical evaluation. We propose a computational design framework for pose-driven automated generation of…

Graphics · Computer Science 2020-07-02 Kurt Leimer , Andreas Winkler , Stefan Ohrhallinger , Przemyslaw Musialski

Learning to represent three dimensional (3D) human pose given a two dimensional (2D) image of a person, is a challenging problem. In order to make the problem less ambiguous it has become common practice to estimate 3D pose in the camera…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Mara Levy , Abhinav Shrivastava

Mesh reconstruction of the cardiac anatomy from medical images is useful for shape and motion measurements and biophysics simulations to facilitate the assessment of cardiac function and health. However, 3D medical images are often acquired…

Image and Video Processing · Electrical Eng. & Systems 2024-10-22 Yihao Luo , Dario Sesia , Fanwen Wang , Yinzhe Wu , Wenhao Ding , Jiahao Huang , Fadong Shi , Anoop Shah , Amit Kaural , Jamil Mayet , Guang Yang , ChoonHwai Yap

The estimation of implicit cross-frame correspondences and the high computational cost have long been major challenges in video semantic segmentation (VSS) for driving scenes. Prior works utilize keyframes, feature propagation, or…

Computer Vision and Pattern Recognition · Computer Science 2024-04-29 Diandian Guo , Deng-Ping Fan , Tongyu Lu , Christos Sakaridis , Luc Van Gool

In the era of generative AI, integrating video generation models into robotics opens new possibilities for the general-purpose robot agent. This paper introduces imitation learning with latent video planning (VILP). We propose a latent…

Robotics · Computer Science 2025-02-05 Zhengtong Xu , Qiang Qiu , Yu She

Prevailing Vision-Language-Action Models (VLAs) for robotic manipulation are built upon vision-language backbones pretrained on large-scale, but disconnected static web data. As a result, despite improved semantic generalization, the policy…

Robotics · Computer Science 2025-12-22 Jonas Pai , Liam Achenbach , Victoriano Montesinos , Benedek Forrai , Oier Mees , Elvis Nava

Masked autoregressive models (MAR) have emerged as a powerful paradigm for image and video generation, combining the flexibility of masked modeling with the expressiveness of continuous tokenizers. However, when sampling individual frames,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Zian Li , Muhan Zhang