English
Related papers

Related papers: MEGA: Masked Generative Autoencoder for Human Mesh…

200 papers

We present TexMesh, a novel approach to reconstruct detailed human meshes with high-resolution full-body texture from RGB-D video. TexMesh enables high quality free-viewpoint rendering of humans. Given the RGB frames, the captured…

Computer Vision and Pattern Recognition · Computer Science 2020-09-22 Tiancheng Zhi , Christoph Lassner , Tony Tung , Carsten Stoll , Srinivasa G. Narasimhan , Minh Vo

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Lingteng Qiu , Xiaodong Gu , Peihao Li , Qi Zuo , Weichao Shen , Junfei Zhang , Kejie Qiu , Weihao Yuan , Guanying Chen , Zilong Dong , Liefeng Bo

We propose Heterogeneous Masked Autoregression (HMA) for modeling action-video dynamics to generate high-quality data and evaluation in scaling robot learning. Building interactive video world models and policies for robotics is difficult…

Robotics · Computer Science 2025-02-07 Lirui Wang , Kevin Zhao , Chaoqi Liu , Xinlei Chen

Masked Autoregressive (MAR) models promise better efficiency in visual generation than autoregressive (AR) models for the ability of parallel generation, yet their acceleration potential remains constrained by the modeling complexity of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Feihong Yan , Peiru Wang , Yao Zhu , Kaiyu Pang , Qingyan Wei , Huiqi Li , Linfeng Zhang

The field of image-to-video generation has made remarkable progress. However, challenges such as human limb twisting and facial distortion persist, especially when generating long videos or modeling intensive motions. Existing human image…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Chang Liu , Mengting Chen , Yixuan Huang , Haoning Wu , Chen Ju , Shuai Xiao , Jinsong Lan , Yanfeng Wang

How do humans recognize an object in a piece of video? Due to the deteriorated quality of single frame, it may be hard for people to identify an occluded object in this frame by just utilizing information within one image. We argue that…

Computer Vision and Pattern Recognition · Computer Science 2020-03-27 Yihong Chen , Yue Cao , Han Hu , Liwei Wang

In recent years, numerous graph generative models (GGMs) have been proposed. However, evaluating these models remains a considerable challenge, primarily due to the difficulty in extracting meaningful graph features that accurately…

Machine Learning · Computer Science 2025-03-18 Chengen Wang , Murat Kantarcioglu

Videos from edited media like movies are a useful, yet under-explored source of information. The rich variety of appearance and interactions between humans depicted over a large temporal context in these films could be a valuable source of…

Computer Vision and Pattern Recognition · Computer Science 2020-12-18 Georgios Pavlakos , Jitendra Malik , Angjoo Kanazawa

The cost and accuracy of simulating complex physical systems using the Finite Element Method (FEM) scales with the resolution of the underlying mesh. Adaptive meshes improve computational efficiency by refining resolution in critical…

Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that advances generative speech enhancement through a compact and…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-16 The Hieu Pham , Tan Dat Nguyen , Phuong Thanh Tran , Joon Son Chung , Duc Dung Nguyen

This paper shows that masked autoencoders (MAE) are scalable self-supervised learners for computer vision. Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels. It is based on two core…

Computer Vision and Pattern Recognition · Computer Science 2021-12-21 Kaiming He , Xinlei Chen , Saining Xie , Yanghao Li , Piotr Dollár , Ross Girshick

Recovering 3D human mesh from monocular images is a popular topic in computer vision and has a wide range of applications. This paper aims to estimate 3D mesh of multiple body parts (e.g., body, hands) with large-scale differences from a…

Computer Vision and Pattern Recognition · Computer Science 2020-10-28 Yu Sun , Qian Bao , Wu Liu , Wenpeng Gao , Yili Fu , Chuang Gan , Tao Mei

In this paper, we define and study a new Cloth2Body problem which has a goal of generating 3D human body meshes from a 2D clothing image. Unlike the existing human mesh recovery problem, Cloth2Body needs to address new and emerging…

Computer Vision and Pattern Recognition · Computer Science 2023-09-29 Lu Dai , Liqian Ma , Shenhan Qian , Hao Liu , Ziwei Liu , Hui Xiong

We present an approach to generating 3D human models from images. The key to our framework is that we predict double-sided orthographic depth maps and color images from a single perspective projected image. Our framework consists of three…

Computer Vision and Pattern Recognition · Computer Science 2022-05-17 Min-Gyu Park , Ju-Mi Kang , Je Woo Kim , Ju Hong Yoon

Recently, implicit neural representation has been widely used to generate animatable human avatars. However, the materials and geometry of those representations are coupled in the neural network and hard to edit, which hinders their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Qifeng Chen , Rengan Xie , Kai Huang , Qi Wang , Wenting Zheng , Rong Li , Yuchi Huo

Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified deep network to learn generative-to-discriminative…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Shiming Ge , Weijia Guo , Chenyu Li , Junzheng Zhang , Yong Li , Dan Zeng

In this age of information, images are a critical medium for storing and transmitting information. With the rapid growth of image data amount, visual compression and visual data perception are two important research topics attracting a lot…

Image and Video Processing · Electrical Eng. & Systems 2024-07-02 Yuefeng Zhang , Chuanmin Jia , Jiannhui Chang , Siwei Ma

Generating facial reactions in a human-human dyadic interaction is complex and highly dependent on the context since more than one facial reactions can be appropriate for the speaker's behaviour. This has challenged existing machine…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Tong Xu , Micol Spitale , Hao Tang , Lu Liu , Hatice Gunes , Siyang Song

Transformer encoder architectures have recently achieved state-of-the-art results on monocular 3D human mesh reconstruction, but they require a substantial number of parameters and expensive computations. Due to the large memory overhead…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Junhyeong Cho , Kim Youwang , Tae-Hyun Oh

Recovering high-fidelity 3D hand geometry from images is a critical task in computer vision, holding significant value for domains such as robotics, animation and VR/AR. Crucially, scalable applications demand both accuracy and deployment…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Yumeng Liu , Xiao-Xiao Long , Marc Habermann , Xuanze Yang , Cheng Lin , Yuan Liu , Yuexin Ma , Wenping Wang , Ligang Liu