English
Related papers

Related papers: GeoMan: Temporally Consistent Human Geometry Estim…

200 papers

Text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models has shown great promise but still suffers from inconsistent 3D geometric structures (Janus problems) and severe artifacts. The aforementioned problems…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Baorui Ma , Haoge Deng , Junsheng Zhou , Yu-Shen Liu , Tiejun Huang , Xinlong Wang

Real-world applications like video gaming and virtual reality often demand the ability to model 3D scenes that users can explore along custom camera trajectories. While significant progress has been made in generating 3D objects from text…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Tianyu Huang , Wangguandong Zheng , Tengfei Wang , Yuhao Liu , Zhenwei Wang , Junta Wu , Jie Jiang , Hui Li , Rynson W. H. Lau , Wangmeng Zuo , Chunchao Guo

Personalized 3D avatars require an animatable representation of digital humans. Doing so instantly from monocular videos offers scalability to broad class of users and wide-scale applications. In this paper, we present a fast, simple, yet…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Pramish Paudel , Anubhav Khanal , Ajad Chhatkuli , Danda Pani Paudel , Jyoti Tandukar

State-of-the-art video generation models produce remarkable photorealism, but they lack the precise control required to align generated content with specific scene requirements. Furthermore, without an underlying explicit geometry, these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Dana Cohen-Bar , Ido Sobol , Raphael Bensadoun , Shelly Sheynin , Oran Gafni , Or Patashnik , Daniel Cohen-Or , Amit Zohar

Intrinsic image decomposition aims to estimate physically based rendering (PBR) parameters such as albedo, roughness, and metallicity from images. While recent methods achieve strong single-view predictions, applying them independently to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Alara Dirik , Stefanos Zafeiriou

Perceiving 3D objects from monocular inputs is crucial for robotic systems, given its economy compared to multi-sensor settings. It is notably difficult as a single image can not provide any clues for predicting absolute depth values.…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Tai Wang , Jiangmiao Pang , Dahua Lin

Human video generation is a dynamic and rapidly evolving task that aims to synthesize 2D human body video sequences with generative models given control conditions such as text, audio, and pose. With the potential for wide-ranging…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Wentao Lei , Jinting Wang , Fengji Ma , Guanjie Huang , Li Liu

Accurately reconstructing human behavior in close-interaction scenarios is crucial for enabling realistic virtual interactions in augmented reality, precise motion analysis in sports, and natural collaborative behavior in human-robot tasks.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Qi Xia , Peishan Cong , Ziyi Wang , Yujing Sun , Qin Sun , Xinge Zhu , Mao Ye , Ruigang Yang , Yuexin Ma

Human motion copy is an intriguing yet challenging task in artificial intelligence and computer vision, which strives to generate a fake video of a target person performing the motion of a source person. The problem is inherently…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Sifan Wu , Zhenguang Liu , Beibei Zhang , Roger Zimmermann , Zhongjie Ba , Xiaosong Zhang , Kui Ren

We present PanoWorld, a panoramic video world model that generates geometry-consistent 360$\degree$ video from a single image and a caption. Existing panoramic video methods optimize primarily for visual realism and do not explicitly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Le Jiang , Xiangyu Bai , Bishoy Galoaa , Shayda Moezzi , Caleb James Lee , Tooba Imtiaz , Edmund Yeh , Jennifer Dy , Yanzhi Wang , Sarah Ostadabbas

We introduce a robust, real-time, high-resolution human video matting method that achieves new state-of-the-art performance. Our method is much lighter than previous approaches and can process 4K at 76 FPS and HD at 104 FPS on an Nvidia GTX…

Computer Vision and Pattern Recognition · Computer Science 2021-08-27 Shanchuan Lin , Linjie Yang , Imran Saleemi , Soumyadip Sengupta

We present a novel approach designed to address the complexities posed by challenging, out-of-distribution data in the single-image depth estimation task. Starting with images that facilitate depth prediction due to the absence of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Fabio Tosi , Pierluigi Zama Ramirez , Matteo Poggi

The appearance of a human in clothing is driven not only by the pose but also by its temporal context, i.e., motion. However, such context has been largely neglected by existing monocular human modeling methods whose neural networks often…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Hansol Lee , Junuk Cha , Yunhoe Ku , Jae Shin Yoon , Seungryul Baek

Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer…

Computer Vision and Pattern Recognition · Computer Science 2021-10-18 Patrick Ruhkamp , Daoyi Gao , Hanzhi Chen , Nassir Navab , Benjamin Busam

We present a novel algorithm for estimating the broad 3D geometric structure of outdoor video scenes. Leveraging spatio-temporal video segmentation, we decompose a dynamic scene captured by a video into geometric classes, based on…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 S. Hussain Raza , Matthias Grundmann , Irfan Essa

In this paper, we tackle the challenging task of learning a generalizable human NeRF model from a monocular video. Although existing generalizable human NeRFs have achieved impressive results, they require muti-view images or videos which…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Chen Li , Jiahao Lin , Gim Hee Lee

We present an algorithm for estimating consistent dense depth maps and camera poses from a monocular video. We integrate a learning-based depth prior, in the form of a convolutional neural network trained for single-image depth estimation,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Johannes Kopf , Xuejian Rong , Jia-Bin Huang

As a crucial task of autonomous driving, 3D object detection has made great progress in recent years. However, monocular 3D object detection remains a challenging problem due to the unsatisfactory performance in depth estimation. Most…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Yinmin Zhang , Xinzhu Ma , Shuai Yi , Jun Hou , Zhihui Wang , Wanli Ouyang , Dan Xu

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

Current texture synthesis methods, which generate textures from fixed viewpoints, suffer from inconsistencies due to the lack of global context and geometric understanding. Meanwhile, recent advancements in video generation models have…

Graphics · Computer Science 2025-06-27 Donggoo Kang , Jangyeong Kim , Dasol Jeong , Junyoung Choi , Jeonga Wi , Hyunmin Lee , Joonho Gwon , Joonki Paik