中文
相关论文

相关论文: SAT: Supervisor Regularization and Animation Augme…

200 篇论文

We address the problem of 3D human pose estimation from 2D input images using only weakly supervised training data. Despite showing considerable success for 2D pose estimation, the application of supervised machine learning to 3D pose…

计算机视觉与模式识别 · 计算机科学 2018-07-31 Matteo Ruggero Ronchi , Oisin Mac Aodha , Robert Eng , Pietro Perona

Recent advances in video diffusion models have enabled realistic and controllable human image animation with temporal coherence. Although generating reasonable results, existing methods often overlook the need for regional supervision in…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Zhongcong Xu , Chaoyue Song , Guoxian Song , Jianfeng Zhang , Jun Hao Liew , Hongyi Xu , You Xie , Linjie Luo , Guosheng Lin , Jiashi Feng , Mike Zheng Shou

The difficulties in both data acquisition and annotation substantially restrict the sample sizes of training datasets for 3D medical imaging applications. As a result, constructing high-performance 3D convolutional neural networks from…

图像与视频处理 · 电气工程与系统科学 2022-01-06 Shu Zhang , Zihao Li , Hong-Yu Zhou , Jiechao Ma , Yizhou Yu

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Haitian Li , Haozhe Xie , Junxiang Xu , Beichen Wen , Fangzhou Hong , Ziwei Liu

We revisit the role of texture in monocular 3D hand reconstruction, not as an afterthought for photorealism, but as a dense, spatially grounded cue that can actively support pose and shape estimation. Our observation is simple: even in…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Giorgos Karvounas , Nikolaos Kyriazis , Iason Oikonomidis , Georgios Pavlakos , Antonis A. Argyros

Sparse-view satellite image surface reconstruction remains highly challenging, fundamentally because the reliability of multi-view matching under satellite imaging conditions is strongly spatially heterogeneous. Affected by large…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Min Chen , Wei Guo , Bin Wang , Wen Li , Tong Fang , Jinbo Zhang , Junqi Zhao , Hong Kuang , Han Hu , Xuming Ge , Qing Zhu , Bo Xu

We present Better Together, a method that simultaneously solves the human pose estimation problem while reconstructing a photorealistic 3D human avatar from multi-view videos. While prior art usually solves these problems separately, we…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Arthur Moreau , Mohammed Brahimi , Richard Shaw , Athanasios Papaioannou , Thomas Tanay , Zhensong Zhang , Eduardo Pérez-Pellitero

Reconstructing high-fidelity, animatable 3D head avatars from effortlessly captured monocular videos is a pivotal yet formidable challenge. Although significant progress has been made in rendering performance and manipulation capabilities,…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Jiawei Zhang , Zijian Wu , Zhiyang Liang , Yicheng Gong , Dongfang Hu , Yao Yao , Xun Cao , Hao Zhu

High-fidelity 4D dynamic facial avatar reconstruction from monocular video is a critical yet challenging task, driven by increasing demands for immersive virtual human applications. While Neural Radiance Fields (NeRF) have advanced scene…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Zhe Chang , Haodong Jin , Ying Sun , Yan Song , Hui Yu

Building high-fidelity digital twins of articulated objects from visual data remains a central challenge. Existing approaches depend on multi-view captures of the object in discrete, static states, which severely constrains their real-world…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Lijun Guo , Haoyu Zhao , Xingyue Zhao , Rong Fu , Linghao Zhuang , Siteng Huang , Zhongyu Li , Hua Zou

We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. We contribute DensePose 3D, a method that can learn such reconstructions in a weakly supervised fashion from 2D image annotations only.…

计算机视觉与模式识别 · 计算机科学 2021-09-02 Roman Shapovalov , David Novotny , Benjamin Graham , Patrick Labatut , Andrea Vedaldi

3D scene reconstruction from multiple views is an important classical problem in computer vision. Deep learning based approaches have recently demonstrated impressive reconstruction results. When training such models, self-supervised…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Arijit Mallick , Jörg Stückler , Hendrik Lensch

Estimating 3D human texture from a single image is essential in graphics and vision. It requires learning a mapping function from input images of humans with diverse poses into the parametric (UV) space and reasonably hallucinating…

计算机视觉与模式识别 · 计算机科学 2023-03-08 Said Fahri Altindis , Adil Meric , Yusuf Dalva , Ugur Gudukbay , Aysegul Dundar

Learning-based image reconstruction models, such as those based on the U-Net, require a large set of labeled images if good generalization is to be guaranteed. In some imaging domains, however, labeled data with pixel- or voxel-level label…

图像与视频处理 · 电气工程与系统科学 2024-01-08 Sean I. Young , Adrian V. Dalca , Enzo Ferrante , Polina Golland , Christopher A. Metzler , Bruce Fischl , Juan Eugenio Iglesias

Monocular Simultaneous Localization and Mapping (SLAM) aims to estimate a robot's pose while simultaneously reconstructing an unknown 3D scene using a single camera. While existing monocular SLAM systems generate detailed 3D geometry…

机器人学 · 计算机科学 2025-11-27 Yuchen Zhou , Haihang Wu

We introduce a novel framework for 3D human avatar generation and personalization, leveraging text prompts to enhance user engagement and customization. Central to our approach are key innovations aimed at overcoming the challenges in…

The modern supervised approaches for human image relighting rely on training data generated from 3D human models. However, such datasets are often small (e.g., Light Stage data with a small number of individuals) or limited to diffuse…

图形学 · 计算机科学 2021-10-18 Daichi Tajima , Yoshihiro Kanamori , Yuki Endo

The human face is central to communication. For immersive applications, the digital presence of a person should mirror the physical reality, capturing the users idiosyncrasies and detailed facial expressions. However, current 3D head avatar…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Jalees Nehvi , Timo Bolkart , Thabo Beeler , Justus Thies

This paper describes how to obtain accurate 3D body models and texture of arbitrary people from a single, monocular video in which a person is moving. Based on a parametric body model, we present a robust processing pipeline achieving 3D…

计算机视觉与模式识别 · 计算机科学 2018-04-17 Thiemo Alldieck , Marcus Magnor , Weipeng Xu , Christian Theobalt , Gerard Pons-Moll

Per-pixel ground-truth depth data is challenging to acquire at scale. To overcome this limitation, self-supervised learning has emerged as a promising alternative for training models to perform monocular depth estimation. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2019-08-20 Clément Godard , Oisin Mac Aodha , Michael Firman , Gabriel Brostow