中文
相关论文

相关论文: STAR: Skeleton-aware Text-based 4D Avatar Generati…

200 篇论文

Text-guided domain adaptation and generation of 3D-aware portraits find many applications in various fields. However, due to the lack of training data and the challenges in handling the high variety of geometry and appearance, the existing…

计算机视觉与模式识别 · 计算机科学 2024-04-15 Biwen Lei , Kai Yu , Mengyang Feng , Miaomiao Cui , Xuansong Xie

High-fidelity head avatar reconstruction plays a crucial role in AR/VR, gaming, and multimedia content creation. Recent advances in 3D Gaussian Splatting (3DGS) have demonstrated effectiveness in modeling complex geometry with real-time…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Shikun Zhang , Cunjian Chen , Yiqun Wang , Qiuhong Ke , Yong Li

We introduce a novel framework for 3D human avatar generation and personalization, leveraging text prompts to enhance user engagement and customization. Central to our approach are key innovations aimed at overcoming the challenges in…

Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor, a novel system necessitating…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Ziyao Huang , Fan Tang , Yong Zhang , Xiaodong Cun , Juan Cao , Jintao Li , Tong-Yee Lee

Although large vision-language models (LVLMs) leverage rich visual token representations to achieve strong performance on multimodal tasks, these tokens also introduce significant computational overhead during inference. Existing…

机器学习 · 计算机科学 2025-05-20 Yichen Guo , Hanze Li , Zonghao Zhang , Jinhao You , Kai Tang , Xiande Huang

Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations, and physical inconsistencies that are largely rooted in the…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Shurui Gui , Deep Anil Patel , Xiner Li , Martin Renqiang Min

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Shangjin Zhai , Zhichao Ye , Jialin Liu , Weijian Xie , Jiaqi Hu , Zhen Peng , Hua Xue , Danpeng Chen , Xiaomeng Wang , Lei Yang , Nan Wang , Haomin Liu , Guofeng Zhang

We introduce a highly robust GAN-based framework for digitizing a normalized 3D avatar of a person from a single unconstrained photo. While the input image can be of a smiling person or taken in extreme lighting conditions, our method can…

计算机视觉与模式识别 · 计算机科学 2021-06-23 Huiwen Luo , Koki Nagano , Han-Wei Kung , Mclean Goldwhite , Qingguo Xu , Zejian Wang , Lingyu Wei , Liwen Hu , Hao Li

Recent advances in Gaussian Splatting have significantly boosted the reconstruction of head avatars, enabling high-quality facial modeling by representing an 3D avatar as a collection of 3D Gaussians. However, existing methods predominantly…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Shiqi Xin , Xiaolin Zhang , Yanbin Liu , Peng Zhang , Caifeng Shan

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Yatian Pang , Bin Zhu , Bin Lin , Mingzhe Zheng , Francis E. H. Tay , Ser-Nam Lim , Harry Yang , Li Yuan

In this paper, we introduce GaussianMotion, a novel human rendering model that generates fully animatable scenes aligned with textual descriptions using Gaussian Splatting. Although existing methods achieve reasonable text-to-3D generation…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Gyumin Shim , Sangmin Lee , Jaegul Choo

Our goal is to create a realistic 3D facial avatar with hair and accessories using only a text description. While this challenge has attracted significant recent interest, existing methods either lack realism, produce unrealistic shapes, or…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Hao Zhang , Yao Feng , Peter Kulits , Yandong Wen , Justus Thies , Michael J. Black

We present a data-driven framework for unsupervised human motion retargeting that animates a target subject with the motion of a source subject. Our method is correspondence-free, requiring neither spatial correspondences between the source…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Rim Rekik , Mathieu Marsot , Anne-Hélène Olivier , Jean-Sébastien Franco , Stefanie Wuhrer

Controllable text-to-image (T2I) diffusion models have shown impressive performance in generating high-quality visual content through the incorporation of various conditions. Current methods, however, exhibit limited performance when guided…

计算机视觉与模式识别 · 计算机科学 2024-11-06 Jiajun Wang , Morteza Ghahremani , Yitong Li , Björn Ommer , Christian Wachinger

Clothed avatar generation has wide applications in virtual and augmented reality, filmmaking, and more. Previous methods have achieved success in generating diverse digital avatars, however, generating avatars with disentangled components…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Weitian Zhang , Yichao Yan , Sijing Wu , Manwen Liao , Xiaokang Yang

We present Better Together, a method that simultaneously solves the human pose estimation problem while reconstructing a photorealistic 3D human avatar from multi-view videos. While prior art usually solves these problems separately, we…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Arthur Moreau , Mohammed Brahimi , Richard Shaw , Athanasios Papaioannou , Thomas Tanay , Zhensong Zhang , Eduardo Pérez-Pellitero

We introduce SimAvatar, a framework designed to generate simulation-ready clothed 3D human avatars from a text prompt. Current text-driven human avatar generation methods either model hair, clothing, and the human body using a unified…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Xueting Li , Ye Yuan , Shalini De Mello , Gilles Daviet , Jonathan Leaf , Miles Macklin , Jan Kautz , Umar Iqbal

Streaming speech-to-avatar synthesis creates real-time animations for a virtual character from audio data. Accurate avatar representations of speech are important for the visualization of sound in linguistics, phonetics, and phonology,…

声音 · 计算机科学 2023-10-26 Tejas S. Prabhune , Peter Wu , Bohan Yu , Gopala K. Anumanchipalli

To advance real-world fashion image editing, we analyze existing two-stage pipelines(mask generation followed by diffusion-based editing)which overly prioritize generator optimization while neglecting mask controllability. This results in…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Yuran Dong , Mang Ye

Real-time animation of virtual characters has traditionally been accomplished by playing short sequences of animations structured in the form of a graph. These methods are time-consuming to set up and scale poorly with the number of motions…

图形学 · 计算机科学 2023-10-10 Jose Luis Ponton