English
Related papers

Related papers: AnyCrowd: Instance-Isolated Identity-Pose Binding …

200 papers

Recent progress in video diffusion models has markedly advanced character animation, which synthesizes motioned videos by animating a static identity image according to a driving video. Explicit methods represent motion using skeleton,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Zhufeng Xu , Xuan Gao , Feng-Lin Liu , Haoxian Zhang , Zhixue Fang , Yu-Kun Lai , Xiaoqiang Liu , Pengfei Wan , Lin Gao

Leveraging Stable Diffusion for the generation of personalized portraits has emerged as a powerful and noteworthy tool, enabling users to create high-fidelity, custom character avatars based on their specific prompts. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Siying Cui , Jia Guo , Xiang An , Jiankang Deng , Yongle Zhao , Xinyu Wei , Ziyong Feng

Producing expressive facial animations from static images is a challenging task. Prior methods relying on explicit geometric priors (e.g., facial landmarks or 3DMM) often suffer from artifacts in cross reenactment and struggle to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Qiang Wang , Mengchao Wang , Fan Jiang , Yaqi Fan , Yonggang Qi , Mu Xu

Whole-body audio-driven avatar pose and expression generation is a critical task for creating lifelike digital humans and enhancing the capabilities of interactive virtual agents, with wide-ranging applications in virtual reality, digital…

Sound · Computer Science 2025-10-15 Tianbao Zhang , Jian Zhao , Yuer Li , Zheng Zhu , Ping Hu , Zhaoxin Fan , Wenjun Wu , Xuelong Li

Character customization, or 'face crafting,' is a vital feature in role-playing games (RPGs), enhancing player engagement by enabling the creation of personalized avatars. Existing automated methods often struggle with generalizability…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Suzhen Wang , Weijie Chen , Wei Zhang , Minda Zhao , Lincheng Li , Rongsheng Zhang , Zhipeng Hu , Xin Yu

Digitizing humans and synthesizing photorealistic avatars with explicit 3D pose and camera controls are central to VR, telepresence, and entertainment. Existing skinning-based workflows require laborious manual rigging or template-based…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Zhilin Guo , Jing Yang , Kyle Fogarty , Jingyi Wan , Boqiao Zhang , Tianhao Wu , Weihao Xia , Chenliang Zhou , Sakar Khattar , Fangcheng Zhong , Cristina Nader Vasconcelos , Cengiz Oztireli

Subject-driven image generation has shown great success in creating personalized content, but its capabilities are largely confined to single subjects in common poses. Current approaches face a fundamental conflict when handling multiple…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Tianze Xia , Zijian Ning , Zonglin Zhao , Mingjia Wang

We address the problem of image-based crowd counting. In particular, we propose a new problem called unlabeled scene-adaptive crowd counting. Given a new target scene, we would like to have a crowd counting model specifically adapted to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Mahesh Kumar Krishna Reddy , Mrigank Rochan , Yiwei Lu , Yang Wang

Learning disentangled representations of data is a fundamental problem in artificial intelligence. Specifically, disentangled latent representations allow generative models to control and compose the disentangled factors in the synthesis…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Yotam Nitzan , Amit Bermano , Yangyan Li , Daniel Cohen-Or

The field of text-to-image (T2I) generation has made significant progress in recent years, largely driven by advancements in diffusion models. Linguistic control enables effective content creation, but struggles with fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Yanan Sun , Yanchen Liu , Yinhao Tang , Wenjie Pei , Kai Chen

We introduce You Only Train Once (YOTO), a dynamic human generation framework, which performs free-viewpoint rendering of different human identities with distinct motions, via only one-time training from monocular videos. Most prior works…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Jaehyeok Kim , Dongyoon Wee , Dan Xu

In this paper, we propose a novel framework named DRL-CPG to learn disentangled latent representation for controllable person image generation, which can produce realistic person images with desired poses and human attributes (e.g., pose,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Wenju Xu , Chengjiang Long , Yongwei Nie , Guanghui Wang

Autoregressive Model (AR) has shown remarkable success in conditional image generation. However, these approaches for multiple reference generation struggle with decoupling different reference identities. In this work, we propose the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Haiyue Sun , Qingdong He , Jinlong Peng , Peng Tang , Jiangning Zhang , Junwei Zhu , Xiaobin Hu , Shuicheng Yan

This study focuses on a novel task in text-to-image (T2I) generation, namely action customization. The objective of this task is to learn the co-existing action from limited data and generalize it to unseen humans or even animals.…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Siteng Huang , Biao Gong , Yutong Feng , Xi Chen , Yuqian Fu , Yu Liu , Donglin Wang

Character image animation has rapidly advanced with the rise of digital humans. However, existing methods rely largely on 2D-rendered pose images for motion guidance, which limits generalization and discards essential 4D information for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yanbo Ding , Xirui Hu , Zhizhi Guo , Yan Zhang , Xinrui Wang , Zhixiang He , Chi Zhang , Yali Wang , Xuelong Li

In human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted portrait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Renshuai Liu , Bowen Ma , Wei Zhang , Zhipeng Hu , Changjie Fan , Tangjie Lv , Yu Ding , Xuan Cheng

Significant progress has been achieved in high-fidelity video synthesis, yet current paradigms often fall short in effectively integrating identity information from multiple subjects. This leads to semantic conflicts and suboptimal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Panwang Pan , Jingjing Zhao , Yuchen Lin , Chenguo Lin , Chenxin Li , Hengyu Liu , Tingting Shen , Yadong MU

Character video generation is a significant real-world application focused on producing high-quality videos featuring specific characters. Recent advancements have introduced various control signals to animate static characters,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Zhao Wang , Hao Wen , Lingting Zhu , Chenming Shang , Yujiu Yang , Qi Dou

Recent advances in garment-centric image generation from text and image prompts based on diffusion models are impressive. However, existing methods lack support for various combinations of attire, and struggle to preserve the garment…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Xinghui Li , Qichao Sun , Pengze Zhang , Fulong Ye , Zhichao Liao , Wanquan Feng , Songtao Zhao , Qian He

Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Zeyu Yang , Nan Song , Wei Li , Xiatian Zhu , Li Zhang , Philip H. S. Torr