中文
相关论文

相关论文: DanceTogether! Identity-Preserving Multi-Person In…

200 篇论文

Synthesizing dynamic appearances of humans in motion plays a central role in applications such as AR/VR and video editing. While many recent methods have been proposed to tackle this problem, handling loose garments with complex textures…

计算机视觉与模式识别 · 计算机科学 2021-11-12 Tuanfeng Y. Wang , Duygu Ceylan , Krishna Kumar Singh , Niloy J. Mitra

Identity-preserving text-to-video (IPT2V) generation aims to create high-fidelity videos with consistent human identity. It is an important task in video generation but remains an open problem for generative models. This paper pushes the…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Shenghai Yuan , Jinfa Huang , Xianyi He , Yunyuan Ge , Yujun Shi , Liuhan Chen , Jiebo Luo , Li Yuan

We present CoMoGen, a controllable video generation framework that generates realistic interactive dynamics from a single binary mask sequence conditioned on an input image. CoMoGen introduces a lightweight MaskAdapter that encodes binary…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Adil Meric , Lin Geng Foo , Mert Kiray , Benjamin Busam , Rishabh Dabral , Christian Theobalt

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zhengwentai Sun , Keru Zheng , Chenghong Li , Hongjie Liao , Xihe Yang , Heyuan Li , Yihao Zhi , Shuliang Ning , Shuguang Cui , Xiaoguang Han

Recent years have seen substantial progress in diffusion-based controllable video generation. However, achieving precise control in complex scenarios, including fine-grained object parts, sophisticated motion trajectories, and coherent…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Haitao Zhou , Chuang Wang , Rui Nie , Jinlin Liu , Dongdong Yu , Qian Yu , Changhu Wang

Generative masked transformers have demonstrated remarkable success across various content generation tasks, primarily due to their ability to effectively model large-scale dataset distributions with high consistency. However, in the…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Yilin Wang , Chuan Guo , Yuxuan Mu , Muhammad Gohar Javed , Xinxin Zuo , Juwei Lu , Hai Jiang , Li Cheng

Gait recognition is a valuable biometric task that enables the identification of individuals from a distance based on their walking patterns. However, it remains limited by the lack of large-scale labeled datasets and the difficulty of…

计算机视觉与模式识别 · 计算机科学 2025-08-20 Sirshapan Mitra , Yogesh S. Rawat

Audio-driven talking face generation has gained significant attention for applications in digital media and virtual avatars. While recent methods improve audio-lip synchronization, they often struggle with temporal consistency, identity…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Fatemeh Nazarieh , Zhenhua Feng , Diptesh Kanojia , Muhammad Awais , Josef Kittler

Learning robotic manipulation from human videos is a promising solution to the data bottleneck in robotics, but the distribution shift between humans and robots remains a critical challenge. Existing approaches often produce entangled…

机器人学 · 计算机科学 2026-05-06 Zhiyuan Li , Wenyan Yang , Wenshuai Zhao , Yue Ma , Yuanpeng Tu , Pekka Marttinen , Joni Pajarinen

Imagine Mr. Bean stepping into Tom and Jerry--can we generate videos where characters interact naturally across different worlds? We study inter-character interaction in text-to-video generation, where the key challenge is to preserve each…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Tingting Liao , Chongjian Ge , Guangyi Liu , Hao Li , Yi Zhou

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

Co-speech gesture generation aims to synthesize realistic body movements that are semantically coherent with speech and faithful to a user-specified gestural style. Existing VQ-VAE based co-speech gesture generation methods improve…

图形学 · 计算机科学 2026-05-11 Junchuan Zhao , Qifan Liang , Ye Wang

Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Xuanchen Wang , Heng Wang , Dongnan Liu , Weidong Cai

Video generation models have made significant progress in generating realistic content, enabling applications in simulation, gaming, and film making. However, current generated videos still contain visual artifacts arising from 3D…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Duolikun Danier , Ge Gao , Steven McDonagh , Changjian Li , Hakan Bilen , Oisin Mac Aodha

High-quality video generation is crucial for many fields, including the film industry and autonomous driving. However, generating videos with spatiotemporal consistencies remains challenging. Current methods typically utilize attention…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Haotian Dong , Xin Wang , Di Lin , Yipeng Wu , Qin Chen , Ruonan Liu , Kairui Yang , Ping Li , Qing Guo

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Zheng Qin , Ruobing Zheng , Yabing Wang , Tianqi Li , Zixin Zhu , Sanping Zhou , Ming Yang , Le Wang

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Ruihao Xi , Xuekuan Wang , Yongcheng Li , Shuhua Li , Zichen Wang , Yiwei Wang , Feng Wei , Cairong Zhao

Video body-swapping aims to replace the body in an existing video with a new body from arbitrary sources, which has garnered more attention in recent years. Existing methods treat video body-swapping as a composite of multiple tasks instead…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Chengshu Zhao , Yunyang Ge , Xinhua Cheng , Bin Zhu , Yatian Pang , Bin Lin , Fan Yang , Feng Gao , Li Yuan

State-of-the-art Text-to-Video (T2V) diffusion models can generate visually impressive results, yet they still frequently fail to compose complex scenes or follow logical temporal instructions. In this paper, we argue that many errors,…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Mariam Hassan , Bastien Van Delft , Wuyang Li , Alexandre Alahi

Recent advances in image-to-video (I2V) generation have achieved remarkable progress in synthesizing high-quality, temporally coherent videos from static images. Among all the applications of I2V, human-centric video generation includes a…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Liao Shen , Wentao Jiang , Yiran Zhu , Jiahe Li , Tiezheng Ge , Zhiguo Cao , Bo Zheng