中文
相关论文

相关论文: TANGO: Co-Speech Gesture Video Reenactment with Hi…

200 篇论文

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

计算机视觉与模式识别 · 计算机科学 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text datasets. Recent…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Xiefan Guo , Jinlin Liu , Miaomiao Cui , Liefeng Bo , Di Huang

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Yangyang Xu , Shengfeng He , Kwan-Yee K. Wong , Ping Luo

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of…

计算机视觉与模式识别 · 计算机科学 2026-02-02 Renjie Lu , Xulong Zhang , Xiaoyang Qu , Jianzong Wang , Shangfei Wang

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Ryan Burgert , Yuancheng Xu , Wenqi Xian , Oliver Pilarski , Pascal Clausen , Mingming He , Li Ma , Yitong Deng , Lingxiao Li , Mohsen Mousavi , Michael Ryoo , Paul Debevec , Ning Yu

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often fail to provide an…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Thomas Ressler-Antal , Frank Fundel , Malek Ben Alaya , Stefan Andreas Baumann , Felix Krause , Ming Gui , Björn Ommer

Face videos accompanied by audio have become integral to our daily lives, while they often suffer from complex degradations. Most face video restoration methods neglect the intrinsic correlations between the visual and audio features,…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Yuqin Cao , Yixuan Gao , Wei Sun , Xiaohong Liu , Yulun Zhang , Xiongkuo Min

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding…

图形学 · 计算机科学 2025-03-31 Yong Xie , Yunlian Sun , Hongwen Zhang , Yebin Liu , Jinhui Tang

Significant advancements have been achieved in the realm of large-scale pre-trained text-to-video Diffusion Models (VDMs). However, previous methods either rely solely on pixel-based VDMs, which come with high computational costs, or on…

计算机视觉与模式识别 · 计算机科学 2025-06-02 David Junhao Zhang , Jay Zhangjie Wu , Jia-Wei Liu , Rui Zhao , Lingmin Ran , Yuchao Gu , Difei Gao , Mike Zheng Shou

Modern text-to-video (T2V) diffusion models can synthesize visually compelling clips, yet they remain brittle at fine-scale structure: even state-of-the-art generators often produce distorted faces and hands, warped backgrounds, and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Tejas Panambur , Ishan Rajendrakumar Dave , Chongjian Ge , Ersin Yumer , Xue Bai

In this work, we build a simple but strong baseline for sounding video generation. Given base diffusion models for audio and video, we integrate them with additional modules into a single model and train it to make the model jointly…

机器学习 · 计算机科学 2025-04-10 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Bing Li , Cheng Zheng , Wenxuan Zhu , Jinjie Mai , Biao Zhang , Peter Wonka , Bernard Ghanem

Current two-stage TTS framework typically integrates an acoustic model with a vocoder -- the acoustic model predicts a low resolution intermediate representation such as Mel-spectrum while the vocoder generates waveform from the…

音频与语音处理 · 电气工程与系统科学 2021-06-23 Jian Cong , Shan Yang , Lei Xie , Dan Su

Text-to-video (T2V) generation technology holds potential to transform multiple domains such as education, marketing, entertainment, and assistive technologies for individuals with visual or reading comprehension challenges, by creating…

图形学 · 计算机科学 2025-10-07 Nilay Kumar , Priyansh Bhandari , G. Maragatham

Human video generation is becoming an increasingly important task with broad applications in graphics, entertainment, and embodied AI. Despite the rapid progress of video diffusion models (VDMs), their use for general-purpose human video…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Hyelin Nam , Hyojun Go , Byeongjun Park , Byung-Hoon Kim , Hyungjin Chung

Recently video diffusion models have emerged as expressive generative tools for high-quality video content creation readily available to general users. However, these models often do not offer precise control over camera poses for video…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Dejia Xu , Weili Nie , Chao Liu , Sifei Liu , Jan Kautz , Zhangyang Wang , Arash Vahdat

Recent studies have demonstrated the efficacy of integrating Group Relative Policy Optimization (GRPO) into flow matching models, particularly for text-to-image and text-to-video generation. However, we find that directly applying these…

计算机视觉与模式识别 · 计算机科学 2026-05-27 Jin Wang , Jianxiang Lu , Guangzheng Xu , Comi Chen , Haoyu Yang , Linqing Wang , Peng Chen , Mingtao Chen , Zhichao Hu , Longhuang Wu , Shuai Shao , Qinglin Lu , Ping Luo

In this paper, we present our framework for neural face/head reenactment whose goal is to transfer the 3D head orientation and expression of a target face to a source face. Previous methods focus on learning embedding networks for identity…

计算机视觉与模式识别 · 计算机科学 2024-02-07 Stella Bounareli , Christos Tzelepis , Vasileios Argyriou , Ioannis Patras , Georgios Tzimiropoulos

Co-Speech Gesture Video Generation aims to generate vivid speech videos from audio-driven still images, which is challenging due to the diversity of body parts in terms of motion amplitude, audio relevance, and detailed features. Relying…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Siyuan Wang , Jiawei Liu , Wei Wang , Yeying Jin , Jinsong Du , Zhi Han