中文
相关论文

相关论文: Text2Video: Text-driven Talking-head Video Synthes…

200 篇论文

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

Recent advances in the diffusion models have significantly improved text-to-image generation. However, generating videos from text is a more challenging task than generating images from text, due to the much larger dataset and higher…

计算机视觉与模式识别 · 计算机科学 2024-12-31 Taegyeong Lee , Soyeong Kwon , Taehwan Kim

We describe an end-to-end speech synthesis system that uses generative adversarial training. We train our Vocoder for raw phoneme-to-audio conversion, using explicit phonetic, pitch and duration modeling. We experiment with several…

机器学习 · 计算机科学 2023-10-17 Tiberiu Boros , Stefan Daniel Dumitrescu , Ionut Mironica , Radu Chivereanu

While recent research has made significant progress in speech-driven talking face generation, the quality of the generated video still lags behind that of real recordings. One reason for this is the use of handcrafted intermediate…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Chenpeng Du , Qi Chen , Tianyu He , Xu Tan , Xie Chen , Kai Yu , Sheng Zhao , Jiang Bian

Most of the existing works in video synthesis focus on generating videos using adversarial learning. Despite their success, these methods often require input reference frame or fail to generate diverse videos from the given data…

图像与视频处理 · 电气工程与系统科学 2020-04-21 Abhishek Aich , Akash Gupta , Rameswar Panda , Rakib Hyder , M. Salman Asif , Amit K. Roy-Chowdhury

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

In this work, we address the task of unconditional head motion generation to animate still human faces in a low-dimensional semantic space from a single reference pose. Different from traditional audio-conditioned talking head generation…

计算机视觉与模式识别 · 计算机科学 2023-04-18 Louis Airale , Xavier Alameda-Pineda , Stéphane Lathuilière , Dominique Vaufreydaz

Text-guided 3D face synthesis has achieved remarkable results by leveraging text-to-image (T2I) diffusion models. However, most existing works focus solely on the direct generation, ignoring the editing, restricting them from synthesizing…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Yunjie Wu , Yapeng Meng , Zhipeng Hu , Lincheng Li , Haoqian Wu , Kun Zhou , Weiwei Xu , Xin Yu

Video generation requires synthesizing consistent and persistent frames with dynamic content over time. This work investigates modeling the temporal relations for composing video with arbitrary length, from a few frames to even infinite,…

计算机视觉与模式识别 · 计算机科学 2022-12-15 Qihang Zhang , Ceyuan Yang , Yujun Shen , Yinghao Xu , Bolei Zhou

Due to the rapid emergence of short videos and the requirement for content understanding and creation, the video captioning task has received increasing attention in recent years. In this paper, we convert traditional video captioning task…

计算机视觉与模式识别 · 计算机科学 2021-03-10 Ziqi Zhang , Zhongang Qi , Chunfeng Yuan , Ying Shan , Bing Li , Ying Deng , Weiming Hu

Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Jiayi He , Xu Wang , Shengeng Tang , Yaxiong Wang , Lechao Cheng , Dan Guo

We propose a novel learned deep prior of body motion for 3D hand shape synthesis and estimation in the domain of conversational gestures. Our model builds upon the insight that body motion and hand gestures are strongly correlated in…

计算机视觉与模式识别 · 计算机科学 2021-04-08 Evonne Ng , Shiry Ginosar , Trevor Darrell , Hanbyul Joo

This paper explores multi-modal controllable Text-to-Speech Synthesis (TTS) where the voice can be generated from face image, and the characteristics of output speech (e.g., pace, noise level, distance, tone, place) can be controllable with…

音频与语音处理 · 电气工程与系统科学 2025-05-27 Minsu Kim , Pingchuan Ma , Honglie Chen , Stavros Petridis , Maja Pantic

The hand plays a pivotal role in human ability to grasp and manipulate objects and controllable grasp synthesis is the key for successfully performing downstream tasks. Existing methods that use human intention or task-level language as…

人工智能 · 计算机科学 2024-04-24 Xiaoyun Chang , Yi Sun

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task,…

计算机视觉与模式识别 · 计算机科学 2022-08-08 Chuan Guo , Xinxin Zuo , Sen Wang , Li Cheng

Faces generated using generative adversarial networks (GANs) have reached unprecedented realism. These faces, also known as "Deep Fakes", appear as realistic photographs with very little pixel-level distortions. While some work has enabled…

计算机视觉与模式识别 · 计算机科学 2023-12-14 Manan Oza , Sukalpa Chanda , David Doermann

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Generating controllable videos conforming to user intentions is an appealing yet challenging topic in computer vision. To enable maneuverable control in line with user intentions, a novel video generation task, named Text-Image-to-Video…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Yaosi Hu , Chong Luo , Zhenzhong Chen

State-of-the-art Text-to-Video (T2V) diffusion models can generate visually impressive results, yet they still frequently fail to compose complex scenes or follow logical temporal instructions. In this paper, we argue that many errors,…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Mariam Hassan , Bastien Van Delft , Wuyang Li , Alexandre Alahi

We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Jiaxin Cheng , Tianjun Xiao , Tong He