中文
相关论文

相关论文: Sketching the Future (STF): Applying Conditional C…

200 篇论文

ControlNets are widely used for adding spatial control to text-to-image diffusion models with different conditions, such as depth maps, scribbles/sketches, and human poses. However, when it comes to controllable video generation,…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Han Lin , Jaemin Cho , Abhay Zala , Mohit Bansal

Recent advances in zero-shot text-to-speech (TTS) synthesis have achieved high-quality speech generation for unseen speakers, but most systems remain unsuitable for real-time applications because of their offline design. Current streaming…

声音 · 计算机科学 2025-06-17 Hui Wang , Yifan Yang , Shujie Liu , Jinyu Li , Lingwei Meng , Yanqing Liu , Jiaming Zhou , Haoqin Sun , Yan Lu , Yong Qin

Recent CLIP-guided 3D optimization methods, such as DreamFields and PureCLIPNeRF, have achieved impressive results in zero-shot text-to-3D synthesis. However, due to scratch training and random initialization without prior knowledge, these…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Jiale Xu , Xintao Wang , Weihao Cheng , Yan-Pei Cao , Ying Shan , Xiaohu Qie , Shenghua Gao

Specifying nuanced and compelling camera motion remains a significant hurdle for non-expert creators using generative tools, creating an "expressive gap" where generic text prompts fail to capture cinematic vision. This barrier limits…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Pooja Guhan , Divya Kothandaraman , Geonsun Lee , Tsung-Wei Huang , Guan-Ming Su , Dinesh Manocha

Multimedia generation approaches occupy a prominent place in artificial intelligence research. Text-to-image models achieved high-quality results over the last few years. However, video synthesis methods recently started to develop. This…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Vladimir Arkhipkin , Zein Shaheen , Viacheslav Vasilev , Elizaveta Dakhova , Andrey Kuznetsov , Denis Dimitrov

Recent advancements in large vision-language models have enabled highly expressive and diverse vector sketch generation. However, state-of-the-art methods rely on a time-consuming optimization process involving repeated feedback from a…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Ellie Arar , Yarden Frenkel , Daniel Cohen-Or , Ariel Shamir , Yael Vinker

Text-to-image generation has witnessed great progress, especially with the recent advancements in diffusion models. Since texts cannot provide detailed conditions like object appearance, reference images are usually leveraged for the…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Zhiqi Huang , Huixin Xiong , Haoyu Wang , Longguang Wang , Zhiheng Li

Recent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of…

计算机视觉与模式识别 · 计算机科学 2022-04-01 Yoad Tewel , Yoav Shalev , Idan Schwartz , Lior Wolf

Text-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., "a woman is drinking water."). Existing TI2V frameworks often require…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Haomiao Ni , Bernhard Egger , Suhas Lohit , Anoop Cherian , Ye Wang , Toshiaki Koike-Akino , Sharon X. Huang , Tim K. Marks

Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precision, while explicit trajectory conditioning imposes…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Songlin Yang , Zhe Wang , Xuyi Yang , Songchun Zhang , Xianghao Kong , Taiyi Wu , Xiaotong Zhao , Ran Zhang , Alan Zhao , Anyi Rao

We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. VoiceCraft employs a…

音频与语音处理 · 电气工程与系统科学 2024-06-17 Puyuan Peng , Po-Yao Huang , Shang-Wen Li , Abdelrahman Mohamed , David Harwath

Real-world videos often have complex dynamics; and methods for generating open-domain video descriptions should be sensitive to temporal structure and allow both input (sequence of frames) and output (sequence of words) of variable length.…

计算机视觉与模式识别 · 计算机科学 2015-10-20 Subhashini Venugopalan , Marcus Rohrbach , Jeff Donahue , Raymond Mooney , Trevor Darrell , Kate Saenko

For artistic applications, video generation requires fine-grained control over both performance and cinematography, i.e., the actor's motion and the camera trajectory. We present ActCam, a zero-shot method for video generation that jointly…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Omar El Khalifi , Thomas Rossi , Oscar Fossey , Thibault Fouque , Ulysse Mizrahi , Philip Torr , Ivan Laptev , Fabio Pizzati , Baptiste Bellot-Gurlet

We present MicroCinema, a straightforward yet effective framework for high-quality and coherent text-to-video generation. Unlike existing approaches that align text prompts with video directly, MicroCinema introduces a Divide-and-Conquer…

Automatically generating textual content with desired attributes is an ambitious task that people have pursued long. Existing works have made a series of progress in incorporating unimodal controls into language models (LMs), whereas how to…

计算与语言 · 计算机科学 2023-06-30 Haoqin Tu , Bowen Yang , Xianfeng Zhao

Recent text-to-video diffusion models have achieved impressive progress. In practice, users often desire the ability to control object motion and camera movement independently for customized video creation. However, current methods lack the…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Shiyuan Yang , Liang Hou , Haibin Huang , Chongyang Ma , Pengfei Wan , Di Zhang , Xiaodong Chen , Jing Liao

Diffusion models have demonstrated remarkable capabilities in text-to-image and text-to-video generation, opening up possibilities for video editing based on textual input. However, the computational cost associated with sequential sampling…

计算机视觉与模式识别 · 计算机科学 2024-11-20 Youyuan Zhang , Xuan Ju , James J. Clark

In the field of media production, video editing techniques play a pivotal role. Recent approaches have had great success at performing novel view image synthesis of static scenes. But adding temporal information adds an extra layer of…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Violeta Menéndez González , Andrew Gilbert , Graeme Phillipson , Stephen Jolly , Simon Hadfield

The development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos,…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Jiayi Gao , Zijin Yin , Changcheng Hua , Yuxin Peng , Kongming Liang , Zhanyu Ma , Jun Guo , Yang Liu

Recent advances in generative video models have enabled the creation of high-quality videos based on natural language prompts. However, these models frequently lack fine-grained temporal control, meaning they do not allow users to specify…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Shira Schiber , Ofir Lindenbaum , Idan Schwartz