中文
相关论文

相关论文: Text-Guided Synthesis of Eulerian Cinemagraphs

200 篇论文

Text-to-image generative models are a new and powerful way to generate visual artwork. However, the open-ended nature of text as interaction is double-edged; while users can input anything and have access to an infinite range of…

人机交互 · 计算机科学 2023-09-29 Vivian Liu , Lydia B. Chilton

Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task of zero-shot text-to-video generation and propose a low-cost approach (without…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Levon Khachatryan , Andranik Movsisyan , Vahram Tadevosyan , Roberto Henschel , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

We introduce a novel and efficient approach for text-based video-to-video editing that eliminates the need for resource-intensive per-video-per-model finetuning. At the core of our approach is a synthetic paired video dataset tailored for…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Jiaxin Cheng , Tianjun Xiao , Tong He

Generating videos of complex human motions such as flips, cartwheels, and martial arts remains challenging for current video diffusion models. Text-only conditioning is temporally ambiguous for fine-grained motion control, while explicit…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Ashkan Taghipour , Morteza Ghahremani , Zinuo Li , Hamid Laga , Farid Boussaid , Mohammed Bennamoun

Scriptwriting has traditionally been text-centric, a modality that only partially conveys the produced audiovisual experience. A formative study with professional writers informed us that connecting textual and audiovisual modalities can…

人机交互 · 计算机科学 2026-04-09 Zhecheng Wang , Jiaju Ma , Eitan Grinspun , Tovi Grossman , Bryan Wang

Generating coherent long-form video sequences from discrete text prompts remains challenging due to difficulties in maintaining temporal coherence, semantic consistency, and scene-action continuity across segments. We propose a novel…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Taewon Kang , Divya Kothandaraman , Ming C. Lin

We propose a new paradigm to automatically generate training data with accurate labels at scale using the text-to-image synthesis frameworks (e.g., DALL-E, Stable Diffusion, etc.). The proposed approach1 decouples training data generation…

计算机视觉与模式识别 · 计算机科学 2023-09-13 Yunhao Ge , Jiashu Xu , Brian Nlong Zhao , Neel Joshi , Laurent Itti , Vibhav Vineet

Despite recent advances in image-to-video generation, better controllability and local animation are less explored. Most existing image-to-video methods are not locally aware and tend to move the entire scene. However, human artists may…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Yue Ma , Yingqing He , Hongfa Wang , Andong Wang , Chenyang Qi , Chengfei Cai , Xiu Li , Zhifeng Li , Heung-Yeung Shum , Wei Liu , Qifeng Chen

Controllable image synthesis models allow creation of diverse images based on text instructions or guidance from a reference image. Recently, denoising diffusion probabilistic models have been shown to generate more realistic imagery than…

计算机视觉与模式识别 · 计算机科学 2022-12-06 Xihui Liu , Dong Huk Park , Samaneh Azadi , Gong Zhang , Arman Chopikyan , Yuxiao Hu , Humphrey Shi , Anna Rohrbach , Trevor Darrell

State-of-the-art Video Scene Graph Generation (VSGG) systems provide structured visual understanding but operate as closed, feed-forward pipelines with no ability to incorporate human guidance. In contrast, promptable segmentation models…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Raphael Ruschel , Hardikkumar Prajapati , Awsafur Rahman , B. S. Manjunath

Existing image-to-video generation methods often produce physically implausible motions and lack precise control over object dynamics. While prior approaches have incorporated physics simulators, they remain confined to 2D planar motions…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Tianyidan Xie , Zhentao Huang , Mingjie Wang , Xin Huang , Jun Zhou , Minglun Gong , Zili Yi

Image animation is a key task in computer vision which aims to generate dynamic visual content from static image. Recent image animation methods employ neural based rendering technique to generate realistic animations. Despite these…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Zuozhuo Dai , Zhenghao Zhang , Yao Yao , Bingxue Qiu , Siyu Zhu , Long Qin , Weizhi Wang

This paper introduces a method for realistic kinetic typography that generates user-preferred animatable 'text content'. We draw on recent advances in guided video diffusion models to achieve visually-pleasing text appearances. To do this,…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Seonmi Park , Inhwan Bae , Seunghyun Shin , Hae-Gon Jeon

Text-to-Image (T2I) models have transformed visual content creation, producing highly realistic images from natural language prompts. However, concerns persist around their potential to replicate and magnify existing societal biases. To…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Sedat Porikli , Vedat Porikli

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

计算机视觉与模式识别 · 计算机科学 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu

We study the problem of synthesizing a number of likely future frames from a single input image. In contrast to traditional methods that have tackled this problem in a deterministic or non-parametric way, we propose to model future frames…

计算机视觉与模式识别 · 计算机科学 2019-08-13 Tianfan Xue , Jiajun Wu , Katherine L. Bouman , William T. Freeman

Synthesizing images from a given text description involves engaging two types of information: the content, which includes information explicitly described in the text (e.g., color, composition, etc.), and the style, which is usually not…

计算机视觉与模式识别 · 计算机科学 2019-08-16 Qicheng Lao , Mohammad Havaei , Ahmad Pesaranghader , Francis Dutil , Lisa Di Jorio , Thomas Fevens

Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single-image-to-3D…

图形学 · 计算机科学 2026-05-21 Xin Zhang , Yabo Chen , Yijie Fang , Wanying Qu , Haibin Huang , Chi Zhang , Feng Xu , Xuelong Li

The techniques for 3D indoor scene capturing are widely used, but the meshes produced leave much to be desired. In this paper, we propose "RoomDreamer", which leverages powerful natural language to synthesize a new room with a different…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Liangchen Song , Liangliang Cao , Hongyu Xu , Kai Kang , Feng Tang , Junsong Yuan , Yang Zhao

Recent works on language-guided image manipulation have shown great power of language in providing rich semantics, especially for face images. However, the other natural information, motions, in language is less explored. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-07-04 Tiankai Hang , Huan Yang , Bei Liu , Jianlong Fu , Xin Geng , Baining Guo