English
Related papers

Related papers: CogCartoon: Towards Practical Story Visualization

200 papers

Diffusion models are widely used for image editing tasks. Existing editing methods often design a representation manipulation procedure by curating an edit direction in the text embedding or score space. However, such a procedure faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Jinqi Luo , Tianjiao Ding , Kwan Ho Ryan Chan , Hancheng Min , Chris Callison-Burch , René Vidal

Animating a still image offers an engaging visual experience. Traditional image animation techniques mainly focus on animating natural scenes with stochastic dynamics (e.g. clouds and fluid) or domain-specific motions (e.g. human hair or…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Jinbo Xing , Menghan Xia , Yong Zhang , Haoxin Chen , Wangbo Yu , Hanyuan Liu , Xintao Wang , Tien-Tsin Wong , Ying Shan

Image-based virtual try-on is an increasingly popular and important task to generate realistic try-on images of the specific person. Recent methods model virtual try-on as image mask-inpaint task, which requires masking the person image and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Xuanpu Zhang , Dan Song , Pengxin Zhan , Tianyu Chang , Jianhao Zeng , Qingguo Chen , Weihua Luo , Anan Liu

We present a method to create storytelling visualization with time series data. Many personal decisions nowadays rely on access to dynamic data regularly, as we have seen during the COVID-19 pandemic. It is thus desirable to construct…

Human-Computer Interaction · Computer Science 2024-02-06 Saiful Khan , Scott Jones , Benjamin Bach , Jaehoon Cha , Min Chen , Julie Meikle , Jonathan C Roberts , Jeyan Thiyagalingam , Jo Wood , Panagiotis D. Ritsos

Cotraining with demonstration data generated both in simulation and on real hardware has emerged as a promising recipe for scaling imitation learning in robotics. This work seeks to elucidate basic principles of this sim-and-real cotraining…

Robotics · Computer Science 2025-08-07 Adam Wei , Abhinav Agarwal , Boyuan Chen , Rohan Bosworth , Nicholas Pfaff , Russ Tedrake

Artistic image stylization aims to render the content provided by text or image with the target style, where content and style decoupling is the key to achieve satisfactory results. However, current methods for content and style…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Ma Zhuoqi , Zhang Yixuan , You Zejun , Tian Long , Liu Xiyang

Despite raw driving videos contain richer information on facial expressions than intermediate representations such as landmarks in the field of portrait animation, they are seldom the subject of research. This is due to two challenges…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Shurong Yang , Huadong Li , Juhao Wu , Minhao Jing , Linze Li , Renhe Ji , Jiajun Liang , Haoqiang Fan

We present StyleClone, a method for training image-to-image translation networks to stylize faces in a specific style, even with limited style images. Our approach leverages textual inversion and diffusion-based guided image generation to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Neeraj Matiyali , Siddharth Srivastava , Gaurav Sharma

Accurate Story visualization requires several necessary elements, such as identity consistency across frames, the alignment between plain text and visual content, and a reasonable layout of objects in images. Most previous works endeavor to…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Yuan Gong , Youxin Pang , Xiaodong Cun , Menghan Xia , Yingqing He , Haoxin Chen , Longyue Wang , Yong Zhang , Xintao Wang , Ying Shan , Yujiu Yang

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Given a single portrait…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Jiapeng Tang , Kai Li , Chengxiang Yin , Liuhao Ge , Fei Jiang , Jiu Xu , Matthias Nießner , Christian Häne , Timur Bagautdinov , Egor Zakharov , Peihong Guo

Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Kiymet Akdemir , Jing Shi , Kushal Kafle , Brian Price , Pinar Yanardag

Recent diffusion-based text-to-image customization methods have achieved significant success in understanding concrete concepts to control generation processes, such as styles and shapes. However, few efforts dive into the realistic yet…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Fan Wu , Cheng Chen , Zhoujie Fu , Jiacheng Wei , Yi Xu , Deheng Ye , Guosheng Lin

Recent advances in AI has made automated analysis of complex media content at scale possible while generating actionable insights regarding character representation along such dimensions as gender and age. Past works focused on quantifying…

Human-Computer Interaction · Computer Science 2025-08-28 Evdoxia Taka , Debadyuti Bhattacharya , Joanne Garde-Hansen , Sanjay Sharma , Tanaya Guha

Visual storytelling is the task of generating stories based on a sequence of images. Inspired by the recent works in neural generation focusing on controlling the form of text, this paper explores the idea of generating these stories in…

Computation and Language · Computer Science 2019-06-18 Shrimai Prabhumoye , Khyathi Raghavi Chandu , Ruslan Salakhutdinov , Alan W Black

Diffusion transformers have shown significant effectiveness in both image and video synthesis at the expense of huge computation costs. To address this problem, feature caching methods have been introduced to accelerate diffusion…

Machine Learning · Computer Science 2025-02-20 Chang Zou , Xuyang Liu , Ting Liu , Siteng Huang , Linfeng Zhang

Characters are integral to long-form narratives, but are poorly understood by existing story analysis and generation systems. While prior work has simplified characters via graph-based methods and brief character descriptions, we aim to…

Computation and Language · Computer Science 2024-12-10 Alexander Gurung , Mirella Lapata

We present a system to convert any set of images (e.g., a video clip or a photo album) into a storyboard. We aim to create multiple pleasing graphic representations of the content at interactive rates, so the user can explore and find the…

Diffusion models trained on large-scale datasets have achieved remarkable progress in image synthesis. However, due to the randomness in the diffusion process, they often struggle with handling diverse low-level tasks that require details…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Yuhao Liu , Zhanghan Ke , Fang Liu , Nanxuan Zhao , Rynson W. H. Lau

Personalized text-to-image models allow users to generate varied styles of images (specified with a sentence) for an object (specified with a set of reference images). While remarkable results have been achieved using diffusion-based…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Fanyue Wei , Wei Zeng , Zhenyang Li , Dawei Yin , Lixin Duan , Wen Li

Storytelling is an open-ended task that entails creative thinking and requires a constant flow of ideas. Natural language generation (NLG) for storytelling is especially challenging because it requires the generated text to follow an…

Computation and Language · Computer Science 2021-09-17 Eden Bensaid , Mauro Martino , Benjamin Hoover , Hendrik Strobelt
‹ Prev 1 3 4 5 6 7 10 Next ›