中文
相关论文

相关论文: MovieFactory: Automatic Movie Creation from Text u…

200 篇论文

This paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term…

计算机视觉与模式识别 · 计算机科学 2017-11-17 Tanya Marwah , Gaurav Mittal , Vineeth N. Balasubramanian

Recent approaches have demonstrated the promise of using diffusion models to generate interactive and explorable worlds. However, most of these methods face critical challenges such as excessively large parameter sizes, reliance on lengthy…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Xiaofeng Mao , Zhen Li , Chuanhao Li , Xiaojie Xu , Kaining Ying , Tong He , Jiangmiao Pang , Yu Qiao , Kaipeng Zhang

We present MagicInfinite, a novel diffusion Transformer (DiT) framework that overcomes traditional portrait animation limitations, delivering high-fidelity results across diverse character types-realistic humans, full-body figures, and…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Hongwei Yi , Tian Ye , Shitong Shao , Xuancheng Yang , Jiantong Zhao , Hanzhong Guo , Terrance Wang , Qingyu Yin , Zeke Xie , Lei Zhu , Wei Li , Michael Lingelbach , Daquan Zhou

With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level…

人工智能 · 计算机科学 2025-04-22 Haotong Yang , Qingyuan Zheng , Yunjian Gao , Yongkun Yang , Yangbo He , Zhouchen Lin , Muhan Zhang

Our objective is the automatic generation of Audio Descriptions (ADs) for edited video material, such as movies and TV series. To achieve this, we propose a two-stage framework that leverages "shots" as the fundamental units of video…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Junyu Xie , Tengda Han , Max Bain , Arsha Nagrani , Eshika Khandelwal , Gül Varol , Weidi Xie , Andrew Zisserman

A large number of annotated training images is crucial for training successful scene text recognition models. However, collecting sufficient datasets can be a labor-intensive and costly process, particularly for low-resource languages. To…

计算机视觉与模式识别 · 计算机科学 2023-06-28 Yangchen Xie , Xinyuan Chen , Hongjian Zhan , Palaiahankote Shivakum , Bing Yin , Cong Liu , Yue Lu

Movie screenplay summarization is challenging, as it requires an understanding of long input contexts and various elements unique to movies. Large language models have shown significant advancements in document summarization, but they often…

计算与语言 · 计算机科学 2024-08-13 Rohit Saxena , Frank Keller

Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zun Wang , Jialu Li , Han Lin , Jaehong Yoon , Mohit Bansal

Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Xindi Yang , Baolu Li , Yiming Zhang , Zhenfei Yin , Lei Bai , Liqian Ma , Zhiyong Wang , Jianfei Cai , Tien-Tsin Wong , Huchuan Lu , Xu Jia

Generating long, cohesive video stories with consistent characters is a significant challenge for current text-to-video AI. We introduce a method that approaches video generation in a filmmaker-like manner. Instead of creating a video in…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Chayan Jain , Rishant Sharma , Archit Garg , Ishan Bhanuka , Pratik Narang , Dhruv Kumar

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see (the video's image…

Generating high-quality videos from textual descriptions poses challenges in maintaining temporal coherence and control over subject motion. We propose VAST (Video As Storyboard from Text), a two-stage framework to address these challenges…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Chi Zhang , Yuanzhi Liang , Xi Qiu , Fangqiu Yi , Xuelong Li

The dominant text generation models compose the output by sequentially selecting words from a fixed vocabulary. In this paper, we formulate text generation as progressively copying text segments (e.g., words or phrases) from an existing…

计算与语言 · 计算机科学 2023-07-17 Tian Lan , Deng Cai , Yan Wang , Heyan Huang , Xian-Ling Mao

Generating realistic and controllable human motions, particularly those involving rich multi-character interactions, remains a significant challenge due to data scarcity and the complexities of modeling inter-personal dynamics. To address…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Ruihao Xi , Xuekuan Wang , Yongcheng Li , Shuhua Li , Zichen Wang , Yiwei Wang , Feng Wei , Cairong Zhao

Generating long videos that can show complex stories, like movie scenes from scripts, has great promise and offers much more than short clips. However, current methods that use autoregression with diffusion models often struggle because…

计算机视觉与模式识别 · 计算机科学 2025-05-28 Guangcong Zheng , Jianlong Yuan , Bo Wang , Haoyang Huang , Guoqing Ma , Nan Duan

Advancements in language foundation models have primarily fueled the recent surge in artificial intelligence. In contrast, generative learning of non-textual modalities, especially videos, significantly trails behind language modeling. This…

计算机视觉与模式识别 · 计算机科学 2024-05-28 Lijun Yu

Recent advancements in large scale text-to-image models have opened new possibilities for guiding the creation of images through human-devised natural language. However, while prior literature has primarily focused on the generation of…

计算机视觉与模式识别 · 计算机科学 2023-02-09 Hyeonho Jeong , Gihyun Kwon , Jong Chul Ye

With the widespread popularity of internet celebrity marketing all over the world, short video production has gradually become a popular way of presenting products information. However, the traditional video production industry usually…

多媒体 · 计算机科学 2024-03-25 Juan Zhang , Jiahao Chen , Cheng Wang , Zhiwang Yu , Tangquan Qi , Can Liu , Di Wu

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

In this work, we implement music production for silent film clips using LLM-driven method. Given the strong professional demands of film music production, we propose the FilmComposer, simulating the actual workflows of professional…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Zhifeng Xie , Qile He , Youjia Zhu , Qiwei He , Mengtian Li
‹ 上一页 1 8 9 10 下一页 ›