中文
相关论文

相关论文: FilmComposer: LLM-Driven Music Production for Sile…

200 篇论文

Controllable video synthesis is a central challenge in computer vision, yet current models struggle with fine grained control beyond textual prompts, particularly for cinematic attributes like camera trajectory and genre. Existing datasets…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Zahra Dehghanian , Morteza Abolghasemi , Hamid Beigy , Hamid R. Rabiee

Text-to-video generation models have shown significant progress in the recent years. However, they still struggle with generating complex dynamic scenes based on compositional text prompts, such as attribute binding for multiple objects,…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Kaiyi Huang , Yukun Huang , Xuefei Ning , Zinan Lin , Yu Wang , Xihui Liu

Music is essential when editing videos, but selecting music manually is difficult and time-consuming. Thus, we seek to automatically generate background music tracks given video input. This is a challenging task since it requires…

计算机视觉与模式识别 · 计算机科学 2023-08-07 Le Zhuo , Zhaokai Wang , Baisen Wang , Yue Liao , Chenxi Bao , Stanley Peng , Songhao Han , Aixi Zhang , Fei Fang , Si Liu

Virtual film production requires intricate decision-making processes, including scriptwriting, virtual cinematography, and precise actor positioning and actions. Motivated by recent advances in automated decision-making with language…

计算与语言 · 计算机科学 2025-01-23 Zhenran Xu , Longyue Wang , Jifang Wang , Zhouyi Li , Senbao Shi , Xue Yang , Yiyu Wang , Baotian Hu , Jun Yu , Min Zhang

Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this…

音频与语音处理 · 电气工程与系统科学 2026-03-24 Jianyi Chen , Rongxiu Zhong , Shilei Zhang , Kun Qian , Jinglei Liu , Yike Guo , Wei Xue

Content creators often draw inspiration from multiple visual sources, combining distinct elements to craft new compositions. Modern computational approaches now aim to emulate this fundamental creative process. Although recent diffusion…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Sara Dorfman , Dana Cohen-Bar , Rinon Gal , Daniel Cohen-Or

Compositional video generation aims to synthesize multiple instances with diverse appearance and motion. However, current approaches mainly focus on binding semantics, neglecting to understand diverse motion categories specified in prompts.…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zixuan Wang , Ziqin Zhou , Feng Chen , Duo Peng , Yixin Hu , Changsheng Li , Yinjie Lei

Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We…

In this paper, we present MovieFactory, a powerful framework to generate cinematic-picture (3072$\times$1280), film-style (multi-scene), and multi-modality (sounding) movies on the demand of natural languages. As the first fully automated…

计算机视觉与模式识别 · 计算机科学 2023-06-13 Junchen Zhu , Huan Yang , Huiguo He , Wenjing Wang , Zixi Tuo , Wen-Huang Cheng , Lianli Gao , Jingkuan Song , Jianlong Fu

Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the motion. In this work,…

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

We present FashionComposer for compositional fashion image generation. Unlike previous methods, FashionComposer is highly flexible. It takes multi-modal input (i.e., text prompt, parametric human model, garment image, and face image) and…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Sihui Ji , Yiyang Wang , Xi Chen , Xiaogang Xu , Hao Luo , Hengshuang Zhao

Recent large vision-language models (LVLMs) for video understanding are primarily fine-tuned with various videos scraped from online platforms. Existing datasets, such as ActivityNet, require considerable human labor for structuring and…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Zhende Song , Chenchen Wang , Jiamu Sheng , Chi Zhang , Shengji Tang , Jiayuan Fan , Tao Chen

Traditional animation generation methods depend on training generative models with human-labelled data, entailing a sophisticated multi-stage pipeline that demands substantial human effort and incurs high training costs. Due to limited…

计算与语言 · 计算机科学 2024-08-20 Yunxin Li , Haoyuan Shi , Baotian Hu , Longyue Wang , Jiashun Zhu , Jinyi Xu , Zhen Zhao , Min Zhang

Most digital music tools emphasize precision and control, but often lack support for tactile, improvisational workflows grounded in environmental interaction. Lumia addresses this by enabling users to "compose through looking"--transforming…

人机交互 · 计算机科学 2025-12-22 Chung-Ta Huang , Connie Cheng , Vealy Lai

We consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods…

声音 · 计算机科学 2023-05-31 Jiashuo Yu , Yaohui Wang , Xinyuan Chen , Xiao Sun , Yu Qiao

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Han Lin , Abhay Zala , Jaemin Cho , Mohit Bansal

Despite their impressive visual fidelity, existing personalized image generators lack interactive control over spatial composition and scale poorly to multiple humans. To address these limitations, we present LayerComposer, an interactive…

In this demo, we present VirtualConductor, a system that can generate conducting video from any given music and a single user's image. First, a large-scale conductor motion dataset is collected and constructed. Then, we propose Audio Motion…

计算机视觉与模式识别 · 计算机科学 2021-08-11 Delong Chen , Fan Liu , Zewen Li , Feng Xu

While recent generative models advance pixel-space video synthesis, they remain limited in producing professional educational videos, which demand disciplinary knowledge, precise visual structures, and coherent transitions, limiting their…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yanzhe Chen , Kevin Qinghong Lin , Mike Zheng Shou