中文
相关论文

相关论文: Video-Robin: Autoregressive Diffusion Planning for…

200 篇论文

We present a framework for learning to generate background music from video inputs. Unlike existing works that rely on symbolic musical annotations, which are limited in quantity and diversity, our method leverages large-scale web videos…

多媒体 · 计算机科学 2024-09-12 Yan-Bo Lin , Yu Tian , Linjie Yang , Gedas Bertasius , Heng Wang

We introduce MusicInfuser, an approach that aligns pre-trained text-to-video diffusion models to generate high-quality dance videos synchronized with specified music tracks. Rather than training a multimodal audio-video or audio-motion…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Susung Hong , Ira Kemelmacher-Shlizerman , Brian Curless , Steven M. Seitz

Video-to-music generation demands both a temporally localized high-quality listening experience and globally aligned video-acoustic signatures. While recent music generation models excel at the former through advanced audio codecs, the…

Video-to-audio (V2A) generation aims to synthesize realistic and semantically aligned audio from silent videos, with potential applications in video editing, Foley sound design, and assistive multimedia. Although the excellent results,…

Breakthroughs in text-to-music generation models are transforming the creative landscape, equipping musicians with innovative tools for composition and experimentation like never before. However, controlling the generation process to…

声音 · 计算机科学 2025-06-19 Teysir Baoueb , Xiaoyu Bie , Xi Wang , Gaël Richard

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have limited generation…

声音 · 计算机科学 2023-07-03 Simian Luo , Chuanhao Yan , Chenxu Hu , Hang Zhao

Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task of zero-shot text-to-video generation and propose a low-cost approach (without…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Levon Khachatryan , Andranik Movsisyan , Vahram Tadevosyan , Roberto Henschel , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Recent advancements in diffusion-based video generation have produced impressive and high-fidelity short videos. To extend these successes to generate coherent long videos, most video diffusion models (VDMs) generate videos in an…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Tianle Cheng , Zeyan Zhang , Kaifeng Gao , Jun Xiao

Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This limitation frequently leads to visual hallucinations and…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Lun Huang , You Xie , Hongyi Xu , Tianpei Gu , Chenxu Zhang , Guoxian Song , Zenan Li , Xiaochen Zhao , Linjie Luo , Guillermo Sapiro

We present the Melody-Guided Music Generation (MG2) model, a novel approach using melody to guide the text-to-music generation that, despite a simple method and limited resources, achieves excellent performance. Specifically, we first align…

声音 · 计算机科学 2024-12-31 Shaopeng Wei , Manzhen Wei , Haoyu Wang , Yu Zhao , Gang Kou

Text-to-video diffusion models have enabled high-quality video synthesis, yet often fail to generate temporally coherent and physically plausible motion. A key reason is the models' insufficient understanding of complex motions that natural…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Aritra Bhowmik , Denis Korzhenkov , Cees G. M. Snoek , Amirhossein Habibian , Mohsen Ghafoorian

Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Ran Galun , Sagie Benaim

With the advance of diffusion models, today's video generation has achieved impressive quality. To extend the generation length and facilitate real-world applications, a majority of video diffusion models (VDMs) generate videos in an…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Kaifeng Gao , Jiaxin Shi , Hanwang Zhang , Chunping Wang , Jun Xiao , Long Chen

Dance-to-music (D2M) generation aims to automatically compose music that is rhythmically and temporally aligned with dance movements. Existing methods typically rely on coarse rhythm embeddings, such as global motion features or binarized…

声音 · 计算机科学 2026-03-03 Jinting Wang , Chenxing Li , Li Liu

Numerous studies in the field of music generation have demonstrated impressive performance, yet virtually no models are able to directly generate music to match accompanying videos. In this work, we develop a generative music AI framework,…

声音 · 计算机科学 2024-06-03 Jaeyong Kang , Soujanya Poria , Dorien Herremans

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on.…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Lei Wang , YuXin Song , Ge Wu , Haocheng Feng , Hang Zhou , Jingdong Wang , Yaxing Wang , jian Yang

Latent Diffusion Models (LDMs) enable high-quality image synthesis while avoiding excessive compute demands by training a diffusion model in a compressed lower-dimensional latent space. Here, we apply the LDM paradigm to high-resolution…

计算机视觉与模式识别 · 计算机科学 2023-12-29 Andreas Blattmann , Robin Rombach , Huan Ling , Tim Dockhorn , Seung Wook Kim , Sanja Fidler , Karsten Kreis

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Existing video generation models struggle to follow complex text prompts and synthesize multiple objects, raising the need for additional grounding input for improved controllability. In this work, we propose to decompose videos into visual…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Weixi Feng , Chao Liu , Sifei Liu , William Yang Wang , Arash Vahdat , Weili Nie

Current open-source diffusion models struggle to generate stable and synchronized audio-visual content, particularly in scenarios demanding complex semantic reasoning. The root cause is that existing methods rely on coarse text embeddings…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Shuyuan Tu , Qi Tian , Zihan Yang , Yue Wu , Xintong Han , Weijie Kong , Jiangfeng Xiong , Jian-Wei Zhang , Zhao Zhong , Liefeng Bo , Zuxuan Wu , Yu-Gang Jiang