中文
相关论文

相关论文: Diff-BGM: A Diffusion Model for Video Background M…

200 篇论文

The recent surge in popularity of diffusion models for image generation has brought new attention to the potential of these models in other areas of media generation. One area that has yet to be fully explored is the application of…

声音 · 计算机科学 2023-02-01 Flavio Schneider

AI-generated content has attracted lots of attention recently, but photo-realistic video synthesis is still challenging. Although many attempts using GANs and autoregressive models have been made in this area, the visual quality and length…

计算机视觉与模式识别 · 计算机科学 2023-03-21 Yingqing He , Tianyu Yang , Yong Zhang , Ying Shan , Qifeng Chen

Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs. However, existing conditioning methods often require additional training and…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Ofir Abramovich , Nadav Z. Cohen , Adi Rosenthal , Ariel Shamir

We introduce Diff4Splat, a feed-forward method that synthesizes controllable and explicit 4D scenes from a single image. Our approach unifies the generative priors of video diffusion models with geometry and motion constraints learned from…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Panwang Pan , Chenguo Lin , Jingjing Zhao , Chenxin Li , Yuchen Lin , Haopeng Li , Honglei Yan , Kairun Wen , Yunlong Lin , Yixuan Yuan , Yadong Mu

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled,…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Ziqi Cai , Taoyu Yang , Zheng Chang , Si Li , Han Jiang , Shuchen Weng , Boxin Shi

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

Recently, the witness of the rapidly growing popularity of short videos on different Internet platforms has intensified the need for a background music (BGM) retrieval system. However, existing video-music retrieval methods only based on…

信息检索 · 计算机科学 2021-08-04 Tingtian Li , Zixun Sun , Haoruo Zhang , Jin Li , Ziming Wu , Hui Zhan , Yipeng Yu , Hengcan Shi

Incorporating a temporal dimension into pretrained image diffusion models for video generation is a prevalent approach. However, this method is computationally demanding and necessitates large-scale video datasets. More critically, the…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Dengsheng Chen , Jie Hu , Xiaoming Wei , Enhua Wu

Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Huanghao Yin , Shenkun Xu , Kanle Shi , Junhai Yong , Bin Wang

We propose a generative model that, given a coarsely edited image, synthesizes a photorealistic output that follows the prescribed layout. Our method transfers fine details from the original image and preserve the identity of its parts.…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Hadi Alzayer , Zhihao Xia , Xuaner Zhang , Eli Shechtman , Jia-Bin Huang , Michael Gharbi

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

机器学习 · 计算机科学 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

Previous studies on music style transfer have mainly focused on one-to-one style conversion, which is relatively limited. When considering the conversion between multiple styles, previous methods required designing multiple modes to…

声音 · 计算机科学 2024-04-24 Hong Huang , Yuyi Wang , Luyao Li , Jun Lin

Text-driven image and video diffusion models have recently achieved unprecedented generation realism. While diffusion models have been successfully applied for image editing, very few works have done so for video editing. We present the…

计算机视觉与模式识别 · 计算机科学 2023-02-03 Eyal Molad , Eliahu Horwitz , Dani Valevski , Alex Rav Acha , Yossi Matias , Yael Pritch , Yaniv Leviathan , Yedid Hoshen

Most existing video diffusion models (VDMs) are limited to mere text conditions. Thereby, they are usually lacking in control over visual appearance and geometry structure of the generated videos. This work presents Moonshot, a new video…

计算机视觉与模式识别 · 计算机科学 2024-01-04 David Junhao Zhang , Dongxu Li , Hung Le , Mike Zheng Shou , Caiming Xiong , Doyen Sahoo

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech and acoustic events.…

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and…

计算机视觉与模式识别 · 计算机科学 2026-02-26 JiaKui Hu , Jialun Liu , Liying Yang , Xinliang Zhang , Kaiwen Li , Shuang Zeng , Yuanwei Li , Haibin Huang , Chi Zhang , Yanye Lu

In this paper, we propose Scene Splatter, a momentum-based paradigm for video diffusion to generate generic scenes from single image. Existing methods, which employ video generation models to synthesize novel views, suffer from limited…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Shengjun Zhang , Jinzhao Li , Xin Fei , Hao Liu , Yueqi Duan

Generating the motion of orchestral conductors from a given piece of symphony music is a challenging task since it requires a model to learn semantic music features and capture the underlying distribution of real conducting motion. Prior…

音频与语音处理 · 电气工程与系统科学 2023-11-14 Zhuoran Zhao , Jinbin Bai , Delong Chen , Debang Wang , Yubo Pan

AI-generated video has advanced rapidly and poses serious challenges to content security and forensic analysis. Existing detectors rely mainly on pixel-level visual cues and generalize poorly to unseen generators. We propose DBINDS, a…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Yanlin Wu , Xiaogang Yuan , Dezhi An

Music profoundly enhances video production by improving quality, engagement, and emotional resonance, sparking growing interest in video-to-music generation. Despite recent advances, existing approaches remain limited in specific scenarios…

多媒体 · 计算机科学 2025-04-11 Xiaohao Liu , Teng Tu , Yunshan Ma , Tat-Seng Chua