English
Related papers

Related papers: Diff-V2M: A Hierarchical Conditional Diffusion Mod…

200 papers

Generating full-length, high-quality songs is challenging, as it requires maintaining long-term coherence both across text and music modalities and within the music modality itself. Existing non-autoregressive (NAR) frameworks, while…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-04 Yuepeng Jiang , Huakang Chen , Ziqian Ning , Jixun Yao , Zerui Han , Di Wu , Meng Meng , Jian Luan , Zhonghua Fu , Lei Xie

Diffusion-based text-to-video generation has witnessed impressive progress in the past year yet still falls behind text-to-image generation. One of the key reasons is the limited scale of publicly available data (e.g., 10M video-text pairs…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Xiang Wang , Shiwei Zhang , Hangjie Yuan , Zhiwu Qing , Biao Gong , Yingya Zhang , Yujun Shen , Changxin Gao , Nong Sang

Given a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as…

Computation and Language · Computer Science 2023-04-05 Gaoxiang Cong , Liang Li , Yuankai Qi , Zhengjun Zha , Qi Wu , Wenyu Wang , Bin Jiang , Ming-Hsuan Yang , Qingming Huang

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source foundational model…

Advances in diffusion-based video generation models, while significantly improving human animation, poses threats of misuse through the creation of fake videos from a specific person's photo and text prompts. Recent efforts have focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Duc Vu , Anh Nguyen , Chi Tran , Anh Tran

Diffusion models excel in noise-to-data generation tasks, providing a mapping from a Gaussian distribution to a more complex data distribution. However they struggle to model translations between complex distributions, limiting their…

Machine Learning · Computer Science 2026-03-27 Viacheslav Vasilev , Arseny Ivanov , Nikita Gushchin , Maria Kovaleva , Alexander Korotin

Text-to-vibration generation converts natural language into haptic feedback, enabling vibration-effect designers to get scenarios-fitted vibrations more efficiently, which shows great potentials in application fields such as metaverse,…

Human-Computer Interaction · Computer Science 2026-05-12 Jiahao Xiong , Fei Wang , Anran Xu , Pinzhi Huang , Tao Wen , Lijia Pan , Cai Chen

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this…

Sound · Computer Science 2026-03-18 Alejandro Paredes La Torre

Perceptual studies demonstrate that conditional diffusion models excel at reconstructing video content aligned with human visual perception. Building on this insight, we propose a video compression framework that leverages conditional…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Fangqiu Yi , Jingyu Xu , Jiawei Shao , Chi Zhang , Xuelong Li

Text-to-image diffusion models have demonstrated an impressive ability to produce high-quality outputs. However, they often struggle to accurately follow fine-grained spatial information in an input text. To this end, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ran Galun , Sagie Benaim

Video Diffusion Models (VDMs) have emerged as powerful generative tools, capable of synthesizing high-quality spatiotemporal content. Yet, their potential goes far beyond mere video generation. We argue that the training dynamics of VDMs,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Pablo Acuaviva , Aram Davtyan , Mariam Hassan , Sebastian Stapf , Ahmad Rahimi , Alexandre Alahi , Paolo Favaro

Recent video diffusion models have demonstrated their great capability in generating visually-pleasing results, while synthesizing the correct physical effects in generated videos remains challenging. The complexity of real-world motions,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Ke Zhang , Cihan Xiao , Jiacong Xu , Yiqun Mei , Vishal M. Patel

Text-to-video diffusion models enable the generation of high-quality videos that follow text instructions, making it easy to create diverse and individual content. However, existing approaches mostly focus on high-quality short video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Roberto Henschel , Levon Khachatryan , Hayk Poghosyan , Daniil Hayrapetyan , Vahram Tadevosyan , Zhangyang Wang , Shant Navasardyan , Humphrey Shi

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Weiming Ren , Huan Yang , Ge Zhang , Cong Wei , Xinrun Du , Wenhao Huang , Wenhu Chen

Long video generation remains a challenging and compelling topic in computer vision. Diffusion based models, among the various approaches to video generation, have achieved state of the art quality with their iterative denoising procedures.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Siyang Zhang , Harry Yang , Ser-Nam Lim

Diffusion models have emerged as powerful generative frameworks by progressively adding noise to data through a forward process and then reversing this process to generate realistic samples. While these models have achieved strong…

Machine Learning · Computer Science 2025-03-04 Xingzhuo Guo , Yu Zhang , Baixu Chen , Haoran Xu , Jianmin Wang , Mingsheng Long

In this paper, we propose a data-driven visual rhythm prediction method, which overcomes the previous works' deficiency that predictions are made primarily by human-crafted hard rules. In our approach, we first extract features including…

Computer Vision and Pattern Recognition · Computer Science 2019-01-30 Yutong Xie , Haiyang Wang , Yan Hao , Zihao Xu

This study introduces a text-conditioned approach to generating drumbeats with Latent Diffusion Models (LDMs). It uses informative conditioning text extracted from training data filenames. By pretraining a text and drumbeat encoder through…

Sound · Computer Science 2024-08-07 Pushkar Jajoria , James McDermott

Video-to-Video synthesis (Vid2Vid) has achieved remarkable results in generating a photo-realistic video from a sequence of semantic maps. However, this pipeline suffers from high computational cost and long inference latency, which largely…

Computer Vision and Pattern Recognition · Computer Science 2022-07-12 Long Zhuo , Guangcong Wang , Shikai Li , Wayne Wu , Ziwei Liu

Diffusion-based text-to-video generation (T2V) or image-to-video (I2V) generation have emerged as a prominent research focus. However, there exists a challenge in integrating the two generative paradigms into a unified model. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Xinyu Xiao , Binbin Yang , Tingtian Li , Yipeng Yu , Sen Lei