English
Related papers

Related papers: SyncFlow: Toward Temporally Aligned Joint Audio-Vi…

200 papers

We study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Yoav Shalev , Lior Wolf

Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence during training, while diffusion methods require multi-step…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Zengwei Yao , Wei Kang , Han Zhu , Liyong Guo , Lingxuan Ye , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Long Lin , Daniel Povey

Flow-based generative models have greatly improved text-to-speech (TTS) synthesis quality, but inference speed remains limited by the iterative sampling process and multiple function evaluations (NFE). The recent MeanFlow model accelerates…

Sound · Computer Science 2025-10-10 Wei Wang , Rong Cao , Yi Guo , Zhengyang Chen , Kuan Chen , Yuanyuan Huo

Generative modeling aims to transform random noise into structured outputs. In this work, we enhance video diffusion models by allowing motion control via structured latent noise sampling. This is achieved by just a change in data: we…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Ryan Burgert , Yuancheng Xu , Wenqi Xian , Oliver Pilarski , Pascal Clausen , Mingming He , Li Ma , Yitong Deng , Lingxiao Li , Mohsen Mousavi , Michael Ryoo , Paul Debevec , Ning Yu

Recent advancements in video diffusion models have shown exceptional abilities in simulating real-world dynamics and maintaining 3D consistency. This progress inspires us to investigate the potential of these models to ensure dynamic…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Jianhong Bai , Menghan Xia , Xintao Wang , Ziyang Yuan , Xiao Fu , Zuozhu Liu , Haoji Hu , Pengfei Wan , Di Zhang

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

Sound · Computer Science 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

Text-guided image generation has witnessed unprecedented progress due to the development of diffusion models. Beyond text and image, sound is a vital element within the sphere of human perception, offering vivid representations and…

Graphics · Computer Science 2023-06-21 Yue Yang , Kaipeng Zhang , Yuying Ge , Wenqi Shao , Zeyue Xue , Yu Qiao , Ping Luo

Recent advancements in video diffusion models have significantly enhanced audio-driven portrait animation. However, current methods still suffer from flickering, identity drift, and poor audio-visual synchronization. These issues primarily…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Zhenjie Liu , Jianzhang Lu , Renjie Lu , Cong Liang , Shangfei Wang

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled…

Many speech segments in movies are re-recorded in a studio during postproduction, to compensate for poor sound quality as recorded on location. Manual alignment of the newly-recorded speech with the original lip movements is a tedious task.…

Computer Vision and Pattern Recognition · Computer Science 2018-08-21 Tavi Halperin , Ariel Ephrat , Shmuel Peleg

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

Autoregressive (AR) models with diffusion heads have recently achieved strong text-to-audio performance, yet their iterative decoding and multi-step sampling process introduce high-latency issues. To address this bottleneck, we propose a…

Recent advances in text-to-image (T2I) diffusion models have enabled impressive image generation capabilities guided by text prompts. However, extending these techniques to video generation remains challenging, with existing text-to-video…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weifeng Chen , Yatai Ji , Jie Wu , Hefeng Wu , Pan Xie , Jiashi Li , Xin Xia , Xuefeng Xiao , Liang Lin

There have been a number of techniques that have demonstrated the generation of multimedia data for one modality at a time using GANs, such as the ability to generate images, videos, and audio. However, so far, the task of multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Vinod K Kurmi , Vipul Bajaj , Badri N Patro , K S Venkatesh , Vinay P Namboodiri , Preethi Jyothi

Compared with ample visual-text pre-training research, few works explore audio-text pre-training, mostly due to the lack of sufficient parallel audio-text data. Most existing methods incorporate the visual modality as a pivot for audio-text…

Sound · Computer Science 2024-03-06 Xuenan Xu , Zhiling Zhang , Zelin Zhou , Pingyue Zhang , Zeyu Xie , Mengyue Wu , Kenny Q. Zhu

Training-free video object editing aims to achieve precise object-level manipulation, including object insertion, swapping, and deletion. However, it faces significant challenges in maintaining fidelity and temporal consistency. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yiyang Chen , Xuanhua He , Xiujun Ma , Yue Ma

We propose DiTFlow, a method for transferring the motion of a reference video to a newly synthesized one, designed specifically for Diffusion Transformers (DiT). We first process the reference video with a pre-trained DiT to analyze…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Alexander Pondaven , Aliaksandr Siarohin , Sergey Tulyakov , Philip Torr , Fabio Pizzati

Joint audio-video generation models have shown that unified generation yields stronger cross-modal coherence than cascaded approaches. However, existing models couple modalities throughout denoising via pervasive attention, treating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhen Ye , Xu Tan , Aoxiong Yin , Hongzhan Lin , Guangyan Zhang , Peiwen Sun , Yiming Li , Chi-Min Chan , Wei Ye , Shikun Zhang , Wei Xue

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal…

Computer Vision and Pattern Recognition · Computer Science 2023-03-27 Ludan Ruan , Yiyang Ma , Huan Yang , Huiguo He , Bei Liu , Jianlong Fu , Nicholas Jing Yuan , Qin Jin , Baining Guo

Diffusion models have significantly improved the quality and diversity of audio generation but are hindered by slow inference speed. Rectified flow enhances inference speed by learning straight-line ordinary differential equation (ODE)…

Sound · Computer Science 2025-05-29 Junqi Zhao , Jinzheng Zhao , Haohe Liu , Yun Chen , Lu Han , Xubo Liu , Mark Plumbley , Wenwu Wang