English
Related papers

Related papers: Kling-Foley: Multimodal Diffusion Transformer for …

200 papers

Recent advancements in foundation models have revolutionized joint audio-video generation. However, existing approaches typically treat human-centric tasks including reference-based audio-video generation (R2AV), video editing (RV2AV) and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Xu Guo , Fulong Ye , Qichao Sun , Liyang Chen , Bingchuan Li , Pengze Zhang , Jiawei Liu , Songtao Zhao , Qian He , Xiangwang Hou

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

Spatial audio has become central to immersive applications such as VR/AR, cinema, and music. Existing generative audio models are largely limited to mono or stereo formats and cannot capture the full 3D localization cues available in…

Sound · Computer Science 2025-09-22 Christian Templin , Yanda Zhu , Hao Wang

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bingqi Ma , Linlong Lang , Ming Zhang , Dailan He , Xingtong Ge , Yi Zhang , Guanglu Song , Yu Liu

We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounter limitations when…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Yiming Zhang , Yicheng Gu , Yanhong Zeng , Zhening Xing , Yuancheng Wang , Zhizheng Wu , Kai Chen

In this work, we present FoleyGRAM, a novel approach to video-to-audio generation that emphasizes semantic conditioning through the use of aligned multimodal encoders. Building on prior advancements in video-to-audio generation, FoleyGRAM…

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jianwen Jiang , Chao Liang , Jiaqi Yang , Gaojie Lin , Tianyun Zhong , Yanbo Zheng

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, and 2) large-scale, high-quality training data. As such, we…

Multimedia · Computer Science 2026-04-16 Zeyue Tian , Zhaoyang Liu , Yizhu Jin , Ruibin Yuan , Liumeng Xue , Xu Tan , Qifeng Chen , Wei Xue , Yike Guo

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational experiences by…

Sound · Computer Science 2025-06-04 Chetwin Low , Weimin Wang

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioning. The key to our…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Moayed Haji-Ali , Willi Menapace , Aliaksandr Siarohin , Ivan Skorokhodov , Alper Canberk , Kwot Sin Lee , Vicente Ordonez , Sergey Tulyakov

Text-based audio generation models have limitations as they cannot encompass all the information in audio, leading to restricted controllability when relying solely on text. To address this issue, we propose a novel model that enhances the…

Sound · Computer Science 2023-12-29 Zhifang Guo , Jianguo Mao , Rui Tao , Long Yan , Kazushige Ouchi , Hong Liu , Xiangdong Wang

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images),…

Sound · Computer Science 2025-08-06 Yan Rong , Jinting Wang , Guangzhi Lei , Shan Yang , Li Liu

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

Recent advances in Diffusion Transformers (DiTs) have enabled high-quality joint audio-video generation, producing videos with synchronized audio within a single model. However, existing controllable generation frameworks are typically…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Liyang Li , Wen Wang , Canyu Zhao , Tianjian Feng , Zhiyue Zhao , Hao Chen , Chunhua Shen

Foley sound presents the background sound for multimedia content and the generation of Foley sound involves computationally modelling sound effects with specialized techniques. In this work, we proposed a system for DCASE 2023 challenge…

Sound · Computer Science 2023-09-18 Yi Yuan , Haohe Liu , Xubo Liu , Xiyuan Kang , Mark D. Plumbley , Wenwu Wang

With the rise of diffusion models, audio-video generation has been revolutionized. However, most existing methods rely on separate modules for each modality, with limited exploration of unified generative architectures. In addition, many…

Multimedia · Computer Science 2025-07-08 Lei Zhao , Linfeng Feng , Dongxu Ge , Rujin Chen , Fangqiu Yi , Chi Zhang , Xiao-Lei Zhang , Xuelong Li

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

Sound · Computer Science 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang