English
Related papers

Related papers: FoleyGen: Visually-Guided Audio Generation

200 papers

Text-to-Image (T2I) generation has made significant advancements with diffusion models, yet challenges persist in handling complex instructions, ensuring fine-grained content control, and maintaining deep semantic consistency. Existing T2I…

Machine Learning · Computer Science 2025-08-08 Xiaoqi Dong , Xiangyu Zhou , Nicholas Evans , Yujia Lin

The content of visual and audio scenes is multi-faceted such that a video can be paired with various audio and vice-versa. Thereby, in video-to-audio generation task, it is imperative to introduce steering approaches for controlling the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Xiulong Liu , Kun Su , Eli Shlizerman

Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high…

Sound · Computer Science 2025-12-30 HaeChun Chung

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Xudong Lin , Gedas Bertasius , Jue Wang , Shih-Fu Chang , Devi Parikh , Lorenzo Torresani

We tackle the task of conditional music generation. We introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MusicGen is comprised…

We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Burak Can Biner , Farrin Marouf Sofian , Umur Berkay Karakaş , Duygu Ceylan , Erkut Erdem , Aykut Erdem

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions…

Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the lack of a systematic sound class diversity framework and a…

Sound · Computer Science 2024-07-19 Baihan Li , Zeyu Xie , Xuenan Xu , Yiwei Guo , Ming Yan , Ji Zhang , Kai Yu , Mengyue Wu

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

Sound · Computer Science 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

Foley sound synthesis is crucial for multimedia production, enhancing user experience by synchronizing audio and video both temporally and semantically. Recent studies on automating this labor-intensive process through video-to-sound…

Sound · Computer Science 2025-09-18 Junwon Lee , Jaekwon Im , Dabin Kim , Juhan Nam

Driving video generation has achieved much progress in controllability, video resolution, and length, but fails to support fine-grained object-level controllability for diverse driving videos, while preserving the spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Li-Heng Chen , Ke Cheng , Yahui Liu , Lei Shi , Shi-Sheng Huang , Hongbo Fu

Video-to-audio (V2A) generation shows great potential in fields such as film production. Despite significant advances, current V2A methods relying on global video information struggle with complex scenes and generating audio tailored to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yingshan Liang , Keyu Fan , Zhicheng Du , Yiran Wang , Qingyang Shi , Xinyu Zhang , Jiasheng Lu , Peiwu Qin

Text-to-audio (T2A) generation has achieved promising results with the recent advances in generative models. However, because of the limited quality and quantity of temporally-aligned audio-text pairs, existing T2A methods struggle to…

Sound · Computer Science 2025-09-19 Yuxuan Jiang , Zehua Chen , Zeqian Ju , Chang Li , Weibei Dou , Jun Zhu

The design of diffusion-based audio generation systems has been investigated from diverse perspectives, such as data space, network architecture, and conditioning techniques, while most of these innovations require model re-training. In…

Sound · Computer Science 2026-04-10 Junyou Wang , Zehua Chen , Binjie Yuan , Kaiwen Zheng , Chang Li , Yuxuan Jiang , Jun Zhu

Language model (LM) based audio generation frameworks, e.g., AudioLM, have recently achieved new state-of-the-art performance in zero-shot audio generation. In this paper, we explore the feasibility of LMs for zero-shot voice conversion. An…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-22 Zhichao Wang , Yuanzhe Chen , Lei Xie , Qiao Tian , Yuping Wang

Recent advances in image, video, text and audio generative techniques, and their use by the general public, are leading to new forms of content generation. Usually, each modality was approached separately, which poses limitations. The…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-05 María Sánchez , Laura Fernández , Julián Arias , Mateo Cámara , Giulia Comini , Adam Gabrys , José Luis Blanco , Juan Ignacio Godino , Luis Alfonso Hernández

Video-conditioned audio generation, including Video-to-Sound (V2S) and Visual Text-to-Speech (VisualTTS), has traditionally been treated as distinct tasks, leaving the potential for a unified generative framework largely underexplored. In…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Xin Cheng , Yuyue Wang , Xihua Wang , Yihan Wu , Kaisi Guan , Yijing Chen , Peng Zhang , Xiaojiang Liu , Meng Cao , Ruihua Song

The ability to envisage the visual of a talking face based just on hearing a voice is a unique human capability. There have been a number of works that have solved for this ability recently. We differ from these approaches by enabling a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-24 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

Audio is indispensable for real-world video, yet generation models have largely overlooked audio components. Current approaches to producing audio-visual content often rely on cascaded pipelines, which increase cost, accumulate errors, and…

The differing representation spaces required for visual understanding and generation pose a challenge in unifying them within the autoregressive paradigm of large language models. A vision tokenizer trained for reconstruction excels at…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Wei Song , Yuran Wang , Zijia Song , Yadong Li , Zenan Zhou , Long Chen , Jianhua Xu , Jiaqi Wang , Kaicheng Yu