English
Related papers

Related papers: SyncFusion: Multimodal Onset-synchronized Video-to…

200 papers

Audio synthesis has broad applications in multimedia. Recent advancements have made it possible to generate relevant audios from inputs describing an audio scene, such as images or texts. However, the immersiveness and expressiveness of the…

Multimedia · Computer Science 2025-08-13 Wei Guo , Heng Wang , Jianbo Ma , Weidong Cai

Acoustic matching aims to re-synthesize an audio clip to sound as if it were recorded in a target acoustic environment. Existing methods assume access to paired training data, where the audio is observed in both source and target…

Multimedia · Computer Science 2023-11-27 Arjun Somayazulu , Changan Chen , Kristen Grauman

Diffusion models have experienced a surge of interest as highly expressive yet efficiently trainable probabilistic models. We show that these models are an excellent fit for synthesising human motion that co-occurs with audio, e.g., dancing…

Machine Learning · Computer Science 2023-05-17 Simon Alexanderson , Rajmund Nagy , Jonas Beskow , Gustav Eje Henter

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strategy for audiovisual…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Liangfa Wei , Jie Zhang , Junfeng Hou , Lirong Dai

We introduce MMAudioSep, a generative model for video/text-queried sound separation that is founded on a pretrained video-to-audio model. By leveraging knowledge about the relationship between video/text and audio learned through a…

Sound · Computer Science 2026-04-20 Akira Takahashi , Shusuke Takahashi , Yuki Mitsufuji

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical…

Sound · Computer Science 2026-04-07 Weiguo Pian , Saksham Singh Kushwaha , Zhimin Chen , Shijian Deng , Kai Wang , Yunhui Guo , Yapeng Tian

Recent advances in text-to-music editing, which employ text queries to modify music (e.g.\ by changing its style or adjusting instrumental components), present unique challenges and opportunities for AI-assisted music creation. Previous…

An abstract sound is defined as a sound that does not disclose identifiable real-world sound events to a listener. Sound fusion aims to synthesize an original sound and a reference sound to generate a novel sound that exhibits auditory…

Sound · Computer Science 2025-08-05 Jing Liu , Enqi Lian , Moyao Deng

Learning to localize the sound source in videos without explicit annotations is a novel area of audio-visual research. Existing work in this area focuses on creating attention maps to capture the correlation between the two modalities to…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Dennis Fedorishin , Deen Dayal Mohan , Bhavin Jawade , Srirangaraj Setlur , Venu Govindaraju

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio…

Machine Learning · Computer Science 2017-04-06 Jesse Engel , Cinjon Resnick , Adam Roberts , Sander Dieleman , Douglas Eck , Karen Simonyan , Mohammad Norouzi

Conventional music visualisation systems rely on handcrafted ad hoc transformations of shapes and colours that offer only limited expressiveness. We propose two novel pipelines for automatically generating music videos from any…

Audio-to-score alignment is a long-standing challenge in music information retrieval and arguably the most widely applicable alignment task for music research. Alignment algorithms match two versions of a piece of music, and for this to…

Sound · Computer Science 2026-05-20 Silvan Peter , Patricia Hu , Gerhard Widmer

As text-based speech editing becomes increasingly prevalent, the demand for unrestricted free-text editing continues to grow. However, existing speech editing techniques encounter significant challenges, particularly in maintaining…

Sound · Computer Science 2024-09-23 Yang Chen , Yuhang Jia , Shiwan Zhao , Ziyue Jiang , Haoran Li , Jiarong Kang , Yong Qin

Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-24 Jianyi Chen , Rongxiu Zhong , Shilei Zhang , Kun Qian , Jinglei Liu , Yike Guo , Wei Xue

Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on storytelling datasets and tasks have paid little attention to…

Multimedia · Computer Science 2023-10-31 Jaeyeon Bae , Seokhoon Jeong , Seokun Kang , Namgi Han , Jae-Yon Lee , Hyounghun Kim , Taehwan Kim

Data synthesis and augmentation are essential for Sound Event Detection (SED) due to the scarcity of temporally labeled data. While augmentation methods like SpecAugment and Mix-up can enhance model performance, they remain constrained by…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Jiarui Hai , Mounya Elhilali

Generating videos for visual storytelling can be a tedious and complex process that typically requires either live-action filming or graphics animation rendering. To bypass these challenges, our key idea is to utilize the abundance of…

Computer Vision and Pattern Recognition · Computer Science 2023-07-14 Yingqing He , Menghan Xia , Haoxin Chen , Xiaodong Cun , Yuan Gong , Jinbo Xing , Yong Zhang , Xintao Wang , Chao Weng , Ying Shan , Qifeng Chen

Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utterance. Compared to…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Muhammad Hamza Mughal , Rishabh Dabral , Ikhsanul Habibie , Lucia Donatelli , Marc Habermann , Christian Theobalt

With the fast development of zero-shot text-to-speech technologies, it is possible to generate high-quality speech signals that are indistinguishable from the real ones. Speech editing, including speech insertion and replacement, appeals to…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-19 Kuan-Yu Chen , Jeng-Lin Li , De-Yan Lu , Jian-Jiun Ding

Audio diffusion models can synthesize a wide variety of sounds. Existing models often operate on the latent domain with cascaded phase recovery modules to reconstruct waveform. This poses challenges when generating high-fidelity audio. In…

Sound · Computer Science 2023-11-21 Ge Zhu , Yutong Wen , Marc-André Carbonneau , Zhiyao Duan
‹ Prev 1 8 9 10 Next ›