中文
相关论文

相关论文: Taming Visually Guided Sound Generation

200 篇论文

Models for audio generation are typically trained on hours of recordings. Here, we illustrate that capturing the essence of an audio source is typically possible from as little as a few tens of seconds from a single training signal.…

声音 · 计算机科学 2021-10-27 Gal Greshler , Tamar Rott Shaham , Tomer Michaeli

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic videos are disregarded.…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Jiawei Liu , Weining Wang , Sihan Chen , Xinxin Zhu , Jing Liu

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

Most GAN(Generative Adversarial Network)-based approaches towards high-fidelity waveform generation heavily rely on discriminators to improve their performance. However, GAN methods introduce much uncertainty into the generation process and…

声音 · 计算机科学 2022-03-22 Shengyuan Xu , Wenxiao Zhao , Jing Guo

Previous audio generation mainly focuses on specified sound classes such as speech or music, whose form and content are greatly restricted. In this paper, we go beyond specific audio generation by using natural language description as a…

声音 · 计算机科学 2023-05-04 Guangwei Li , Xuenan Xu , Lingfeng Dai , Mengyue Wu , Kai Yu

Previous works (Donahue et al., 2018a; Engel et al., 2019a) have found that generating coherent raw audio waveforms with GANs is challenging. In this paper, we show that it is possible to train GANs reliably to generate high quality…

音频与语音处理 · 电气工程与系统科学 2019-12-10 Kundan Kumar , Rithesh Kumar , Thibault de Boissiere , Lucas Gestin , Wei Zhen Teoh , Jose Sotelo , Alexandre de Brebisson , Yoshua Bengio , Aaron Courville

Recent developments in generative models have shown that deep learning combined with traditional digital signal processing (DSP) techniques could successfully generate convincing violin samples [1], that source-excitation combined with…

音频与语音处理 · 电气工程与系统科学 2020-08-07 Ollie McCarthy , Zohaib Ahmed

Deep learning based visual to sound generation systems essentially need to be developed particularly considering the synchronicity aspects of visual and audio features with time. In this research we introduce a novel task of guiding a class…

机器学习 · 计算机科学 2021-07-21 Sanchita Ghose , John J. Prevost

Diffusion-based audio and music generation models commonly perform generation by constructing an image representation of audio (e.g., a mel-spectrogram) and then convert it to audio using a phase reconstruction model or vocoder. Typical…

声音 · 计算机科学 2024-10-08 Ge Zhu , Juan-Pablo Caceres , Zhiyao Duan , Nicholas J. Bryan

In recent years, with the realistic generation results and a wide range of personalized applications, diffusion-based generative models gain huge attention in both visual and audio generation areas. Compared to the considerable advancements…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Shiqi Yang , Zhi Zhong , Mengjie Zhao , Shusuke Takahashi , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Audio signals are sampled at high temporal resolutions, and learning to synthesize audio requires capturing structure across a range of timescales. Generative adversarial networks (GANs) have seen wide success at generating images that are…

声音 · 计算机科学 2019-02-12 Chris Donahue , Julian McAuley , Miller Puckette

Recent advances in generative models that iteratively synthesize audio clips sparked great success to text-to-audio synthesis (TTA), but with the cost of slow synthesis speed and heavy computation. Although there have been attempts to…

In this paper we introduce StyleWaveGAN, a style-based drum sound generator that is a variation of StyleGAN, a state-of-the-art image generator. By conditioning StyleWaveGAN on both the type of drum and several audio descriptors, we are…

声音 · 计算机科学 2022-08-29 Antoine Lavault , Axel Roebel , Matthieu Voiry

Single-image generative adversarial networks learn from the internal distribution of a single training example to generate variations of it, removing the need of a large dataset. In this paper we introduce SpecSinGAN, an unconditional…

声音 · 计算机科学 2022-04-06 Adrián Barahona-Ríos , Tom Collins

Text-to-audio (TTA) generation can significantly benefit the media industry by reducing production costs and enhancing work efficiency. However, most current TTA models (primarily diffusion-based) suffer from slow inference speeds and high…

声音 · 计算机科学 2025-12-30 HaeChun Chung

Video to sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls or specializations of the…

多媒体 · 计算机科学 2022-11-22 Chenye Cui , Yi Ren , Jinglin Liu , Rongjie Huang , Zhou Zhao

Videos show continuous events, yet most $-$ if not all $-$ video synthesis frameworks treat them discretely in time. In this work, we think of videos of what they should be $-$ time-continuous signals, and extend the paradigm of neural…

计算机视觉与模式识别 · 计算机科学 2022-06-02 Ivan Skorokhodov , Sergey Tulyakov , Mohamed Elhoseiny
‹ 上一页 1 2 3 10 下一页 ›