中文
相关论文

相关论文: Generating Realistic Images from In-the-wild Sound…

200 篇论文

The recent success of the generative model shows that leveraging the multi-modal embedding space can manipulate an image using text information. However, manipulating an image with other sources rather than text, such as sound, is not easy…

图形学 · 计算机科学 2021-12-02 Seung Hyun Lee , Wonseok Roh , Wonmin Byeon , Sang Ho Yoon , Chan Young Kim , Jinkyu Kim , Sangpil Kim

In this paper we propose a deep learning method for performing attributed-based music-to-image translation. The proposed method is applied for synthesizing visual stories according to the sentiment expressed by songs. The generated images…

计算机视觉与模式识别 · 计算机科学 2019-12-13 Nikolaos Passalis , Stavros Doropoulos

Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on…

声音 · 计算机科学 2025-12-11 Hao Zhou , Xiaobao Guo , Yuzhe Zhu , Adams Wai-Kin Kong

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting sounds that may not…

声音 · 计算机科学 2023-11-10 David Chuan-En Lin , Nikolas Martelaro

In this paper, we introduce a novel approach to address the task of synthesizing speech from silent videos of any in-the-wild speaker solely based on lip movements. The traditional approach of directly generating speech from lip videos…

多媒体 · 计算机科学 2024-03-05 Sindhu Hegde , Rudrabha Mukhopadhyay , C. V. Jawahar , Vinay Namboodiri

Can we determine someone's geographic location purely from the sounds they hear? Are acoustic signals enough to localize within a country, state, or even city? We tackle the challenge of global-scale audio geolocation, formalize the…

声音 · 计算机科学 2025-07-23 Mustafa Chasmai , Wuao Liu , Subhransu Maji , Grant Van Horn

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

声音 · 计算机科学 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

Localizing visual sounds consists on locating the position of objects that emit sound within an image. It is a growing research area with potential applications in monitoring natural and urban environments, such as wildlife migration and…

声音 · 计算机科学 2022-04-12 Ho-Hsiang Wu , Magdalena Fuentes , Prem Seetharaman , Juan Pablo Bello

Previous audio generation mainly focuses on specified sound classes such as speech or music, whose form and content are greatly restricted. In this paper, we go beyond specific audio generation by using natural language description as a…

声音 · 计算机科学 2023-05-04 Guangwei Li , Xuenan Xu , Lingfeng Dai , Mengyue Wu , Kai Yu

We introduce the visual acoustic matching task, in which an audio clip is transformed to sound like it was recorded in a target environment. Given an image of the target environment and a waveform for the source audio, the goal is to…

计算机视觉与模式识别 · 计算机科学 2022-06-15 Changan Chen , Ruohan Gao , Paul Calamia , Kristen Grauman

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains…

多媒体 · 计算机科学 2025-09-09 Jorge E. León , Miguel Carrasco

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that…

计算机视觉与模式识别 · 计算机科学 2024-01-30 Vladimir Iashin , Weidi Xie , Esa Rahtu , Andrew Zisserman

Implicit Neural Representations (INRs) are nowadays used to represent multimedia signals across various real-life applications, including image super-resolution, image compression, or 3D rendering. Existing methods that leverage INRs are…

机器学习 · 计算机科学 2023-06-21 Filip Szatkowski , Karol J. Piczak , Przemysław Spurek , Jacek Tabor , Tomasz Trzciński

We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a set of components that represents the sound from each pixel.…

计算机视觉与模式识别 · 计算机科学 2018-10-16 Hang Zhao , Chuang Gan , Andrew Rouditchenko , Carl Vondrick , Josh McDermott , Antonio Torralba

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion blur or occlusions. In…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Daniel Adebi , Sagnik Majumder , Kristen Grauman

Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited research on enhancing the diversity of generated audio,…

声音 · 计算机科学 2024-03-05 Zeyu Xie , Baihan Li , Xuenan Xu , Mengyue Wu , Kai Yu

Motivated by the recent progress in generative models, we introduce a model that generates images from natural language descriptions. The proposed model iteratively draws patches on a canvas, while attending to the relevant words in the…

机器学习 · 计算机科学 2016-03-01 Elman Mansimov , Emilio Parisotto , Jimmy Lei Ba , Ruslan Salakhutdinov

The rise of deep learning algorithms has led many researchers to withdraw from using classic signal processing methods for sound generation. Deep learning models have achieved expressive voice synthesis, realistic sound textures, and…

声音 · 计算机科学 2022-01-10 Anastasia Natsiou , Sean O'Leary

Video editing and generation methods often rely on pre-trained image-based diffusion models. During the diffusion process, however, the reliance on rudimentary noise sampling techniques that do not preserve correlations present in…

计算机视觉与模式识别 · 计算机科学 2025-04-07 Pascal Chang , Jingwei Tang , Markus Gross , Vinicius C. Azevedo

Audio denoising has been explored for decades using both traditional and deep learning-based methods. However, these methods are still limited to either manually added artificial noise or lower denoised audio quality. To overcome these…

声音 · 计算机科学 2022-10-20 Youshan Zhang , Jialu Li