English
Related papers

Related papers: Learning from Silence and Noise for Visual Sound S…

200 papers

Discriminatively localizing sounding objects in cocktail-party, i.e., mixed sound scenes, is commonplace for humans, but still challenging for machines. In this paper, we propose a two-stage learning framework to perform self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Di Hu , Rui Qian , Minyue Jiang , Xiao Tan , Shilei Wen , Errui Ding , Weiyao Lin , Dejing Dou

There has been a growing interest in the task of generating sound for silent videos, primarily because of its practicality in streamlining video post-production. However, existing methods for video-sound generation attempt to directly…

Multimedia · Computer Science 2024-04-04 Zhifeng Xie , Shengye Yu , Qile He , Mengtian Li

Background sound is an informative form of art that is helpful in providing a more immersive experience in real-application voice conversion (VC) scenarios. However, prior research about VC, mainly focusing on clean voices, pay rare…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Jixun Yao , Yi Lei , Qing Wang , Pengcheng Guo , Ziqian Ning , Lei Xie , Hai Li , Junhui Liu , Danming Xie

Low signal-to-noise ratio videos -- such as those from underwater sonar, ultrasound, and microscopy -- pose significant challenges for computer vision models, particularly when paired clean imagery is unavailable. We present Spatiotemporal…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Suzanne Stathatos , Michael Hobley , Pietro Perona , Markus Marks

Source localization and spectral estimation are among the most fundamental problems in statistical and array signal processing. Methods which rely on the orthogonality of the signal and noise subspaces, such as Pisarenko's method, MUSIC,…

Information Theory · Computer Science 2019-04-16 Matthew W. Morency , Sergiy A. Vorobyov , Geert Leus

Target-Speaker Voice Activity Detection (TS-VAD) is the task of detecting the presence of speech from a known target-speaker in an audio frame. Recently, deep neural network-based models have shown good performance in this task. However,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-07 Holger Severin Bovbjerg , Jan Østergaard , Jesper Jensen , Zheng-Hua Tan

Predicting audio quality in voice synthesis and conversion systems is a critical yet challenging task, especially when traditional methods like Mean Opinion Scores (MOS) are cumbersome to collect at scale. This paper addresses the gap in…

Sound · Computer Science 2023-12-27 Aditya Ravuri , Erica Cooper , Junichi Yamagishi

Constructing an embedding space for musical instrument sounds that can meaningfully represent new and unseen instruments is important for downstream music generation tasks such as multi-instrument synthesis and timbre transfer. The…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-28 Xuan Shi , Erica Cooper , Junichi Yamagishi

The task of audio-visual sound source localization has been well studied under constrained scenes, where the audio recordings are clean. However, in real-world scenarios, audios are usually contaminated by off-screen sound and background…

Computer Vision and Pattern Recognition · Computer Science 2022-02-15 Xian Liu , Rui Qian , Hang Zhou , Di Hu , Weiyao Lin , Ziwei Liu , Bolei Zhou , Xiaowei Zhou

Contrastive representation learning has proven to be an effective self-supervised learning method for images and videos. Most successful approaches are based on Noise Contrastive Estimation (NCE) and use different views of an instance as…

Computer Vision and Pattern Recognition · Computer Science 2023-09-27 Julien Denize , Jaonary Rabarisoa , Astrid Orcesi , Romain Hérault

Pseudo-label learning methods have been widely applied in weakly-supervised temporal action localization. Existing works directly utilize weakly-supervised base model to generate instance-level pseudo-labels for training the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Quan Zhang , Yuxin Qi , Xi Tang , Rui Yuan , Xi Lin , Ke Zhang , Chun Yuan

In this paper we propose a novel learning framework called Supervised and Weakly Supervised Learning where the goal is to learn simultaneously from weakly and strongly labeled data. Strongly labeled data can be simply understood as fully…

Machine Learning · Computer Science 2017-02-21 Anurag Kumar , Bhiksha Raj

Video quality significantly affects video classification. We found this problem when we classified Mild Cognitive Impairment well from clear videos, but worse from blurred ones. From then, we realized that referring to Video Quality…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Jian Sun , Mohammad H. Mahoor

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a…

Sound · Computer Science 2022-02-15 Ke Chen , Xingjian Du , Bilei Zhu , Zejun Ma , Taylor Berg-Kirkpatrick , Shlomo Dubnov

The problem of source localization with ad hoc microphone networks in noisy and reverberant enclosures, given a training set of prerecorded measurements, is addressed in this paper. The training set is assumed to consist of a limited number…

Sound · Computer Science 2016-10-18 Bracha Laufer-Goldshtein , Ronen Talmon , Sharon Gannot

Conventional voice conversion modifies voice characteristics from a source speaker to a target speaker, relying on audio input from both sides. However, this process becomes infeasible when clean audio is unavailable, such as in silent…

Sound · Computer Science 2025-08-05 Yifan Liu , Yu Fang , Zhouhan Lin

The hearing sense on a mobile robot is important because it is omnidirectional and it does not require direct line-of-sight with the sound source. Such capabilities can nicely complement vision to help localize a person or an interesting…

Robotics · Computer Science 2016-02-29 Jean-Marc Valin , François Michaud , Jean Rouat , Dominic Létourneau

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-22 Hengyi Hong , Qing Wang , Jun Du , Ruoyu Wei , Mingqi Cai , Xin Fang

Localizing a moving sound source in the real world involves determining its direction-of-arrival (DOA) and distance relative to a microphone. Advancements in DOA estimation have been facilitated by data-driven methods optimized with large…

Sound · Computer Science 2023-09-19 Saksham Singh Kushwaha , Iran R. Roman , Magdalena Fuentes , Juan Pablo Bello

The ability to localize and track acoustic events is a fundamental prerequisite for equipping machines with the ability to be aware of and engage with humans in their surrounding environment. However, in realistic scenarios, audio signals…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-22 Christine Evers , Heinrich Loellmann , Heinrich Mellmann , Alexander Schmidt , Hendrik Barfuss , Patrick Naylor , Walter Kellermann