English
Related papers

Related papers: Voice activity detection in the wild via weakly su…

200 papers

Weakly supervised video anomaly detection (WSVAD) is a challenging task. Generating fine-grained pseudo-labels based on weak-label and then self-training a classifier is currently a promising solution. However, since the existing methods…

Computer Vision and Pattern Recognition · Computer Science 2024-04-15 Zhiwei Yang , Jing Liu , Peng Wu

Voice Type Discrimination (VTD) refers to discrimination between regions in a recording where speech was produced by speakers that are physically within proximity of the recording device ("Live Speech") from speech and other types of audio…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-12 Tyler Vuong , Yangyang Xia , Richard Stern

Pseudo-label learning methods have been widely applied in weakly-supervised temporal action localization. Existing works directly utilize weakly-supervised base model to generate instance-level pseudo-labels for training the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Quan Zhang , Yuxin Qi , Xi Tang , Rui Yuan , Xi Lin , Ke Zhang , Chun Yuan

Isolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate…

Computer Vision and Pattern Recognition · Computer Science 2018-02-13 Aviv Gabbay , Ariel Ephrat , Tavi Halperin , Shmuel Peleg

In this paper, we propose a multi-level attention model to solve the weakly labelled audio classification problem. The objective of audio classification is to predict the presence or absence of audio events in an audio clip. Recently,…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-08 Changsong Yu , Karim Said Barsim , Qiuqiang Kong , Bin Yang

Deep neural networks (DNNs) have a high capacity to completely memorize noisy labels given sufficient training time, and its memorization, unfortunately, leads to performance degradation. Recently, virtual adversarial training (VAT)…

Computation and Language · Computer Science 2022-06-24 Do-Myoung Lee , Yeachan Kim , Chang-gyun Seo

Analyses for biodiversity monitoring based on passive acoustic monitoring (PAM) recordings is time-consuming and challenged by the presence of background noise in recordings. Existing models for sound event detection (SED) worked only on…

Automatic speaker verification (ASV) technology is recently finding its way to end-user applications for secure access to personal data, smart services or physical facilities. Similar to other biometric technologies, speaker verification is…

Sound · Computer Science 2016-09-16 Cemal Hanilci , Tomi Kinnunen , Md Sahidullah , Aleksandr Sizov

Training-free video anomaly detection (VAD) has recently emerged as a scalable alternative to supervised approaches, yet existing methods largely rely on static prompting and geometry-agnostic feature fusion. As a result, anomaly inference…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ali Zia , Usman Ali , Muhammad Umer Ramzan , Hamza Abid , Abdul Rehman , Wei Xiang

In audio-visual navigation (AVN), an intelligent agent needs to navigate to a constantly sound-making object in complex 3D environments based on its audio and visual perceptions. While existing methods attempt to improve the navigation…

Sound · Computer Science 2022-06-02 Shunqi Mao , Chaoyi Zhang , Heng Wang , Weidong Cai

Acoustic event detection for content analysis in most cases relies on lots of labeled data. However, manually annotating data is a time-consuming task, which thus makes few annotated resources available so far. Unlike audio event detection,…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Yong Xu , Qiang Huang , Wenwu Wang , Philip J. B. Jackson , Mark D. Plumbley

The latest research in the field of voice anti-spoofing (VAS) shows that deep neural networks (DNN) outperform classic approaches like GMM in the task of presentation attack detection. However, DNNs require a lot of data to converge, and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-02 Ivan Yakovlev , Mikhail Melnikov , Nikita Bukhal , Rostislav Makarov , Alexander Alenin , Nikita Torgashov , Anton Okhotnikov

Spatio-temporal action detection (STAD) is an important fine-grained video understanding task. Current methods require box and label supervision for all action classes in advance. However, in real-world applications, it is very likely to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Tao Wu , Shuqiu Ge , Jie Qin , Gangshan Wu , Limin Wang

We propose noise-robust voice conversion (VC) which takes into account the recording quality and environment of noisy source speech. Conventional denoising training improves the noise robustness of a VC model by learning noisy-to-clean VC…

Visual detection of Micro Air Vehicles (MAVs) has attracted increasing attention in recent years due to its important application in various tasks. The existing methods for MAV detection assume that the training set and testing set have the…

Robotics · Computer Science 2024-03-26 Yin Zhang , Jinhong Deng , Peidong Liu , Wen Li , Shiyu Zhao

Passive acoustic monitoring is a sustainable method of monitoring wildlife and environments that leads to the generation of large datasets and, currently, a processing backlog. Academic research into automating this process is focused on…

Sound · Computer Science 2025-11-18 Wendy Lomas , Andrew Gascoyne , Colin Dubreuil , Stefano Vaglio , Liam Naughton

Semi- and weakly-supervised learning have recently attracted considerable attention in the object detection literature since they can alleviate the cost of annotation needed to successfully train deep learning models. State-of-art…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Akhil Meethal , Marco Pedersoli , Zhongwen Zhu , Francisco Perdigon Romero , Eric Granger

Humans have the ability to utilize visual cues, such as lip movements and visual scenes, to enhance auditory perception, particularly in noisy environments. However, current Automatic Speech Recognition (ASR) or Audio-Visual Speech…

Computation and Language · Computer Science 2025-04-11 Lakshmipathi Balaji , Karan Singla

This study explores the design and application of Complex-Valued Convolutional Neural Networks (CVCNNs) in audio signal processing, with a focus on preserving and utilizing phase information often neglected in real-valued networks. We begin…

Machine Learning · Computer Science 2025-10-14 Naman Agrawal

With recent deep learning based approaches showing promising results in removing noise from images, the best denoising performance has been reported in a supervised learning setup that requires a large set of paired noisy images and ground…

Image and Video Processing · Electrical Eng. & Systems 2022-09-20 Rihuan Ke