English
Related papers

Related papers: Voice activity detection in the wild via weakly su…

200 papers

Recently, an audio-visual speech generative model based on variational autoencoder (VAE) has been proposed, which is combined with a nonnegative matrix factorization (NMF) model for noise variance to perform unsupervised speech enhancement.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-12 Mostafa Sadeghi , Xavier Alameda-Pineda

Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process free-text sound…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Jiarui Hai , Helin Wang , Weizhe Guo , Mounya Elhilali

Sound event detection (SED) is a task to detect sound events in an audio recording. One challenge of the SED task is that many datasets such as the Detection and Classification of Acoustic Scenes and Events (DCASE) datasets are weakly…

Sound · Computer Science 2020-08-25 Qiuqiang Kong , Yong Xu , Wenwu Wang , Mark D. Plumbley

Although deep learning based multi-channel speech enhancement has achieved significant advancements, its practical deployment is often limited by constrained computational resources, particularly in low signal-to-noise ratio (SNR)…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-27 Zheng Wang , Xiaobin Rong , Yu Sun , Tianchi Sun , Zhibin Lin , Jing Lu

Generative Face Video Coding (GFVC) achieves superior rate-distortion performance by leveraging the strong inference capabilities of deep generative models. However, its practical deployment is hindered by large model parameters and high…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Zihan Zhang , Shanzhi Yin , Bolin Chen , Ru-Ling Liao , Shiqi Wang , Yan Ye

Polyphonic Sound Event Detection (SED) in real-world recordings is a challenging task because of the dynamic polyphony level, intensity, and duration of sound events. Current polyphonic SED systems fail to model the temporal structure of…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-02 Arjun Pankajakshan , Helen L. Bear , Emmanouil Benetos

In this work, we present a novel audio-visual dataset for active speaker detection in the wild. A speaker is considered active when his or her face is visible and the voice is audible simultaneously. Although active speaker detection is a…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 You Jin Kim , Hee-Soo Heo , Soyeon Choe , Soo-Whan Chung , Yoohwan Kwon , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

Video anomaly detection (VAD) plays a vital role in real-world applications such as security surveillance, autonomous driving, and industrial monitoring. Recent advances in large pre-trained models have opened new opportunities for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 He Huang , Zixuan Hu , Dongxiao Li , Yao Xiao , Ling-Yu Duan

Video Anomaly Detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Yang Liu , Siao Liu , Xiaoguang Zhu , Jielin Li , Hao Yang , Liangyu Teng , Juncen Guo , Yan Wang , Dingkang Yang , Jing Liu

The deep learning-based speech enhancement (SE) methods always take the clean speech's waveform or time-frequency spectrum feature as the learning target, and train the deep neural network (DNN) by reducing the error loss between the DNN's…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-02 Yuewei Zhang , Huanbin Zou , Jie Zhu

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

The rise of advanced large language models such as GPT-4, GPT-4o, and the Claude family has made fake audio detection increasingly challenging. Traditional fine-tuning methods struggle to keep pace with the evolving landscape of synthetic…

Sound · Computer Science 2024-08-14 Xiaohui Zhang , Jiangyan Yi , Jianhua Tao

The use of large-scale vision-language datasets is limited for object detection due to the negative impact of label noise on localization. Prior methods have shown how such large-scale datasets can be used for pretraining, which can provide…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Arushi Rai , Adriana Kovashka

We introduce AudioScopeV2, a state-of-the-art universal audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify…

Sound · Computer Science 2022-07-22 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

The audio segmentation mismatch between training data and those seen at run-time is a major problem in direct speech translation. Indeed, while systems are usually trained on manually segmented corpora, in real use cases they are often…

Sound · Computer Science 2021-10-15 Marco Gaido , Matteo Negri , Mauro Cettolo , Marco Turchi

We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model…

Computation and Language · Computer Science 2019-09-12 Steffen Schneider , Alexei Baevski , Ronan Collobert , Michael Auli

The speech representations learned from large-scale unlabeled data have shown better generalizability than those from supervised learning and thus attract a lot of interest to be applied for various downstream tasks. In this paper, we…

Sound · Computer Science 2022-01-25 Zhengyang Chen , Sanyuan Chen , Yu Wu , Yao Qian , Chengyi Wang , Shujie Liu , Yanmin Qian , Michael Zeng

Speech Emotion Recognition (SER) often operates on speech segments detected by a Voice Activity Detection (VAD) model. However, VAD models may output flawed speech segments, especially in noisy environments, resulting in degraded…

Sound · Computer Science 2024-10-18 Natsuo Yamashita , Masaaki Yamamoto , Yohei Kawaguchi

Video Anomaly Detection (VAD), which aims to detect anomalies that deviate from expectation, has attracted increasing attention in recent years. Existing advancements in VAD primarily focus on model architectures and training strategies,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Zihao Liu , Xiaoyu Wu , Wenna Li , Linlin Yang , Shengjin Wang

Sound recordings are used in various ecological studies, including acoustic wildlife monitoring. Such surveys require automatic detection of target sound events. However, current detectors, especially those relying on band-limited energy,…

Applications · Statistics 2021-10-13 Julius Juodakis , Stephen Marsland