English
Related papers

Related papers: Unsupervised Voice Activity Detection by Modeling …

200 papers

Rapid advances in speech synthesis and audio editing have made realistic forgeries increasingly accessible, yet existing detection methods remain vulnerable to tampering or depend on visual/wearable sensors. In this paper, we present…

Human-Computer Interaction · Computer Science 2026-03-31 Mingda Han , Huanqi Yang , Chaoqun Li , Wenhao Li , Guoming Zhang , Yanni Yang , Yetong Cao , Weitao Xu , Pengfei Hu

Keyword spotting (KWS) and speaker verification (SV) have been studied independently although it is known that acoustic and speaker domains are complementary. In this paper, we propose a multi-task network that performs KWS and SV…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Myunghun Jung , Youngmoon Jung , Jahyun Goo , Hoirin Kim

Deep neural networks (DNNs) remain challenged by distribution shifts in complex open-world domains like automated driving (AD): Robustness against yet unknown novel objects (semantic shift) or styles like lighting conditions (covariate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Mert Keser , Halil Ibrahim Orhan , Niki Amini-Naieni , Gesina Schwalbe , Alois Knoll , Matthias Rottmann

Recently, pioneer research works have proposed a large number of acoustic features (log power spectrogram, linear frequency cepstral coefficients, constant Q cepstral coefficients, etc.) for audio deepfake detection, obtaining good…

An objective understanding of media depictions, such as inclusive portrayals of how much someone is heard and seen on screen such as in film and television, requires the machines to discern automatically who, when, how, and where someone is…

Computer Vision and Pattern Recognition · Computer Science 2022-12-26 Rahul Sharma , Krishna Somandepalli , Shrikanth Narayanan

Recent advances in generative speech have increased the need for automatic detection of obviously failed synthetic outputs. This is particularly important in clinical settings such as AVATAR therapy, in which schizophrenia patients engage…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Jana Shokr , Minos Papadopoulos , Jeremy Cooperstock , Pavo Orepic

In this work, we investigate the effectiveness of two techniques for improving variational autoencoder (VAE) based voice conversion (VC). First, we reconsider the relationship between vocoder features extracted using the high quality…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-09 Wen-Chin Huang , Yi-Chiao Wu , Chen-Chou Lo , Patrick Lumban Tobing , Tomoki Hayashi , Kazuhiro Kobayashi , Tomoki Toda , Yu Tsao , Hsin-Min Wang

Due to the superior modeling ability of deep neural network (DNN), it is widely used in voice activity detection (VAD). However, the performance may degrade if no sufficient data especially for practical data could be used for training,…

Sound · Computer Science 2020-05-19 Lu Ma , Xiaomeng Zhang , Pei Zhao , Tengrong Su

Overlapping Speech Detection (OSD) aims to identify regions where multiple speakers overlap in a conversation, a critical challenge in multi-party speech processing. This work proposes a speaker-aware progressive OSD model that leverages a…

Sound · Computer Science 2025-05-30 Zhaokai Sun , Li Zhang , Qing Wang , Pan Zhou , Lei Xie

Environmental sound recordings often contain intelligible speech, raising privacy concerns that limit analysis, sharing and reuse of data. In this paper, we introduce a method that renders speech unintelligible while preserving both the…

Sound · Computer Science 2025-07-14 Modan Tailleur , Mathieu Lagrange , Pierre Aumond , Vincent Tourre

A state transition model (STM) based on chunk-wise classification was proposed for end-point detection (EPD). In general, EPD is developed using frame-wise voice activity detection (VAD) with additional STM, in which the state transition is…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-24 Juntae Kim , Jaesung Bae , Minsoo Hahn

Voice Conversion (VC) is a technique that aims to transform the non-linguistic information of a source utterance to change the perceived identity of the speaker. While there is a rich literature on VC, most proposed methods are trained and…

We propose a fully unsupervised algorithm that detects from encephalography (EEG) recordings when a subject actively listens to sound, versus when the sound is ignored. This problem is known as absolute auditory attention decoding (aAAD).…

Signal Processing · Electrical Eng. & Systems 2025-04-25 Nicolas Heintz , Tom Francart , Alexander Bertrand

Unsupervised blind source separation methods do not require a training phase and thus cannot suffer from a train-test mismatch, which is a common concern in neural network based source separation. The unsupervised techniques can be…

Sound · Computer Science 2021-06-11 Christoph Boeddeker , Frederik Rautenberg , Reinhold Haeb-Umbach

We propose an algorithm to extract noise-robust acoustic features from noisy speech. We use Total Variability Modeling in combination with Non-negative Matrix Factorization (NMF) to learn a total variability subspace and adapt NMF…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-17 Kunal Dhawan , Colin Vaz , Ruchir Travadi , Shrikanth Narayanan

Audiovisual segmentation (AVS) aims to identify visual regions corresponding to sound sources, playing a vital role in video understanding, surveillance, and human-computer interaction. Traditional AVS methods depend on large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Seung-jae Lee , Paul Hongsuck Seo

We present a cross-modal unsupervised framework for active speaker detection in media content such as TV shows and movies. Machine learning advances have enabled impressive performance in identifying individuals from speech and facial…

Image and Video Processing · Electrical Eng. & Systems 2022-09-27 Rahul Sharma , Shrikanth Narayanan

Rapid advances in singing voice synthesis have increased unauthorized imitation risks, creating an urgent need for better Singing Voice Deepfake (SingFake) Detection, also known as SVDD. Unlike speech, singing contains complex pitch, wide…

Sound · Computer Science 2026-04-07 Xuanjun Chen , Chia-Yu Hu , Sung-Feng Huang , Haibin Wu , Hung-yi Lee , Jyh-Shing Roger Jang

Most cross-domain unsupervised Video Anomaly Detection (VAD) works assume that at least few task-relevant target domain training data are available for adaptation from the source to the target domain. However, this requires laborious…

Computer Vision and Pattern Recognition · Computer Science 2022-12-15 Abhishek Aich , Kuan-Chuan Peng , Amit K. Roy-Chowdhury

Voice conversion is a task to convert a non-linguistic feature of a given utterance. Since naturalness of speech strongly depends on its pitch pattern, in some applications, it would be desirable to keep the original rise/fall pitch pattern…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-21 Chihiro Watanabe , Hirokazu Kameoka