English
Related papers

Related papers: Waveform-based Voice Activity Detection Exploiting…

200 papers

Complex-valued processing has brought deep learning-based speech enhancement and signal extraction to a new level. Typically, the process is based on a time-frequency (TF) mask which is applied to a noisy spectrogram, while complex masks…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-02 Hendrik Schröter , Alberto N. Escalante-B. , Tobias Rosenkranz , Andreas Maier

This article describes the development of a novel U-Net-enhanced Wavelet Neural Operator (U-WNO),which combines wavelet decomposition, operator learning, and an encoder-decoder mechanism. This approach harnesses the superiority of the…

Image and Video Processing · Electrical Eng. & Systems 2024-11-27 Pranava Seth , Deepak Mishra , Veena Iyer

Chord recognition systems depend on robust feature extraction pipelines. While these pipelines are traditionally hand-crafted, recent advances in end-to-end machine learning have begun to inspire researchers to explore data-driven methods…

Machine Learning · Computer Science 2016-12-16 Filip Korzeniowski , Gerhard Widmer

Most deep learning-based models for speech enhancement have mainly focused on estimating the magnitude of spectrogram while reusing the phase from noisy speech for reconstruction. This is due to the difficulty of estimating the phase of…

Sound · Computer Science 2019-04-03 Hyeong-Seok Choi , Jang-Hyun Kim , Jaesung Huh , Adrian Kim , Jung-Woo Ha , Kyogu Lee

Direct speech-to-text translation (ST) models are usually trained on corpora segmented at sentence level, but at inference time they are commonly fed with audio split by a voice activity detector (VAD). Since VAD segmentation is not…

Computation and Language · Computer Science 2020-08-06 Marco Gaido , Mattia Antonino Di Gangi , Matteo Negri , Mauro Cettolo , Marco Turchi

Audio-visual speech enhancement system is regarded to be one of promising solutions for isolating and enhancing speech of desired speaker. Conventional methods focus on predicting clean speech spectrum via a naive convolution neural network…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-28 Xinmeng Xu , Jianjun Hao

This paper focuses on explaining the timbre conveyed by speech signals and introduces a task termed voice timbre attribute detection (vTAD). In this task, voice timbre is explained with a set of sensory attributes describing its human…

Sound · Computer Science 2025-06-24 Jinghao He , Zhengyan Sheng , Liping Chen , Kong Aik Lee , Zhen-Hua Ling

Active speaker detection plays a vital role in human-machine interaction. Recently, a few end-to-end audiovisual frameworks emerged. However, these models' inference time was not explored and are not applicable for real-time applications…

Sound · Computer Science 2022-11-24 Fiseha B. Tesema , Zheyuan Lin , Shiqiang Zhu , Wei Song , Jason Gu , Hong Wu

This article focuses on overlapped speech and gender detection in order to study interactions between women and men in French audiovisual media (Gender Equality Monitoring project). In this application context, we need to automatically…

Sound · Computer Science 2022-09-12 Martin Lebourdais , Marie Tahon , Antoine Laurent , Sylvain Meignier

Open-vocabulary change detection aims to identify semantic changes in bi-temporal remote sensing images without predefined categories. Recent methods combine foundation models such as SAM, DINO and CLIP, but typically process each timestamp…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Zuzheng Kuang , Honghao Chang , Boqiang Liang , Haoqian Wang , Lijun He , Fan Li , Haixia Bi

Supervised multi-channel audio source separation requires extracting useful spectral, temporal, and spatial features from the mixed signals. The success of many existing systems is therefore largely dependent on the choice of features used…

Sound · Computer Science 2018-03-05 Emad M. Grais , Dominic Ward , Mark D. Plumbley

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with…

Sound · Computer Science 2024-09-17 Anna Wang , Da Liu , Zhiyu Zhang , Shengqiang Liu , Jie Gao , Yali Li

The history of computing started with analog computers consisting of physical devices performing specialized functions such as predicting the trajectory of cannon balls. In modern times, this idea has been extended, for example, to…

Image and Video Processing · Electrical Eng. & Systems 2022-08-29 Callen MacPhee , Bahram Jalali

Video Anomaly Detection (VAD) is a fundamental challenge in computer vision, particularly due to the open-set nature of anomalies. While recent training-free approaches utilizing Vision-Language Models (VLMs) have shown promise, they…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Lokman Bekit , Hamza Karim , Nghia T Nguyen , Yasin Yilmaz

Overlapped speech detection (OSD) is critical for speech applications in scenario of multi-party conversion. Despite numerous research efforts and progresses, comparing with speech activity detection (VAD), OSD remains an open challenge and…

Sound · Computer Science 2022-09-27 Ziqing Du , Kai Liu , Xucheng Wan , Huan Zhou

This paper aims at eliminating the interfering speakers' speech, additive noise, and reverberation from the noisy multi-talker speech mixture that benefits automatic speech recognition (ASR) backend. While the recently proposed Weighted…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-19 Zhaoheng Ni , Yong Xu , Meng Yu , Bo Wu , Shixiong Zhang , Dong Yu , Michael I Mandel

This study investigated the waveform representation for audio signal classification. Recently, many studies on audio waveform classification such as acoustic event detection and music genre classification have been published. Most studies…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-19 Masaki Okawa , Takuya Saito , Naoki Sawada , Hiromitsu Nishizaki

Vector Quantized Variational AutoEncoders (VQ-VAE) are a powerful representation learning framework that can discover discrete groups of features from a speech signal without supervision. Until now, the VQ-VAE architecture has previously…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Yi Zhao , Haoyu Li , Cheng-I Lai , Jennifer Williams , Erica Cooper , Junichi Yamagishi

Environmental sound recordings often contain intelligible speech, raising privacy concerns that limit analysis, sharing and reuse of data. In this paper, we introduce a method that renders speech unintelligible while preserving both the…

Sound · Computer Science 2025-07-14 Modan Tailleur , Mathieu Lagrange , Pierre Aumond , Vincent Tourre

2D convolution is widely used in sound event detection (SED) to recognize two dimensional time-frequency patterns of sound events. However, 2D convolution enforces translation equivariance on sound events along both time and frequency axis…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-05 Hyeonuk Nam , Seong-Hu Kim , Byeong-Yun Ko , Yong-Hwa Park
‹ Prev 1 8 9 10 Next ›