English
Related papers

Related papers: Spatial-temporal Graph Based Multi-channel Speaker…

200 papers

This paper investigates the application of environmental feature representations for room verification tasks and acoustic meta-data estimation. Audio recordings contain both speaker and non-speaker information. We refer to the…

Sound · Computer Science 2022-03-10 Desmond Caulley

Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Michael Neri , Tuomas Virtanen

Recently, speech enhancement technologies that are based on deep learning have received considerable research attention. If the spatial information in microphone signals is exploited, microphone arrays can be advantageous under some adverse…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-19 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

A three-stage approach is proposed for speaker counting and speech separation in noisy and reverberant environments. In the spatial feature extraction, a spatial coherence matrix (SCM) is computed using whitened relative transfer functions…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Yicheng Hsu , Mingsian Bai

The convolutional neural network (CNN) based approaches have shown great success for speaker verification (SV) tasks, where modeling long temporal context and reducing information loss of speaker characteristics are two important challenges…

Sound · Computer Science 2021-08-31 Yanfeng Wu , Chenkai Guo , Junan Zhao , Xiao Jin , Jing Xu

In recent decades, neural network based methods have significantly improved the performace of speech enhancement. Most of them estimate time-frequency (T-F) representation of target speech directly or indirectly, then resynthesize waveform…

Sound · Computer Science 2020-02-06 Jingdong Li , Hui Zhang , Xueliang Zhang , Changliang Li

In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of…

Computer Vision and Pattern Recognition · Computer Science 2018-06-08 Amirsina Torfi , Jeremy Dawson , Nasser M. Nasrabadi

Convolutional neural networks (CNN) are one of the best-performing neural network architectures for environmental sound classification (ESC). Recently, temporal attention mechanisms have been used in CNN to capture the useful information…

Sound · Computer Science 2020-05-22 Helin Wang , Yuexian Zou , Dading Chong , Wenwu Wang

Automatic Modulation Recognition (AMR) is an essential part of Intelligent Transportation System (ITS) dynamic spectrum allocation. However, current deep learning-based AMR (DL-AMR) methods are challenged to extract discriminative and…

Signal Processing · Electrical Eng. & Systems 2025-08-19 Mingyuan Shao , Zhengqiu Fu , Dingzhao Li , Fuqing Zhang , Yilin Cai , Shaohua Hong , Lin Cao , Yuan Peng , Jie Qi

Audio deepfake detection is increasingly important as synthetic speech becomes more realistic and accessible. Recent methods, including those using graph neural networks (GNNs) to model frequency and temporal dependencies, show strong…

Sound · Computer Science 2026-01-13 Falih Gozi Febrinanto , Kristen Moore , Chandra Thapa , Jiangang Ma , Vidya Saikrishna

Audio tagging aims at predicting sound events occurred in a recording. Traditional models require enormous laborious annotations, otherwise performance degeneration will be the norm. Therefore, we investigate robust audio tagging models in…

Sound · Computer Science 2021-10-05 Zhiling Zhang , Zelin Zhou , Haifeng Tang , Guangwei Li , Mengyue Wu , Kenny Q. Zhu

Single-channel speech separation in time domain and frequency domain has been widely studied for voice-driven applications over the past few years. Most of previous works assume known number of speakers in advance, however, which is not…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-02 Yiming Xiao , Haijian Zhang

Utilizing the large-scale unlabeled data from the target domain via pseudo-label clustering algorithms is an important approach for addressing domain adaptation problems in speaker verification tasks. In this paper, we propose a novel…

Sound · Computer Science 2023-05-23 Zhuo Li , Jingze Lu , Zhenduo Zhao , Wenchao Wang , Pengyuan Zhang

This paper proposes a novel framework for lung sound event detection, segmenting continuous lung sound recordings into discrete events and performing recognition on each event. Exploiting the lightweight nature of Temporal Convolution…

Distributed Microphone Arrays (DMAs) present many challenges with respect to centralized microphone arrays. An important requirement of applications on these arrays is handling a variable number of input channels. We consider the use of…

Sound · Computer Science 2023-06-29 Eric Grinstein , Mike Brookes , Patrick A. Naylor

In speech enhancement, an end-to-end deep neural network converts a noisy speech signal to a clean speech directly in time domain without time-frequency transformation or mask estimation. However, aggregating contextual information from a…

Sound · Computer Science 2020-02-10 Kai Zhen , Mi Suk Lee , Minje Kim

In speaker verification, traditional models often emphasize modeling long-term contextual features to capture global speaker characteristics. However, this approach can neglect fine-grained voiceprint information, which contains highly…

Sound · Computer Science 2025-05-07 Ya Li , Bin Zhou , Bo Hu

This paper introduces an online speaker diarization system that can handle long-time audio with low latency. We enable Agglomerative Hierarchy Clustering (AHC) to work in an online fashion by introducing a label matching algorithm. This…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-27 Yucong Zhang , Qinjian Lin , Weiqing Wang , Lin Yang , Xuyang Wang , Junjie Wang , Ming Li

To improve speaker verification in real scenarios with interference speakers, noise, and reverberation, we propose to bring together advancements made in multi-channel speech features. Specifically, we combine spectral, spatial, and…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-12 Saurabh Kataria , Shi-Xiong Zhang , Dong Yu

Artefacts that serve to distinguish bona fide speech from spoofed or deepfake speech are known to reside in specific subbands and temporal segments. Various approaches can be used to capture and model such artefacts, however, none works…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-24 Hemlata Tak , Jee-weon Jung , Jose Patino , Madhu Kamble , Massimiliano Todisco , Nicholas Evans