English
Related papers

Related papers: Advanced Signal Analysis in Detecting Replay Attac…

200 papers

Speaker extraction aims to extract target speech signal from a multi-talker environment with interference speakers and surrounding noise, given the target speaker's reference information. Most speaker extraction systems achieve satisfactory…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-12 Chengyun Deng , Shiqian Ma , Yi Zhang , Yongtao Sha , Hui Zhang , Hui Song , Xiangang Li

Recurrence quantification analysis (RQA) is a well established method of nonlinear data analysis. In this work we present a new strategy for an almost parameter-free RQA. The approach finally omits the choice of the threshold parameter by…

Data Analysis, Statistics and Probability · Physics 2022-11-02 Radim Pánis , Karel Adámek , Norbert Marwan

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

It is now well-known that automatic speaker verification (ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofing countermeasure…

Cryptography and Security · Computer Science 2024-01-30 Xuechen Liu , Md Sahidullah , Kong Aik Lee , Tomi Kinnunen

This article develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally…

Applications · Statistics 2011-08-25 Daniel Rudoy , Thomas F. Quatieri , Patrick J. Wolfe

Audio analysis is useful in many application scenarios. The state-of-the-art audio analysis approaches assume the data distribution at training and deployment time will be the same. However, due to various real-life challenges, the data may…

Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Tamás Gábor Csapó

Automatic Speech Recognition (ASR) systems suffer considerably when source speech is corrupted with noise or room impulse responses (RIR). Typically, speech enhancement is applied in both mismatched and matched scenario training and…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-26 Shashi Kumar , Shakti P. Rath , Abhishek Pandey

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and…

Sound · Computer Science 2022-04-05 Ruiteng Zhang , Jianguo Wei , Wenhuan Lu , Lin Zhang , Yantao Ji , Junhai Xu , Xugang Lu

The performance of deep learning methods critically depends on the quality and quantity of the available training data. This is especially the case for physiological time series, which are both noisy and scarce, which calls for data…

Signal Processing · Electrical Eng. & Systems 2025-04-08 Nina Moutonnet , Gregory Scott , Danilo P. Mandic

Speech emotion recognition (SER) has drawn increasing attention for its applications in human-machine interaction. However, existing SER methods ignore the information gap between the pre-training speech recognition task and the downstream…

Sound · Computer Science 2023-10-03 Dongyuan Li , Yusong Wang , Kotaro Funakoshi , Manabu Okumura

Speech signal analysis poses significant challenges, particularly in tasks such as speech quality evaluation and profiling, where the goal is to predict multiple perceptual and objective metrics. For instance, metrics like PESQ (Perceptual…

The performances of automatic speech recognition (ASR) systems are usually evaluated by the metric word error rate (WER) when the manually transcribed data are provided, which are, however, expensively available in the real scenario. In…

Computation and Language · Computer Science 2020-09-01 Kai Fan , Jiayi Wang , Bo Li , Shiliang Zhang , Boxing Chen , Niyu Ge , Zhijie Yan

In recent years, the rapid progress in speaker verification (SV) technology has been driven by the extraction of speaker representations based on deep learning. However, such representations are still vulnerable to emotion variability. To…

Sound · Computer Science 2025-05-27 Jingguang Tian , Xinhui Hu , Xinkang Xu

The notion of recurrence over continuous or dense time, as required for expressing Analog and Mixed-Signal (AMS) behaviours, is fundamentally different from what is offered by the recurrence operators of SystemVerilog Assertions (SVA). This…

Formal Languages and Automata Theory · Computer Science 2020-11-18 Sayandeep Sanyal , Antonio Anastasio Bruto da Costa , Pallab Dasgupta

Large Audio Language Models (LALMs) demonstrate impressive general audio understanding, but once deployed, they are static and fail to improve with new real-world audio data. As traditional supervised fine-tuning is costly, we introduce a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-23 Haoyu Zhang , Jiaxian Guo , Yusuke Iwasawa , Yutaka Matsuo

ASV (automatic speaker verification) systems are intrinsically required to reject both non-target (e.g., voice uttered by different speaker) and spoofed (e.g., synthesised or converted) inputs. However, there is little consideration for how…

We present a novel speaker-independent acoustic-to-articulatory inversion (AAI) model, overcoming the limitations observed in conventional AAI models that rely on acoustic features derived from restricted datasets. To address these…

Signal Processing · Electrical Eng. & Systems 2024-06-26 Woo-Jin Chung , Hong-Goo Kang

Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community. In particular, speech representations learned by SSL models have been shown to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-11 Eesung Kim , Jae-Jin Jeon , Hyeji Seo , Hoon Kim

In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-16 Rajeev Rikhye , Quan Wang , Qiao Liang , Yanzhang He , Ding Zhao , Yiteng , Huang , Arun Narayanan , Ian McGraw