English
Related papers

Related papers: Maximum Voiced Frequency Estimation: Exploiting Am…

200 papers

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from the last layer of a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Youngmoon Jung , Seong Min Kye , Yeunju Choi , Myunghun Jung , Hoirin Kim

This article focuses on estimating relative transfer functions (RTFs) for beamforming applications. Traditional methods often assume that spectra are uncorrelated, an assumption that is often violated in practical scenarios due to factors…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Giovanni Bologni , Richard C. Hendriks , Richard Heusdens

The advancements of AI-synthesized human voices have introduced a growing threat of impersonation and disinformation. It is therefore of practical importance to developdetection methods for synthetic human voices. This work proposes a new…

Sound · Computer Science 2023-04-28 Chengzhe Sun , Shan Jia , Shuwei Hou , Ehab AlBadawy , Siwei Lyu

Singing voice synthesis (SVS), as a specific task for generating the vocal singing voice from a music score, has drawn much attention in recent years. SVS faces the challenge that the singing has various pronunciation flexibility…

Sound · Computer Science 2023-03-16 Yuning Wu , Jiatong Shi , Tao Qian , Dongji Gao , Qin Jin

Singing voice separation (SVS) is a task that separates singing voice audio from its mixture with instrumental audio. Previous SVS studies have mainly employed the spectrogram masking method which requires a large dimensionality in…

Sound · Computer Science 2022-11-30 Jaekwon Im , Soonbeom Choi , Sangeon Yong , Juhan Nam

This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction, the BiVocoder takes…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Hui-Peng Du , Ye-Xin Lu , Yang Ai , Zhen-Hua Ling

Recent advances in generative models have amplified the risk of malicious misuse of speech synthesis technologies, enabling adversaries to impersonate target speakers and access sensitive resources. Although speech deepfake detection has…

Speech generated by parametric synthesizers generally suffers from a typical buzziness, similar to what was encountered in old LPC-like vocoders. In order to alleviate this problem, a more suited modeling of the excitation should be…

Sound · Computer Science 2020-01-06 Thomas Drugman , Geoffrey Wilfart , Thierry Dutoit

We present an automated pipeline for estimating Verb Frame Frequencies (VFFs), the frequency with which a verb appears in particular syntactic frames. VFFs provide a powerful window into syntax in both human and machine language systems,…

Computation and Language · Computer Science 2025-07-31 Adam M. Morgan , Adeen Flinker

Usually, hearing impaired people use hearing aids which are implemented with speech enhancement algorithms. Estimation of speech and estimation of nose are the components in single channel speech enhancement system. The main objective of…

Sound · Computer Science 2014-11-10 M. Ravichandra Kumar , B. Ravi Teja

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-05 Meng Ge , Chenglin Xu , Longbiao Wang , Eng Siong Chng , Jianwu Dang , Haizhou Li

The purpose of this note is to show how the method of maximum entropy in the mean (MEM) may be used to improve parametric estimation when the measurements are corrupted by large level of noise. The method is developed in the context on a…

Machine Learning · Computer Science 2021-08-23 Henryk Gzyl , Enrique ter Horst

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Automatic speaker verification (ASV) is the process to recognize persons using voice as biometric. The ASV systems show considerable recognition performance with sufficient amount of speech from matched condition. One of the crucial…

Multimedia · Computer Science 2018-12-04 Arnab Poddar , Md Sahidullah , Goutam Saha

In this work, we propose a novel consistency-preserving loss function for recovering the phase information in the context of phase reconstruction (PR) and speech enhancement (SE). Different from conventional techniques that directly…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-25 Pin-Jui Ku , Chun-Wei Ho , Hao Yen , Sabato Marco Siniscalchi , Chin-Hui Lee

In this contribution, we present a novel online approach to multichannel speech enhancement. The proposed method estimates the enhanced signal through a filter-and-sum framework. More specifically, complex-valued masks are estimated by a…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-09 Mhd Modar Halimeh , Walter Kellermann

For the difficulty and large computational complexity of modeling more frequency bands, full-band speech enhancement based on deep neural networks is still challenging. Previous studies usually adopt compressed full-band speech features in…

Sound · Computer Science 2022-08-02 Guochen Yu , Yuansheng Guan , Weixin Meng , Chengshi Zheng , Hui Wang

Phase serves as a critical component of speech that influences the quality and intelligibility. Current speech enhancement algorithms are beginning to address phase distortions, but the algorithms focus on normal-hearing (NH) listeners. It…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-30 Zhuohuang Zhang , Donald S. Williamson , Yi Shen

Automatic speech quality assessment plays a crucial role in the development of speech synthesis systems, but existing models exhibit significant performance variations across different granularity levels of prediction tasks. This paper…

Sound · Computer Science 2025-07-09 Xintong Hu , Yixuan Chen , Rui Yang , Wenxiang Guo , Changhao Pan

The estimation of glottal flow from a speech waveform is a key method for speech analysis and parameterization. Significant research effort has been made to dissociate the first vocal tract resonance from the glottal formant (the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-09 Olivier Perrotin , Ian Vince McLoughlin