English
Related papers

Related papers: On the Use of a Spectral Glottal Model for the Sou…

200 papers

We compare the Frequency-Resolved Frozen Phonon Multislice (FRFPMS) method, introduced in Phys. Rev. Lett. 124, 025501 (2020), with other theoretical approaches used to account for the inelastic scattering of high energy electrons, namely…

Materials Science · Physics 2023-04-27 Paul M. Zeiger , Juri Barthel , Leslie J. Allen , Ján Rusz

Speaker Verification (SV) systems trained on adults speech often underperform on children's SV due to the acoustic mismatch, and limited children speech data makes fine-tuning not very effective. In this paper, we propose an innovative…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Vishwas M. Shetty , Jiusi Zheng , Abeer Alwan

A primary challenge in developing synthetic spatial hearing systems, particularly underwater, is accurately modeling sound scattering. Biological organisms achieve 3D spatial hearing by exploiting sound scattering off their bodies to…

Sound · Computer Science 2026-03-03 Siminfar Samakoush Galougah , Pranav Pulijala , Ramani Duraiswami

The binaural minimum-variance distortionless-response (BMVDR) beamformer is a well-known noise reduction algorithm that can be steered using the relative transfer function (RTF) vector of the desired speech source. Exploiting the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-22 Nico Gößling , Wiebke Middelberg , Simon Doclo

Adaptive Local Iterative Filtering (ALIF) is a currently proposed novel time-frequency analysis tool. It has been empirically shown that ALIF is able to separate components and overcome the mode-mixing problem. However, so far its…

Numerical Analysis · Mathematics 2020-05-12 Antonio Cicone , Hau-Tieng Wu

Spectral subtraction, widely used for its simplicity, has been employed to address the Robot Ego Speech Filtering (RESF) problem for detecting speech contents of human interruption from robot's single-channel microphone recordings when it…

Robotics · Computer Science 2024-09-11 Yue Li , Koen V. Hindriks , Florian A. Kunneman

Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-01 Saba Tabatabaee , Suzanne Boyce , Liran Oren , Mark Tiede , Carol Espy-Wilson

We study permutation invariant training (PIT), which targets at the permutation ambiguity problem for speaker independent source separation models. We extend two state-of-the-art PIT strategies. First, we look at the two-stage speaker…

Sound · Computer Science 2021-04-06 Xiaoyu Liu , Jordi Pons

Disorders of voice production have severe effects on the quality of life of the affected individuals. A simulation approach is used to investigate the cause-effect chain in voice production showing typical characteristics of voice such as…

Sound · Computer Science 2022-07-20 Florian Kraxberger , Andreas Wurzinger , Stefan Schoder

Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-10 Hongyu Wang , Hui Li , Bo Li

Score-based generative models (SGMs) have recently shown impressive results for difficult generative tasks such as the unconditional and conditional generation of natural images and audio signals. In this work, we extend these models to the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-08 Simon Welker , Julius Richter , Timo Gerkmann

Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-21 Tsun-An Hsieh , Minje Kim

Audio-visual speech separation aims to isolate each speaker's clean voice from mixtures by leveraging visual cues such as lip movements and facial features. While visual information provides complementary semantic guidance, existing methods…

Sound · Computer Science 2025-10-13 Ke Xue , Rongfei Fan , Lixin , Dawei Zhao , Chao Zhu , Han Hu

Time-frequency analysis for non-linear and non-stationary signals is extraordinarily challenging. To capture features in these signals, it is necessary for the analysis methods to be local, adaptive and stable. In recent years,…

Numerical Analysis · Mathematics 2015-10-26 Antonio Cicone , Jingfang Liu , Haomin Zhou

Diffusion models are receiving a growing interest for a variety of signal generation tasks such as speech or music synthesis. WaveGrad, for example, is a successful diffusion model that conditionally uses the mel spectrogram to guide a…

Sound · Computer Science 2024-02-27 Haocheng Liu , Teysir Baoueb , Mathieu Fontaine , Jonathan Le Roux , Gael Richard

It was recently shown that complex cepstrum can be effectively used for glottal flow estimation by separating the causal and anticausal components of speech. In order to guarantee a correct estimation, some constraints on the window have…

Sound · Computer Science 2020-05-12 Thomas Drugman , Thierry Dutoit

In this work, we present a method for learning interpretable music signal representations directly from waveform signals. Our method can be trained using unsupervised objectives and relies on the denoising auto-encoder model that uses a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Stylianos I. Mimilakis , Konstantinos Drossos , Gerald Schuller

Diffusion models have shown promising results in speech enhancement, using a task-adapted diffusion process for the conditional generation of clean speech given a noisy mixture. However, at test time, the neural network used for score…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-17 Bunlong Lay , Jean-Marie Lemercier , Julius Richter , Timo Gerkmann

Iterative Filtering (IF) is an alternative technique to the Empirical Mode Decomposition (EMD) algorithm for the decomposition of non-stationary and non-linear signals. Recently in [1] IF has been proved to be convergent for any $L^2$…

Numerical Analysis · Mathematics 2015-07-28 Antonio Cicone , Haomin Zhou

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono