English
Related papers

Related papers: Causal-Anticausal Decomposition of Speech using Co…

200 papers

This paper presents an Expert Decision Support System for the identification of time-invariant, aeroacoustic source types. The system comprises two steps: first, acoustic properties are calculated based on spectral and spatial information.…

Sound · Computer Science 2022-03-09 Armin Goudarzi , Carsten Spehr , Steffen Herbold

The increasing size and complexity of medical imaging datasets, particularly in 3D formats, present significant barriers to collaborative research and transferability. This study investigates whether the ZFP compression technique can…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Shimaa Elbana , Ahmad Kamal , Shahd Ahmed Ali , Ahmad Al-Kabbany

Glottal Closure Instants (GCIs) correspond to the temporal locations of significant excitation to the vocal tract occurring during the production of voiced speech. GCI detection from speech signals is a well-studied problem given its…

Sound · Computer Science 2019-07-11 Mohit Goyal , Varun Srivastava , Prathosh A. P

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

Sound · Computer Science 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

Sonar systems are frequently used to classify objects at a distance by using the structure of the echoes of acoustic waves as a proxy for the object's shape and composition. Traditional synthetic aperture processing is highly effective in…

Computational Engineering, Finance, and Science · Computer Science 2021-12-14 Michael Robinson

Glottal Closure Instants (GCI) detection consists in automatically detecting temporal locations of most significant excitation of the vocal tract from the speech signal. It is used in many speech analysis and processing applications, and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-21 Luc Ardaillon , Axel Roebel

We introduce DeCaFlow, a deconfounding causal generative model. Training once per dataset using just observational data and the underlying causal graph, DeCaFlow enables accurate causal inference on continuous variables under the presence…

Machine Learning · Computer Science 2025-10-27 Alejandro Almodóvar , Adrián Javaloy , Juan Parras , Santiago Zazo , Isabel Valera

Current accent conversion (AC) systems do not disentangle the two main sources of non-native accent: segmental and prosodic characteristics. Being able to manipulate a non-native speaker's segmental and/or prosodic channels independently is…

Computation and Language · Computer Science 2024-08-21 Waris Quamer , Ricardo Gutierrez-Osuna

Formant tracking is one of the most fundamental problems in speech processing. Traditionally, formants are estimated using signal processing methods. Recent studies showed that generic convolutional architectures can outperform recurrent…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-11 Wang Dai , Jinsong Zhang , Yingming Gao , Wei Wei , Dengfeng Ke , Binghuai Lin , Yanlu Xie

This paper introduces a computationally efficient technique for estimating high-resolution Doppler blood flow from an ultrafast ultrasound image sequence. More precisely, it consists in a new fast alternating minimization algorithm that…

Image and Video Processing · Electrical Eng. & Systems 2020-11-04 Duong-Hung Pham , Adrian Basarab , Jean-Pierre Remenieras , Paul Rodríguez , Denis Kouamé

We propose to combine cepstrum and nonlinear time-frequency (TF) analysis to study mutiple component oscillatory signals with time-varying frequency and amplitude and with time-varying non-sinusoidal oscillatory pattern. The concept of…

Data Analysis, Statistics and Probability · Physics 2016-11-23 Chen-Yun Lin , Li Su , Hau-tieng Wu

The effectiveness of zero-shot classification in large vision-language models (VLMs), such as Contrastive Language-Image Pre-training (CLIP), depends on access to extensive, well-aligned text-image datasets. In this work, we introduce two…

Computer Vision and Pattern Recognition · Computer Science 2025-02-27 Anju Rani , Daniel O. Arroyo , Petar Durdevic

Zero-resource speech technology is a growing research area that aims to develop methods for speech processing in the absence of transcriptions, lexicons, or language modelling text. Early term discovery systems focused on identifying…

Computation and Language · Computer Science 2017-09-19 Herman Kamper , Aren Jansen , Sharon Goldwater

While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign…

Sound · Computer Science 2026-02-06 Jinzuomu Zhong , Korin Richmond , Zhiba Su , Siqi Sun

Traditional speech enhancement techniques modify the magnitude of a speech in time-frequency domain, and use the phase of a noisy speech to resynthesize a time domain speech. This work proposes a complex-valued Gaussian process latent…

Sound · Computer Science 2017-01-02 Sih-Huei Chen , Yuan-Shan Lee , Jia-Ching Wang

Music source separation is important for applications such as karaoke and remixing. Much of previous research focuses on estimating short-time Fourier transform (STFT) magnitude and discarding phase information. We observe that, for singing…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Yixuan Zhang , Yuzhou Liu , DeLiang Wang

Flow cytometry mainly used for detecting the characteristics of a number of biochemical substances based on the expression of specific markers in cells. It is particularly useful for detecting membrane surface receptors, antigens, ions, or…

Machine Learning · Computer Science 2023-03-17 Yanhua Xu

Purpose: This work explores the use of external phrase break prediction models to enhance listener comprehension in End-to-End Text-to-Speech (TTS) systems. Methods: The effectiveness of these models is evaluated based on listener…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-31 Anandaswarup Vadapalli

This paper introduces GlOttal-flow LPC Filter (GOLF), a novel method for singing voice synthesis (SVS) that exploits the physical characteristics of the human voice using differentiable digital signal processing. GOLF employs a glottal…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-21 Chin-Yun Yu , György Fazekas

Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…

Sound · Computer Science 2007-05-23 Friedhelm R. Drepper