English
Related papers

Related papers: Inferring Pitch from Coarse Spectral Features

200 papers

We aim to explain whether a stress memory task has a significant impact on tonal coarticulation. We contribute a novel approach to analyse tonal coarticulation in phonetics, where several f0 contours are compared with respect to their…

Applications · Statistics 2024-09-10 Valentina Masarotto , Yiya Chen

New formulae for the resonant scattering and the production amplitudes near an inelastic threshold are derived. It is shown that the Flatte formula, frequently used in experimental analyses, is not sufficiently accurate. Its application to…

High Energy Physics - Phenomenology · Physics 2009-11-13 L. Lesniak

Learning a new language involves constantly comparing speech productions with reference productions from the environment. Early in speech acquisition, children make articulatory adjustments to match their caregivers' speech. Grownup…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Talia Ben-Simon , Felix Kreuk , Faten Awwad , Jacob T. Cohen , Joseph Keshet

In this work, we present the misspecified Gaussian Cram\'er-Rao lower bound for the parameters of a harmonic signal, or pitch, when signal measurements are collected from an almost, but not quite, harmonic model. For the asymptotic case of…

Signal Processing · Electrical Eng. & Systems 2019-10-29 Filip Elvander , Jie Ding , Andreas Jakobsson

A large number of works view the automatic assessment of speech from an utterance- or system-level perspective. While such approaches are good in judging overall quality, they cannot adequately explain why a certain score was assigned to an…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-30 Michael Kuhlmann , Alexander Werning , Thilo von Neumann , Reinhold Haeb-Umbach

Frequency is a central concept in Mathematics, Physics, and Signal Processing. It is the main tool for describing the oscillatory behavior of signals, which is usually argued to be the manifestation of some of their key features, depending…

Signal Processing · Electrical Eng. & Systems 2021-05-28 Móises Soto-Bajo , Andrés Fraguela Collar , Javier Herrera Vega , Raúl Felipe-Sosa

Speaker verification is the process by which a speakers claim of identity is tested against a claimed speaker by his or her voice. Speaker verification is done by the use of some parameters (features) from the speakers voice which can be…

Sound · Computer Science 2019-08-16 Bhavana V. S , Pradip K. Das

The present study has two goals relating to the grammar of prosody, understood as the rhythms and melodies of speech. First, an overview is provided of the computable grammatical and phonetic approaches to prosody analysis which use…

Computation and Language · Computer Science 2019-12-17 Dafydd Gibbon

Automatic speech recognition (ASR) systems are known to be sensitive to the sociolinguistic variability of speech data, in which gender plays a crucial role. This can result in disparities in recognition accuracy between male and female…

Computation and Language · Computer Science 2023-10-11 Dennis Fucci , Marco Gaido , Matteo Negri , Mauro Cettolo , Luisa Bentivogli

We propose the product-of-filters (PoF) model, a generative model that decomposes audio spectra as sparse linear combinations of "filters" in the log-spectral domain. PoF makes similar assumptions to those used in the classic homomorphic…

Machine Learning · Statistics 2014-11-27 Dawen Liang , Matthew D. Hoffman , Gautham J. Mysore

The success of non-linear optics relies largely on pulse-to-pulse consistency. In contrast, covariance based techniques used in photoionization electron spectroscopy and mass spectrometry have shown that wealth of information can be…

Whisper, as a form of speech, is not sufficiently addressed by mainstream speech applications. This is due to the fact that systems built for normal speech do not work as expected for whispered speech. A first step to building a speech…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 S. Johanan Joysingh , P. Vijayalakshmi , T. Nagarajan

Functional linear regression is an important topic in functional data analysis. It is commonly assumed that samples of the functional predictor are independent realizations of an underlying stochastic process, and are observed over a grid…

Methodology · Statistics 2020-09-15 Cheng Chen , Shaojun Guo , Xinghao Qiao

Given data $y$ and $k$ covariates $x$ the problem is to decide which covariates to include when approximating $y$ by a linear function of the covariates. The decision is based on replacing subsets of the covariates by i.i.d. normal random…

Statistics Theory · Mathematics 2016-05-09 Laurie Davies

Recent advances in machine learning and the availability of articulatory datasets allow vocal tract synthesis to be conditioned on phonetic sequences, a primary task of articulatory speech synthesis. However, quality assessment needs a…

Computation and Language · Computer Science 2026-05-21 Vinicius Ribeiro , Yves Laprie

Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…

Sound · Computer Science 2007-05-23 Friedhelm R. Drepper

Drawing causal inference with observational studies is the central pillar of many disciplines. One sufficient condition for identifying the causal effect is that the treatment-outcome relationship is unconfounded conditional on the observed…

Statistics Theory · Mathematics 2017-01-17 Peng Ding , Tyler VanderWeele , James Robins

Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio, different people may perceive the same audio differently, resulting in caption disparities (i.e., one audio may correlate to…

Sound · Computer Science 2022-04-19 Yiming Zhang , Hong Yu , Ruoyi Du , Zhanyu Ma , Yuan Dong

This paper concerns the inverse source scattering problems of recovering random sources for acoustic and elastic waves. The underlying sources are assumed to be random functions driven by an additive white noise. The inversion process aims…

Numerical Analysis · Mathematics 2024-12-10 Yan Chang , Yukun Guo , Zhipeng Yang , Yue Zhao

Partial audio deepfake localization poses unique challenges and remain underexplored compared to full-utterance spoofing detection. While recent methods report strong in-domain performance, their real-world utility remains unclear. In this…

Sound · Computer Science 2025-09-01 Hieu-Thi Luong , Inbal Rimon , Haim Permuter , Kong Aik Lee , Eng Siong Chng