English
Related papers

Related papers: Inferring Pitch from Coarse Spectral Features

200 papers

Contamination of covariates by measurement error is a classical problem in multivariate regression, where it is well known that failing to account for this contamination can result in substantial bias in the parameter estimators. The nature…

Methodology · Statistics 2017-12-13 Anirvan Chakraborty , Victor M. Panaretos

Accurate and real-time monophonic pitch estimation in noisy conditions, particularly on resource-constrained devices, remains an open challenge in audio processing. We present \emph{SwiftF0}, a novel, lightweight neural model that sets a…

Sound · Computer Science 2025-08-27 Lars Nieradzik

Despite rapid advancement in recent years, current speech enhancement models often produce speech that differs in perceptual quality from real clean speech. We propose a learning objective that formalizes differences in perceptual quality,…

In music and speech, meaning is derived at multiple levels of context. Affect, for example, can be inferred both by a short sound token and by sonic patterns over a longer temporal window such as an entire recording. In this letter, we…

Sound · Computer Science 2022-09-12 Camille Noufi , Prateek Verma

Reduced articulatory precision is common in speech, but for dialog its acoustic properties and pragmatic functions have been little studied. We here try to remedy this gap. This technical report contains content that was omitted from the…

Computation and Language · Computer Science 2024-05-03 Nigel G. Ward , Carlos A. Ortega

This paper presents a new method of singing voice analysis that performs mutually-dependent singing voice separation and vocal fundamental frequency (F0) estimation. Vocal F0 estimation is considered to become easier if singing voices can…

Sound · Computer Science 2016-11-29 Yukara Ikemiya , Katsutoshi Itoyama , Kazuyoshi Yoshii

We have developed a sparse mathematical representation of speech that minimizes the number of active model neurons needed to represent typical speech sounds. The model learns several well-known acoustic features of speech such as harmonic…

Neurons and Cognition · Quantitative Biology 2012-09-25 Nicole L. Carlson , Vivienne L. Ming , Michael R. DeWeese

Covariance regression analysis is an approach to linking the covariance of responses to a set of explanatory variables $X$, where $X$ can be a vector, matrix, or tensor. Most of the literature on this topic focuses on the "Fixed-$X$"…

Statistics Theory · Mathematics 2025-01-08 Tao Zou , Wei Lan , Runze Li , Chih-Ling Tsai

The goal of this work is to recover articulatory information from the speech signal by acoustic-to-articulatory inversion. One of the main difficulties with inversion is that the problem is underdetermined and inversion methods generally…

Computation and Language · Computer Science 2007-05-23 Blaise Potard , Yves Laprie

Two and a half millennia ago Pythagoras initiated the scientific study of the pitch of sounds; yet our understanding of the mechanisms of pitch perception remains incomplete. Physical models of pitch perception try to explain from…

Chaotic Dynamics · Physics 2012-07-24 Julyan H. E. Cartwright , Diego L. Gonzalez , Oreste Piro

Acoustic context effects, where surrounding changes in pitch, rate or timbre influence the perception of a sound, are well documented in speech perception, but how they interact with language background remains unclear. Using a…

Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-10 Aníbal J. S. Ferreira , Luis M. T. Jesus , Laurentino M. M. Leal , Jorge E. F. Spratley

In this paper I report on an investigation into the problem of assigning tones to pitch contours. The proposed model is intended to serve as a tool for phonologists working on instrumentally obtained pitch data from tone languages.…

cmp-lg · Computer Science 2008-02-03 Steven Bird

In this paper, a method of pitch tracking based on variance minimization of locally periodic subsamples of an acoustic signal is presented. Replicates along the length of the periodically sampled data of the signal vector are taken and…

Sound · Computer Science 2009-09-29 Roudra Chakraborty , Debapriya Sengupta , Sagnik Sinha

Pitch is a foundational aspect of our perception of audio signals. Pitch contours are commonly used to analyze speech and music signals and as input features for many audio tasks, including music transcription, singing voice synthesis, and…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-13 Max Morrison , Caedon Hsieh , Nathan Pruyne , Bryan Pardo

This chapter reconsiders the concept of pitch in contemporary popular music (CPM), particularly in electronic contexts where traditional assumptions may fail. Drawing on phenomenological and inductive methods, it argues that pitch is not an…

Sound · Computer Science 2025-07-08 Emmanuel Deruty

This study focuses on generating fundamental frequency (F0) curves of singing voice from musical scores stored in a midi-like notation. Current statistical parametric approaches to singing F0 modeling meet difficulties in reproducing…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-13 Kanru Hua

Unlike text, speech conveys information about the speaker, such as gender, through acoustic cues like pitch. This gives rise to modality-specific bias concerns. For example, in speech translation (ST), when translating from languages with…

Computation and Language · Computer Science 2026-04-29 Lina Conti , Dennis Fucci , Marco Gaido , Matteo Negri , Guillaume Wisniewski , Luisa Bentivogli

Cough is a primary symptom of most respiratory diseases, and changes in cough characteristics provide valuable information for diagnosing respiratory diseases. The characterization of cough sounds still lacks concrete evidence, which makes…

Sound · Computer Science 2023-08-08 Naveenkumar Vodnala , Pratap Reddy Lankireddy , Padmasai Yarlagadda

Recently it was shown that within the Silent Speech Interface (SSI) field, the prediction of F0 is possible from Ultrasound Tongue Images (UTI) as the articulatory input, using Deep Neural Networks for articulatory-to-acoustic mapping.…