English
Related papers

Related papers: Inferring Pitch from Coarse Spectral Features

200 papers

Jitter and shimmer measurements have shown to be carriers of voice quality and prosodic information which enhance the performance of tasks like speaker recognition, diarization or automatic speech recognition (ASR). However, such features…

Computation and Language · Computer Science 2021-12-22 Guillermo Cámbara , Jordi Luque , Mireia Farrús

Fundamental frequency (f0) estimation from polyphonic music includes the tasks of multiple-f0, melody, vocal, and bass line estimation. Historically these problems have been approached separately, and only recently, using learning-based…

Sound · Computer Science 2018-09-05 Rachel M. Bittner , Brian McFee , Juan P. Bello

We introduce a simple and linear SNR (strictly speaking, periodic to random power ratio) estimator (0dB to 80dB without additional calibration/linearization) for providing reliable descriptions of aperiodicity in speech corpus. The main…

Audio and Speech Processing · Electrical Eng. & Systems 2018-07-06 Hideki Kawahara , Ken-Ichi Sakakibara , Masanori Morise , Hideki Banno , Tomoki Toda

Speech enhancement (SE) performance is known to depend on noise characteristics and signal to noise ratio (SNR), yet intrinsic properties of the clean speech signal itself remain an underexplored factor. In this work, we systematically…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Mingchi Hou , Ina Kodrasi

In this work, we consider the modeling of signals that are almost, but not quite, harmonic, i.e., composed of sinusoids whose frequencies are close to being integer multiples of a common frequency. Typically, in applications, such signals…

Signal Processing · Electrical Eng. & Systems 2020-12-16 Filip Elvander , Andreas Jakobsson

This paper presents a simple Fourier-matching method to rigorously study resonance frequencies of a sound-hard slab with a finite number of arbitrarily shaped cylindrical holes of diameter ${\cal O}(h)$ for $h\ll1$. Outside the holes, a…

Analysis of PDEs · Mathematics 2021-04-07 Wangtao Lu , Wei Wang , Jiaxin Zhou

In this paper, a new statistic feature of the discrete short-time amplitude spectrum is discovered by experiments for the signals of unvoiced pronunciation. For the random-varying short-time spectrum, this feature reveals the relationship…

Sound · Computer Science 2016-12-22 Xiaodong Zhuang

In this work, we propose a novel method for modeling numerous speakers, which enables expressing the overall characteristics of speakers in detail like a trained multi-speaker model without additional training on the target speaker's…

Sound · Computer Science 2024-06-03 Jungil Kong , Junmo Lee , Jeongmin Kim , Beomjeong Kim , Jihoon Park , Dohee Kong , Changheon Lee , Sangjin Kim

Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on emotion similarity…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Yun-Shao Tsai , Yi-Cheng Lin , Huang-Cheng Chou , Tzu-Wen Hsu , Yun-Man Hsu , Chun Wei Chen , Shrikanth Narayanan , Hung-yi Lee

Standard sparse pseudo-input approximations to the Gaussian process (GP) cannot handle complex functions well. Sparse spectrum alternatives attempt to answer this but are known to over-fit. We suggest the use of variational inference for…

Machine Learning · Statistics 2015-03-23 Yarin Gal , Richard Turner

Although regression analysis has a great history, we consider that it has always continued being confused. For example, the fundamental terms in regression analysis (e.g., "regression", "least-squares method", "explanatory variable",…

Statistics Theory · Mathematics 2014-03-04 Shiro Ishikawa

Phonons, the quantum mechanical representation of lattice vibrations, and their coupling to the electronic degrees of freedom are important for understanding thermal and electric properties of materials. For the first time, phonons have…

Strongly Correlated Electrons · Physics 2011-01-25 H. Yavas , M. van Veenendaal , J. van den Brink , L. J. P. Ament , A. Alatas , B. M. Leu , M. -O. Apostu , N. Wizent , G. Behr , W. Sturhahn , H. Sinn , E. E. Alp

Spectral estimators are fundamental in lowrank matrix models and arise throughout machine learning and statistics, with applications including network analysis, matrix completion and PCA. These estimators aim to recover the leading…

Statistics Theory · Mathematics 2025-02-17 Hao Yan , Keith Levin

The spectrum and coherency are useful quantities for characterizing the temporal correlations and functional relations within and between point processes. This paper begins with a review of these quantities, their interpretation and how…

Biological Physics · Physics 2007-05-23 M. R. Jarvis , P. P. Mitra

We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music…

Machine Learning · Computer Science 2024-02-12 Xuanjie Liu , Daniel Chin , Yichen Huang , Gus Xia

This paper presents a geometric approach to pitch estimation (PE)-an important problem in Music Information Retrieval (MIR), and a precursor to a variety of other problems in the field. Though there exist a number of highly-accurate…

Sound · Computer Science 2020-12-09 Tom Goodman , Karoline van Gemst , Peter Tino

An inversion of the speech polarity may have a dramatic detrimental effect on the performance of various techniques of speech processing. An automatic method for determining the speech polarity (which is dependent upon the recording setup)…

Sound · Computer Science 2020-05-19 Thomas Drugman , Thierry Dutoit

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data between different…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Kun Zhou , Berrak Sisman , Haizhou Li

This paper introduces a novel technique for reconstructing the phase of modified spectrograms of audio signals. From the analysis of mixtures of sinusoids we obtain relationships between phases of successive time frames in the…

Sound · Computer Science 2016-05-25 Paul Magron , Roland Badeau , Bertrand David
‹ Prev 1 4 5 6 7 8 10 Next ›