Related papers: Measuring pitch extractors' response to frequency-…
In this paper, robust detection, tracking and geometry estimation methods are developed and combined into a system for estimating time-difference estimates, microphone localization and sound source movement. No assumptions on the 3D…
Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based…
This paper presents a polyphonic pitch tracking system able to extract both framewise and note-based estimates from audio. The system uses several artificial neural networks in a deep layered learning setup. First, cascading networks are…
The objective of the present study is exploratory: to introduce and apply a new theory of speech rhythm zones or rhythm formants (R-formants). R-formants are zones of high magnitude frequencies in the low frequency (LF) long-term spectrum…
Perceptually-inspired objective functions such as the perceptual evaluation of speech quality (PESQ), signal-to-distortion ratio (SDR), and short-time objective intelligibility (STOI), have recently been used to optimize performance of…
In audio processing applications, the generation of expressive sounds based on high-level representations demonstrates a high demand. These representations can be used to manipulate the timbre and influence the synthesis of creative…
The traditional approach to morphological inflection (the task of modifying a base word (lemma) to express grammatical categories) has been, for decades, to consider lexical entries of lemma-tag-form triples uniformly, lacking any…
We present an analytical calculation of the response of a driven Duffing oscillator to low-frequency fluctuations in the resonance frequency and damping. We find that fluctuations in these parameters manifest themselves distinctively,…
Emotions lie on a continuum, but current models treat emotions as a finite valued discrete variable. This representation does not capture the diversity in the expression of emotion. To better represent emotions we propose the use of natural…
This paper discusses a fast algorithm for analyzing the phases of a multi-tone phase calibration signal. The multi-tone method is superior to its single-tone counterpart. The PCal signal is a wide-band frequency comb derived from an…
Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific…
In mobile speech communication applications, wind noise can lead to a severe reduction of speech quality and intelligibility. Since the performance of speech enhancement algorithms using acoustic microphones tends to substantially degrade…
The Frequency Following Response (FFR) reflects the brain's neural encoding of auditory stimuli including speech. Because the fundamental frequency (F0), a physical correlate of pitch, is one of the essential features of speech, there has…
We present a pressure sensor based on a Michelson interferometer, for use in photoacoustic tomography. Quadrature phase detection is employed allowing measurement at any point on the mirror surface without having to retune the…
We consider the problem of obtaining information about an inaccessible half-space from acoustic measurements made in the accessible half-space. If the measurements are of limited precision, some scatterers will be undetectable because their…
The size distribution of aerosol droplets is a key parameter in a myriad of processes, and it is typically measured with optical aids (e.g., lasers or cameras) that require sophisticated calibration, thus making the measurement cost…
We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the generated speech can be…
Many important phenomena in quantum devices are dynamic, meaning that they cannot be studied using time-averaged measurements alone. Experiments that measure such transient effects are collectively known as fast readout. One of the most…
Under the excitation of strings, the wooden structure of string instruments is generally assumed to undergo linear vibrations. As an alternative to the direct measurement of the distortion rate at several vibration levels and frequencies,…
Speaker verification is to judge the similarity between two unknown voices in an open set, where the ideal speaker embedding should be able to condense discriminant information into a compact utterance-level representation that has small…