Related papers: Using phonetic constraints in acoustic-to-articula…
This paper analyzes the reconstruction of diffusion and absorption parameters in an elliptic equation from knowledge of internal data. In the application of photo-acoustics, the internal data are the amount of thermal energy deposited by…
The unique conduction properties of condensed matter systems with topological order have recently inspired a quest for similar effects in classical wave phenomena. Acoustic topological insulators, in particular, hold the promise to…
We aim to explain whether a stress memory task has a significant impact on tonal coarticulation. We contribute a novel approach to analyse tonal coarticulation in phonetics, where several f0 contours are compared with respect to their…
Articulatory information has been shown to be effective in improving the performance of HMM-based and DNN-based text-to-speech synthesis. Speech synthesis research focuses traditionally on text-to-speech conversion, when the input is text…
In this work, we propose a new mathematical vocoder algorithm(modified spectral inversion) that generates a waveform from acoustic features without phase estimation. The main benefit of using our proposed method is that it excludes the…
The acoustic inverse obstacle scattering problem consists of determining the shape of a domain from measurements of the scattered far field due to some set of incident fields (probes). For a penetrable object with known sound speed, this…
Speech is produced through the coordination of vocal tract constricting organs: lips, tongue, velum, and glottis. Previous works developed Speech Inversion (SI) systems to recover acoustic-to-articulatory mappings for lip and tongue…
Airborne electromagnetic surveys may consist of hundreds of thousands of soundings. In most cases, this makes 3D inversions unfeasible even when the subsurface is characterized by a high level of heterogeneity. Instead, approaches based on…
This study experimentally validates a numerical model of electromagnetic propagation through the human head during the pronunciation of different vowels, with the goal of improving our understanding of the underlying physical phenomena. A…
We investigate the potential of stochastic neural networks for learning effective waveform-based acoustic models. The waveform-based setting, inherent to fully end-to-end speech recognition systems, is motivated by several comparative…
This paper concerns the inverse source scattering problems of recovering random sources for acoustic and elastic waves. The underlying sources are assumed to be random functions driven by an additive white noise. The inversion process aims…
Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…
Consider the scattering of a time-harmonic acoustic plane wave by a bounded elastic obstacle which is immersed in a homogeneous acoustic medium. This paper concerns an inverse acoustic-elastic interaction problem, which is to determine the…
Phonetic ambiguity and confusibility are bugbears for any form of bottom-up or data-driven approach to language processing. The question of when an input is ``close enough'' to a target word pervades the entire problem spaces of speech…
The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneous acoustic…
Recent studies have shown that frame-level deep speaker features can be derived from a deep neural network with the training target set to discriminate speakers by a short speech segment. By pooling the frame-level features, utterance-level…
Automatic speech recognition (ASR) is a relevant area in multiple settings because it provides a natural communication mechanism between applications and users. ASRs often fail in environments that use language specific to particular…
While Word2Vec represents words (in text) as vectors carrying semantic information, audio Word2Vec was shown to be able to represent signal segments of spoken words as vectors carrying phonetic structure information. Audio Word2Vec can be…
This paper addresses the problem of speech separation and enhancement from multichannel convolutive and noisy mixtures, \emph{assuming known mixing filters}. We propose to perform the speech separation and enhancement task in the short-time…
This work describes and analyzes the domain derivative for a time-dependent acoustic scattering problem. We study the nonlinear operator that maps a sound-soft scattering object to the solution of the time-dependent wave equation evaluated…