Related papers: Speech Polarity Detection Using Hilbert Phase Info…
Imbalanced data commonly exists in real world, espacially in sentiment-related corpus, making it difficult to train a classifier to distinguish latent sentiment in text data. We observe that humans often express transitional emotion between…
Speech sound disorders are a common communication impairment in childhood. Because speech disorders can negatively affect the lives and the development of children, clinical intervention is often recommended. To help with diagnosis and…
The hyperanalytic signal is the straight forward generalization of the classical analytic signal. It is defined by a complexification of two canonical complex signals, which can be considered as an inverse operation of the Cayley-Dickson…
With the forthcoming release of high precision polarization measurements, such as from the Planck satellite, the metrology of polarization needs to improve. In particular, it is crucial to take into account full knowledge of the noise…
We describe a novel language-independent approach to the task of determining the polarity, positive or negative, of the author's opinion on a specific topic in natural language text. In particular, weights are assigned to attributes,…
State-of-the-art Deep Learning systems for speaker verification are commonly based on speaker embedding extractors. These architectures are usually composed of a feature extractor front-end together with a pooling layer to encode…
A non-invasive method for the monitoring of heart activity can help to reduce the deaths caused by heart disorders such as stroke, arrhythmia and heart attack. The human voice can be considered as a biometric data that can be used for…
We develop two algorithms, based on maximum likelihood (ML) inference, for estimating the parameters of polarized radio sources which emit at a single rotation measure (RM), e.g., pulsars. These algorithms incorporate the flux density…
Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement.…
The study and application of signal detection techniques based on cross-correlation method for acoustic transient signals in noisy and reverberant environments are presented. These techniques are shown to provide high signal to noise ratio,…
Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community. In particular, speech representations learned by SSL models have been shown to…
Conventional optical coherent receivers capture the full electrical field, including amplitude and phase, of a signal waveform by measuring its interference against a stable continuous-wave local oscillator (LO). In optical coherent…
We propose a variation to the commonly used Word Error Rate (WER) metric for speech recognition evaluation which incorporates the alignment of phonemes, in the absence of time boundary information. After computing the Levenshtein alignment…
Online polarization poses a growing challenge for democratic discourse, yet most computational social science research remains monolingual, culturally narrow, or event-specific. We introduce POLAR, a multilingual, multicultural, and…
The synchrotron radiation is commonly known to be completely linearly polarized when observed in the orbital plane of the synchrotron motion. Under actual experimental conditions, however, the degree of polarization of the synchrotron…
This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The…
Detection of transitions between broad phonetic classes in a speech signal is an important problem which has applications such as landmark detection and segmentation. The proposed hierarchical method detects silence to non-silence…
Given a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain,…
This paper explores the quantum detection of Phase-Shift Keying (PSK)-coded coherent states through the lens of active hypothesis testing, focusing on a Dolinar-like receiver with constraints on displacement amplitude and energy. With…
Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with…