English
Related papers

Related papers: Vocal wow in an adapted reflex resonance model

200 papers

This paper investigates the temporal excitation patterns of creaky voice. Creaky voice is a voice quality frequently used as a phrase-boundary marker, but also as a means of portraying attitude, affective states and even social status.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-02 Thomas Drugman , John Kane , Christer Gobl

Linear systems such as room acoustics and string oscillations may be modeled as the sum of mode responses, each characterized by a frequency, damping and amplitude. Here, we consider finding the mode parameters from impulse response…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-24 Orchisama Das , Jonathan S. Abel

Modeling and estimation of the vocal tract and glottal source parameters of vowels from raw speech can be typically done by using the Auto-Regressive with eXogenous input (ARX) model and Liljencrants-Fant (LF) model with an iteration-based…

Sound · Computer Science 2024-10-08 Kai Lia , Masato Akagia , Yongwei Lib , Masashi Unokia

Time-varying covariates in longitudinal studies frequently evolve through reciprocal feedback, undergo role reversal, and reflect unobserved individual heterogeneity. Standard statistical frameworks often assume fixed covariate roles and…

Methodology · Statistics 2026-02-27 Niloofar Ramezani , Pascal Nitiema , Jeffrey R. Wilson

Voice conversion is a task to convert a non-linguistic feature of a given utterance. Since naturalness of speech strongly depends on its pitch pattern, in some applications, it would be desirable to keep the original rise/fall pitch pattern…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-21 Chihiro Watanabe , Hirokazu Kameoka

During slow-wave sleep, the brain produces traveling waves of slow oscillations (SOs; $\leq 2$ Hz), characterized by the propagation of alternating high- and low-activity states. The question of internal mechanisms that modulate traveling…

Adaptation and Self-Organizing Systems · Physics 2025-10-16 Ronja Strömsdörfer , Klaus Obermayer

Audio-Visual Question Answering (AVQA) is a complex multi-modal reasoning task, demanding intelligent systems to accurately respond to natural language queries based on audio-video input pairs. Nevertheless, prevalent AVQA approaches are…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Jie Ma , Min Hu , Pinghui Wang , Wangchun Sun , Lingyun Song , Hongbin Pei , Jun Liu , Youtian Du

We introduce a novel, perceptually derived metric (P-Reverb) that relates the just-noticeable difference (JND) of the early sound field(also called early reflections) to the late sound field (known as late reflections or reverberation).…

Sound · Computer Science 2019-02-20 Atul Rungta , Nicholas Rewkowski , Roberta Klatzky , Dinesh Manocha

In this paper, a multi-channel time-varying covariance matrix model for late reverberation reduction is proposed. Reflecting that variance of the late reverberation is time-varying and it depends on past speech source variance, the proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-22 Masahito Togami

Current ML models for music emotion recognition, while generally working quite well, do not give meaningful or intuitive explanations for their predictions. In this work, we propose a 2-step procedure to arrive at spectrogram-level…

Sound · Computer Science 2019-05-29 Verena Haunschmid , Shreyan Chowdhury , Gerhard Widmer

The concept of a common modulated oscillation spanning multiple time series is formalized, a method for the recovery of such a signal from potentially noisy observations is proposed, and the time-varying bias properties of the recovery…

Methodology · Statistics 2015-05-27 Jonathan M. Lilly , Sofia C. Olhede

Referring expressions are natural language constructions used to identify particular objects within a scene. In this paper, we propose a unified framework for the tasks of referring expression comprehension and generation. Our model is…

Computer Vision and Pattern Recognition · Computer Science 2017-04-19 Licheng Yu , Hao Tan , Mohit Bansal , Tamara L. Berg

Voice Onset Time (VOT), a key measurement of speech for basic research and applied medical studies, is the time between the onset of a stop burst and the onset of voicing. When the voicing onset precedes burst onset the VOT is negative; if…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-30 Yosi Shrem , Matthew Goldrick , Joseph Keshet

A recurrent neural network model of phonological pattern learning is proposed. The model is a relatively simple neural network with one recurrent layer, and displays biases in learning that mimic observed biases in human learning.…

Computation and Language · Computer Science 2024-05-31 Amanda Doucette

Pre-trained model representations have demonstrated state-of-the-art performance in speech recognition, natural language processing, and other applications. Speech models, such as Bidirectional Encoder Representations from Transformers…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-07 Vikramjit Mitra , Vasudha Kowtha , Hsiang-Yun Sherry Chien , Erdrin Azemi , Carlos Avendano

Nasalization of vowels is a phenomenon where oral and nasal tracts participate simultaneously for the production of speech. Acoustic coupling of oral and nasal tracts results in a complex production system, which is subjected to a…

Sound · Computer Science 2020-09-15 RaviShankar Prasad , B. Yegnanarayana

Singing voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated into the deep learning optimization. Consequently, most of…

This work explores the effect of gender and linguistic-based vocal variations on the accuracy of emotive expression classification. Emotive expressions are considered from the perspective of spectral features in speech (Mel-frequency…

Sound · Computer Science 2022-10-28 Zachary Dair , Ryan Donovan , Ruairi O'Reilly

We present NeuroVoc, a flexible model-agnostic vocoder framework that reconstructs acoustic waveforms from simulated neural activity patterns using an inverse Fourier transform. The system applies straightforward signal processing to…

The proposed model is only for the audio module. All videos in the OMG Emotion Dataset are converted to WAV files. The proposed model makes use of semi-supervised learning for the emotion recognition. A GAN is trained with unsupervised…

Sound · Computer Science 2018-05-07 Ingryd Pereira , Diego Santos
‹ Prev 1 4 5 6 7 8 10 Next ›