English
Related papers

Related papers: Vocal wow in an adapted reflex resonance model

200 papers

Acoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio…

Sound · Computer Science 2022-04-05 Jianrong Wang , Jinyu Liu , Longxuan Zhao , Shanyu Wang , Ruiguo Yu , Li Liu

Using the model of a generalized Van der Pol oscillator in the regime of subcritical Hopf bifurcation we investigate the influence of time delay on noise-induced oscillations. It is shown that for appropriate choices of time delay either…

Chaotic Dynamics · Physics 2015-06-23 V. Semenov , A. Feoktistov , T. Vadivasova , E. Schöll , A. Zakharova

Poor laryngeal muscle coordination that results in abnormal glottal posturing is believed to be a primary etiologic factor in common voice disorders such as non-phonotraumatic vocal hyperfunction. Abnormal activity of antagonistic laryngeal…

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

A mathematical model describing the coupling between two independent amplification mechanisms in auditory hair cells is proposed and analyzed. Hair cells are cells in the inner ear responsible for translating sound-induced mechanical…

Pattern Formation and Solitons · Physics 2009-11-13 K. A. Montgomery , M. Silber , S. A. Solla

Affective responses to music are highly personal. Despite consensus that idiosyncratic factors play a key role in regulating how listeners emotionally respond to music, precisely measuring the marginal effects of these variables has proved…

Computation and Language · Computer Science 2022-10-19 Sky CH-Wang , Evan Li , Oliver Li , Smaranda Muresan , Zhou Yu

Expressive singing voice correction is an appealing but challenging problem. A robust time-warping algorithm which synchronizes two singing recordings can provide a promising solution. We thereby propose to address the problem by canonical…

Audio and Speech Processing · Electrical Eng. & Systems 2017-11-27 Yin-Jyun Luo , Ming-Tso Chen , Tai-Shih Chi , Li Su

Models of physical systems are used to explain and predict experimental results and observations. When students encounter discrepancies between the actual and expected behavior of a system, they revise their models to include the newly…

Physics Education · Physics 2022-07-06 Laura Ríos , Benjamin Pollard , Dimitri R. Dounas-Frazer , H. J. Lewandowski

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

Several audio-visual speech recognition models have been recently proposed which aim to improve the robustness over audio-only models in the presence of noise. However, almost all of them ignore the impact of the Lombard effect, i.e., the…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-10 Pingchuan Ma , Stavros Petridis , Maja Pantic

We analyze a stimulated revival (echo) effect for the breathing modes of the atomic oscillations in optical lattices. The effect arises from the dephasing due to the weak anharmonicity being partly reversed in time by means of additional…

Condensed Matter · Physics 2009-10-30 A. Bulatov , A. Kuklov , B. E. Vugmeister , H. Rabitz

Representations in the auditory cortex might be based on mechanisms similar to the visual ventral stream; modules for building invariance to transformations and multiple layers for compositionality and selectivity. In this paper we propose…

In the field of affective computing, traditional methods for generating emotions predominantly rely on deep learning techniques and large-scale emotion datasets. However, deep learning techniques are often complex and difficult to…

Human-Computer Interaction · Computer Science 2025-03-24 Haidong Wang , Qia Shan , JianHua Zhang , PengFei Xiao , Ao Liu

In recent years, a new method for experimental nonlinear modal analysis has been developed, which is based on the extended periodic motion concept. The method is well suited to experimentally obtain amplitude-dependent modal properties…

Systems and Control · Electrical Eng. & Systems 2021-08-16 Maren Scheel

Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-25 Alkis Koudounas , Moreno La Quatra , Marco Sabato Siniscalchi , Elena Baralis

How important are different temporal speech modulations for speech recognition? We answer this question from two complementary perspectives. Firstly, we quantify the amount of phonetic \textit{information} in the modulation spectrum of…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Samik Sadhu , Hynek Hermansky

Speech requires programming the sequence of vocal gestures that produce the sounds of words. Here we explored the timing of this program by asking our participants to pronounce, as quickly as possible, a sequence of…

Neurons and Cognition · Quantitative Biology 2018-05-23 Alan Taitz , Diego E Shalom , Marcos A Trevisan

Dynamic Range Compression (DRC) is a widely used audio effect that adjusts signal dynamics for applications in music production, broadcasting, and speech processing. Inverting DRC is of broad importance for restoring the original dynamics,…

Sound · Computer Science 2025-09-11 Haoran Sun , Dominique Fourer , Hichem Maaref

Acoustic velocity vectors (AVVs) are related to the human's perception of sound at low frequencies and are widely used in Ambisonics. This paper proposes a spatial sound field reproduction algorithm called velocity matching, which…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-06 Jiarui Wang , Thushara Abhayapala , Jihui Aimee Zhang , Prasanga Samarasinghe