English
Related papers

Related papers: An acoustic glottal source for vocal tract physica…

200 papers

Infant speech perception and learning is modeled using Echo State Network classification and Reinforcement Learning. Ambient speech for the modeled infant learner is created using the speech synthesizer Vocaltractlab. An auditory system is…

Sound · Computer Science 2016-10-21 Philip Zurbuchen

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-04 Yiwei Guo , Chenpeng Du , Ziyang Ma , Xie Chen , Kai Yu

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

Image and Video Processing · Electrical Eng. & Systems 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken discourse is also…

Computation and Language · Computer Science 2021-01-05 Rajaswa Patil , Yaman Kumar Singla , Rajiv Ratn Shah , Mika Hama , Roger Zimmermann

At present emotion extraction from speech is a very important issue due to its diverse applications. Hence, it becomes absolutely necessary to obtain models that take into consideration the speaking styles of a person, vocal tract…

Voiced segments of speech are assumed to be composed of non-stationary acoustic objects which can be described as stationary response of a non-stationary fundamental drive (FD) process and which are furthermore suited to reconstruct the…

Sound · Computer Science 2007-05-23 Friedhelm R. Drepper

In this paper, a novel time domain sampling method based on the initial arrival time of waves is proposed to reconstruct acoustic sources, including point sources, curve sources, surface sources and block sources. The uniqueness of…

Mathematical Physics · Physics 2025-09-30 Qiuyi Li , Bo Chen , Peng Gao , Yu Sun , Yao Sun

The solar acoustic oscillations are likely stochastically excited by convective dynamics in the solar photosphere, though few direct observations of individual source events have been made and their detailed characteristics are still…

Solar and Stellar Astrophysics · Physics 2021-07-07 Shah Mohammad Bahauddin , Mark Peter Rast

A method for analyzing sampling jitter in audio equipment is proposed. The method is based on the time-domain analysis where the time fluctuations of zero-crossing points in recorded sinusoidal waves are employed to characterize jitter.…

Sound · Computer Science 2023-05-09 Makoto Takeuchi , Haruo Saito

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

Sound · Computer Science 2022-06-22 Yuan Gong , Jin Yu , James Glass

Acoustic measurements on a vortex whistle corroborates earlier findings that the frequency increases linearly with the velocity of the air flowing into the whistle. Measurements in a reverberant chamber shows that the acoustic power…

Fluid Dynamics · Physics 2016-04-11 Ulf Kristiansen , Muriel Amielh

Understanding how sound propagates through different media is fundamental to both science and technology. While sound plays a critical role in natural navigation and underlies a wide range of applications - from medical ultrasound to sonar…

Physics Education · Physics 2025-07-30 Helio Takai , Tom Tomaszewski , Jeremy Tomaszewski , Joe Sundermier

Assessment of voice signals has long been performed with the assumption of periodicity as this facilitates analysis. Near periodicity of normal voice signals makes short-time harmonic modeling an appealing choice to extract vocal feature…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-10 Takeshi Ikuma , Andrew J. McWhorter , Lacey Adkins , Melda Kunduk

Computer voice is experiencing a renaissance through the growing popularity of voice-based interfaces, agents, and environments. Yet, how to measure the user experience (UX) of voice-based systems remains an open and urgent question,…

Human-Computer Interaction · Computer Science 2021-03-15 Katie Seaborn , Jacqueline Urakami

This paper investigates the differences occuring in the excitation for different voice qualities. Its goal is two-fold. First a large corpus containing three voice qualities (modal, soft and loud) uttered by the same speaker is analyzed and…

Sound · Computer Science 2020-01-06 Thomas Drugman , Thierry Dutoit , Baris Bozkurt

In this work we discuss the problem of identifying sound sources from pressure measurements with a Bayesian approach. The acoustics are modelled by the Helmholtz equation and the goal is to get information about the number, strength and…

Numerical Analysis · Mathematics 2019-01-16 Sebastian Engel , Dominik Hafemeyer , Christian Münch , Daniel Schaden

This paper describes an original experimental procedure to measure the mechanical interaction between the tongue and teeth and palate during speech production. It consists in using edentulous people as subjects and to insert pressure…

Medical Physics · Physics 2016-08-16 Christophe Jeannin , Pascal Perrier , Yohan Payan , André Dittmar , Brigitte Grosgogeat

In this thesis, we propose an artificial auditory system that gives a robot the ability to locate and track sounds, as well as to separate simultaneous sound sources and recognising simultaneous speech. We demonstrate that it is possible to…

Robotics · Computer Science 2016-02-23 Jean-Marc Valin

We address the problem of human-in-the-loop control for generating prosody in the context of text-to-speech synthesis. Controlling prosody is challenging because existing generative models lack an efficient interface through which users can…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-17 Dan Andrei Iliescu , Devang Savita Ram Mohan , Tian Huey Teh , Zack Hodari

As audio-visual systems increasingly bring immersive and interactive capabilities into our work and leisure activities, so the need for naturalistic test material grows. New volumetric datasets have captured high-quality 3D video, but…

Multimedia · Computer Science 2021-05-04 Hanne Stenzel , Davide Berghi , Marco Volino , Philip J. B. Jackson