中文
相关论文

相关论文: Automatic measurement of vowel duration via struct…

200 篇论文

We describe an arrangement for simultaneous recording of speech and geometry of vocal tract in patients undergoing surgery involving this area. Experimental design is considered from an articulatory phonetic point of view. The speech and…

Duration modelling has become an important research problem once more with the rise of non-attention neural text-to-speech systems. The current approaches largely fall back to relying on previous statistical parametric speech synthesis…

音频与语音处理 · 电气工程与系统科学 2022-06-29 Ammar Abbas , Thomas Merritt , Alexis Moinet , Sri Karlapati , Ewa Muszynska , Simon Slangen , Elia Gatti , Thomas Drugman

Self-talk-an internal dialogue that can occur silently or be spoken aloud-plays a crucial role in emotional regulation, cognitive processing, and motivation, yet has remained largely invisible and unmeasurable in everyday life. In this…

声音 · 计算机科学 2025-11-12 Euihyeok Lee , Seonghyeon Kim , SangHun Im , Heung-Seon Oh , Seungwoo Kang

We propose the first method to adaptively modify the duration of a given speech signal. Our approach uses a Bayesian framework to define a latent attention map that links frames of the input and target utterances. We train a masked…

音频与语音处理 · 电气工程与系统科学 2021-07-13 Ravi Shankar , Archana Venkataraman

The task of quantifying human behavior by observing interaction cues is an important and useful one across a range of domains in psychological research and practice. Machine learning-based approaches typically perform this task by first…

计算与语言 · 计算机科学 2020-08-28 Sandeep Nallan Chakravarthula , Brian Baucom , Shrikanth Narayanan , Panayiotis Georgiou

Metric functions for phoneme perception capture the similarity structure among phonemes in a given language and therefore play a central role in phonology and psycho-linguistics. Various phenomena depend on phoneme similarity, such as…

机器学习 · 计算机科学 2018-09-24 Yair Lakretz , Gal Chechik , Evan-Gary Cohen , Alessandro Treves , Naama Friedmann

Automatic inference of important paralinguistic information such as age from speech is an important area of research with numerous spoken language technology based applications. Speaker age estimation has applications in enabling…

音频与语音处理 · 电气工程与系统科学 2024-06-19 Prashanth Gurunath Shivakumar , Somer Bishop , Catherine Lord , Shrikanth Narayanan

Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization criterion. On the other hand, automated metrics are efficient…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Pranay Manocha , Adam Finkelstein , Richard Zhang , Nicholas J. Bryan , Gautham J. Mysore , Zeyu Jin

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Pablo M. Delgado , Jürgen Herre

Vowels are primarily characterized by tongue position. Humans have discovered these features of vowel articulation through their own experience and explicit objective observation such as using MRI. With this knowledge and our experience, we…

计算与语言 · 计算机科学 2025-01-30 Haruki Sakajo , Yusuke Sakai , Hidetaka Kamigaito , Taro Watanabe

Spectro-temporal dynamics of consonant-vowel (CV) transition regions are considered to provide robust cues related to articulation. In this work, we propose an objective measure of precise articulation, dubbed the objective articulation…

音频与语音处理 · 电气工程与系统科学 2022-03-21 Vikram C. Mathad , Julie M. Liss , Kathy Chapman , Nancy Scherer , Visar Berisha

Text does not fully specify the spoken form, so text-to-speech models must be able to learn from speech data that vary in ways not explained by the corresponding text. One way to reduce the amount of unexplained variation in training data…

This paper is concerned with automatic continuous speech recognition using trainable systems. The aim of this work is to build acoustic models for spoken Swedish. This is done employing hidden Markov models and using the SpeechDat database…

音频与语音处理 · 电气工程与系统科学 2024-04-26 Giampiero Salvi

Traditional approaches for understanding phonological learning have predominantly relied on curated text data. Although insightful, such approaches limit the knowledge captured in textual representations of the spoken language. To overcome…

计算与语言 · 计算机科学 2024-07-10 Sneha Ray Barman , Shakuntala Mahanta , Neeraj Kumar Sharma

This paper proposes a model for automatic prosodic label annotation, where the predicted labels can be used for training a prosody-controllable text-to-speech model. The proposed model utilizes not only rich acoustic features extracted by a…

音频与语音处理 · 电气工程与系统科学 2025-07-08 Tomoki Koriyama

This work presents a novel methodology for calculating the phonetic similarity between words taking motivation from the human perception of sounds. This metric is employed to learn a continuous vector embedding space that groups similar…

计算与语言 · 计算机科学 2021-10-01 Rahul Sharma , Kunal Dhawan , Balakrishna Pailla

A sound source was proposed for acoustic measurements of physical models of the human vocal tract. The physical models are produced by Fast Prototyping, based on Magnetic Resonance Imaging during prolonged vowel production. The sound…

仪器与探测器 · 物理学 2017-11-22 Antti Hannukainen , Juha Kuortti , Jarmo Malinen , Antti Ojalammi

Growing digital archives and improving algorithms for automatic analysis of text and speech create new research opportunities for fundamental research in phonetics. Such empirical approaches allow statistical evaluation of a much larger set…

计算与语言 · 计算机科学 2017-06-05 Elodie Gauthier , Laurent Besacier , Sylvie Voisin

Textless spoken language models (SLMs) are generative models of speech that do not rely on text supervision. Most textless SLMs learn to predict the next semantic token, a discrete representation of linguistic content, and rely on a…

计算与语言 · 计算机科学 2025-10-23 Ju-Chieh Chou , Jiawei Zhou , Karen Livescu

Audio-recordings collected with a child-worn device are a fundamental tool in child language research. Long-form recordings collected over whole days promise to capture children's input and production with minimal observer bias, and…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Loann Peurey , Marvin Lavechin , Tarek Kunze , Manel Khentout , Lucas Gautheron , Emmanuel Dupoux , Alejandrina Cristia