中文
相关论文

相关论文: Silent Speech and Emotion Recognition from Vocal T…

200 篇论文

Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven…

图像与视频处理 · 电气工程与系统科学 2024-09-25 Hong Nguyen , Sean Foley , Kevin Huang , Xuan Shi , Tiantian Feng , Shrikanth Narayanan

Vocal tract configurations play a vital role in generating distinguishable speech sounds, by modulating the airflow and creating different resonant cavities in speech production. They contain abundant information that can be utilized to…

声音 · 计算机科学 2018-07-31 Pramit Saha , Praneeth Srungarapu , Sidney Fels

Real-time Magnetic Resonance Imaging (rtMRI) visualizes vocal tract action, offering a comprehensive window into speech articulation. However, its signals are high dimensional and noisy, hindering interpretation. We investigate compact…

图像与视频处理 · 电气工程与系统科学 2026-01-30 Jay Park , Hong Nguyen , Sean Foley , Jihwan Lee , Yoonjeong Lee , Dani Byrd , Shrikanth Narayanan

Accurate modeling of the vocal tract is necessary to construct articulatory representations for interpretable speech processing and linguistics. However, vocal tract modeling is challenging because many internal articulators are occluded…

计算机视觉与模式识别 · 计算机科学 2024-06-25 Rishi Jain , Bohan Yu , Peter Wu , Tejas Prabhune , Gopala Anumanchipalli

Articulatory-to-acoustic (forward) mapping is a technique to predict speech using various articulatory acquisition techniques (e.g. ultrasound tongue imaging, lip video). Real-time MRI (rtMRI) of the vocal tract has not been used before for…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Tamás Gábor Csapó

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisitions may provide…

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

Previous real-time MRI (rtMRI)-based speech synthesis models depend heavily on noisy ground-truth speech. Applying loss directly over ground truth mel-spectrograms entangles speech content with MRI noise, resulting in poor intelligibility.…

声音 · 计算机科学 2025-01-20 Neil Shah , Ayan Kashyap , Shirish Karande , Vineet Gandhi

Articulatory acoustic inversion aims to reconstruct the complete geometry of the vocal tract from the speech signal. In this paper, we present a comparative study of several levels of phonetic segmentation accuracy, together with a…

音频与语音处理 · 电气工程与系统科学 2026-03-13 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

The various speech sounds of a language are obtained by varying the shape and position of the articulators surrounding the vocal tract. Analyzing their variations is crucial for understanding speech production, diagnosing speech disorders…

图像与视频处理 · 电气工程与系统科学 2020-02-04 Mohammad Eslami , Christiane Neuschaefer-Rube , Antoine Serrurier

Real-time magnetic resonance imaging (rtMRI) of speech production enables non-invasive visualization of dynamic vocal-tract motion and is valuable for speech science and clinical assessment. However, rtMRI is fundamentally constrained by…

Encouraged by the success of deep neural networks on a variety of visual tasks, much theoretical and experimental work has been aimed at understanding and interpreting how vision networks operate. Meanwhile, deep neural networks have also…

Real-time Magnetic Resonance Imaging (rtMRI) is frequently used in speech production studies as it provides a complete view of the vocal tract during articulation. This study investigates the effectiveness of rtMRI in analyzing vocal tract…

音频与语音处理 · 电气工程与系统科学 2025-03-27 Masoud Thajudeen Tholan , Vinayaka Hegde , Chetan Sharma , Prasanta Kumar Ghosh

We present a multilinear statistical model of the human tongue that captures anatomical and tongue pose related shape variations separately. The model is derived from 3D magnetic resonance imaging data of 11 speakers sustaining speech…

计算机视觉与模式识别 · 计算机科学 2018-04-18 Alexander Hewer , Stefanie Wuhrer , Ingmar Steiner , Korin Richmond

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

音频与语音处理 · 电气工程与系统科学 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components:…

音频与语音处理 · 电气工程与系统科学 2026-01-27 Jihwan Lee , Sean Foley , Thanathai Lertpetchpun , Kevin Huang , Yoonjeong Lee , Tiantian Feng , Louis Goldstein , Dani Byrd , Shrikanth Narayanan

In this paper the task of emotion recognition from speech is considered. Proposed approach uses deep recurrent neural network trained on a sequence of acoustic features calculated over small speech intervals. At the same time special…

计算与语言 · 计算机科学 2018-07-06 Vladimir Chernykh , Pavel Prikhodko

Infants, adults, non-human primates and non-primates all learn patterns implicitly, and they do so across modalities. The biological evidence supports the hypothesis that the mechanism for this learning is general but computationally local.…

神经元与认知 · 定量生物学 2021-08-16 John Rohrlich , Randall C. O'Reilly

Articulatory-to-acoustic inversion strongly depends on the type of data used. While most previous studies rely on EMA, which is limited by the number of sensors and restricted to accessible articulators, we propose an approach aiming at a…

音频与语音处理 · 电气工程与系统科学 2026-03-31 Sofiane Azzouz , Pierre-André Vuissoz , Yves Laprie

We release the USC Long Single-Speaker (LSS) dataset containing real-time MRI video of the vocal tract dynamics and simultaneous audio obtained during speech production. This unique dataset contains roughly one hour of video and audio data…

‹ 上一页 1 2 3 10 下一页 ›