中文
相关论文

相关论文: Phone Duration Modeling for Speaker Age Estimation…

200 篇论文

Vocal aging, a universal process of human aging, can largely affect one's language use, possibly including some subtle acoustic features of one's utterances like Voice Onset Time. To figure out the time effects, Queen Elizabeth's Christmas…

声音 · 计算机科学 2018-10-17 Xuanda Chen , Ziyu Xiong , Jian Hu

With the advent of high-quality speech synthesis, there is a lot of interest in controlling various prosodic attributes of speech. Speaking rate is an essential attribute towards modelling the expressivity of speech. In this work, we…

音频与语音处理 · 电气工程与系统科学 2023-10-16 Jesuraj Bandekar , Sathvik Udupa , Abhayjeet Singh , Anjali Jayakumar , Deekshitha G , Sandhya Badiger , Saurabh Kumar , Pooja VH , Prasanta Kumar Ghosh

It takes several years for the developing brain of a baby to fully master word repetition-the task of hearing a word and repeating it aloud. Repeating a new word, such as from a new language, can be a challenging task also for adults.…

计算与语言 · 计算机科学 2025-06-17 Daniel Dager , Robin Sobczyk , Emmanuel Chemla , Yair Lakretz

We address the problem of inferring a speaker's level of certainty based on prosodic information in the speech signal, which has application in speech-based dialogue systems. We show that using phrase-level prosodic features centered around…

计算与语言 · 计算机科学 2011-03-11 Heather Pon-Barry , Stuart M. Shieber

Speech is a rich biomarker that encodes substantial information about the health of a speaker, and thus it has been proposed for the detection of numerous diseases, achieving promising results. However, questions remain about what the…

音频与语音处理 · 电气工程与系统科学 2024-09-17 Catarina Botelho , Alberto Abad , Tanja Schultz , Isabel Trancoso

We address the problem of detecting who spoke when in child-inclusive spoken interactions i.e., automatic child-adult speaker classification. Interactions involving children are richly heterogeneous due to developmental differences. The…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Rimita Lahiri , Tiantian Feng , Rajat Hebbar , Catherine Lord , So Hyun Kim , Shrikanth Narayanan

Self-supervised techniques for learning speech representations have been shown to develop linguistic competence from exposure to speech without the need for human labels. In order to fully realize the potential of these approaches and…

Speech and language biomarkers have the potential to be regular, objective assessments of symptom severity in several health conditions, both in-clinic and remotely using mobile devices. However, the complex nature of speech and often…

We present a speaker-aware approach for simulating multi-speaker conversations that captures temporal consistency and realistic turn-taking dynamics. Prior work typically models aggregate conversational statistics under an independence…

声音 · 计算机科学 2026-05-25 Máté Gedeon , Péter Mihajlik

The speech signal is a consummate example of time-series data. The acoustics of the signal change over time, sometimes dramatically. Yet, the most common type of comparison we perform in phonetics is between instantaneous acoustic…

音频与语音处理 · 电气工程与系统科学 2023-04-18 Matthew C. Kelley

A transversal study of the pitch variability of parkinsonian voices in read speech is presented. 30 patients suffering from Parkinson's disease (PD) and 32 healthy speakers were recorded while reading a text without voiceless phonemes. The…

音频与语音处理 · 电气工程与系统科学 2024-02-12 Pablo Rodriguez-Perez , Ruben Fraile , Miguel Garcia-Escrig , Nicolas Saenz-Lechon , Juana M. Gutierrez-Arriola , Victor Osma-Ruiz

Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Munender Varshney , Ravindra Yadav , Vinay P. Namboodiri , Rajesh M Hegde

Children's automatic speech recognition (ASR) often underperforms compared to that of adults due to a confluence of interdependent factors: physiological (e.g., smaller vocal tracts), cognitive (e.g., underdeveloped pronunciation), and…

计算与语言 · 计算机科学 2025-06-03 Vishwanath Pratap Singh , Md. Sahidullah , Tomi Kinnunen

The ability to automatically determine the age audience of a novel provides many opportunities for the development of information retrieval tools. Firstly, developers of book recommendation systems and electronic libraries may be interested…

计算与语言 · 计算机科学 2021-08-30 Anna Glazkova , Yury Egorov , Maksim Glazkov

Disentangling speaker and content attributes of a speech signal into separate latent representations followed by decoding the content with an exchanged speaker representation is a popular approach for voice conversion, which can be trained…

音频与语音处理 · 电气工程与系统科学 2022-09-07 Michael Kuhlmann , Fritz Seebauer , Janek Ebbers , Petra Wagner , Reinhold Haeb-Umbach

An attempt is made to develop a smart toy to help the children suffering with communication disorders. The children suffering with such disorders need additional attention and guidance to understand different types of social events and life…

人机交互 · 计算机科学 2019-06-12 Amr Jadi

The diagnosis of autism spectrum disorder (ASD) is a complex, challenging task as it depends on the analysis of interactional behaviors by psychologists rather than the use of biochemical diagnostics. In this paper, we present a modeling…

音频与语音处理 · 电气工程与系统科学 2024-01-19 Tahiya Chowdhury , Veronica Romero , Amanda Stent

Eavesdropping from the user's smartphone is a well-known threat to the user's safety and privacy. Existing studies show that loudspeaker reverberation can inject speech into motion sensor readings, leading to speech eavesdropping. While…

声音 · 计算机科学 2022-12-26 Ahmed Tanvir Mahdad , Cong Shi , Zhengkun Ye , Tianming Zhao , Yan Wang , Yingying Chen , Nitesh Saxena

While high-performing language models are typically trained on hundreds of billions of words, human children become fluent language users with a much smaller amount of data. What are the features of the data they receive, and how do these…

计算与语言 · 计算机科学 2024-10-10 Steven Y. Feng , Noah D. Goodman , Michael C. Frank

Capturing high-level structure in audio waveforms is challenging because a single second of audio spans tens of thousands of timesteps. While long-range dependencies are difficult to model directly in the time domain, we show that they can…

音频与语音处理 · 电气工程与系统科学 2019-06-05 Sean Vasquez , Mike Lewis