中文
相关论文

相关论文: Phone Duration Modeling for Speaker Age Estimation…

200 篇论文

This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The…

音频与语音处理 · 电气工程与系统科学 2022-06-30 Si-Ioi Ng , Cymie Wing-Yee Ng , Jiarui Wang , Tan Lee

We present a syntactic dependency treebank for naturalistic child and child-directed speech in English (MacWhinney, 2000). Our annotations largely followed the guidelines of the Universal Dependencies project (UD (Zeman et al., 2022)), with…

计算与语言 · 计算机科学 2022-09-29 Zoey Liu , Emily Prud'hommeaux

Voice disguise, purposeful modification of one's speaker identity with the aim of avoiding being identified as oneself, is a low-effort way to fool speaker recognition, whether performed by a human or an automatic speaker verification (ASV)…

声音 · 计算机科学 2018-05-29 Rosa González Hautamäki , Anssi Kanervisto , Ville Hautamäki , Tomi Kinnunen

Speaker diarization, or the task of finding "who spoke and when", is now used in almost every speech processing application. Nevertheless, its fairness has not yet been evaluated because there was no protocol to study its biases one by one.…

声音 · 计算机科学 2023-02-21 Yannis Tevissen , Jérôme Boudy , Gérard Chollet , Frédéric Petitpont

The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each person…

VoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by…

机器学习 · 计算机科学 2021-12-21 Khaled Hechmi , Trung Ngo Trong , Ville Hautamaki , Tomi Kinnunen

This study investigates the explainability of embedding representations, specifically those used in modern audio spoofing detection systems based on deep neural networks, known as spoof embeddings. Building on established work in speaker…

声音 · 计算机科学 2024-12-25 Xuechen Liu , Junichi Yamagishi , Md Sahidullah , Tomi kinnunen

Modern language models (LMs) must be trained on many orders of magnitude more words of training data than human children receive before they begin to produce useful behavior. Assessing the nature and origins of this "data gap" requires…

计算与语言 · 计算机科学 2026-04-01 Steven Y. Feng , Alvin W. M. Tan , Michael C. Frank

When faced with self-regulation challenges, children have been known the use their language to inhibit their emotions and behaviors. Yet, to date, there has been a critical lack of evidence regarding what patterns in their speech children…

计算与语言 · 计算机科学 2021-11-01 Arnav Bhakta , Yeunjoo Kim , Pamela Cole

Interactions involving children span a wide range of important domains from learning to clinical diagnostic and therapeutic contexts. Automated analyses of such interactions are motivated by the need to seek accurate insights and offer…

音频与语音处理 · 电气工程与系统科学 2025-06-13 Anfeng Xu , Kevin Huang , Tiantian Feng , Helen Tager-Flusberg , Shrikanth Narayanan

Accurate control of the total duration of generated speech by adjusting the speech rate is crucial for various text-to-speech (TTS) applications. However, the impact of adjusting the speech rate on speech quality, such as intelligibility…

音频与语音处理 · 电气工程与系统科学 2024-06-07 Sefik Emre Eskimez , Xiaofei Wang , Manthan Thakker , Chung-Hsien Tsai , Canrun Li , Zhen Xiao , Hemin Yang , Zirun Zhu , Min Tang , Jinyu Li , Sheng Zhao , Naoyuki Kanda

We investigate the automatic processing of child speech therapy sessions using ultrasound visual biofeedback, with a specific focus on complementing acoustic features with ultrasound images of the tongue for the tasks of speaker diarization…

音频与语音处理 · 电气工程与系统科学 2019-08-16 Manuel Sam Ribeiro , Aciel Eshky , Korin Richmond , Steve Renals

Machine learning models for speech-based depression classification offer promise for health care applications. Despite growing work on depression classification, little is understood about how the length of speech-input impacts model…

计算与语言 · 计算机科学 2025-01-03 Tomasz Rutowski , Amir Harati , Yang Lu , Elizabeth Shriberg

Speech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named…

音频与语音处理 · 电气工程与系统科学 2020-06-26 Jing Shi , Jiaming Xu , Yusuke Fujita , Shinji Watanabe , Bo Xu

Ear recognition as a biometric modality is becoming increasingly popular, with promising broader application areas. While current applications involve adults, one of the challenges in ear recognition for children is the rapid structural…

计算机视觉与模式识别 · 计算机科学 2024-08-06 Afzal Hossain , Tipu Sultan , Stephanie Schuckers

Mobile phone metadata is increasingly used for humanitarian purposes in developing countries as traditional data is scarce. Basic demographic information is however often absent from mobile phone datasets, limiting the operational impact of…

机器学习 · 计算机科学 2017-11-16 Bjarke Felbo , Pål Sundsøy , Alex 'Sandy' Pentland , Sune Lehmann , Yves-Alexandre de Montjoye

The time domain waveform of a speech signal carries all of the auditory information. From the phonological point of view, it little can be said on the basis of the waveform itself. However, past research in mathematics, acoustics, and…

声音 · 计算机科学 2013-05-07 Urmila Shrawankar , V M Thakare

Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Gasser Elbanna

Ultrasound tongue imaging is widely used for speech production research, and it has attracted increasing attention as its potential applications seem to be evident in many different fields, such as the visual biofeedback tool for second…

计算机视觉与模式识别 · 计算机科学 2021-01-28 Kele Xu , Tamas Gábor Csapó , Ming Feng

We develop automatic speech recognition (ASR) systems for stories told by Afrikaans and isiXhosa preschool children. Oral narratives provide a way to assess children's language development before they learn to read. We consider a range of…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Christiaan Jacobs , Annelien Smith , Daleen Klop , Ondřej Klejch , Febe de Wet , Herman Kamper