中文
相关论文

相关论文: Children's Speech Recognition through Discrete Tok…

200 篇论文

Discrete speech tokens have gained attention for their storage efficiency and integration with Large Language Models (LLMs). They are commonly categorized into acoustic and semantic tokens, with the latter being more advantageous for…

音频与语音处理 · 电气工程与系统科学 2025-12-04 Mohan Shi , Natarajan Balaji Shankar , Kaiyuan Zhang , Zilai Wang , Abeer Alwan

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ…

音频与语音处理 · 电气工程与系统科学 2023-12-15 Yifan Yang , Feiyu Shen , Chenpeng Du , Ziyang Ma , Kai Yu , Daniel Povey , Xie Chen

Self-supervised learning (SSL) of speech has shown impressive results in speech-related tasks, particularly in automatic speech recognition (ASR). While most methods employ the output of intermediate layers of the SSL model as real-valued…

声音 · 计算机科学 2023-05-30 Xuankai Chang , Brian Yan , Yuya Fujita , Takashi Maekaku , Shinji Watanabe

In this study, we gained insight that contributes to achieving accent-robust ASR using only native speech data. In human perception of non-native speech, the phenomenon known as "interlanguage speech intelligibility benefit" (ISIB) is…

声音 · 计算机科学 2025-05-23 Kentaro Onda , Keisuke Imoto , Satoru Fukayama , Daisuke Saito , Nobuaki Minematsu

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…

计算与语言 · 计算机科学 2025-10-30 Shreyas Gopal , Ashutosh Anshul , Haoyang Li , Yue Heng Yeo , Hexin Liu , Eng Siong Chng

Children are one of the most under-represented groups in speech technologies, as well as one of the most vulnerable in terms of privacy. Despite this, anonymization techniques targeting this population have received little attention. In…

计算机与社会 · 计算机科学 2025-06-05 Ajinkya Kulkarni , Francisco Teixeira , Enno Hermann , Thomas Rolland , Isabel Trancoso , Mathew Magimai Doss

Recent advancements in Automatic Speech Recognition (ASR) systems, exemplified by Whisper, have demonstrated the potential of these systems to approach human-level performance given sufficient data. However, this progress doesn't readily…

音频与语音处理 · 电气工程与系统科学 2024-05-16 Ahmed Adel Attia , Jing Liu , Wei Ai , Dorottya Demszky , Carol Espy-Wilson

Recent studies have highlighted the potential of discrete tokens derived from self-supervised learning (SSL) models for various speech-related tasks. These tokens serve not only as substitutes for text in language modeling but also as…

声音 · 计算机科学 2025-05-23 Kentaro Onda , Yosuke Kashiwagi , Emiru Tsunoo , Hayato Futami , Shinji Watanabe

Automatic Speech Recognition (ASR) systems often struggle with transcribing child speech due to the lack of large child speech datasets required to accurately train child-friendly ASR models. However, there are huge amounts of annotated…

音频与语音处理 · 电气工程与系统科学 2023-07-26 Rishabh Jain , Andrei Barcovschi , Mariam Yiwere , Peter Corcoran , Horia Cucu

With the advancement of Self-supervised Learning (SSL) in speech-related tasks, there has been growing interest in utilizing discrete tokens generated by SSL for automatic speech recognition (ASR), as they offer faster processing…

计算与语言 · 计算机科学 2024-09-16 Mingyu Cui , Daxin Tan , Yifan Yang , Dingdong Wang , Huimeng Wang , Xiao Chen , Xie Chen , Xunying Liu

Automatic speech recognition (ASR) has the potential to substantially reduce manual annotation effort in child speech research by generating automatic transcriptions. However, obtaining reliably high-quality ASR transcriptions for child…

计算与语言 · 计算机科学 2026-05-29 Gus Lathouwers , Lingyun Gao , Catia Cucchiarini , Helmer Strik

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising…

声音 · 计算机科学 2025-07-18 Luca Della Libera , Cem Subakan , Mirco Ravanelli

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

Discretized representations of speech signals are efficient alternatives to continuous features for various speech applications, including automatic speech recognition (ASR) and speech language models. However, these representations, such…

声音 · 计算机科学 2026-02-05 Takanori Ashihara , Shota Horiguchi , Kohei Matsuura , Tsubasa Ochiai , Marc Delcroix

Whispering is a distinct form of speech known for its soft, breathy, and hushed characteristics, often used for private communication. The acoustic characteristics of whispered speech differ substantially from normally phonated speech and…

音频与语音处理 · 电气工程与系统科学 2024-02-08 Zhaofeng Lin , Tanvina Patel , Odette Scharenborg

The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is expected when applying an automatic speech recognition (ASR)…

音频与语音处理 · 电气工程与系统科学 2022-05-26 Wei Liu , Jingyu Li , Tan Lee

The emergence of voice-assistant devices ushers in delightful user experiences not just on the smart home front, but also in diverse educational environments from classrooms to personalized-learning/tutoring. However, the use of voice as an…

音频与语音处理 · 电气工程与系统科学 2021-04-23 Mohammad Niknazar , Aditya Vempaty , Ravi Kokku

Discrete Speech Representation Tokens (DSRTs) have become a foundational component in speech generation. While prior work has extensively studied phonetic and speaker information in DSRTs, how accent information is encoded in DSRTs remains…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Jinzuomu Zhong , Yi Wang , Korin Richmond , Peter Bell

A key desiderata for inclusive and accessible speech recognition technology is ensuring its robust performance to children's speech. Notably, this includes the rapidly advancing neural network based end-to-end speech recognition systems.…

音频与语音处理 · 电气工程与系统科学 2021-02-22 Prashanth Gurunath Shivakumar , Shrikanth Narayanan

Pre-trained models, especially self-supervised learning (SSL) models, have demonstrated impressive results in automatic speech recognition (ASR) task. While most applications of SSL models focus on leveraging continuous representations as…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Zehan Li , Yan Yang , Xueqing Li , Jian Kang , Xiao-Lei Zhang , Jie Li
‹ 上一页 1 2 3 10 下一页 ›