中文
相关论文

相关论文: Pitch-Aware RNN-T for Mandarin Chinese Mispronunci…

200 篇论文

Due to the absence of clean reference signals and spatial cues, monaural unsupervised speech dereverberation is a challenging ill-posed inverse problem. To realize it, we propose augmented reverberant-target training (ARTT), which consists…

音频与语音处理 · 电气工程与系统科学 2026-03-20 Siqi Song , Fulin Wu , Zhong-Qiu Wang

The front-end module in multi-channel automatic speech recognition (ASR) systems mainly use microphone array techniques to produce enhanced signals in noisy conditions with reverberation and echos. Recently, neural network (NN) based…

声音 · 计算机科学 2020-11-19 Yuxiang Kong , Jian Wu , Quandong Wang , Peng Gao , Weiji Zhuang , Yujun Wang , Lei Xie

The disparity in phonology between learner's native (L1) and target (L2) language poses a significant challenge for mispronunciation detection and diagnosis (MDD) systems. This challenge is further intensified by lack of annotated L2 data.…

声音 · 计算机科学 2023-08-08 Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali

End-To-End speech recognition have become increasingly popular in mandarin speech recognition and achieved delightful performance. Mandarin is a tonal language which is different from English and requires special treatment for the acoustic…

计算与语言 · 计算机科学 2018-05-15 Wei Zou , Dongwei Jiang , Shuaijiang Zhao , Xiangang Li

Visual speech recognition (VSR) aims to transcribe spoken content from silent lip-motion videos and is particularly challenging in Mandarin due to severe viseme ambiguity and pervasive homophones. We propose VALLR-Pin, a two-stage Mandarin…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Chang Sun , Dongliang Xie , Wanpeng Xie , Bo Qin , Hong Yang

The word error rate (WER) of an automatic speech recognition (ASR) system increases when a mismatch occurs between the training and the testing conditions due to the noise, etc. In this case, the acoustic information can be less reliable.…

计算与语言 · 计算机科学 2020-11-03 Dominique Fohr , Irina Illina

Grapheme-to-phoneme (G2P) conversion serves as an essential component in Chinese Mandarin text-to-speech (TTS) system, where polyphone disambiguation is the core issue. In this paper, we propose an end-to-end framework to predict the…

音频与语音处理 · 电气工程与系统科学 2025-01-03 Dongyang Dai , Zhiyong Wu , Shiyin Kang , Xixin Wu , Jia Jia , Dan Su , Dong Yu , Helen Meng

As false information continues to proliferate across social media platforms, effective rumor detection has emerged as a pressing challenge in natural language processing. This paper proposes RAGAT-Mind, a multi-granular modeling approach…

计算与语言 · 计算机科学 2025-04-25 Zhenkai Qin , Guifang Yang , Dongze Wu

Chinese dialects are different variations of Chinese and can be considered as different languages in the same language family with Mandarin. Though they all use Chinese characters, the pronunciations, grammar and idioms can vary…

计算与语言 · 计算机科学 2022-12-13 Junhui Zhang , Wudi Bao , Junjie Pan , Xiang Yin , Zejun Ma

Grapheme-to-phoneme (G2P) conversion is an indispensable part of the Chinese Mandarin text-to-speech (TTS) system, and the core of G2P conversion is to solve the problem of polyphone disambiguation, which is to pick up the correct…

音频与语音处理 · 电气工程与系统科学 2022-07-26 Song Zhang , Ken Zheng , Xiaoxu Zhu , Baoxiang Li

Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT,…

计算与语言 · 计算机科学 2021-07-01 Zijun Sun , Xiaoya Li , Xiaofei Sun , Yuxian Meng , Xiang Ao , Qing He , Fei Wu , Jiwei Li

In this article, we present an approach for non native automatic speech recognition (ASR). We propose two methods to adapt existing ASR systems to the non-native accents. The first method is based on the modification of acoustic models…

计算与语言 · 计算机科学 2007-11-08 Ghazi Bouselmi , Dominique Fohr , Irina Illina , Jean-Paul Haton

Artificial intelligence (AI) has significantly advanced speech recognition applications. However, many existing neural network-based methods struggle with noise, reducing accuracy in real-world environments. This study addresses isolated…

声音 · 计算机科学 2025-02-12 Ali Nasr-Esfahani , Mehdi Bekrani , Roozbeh Rajabi

Training speech recognizers with unpaired speech and text -- known as unsupervised speech recognition (UASR) -- is a crucial step toward extending ASR to low-resource languages in the long-tail distribution and enabling multimodal learning…

计算与语言 · 计算机科学 2025-10-07 Liming Wang , Junrui Ni , Kai-Wei Chang , Saurabhchand Bhati , David Harwath , Mark Hasegawa-Johnson , James R. Glass

Multilingual models can improve language processing, particularly for low resource situations, by sharing parameters across languages. Multilingual acoustic models, however, generally ignore the difference between phonemes (sounds that can…

One of the challenges in automatic speech recognition is foreign words recognition. It is observed that a speaker's pronunciation of a foreign word is influenced by his native language knowledge, and such phenomenon is known as the effect…

计算与语言 · 计算机科学 2022-10-10 Lei Wang , Rong Tong

State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose…

声音 · 计算机科学 2025-03-13 Jiaming Zhou , Shiwan Zhao , Jiabei He , Hui Wang , Wenjia Zeng , Yong Chen , Haoqin Sun , Aobo Kong , Yong Qin

Machine reading comprehension (MRC) is an important area of conversation agents and draws a lot of attention. However, there is a notable limitation to current MRC benchmarks: The labeled answers are mostly either spans extracted from the…

计算与语言 · 计算机科学 2023-10-10 Nuo Chen , Hongguang Li , Yinan Bao , Baoyuan Wang , Jia Li

This paper proposes a modification to RNN-Transducer (RNN-T) models for automatic speech recognition (ASR). In standard RNN-T, the emission of a blank symbol consumes exactly one input frame; in our proposed method, we introduce additional…

音频与语音处理 · 电气工程与系统科学 2024-04-15 Hainan Xu , Fei Jia , Somshubra Majumdar , Shinji Watanabe , Boris Ginsburg

Bidirectional Encoder Representations from Transformers (BERT) has shown marvelous improvements across various NLP tasks, and its consecutive variants have been proposed to further improve the performance of the pre-trained language models.…

计算与语言 · 计算机科学 2021-11-29 Yiming Cui , Wanxiang Che , Ting Liu , Bing Qin , Ziqing Yang