English
Related papers

Related papers: g2pW: A Conditional Weighted Softmax BERT for Poly…

200 papers

In this paper, we propose a novel multi-modal multi-task encoder-decoder pre-training framework (MMSpeech) for Mandarin automatic speech recognition (ASR), which employs both unlabeled speech and text data. The main difficulty in…

Multimedia · Computer Science 2022-12-02 Xiaohuan Zhou , Jiaming Wang , Zeyu Cui , Shiliang Zhang , Zhijie Yan , Jingren Zhou , Chang Zhou

This paper introduces PnG BERT, a new encoder model for neural TTS. This model is augmented from the original BERT model, by taking both phoneme and grapheme representations of text as input, as well as the word-level alignment between…

Computation and Language · Computer Science 2021-06-08 Ye Jia , Heiga Zen , Jonathan Shen , Yu Zhang , Yonghui Wu

Lip-to-speech (L2S) synthesis for Mandarin is a significant challenge, hindered by complex viseme-to-phoneme mappings and the critical role of lexical tones in intelligibility. To address this issue, we propose Lexical Tone-Aware…

Sound · Computer Science 2025-10-01 Kang Yang , Yifan Liang , Fangkun Liu , Zhenping Xie , Chengshi Zheng

BERT-based models have shown a remarkable ability in the Chinese Spelling Check (CSC) task recently. However, traditional BERT-based methods still suffer from two limitations. First, although previous works have identified that explicit…

Computation and Language · Computer Science 2023-12-29 Yongchang Cao , Liang He , Zhen Wu , Xinyu Dai

Although end-to-end text-to-speech (TTS) models can generate natural speech, challenges still remain when it comes to estimating sentence-level phonetic and prosodic information from raw text in Japanese TTS systems. In this paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-25 Rem Hida , Masaki Hamada , Chie Kamada , Emiru Tsunoo , Toshiyuki Sekiya , Toshiyuki Kumakura

Can pre-trained BERT for one language and GPT for another be glued together to translate texts? Self-supervised training using only monolingual data has led to the success of pre-trained (masked) language models in many NLP tasks. However,…

Computation and Language · Computer Science 2021-09-14 Zewei Sun , Mingxuan Wang , Lei Li

Text-to-Text Transfer Transformer (T5) has recently been considered for the Grapheme-to-Phoneme (G2P) transduction. As a follow-up, a tokenizer-free byte-level model based on T5 referred to as ByT5, recently gave promising results on…

Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-20 Ziyue Jiang , Zhe Su , Zhou Zhao , Qian Yang , Yi Ren , Jinglin Liu , Zhenhui Ye

As a key component of automated speech recognition (ASR) and the front-end in text-to-speech (TTS), grapheme-to-phoneme (G2P) plays the role of converting letters to their corresponding pronunciations. Existing methods are either slow or…

Computation and Language · Computer Science 2023-03-03 Chunfeng Wang , Peisong Huang , Yuxiang Zou , Haoyu Zhang , Shichao Liu , Xiang Yin , Zejun Ma

Word meaning, representation, and interpretation play fundamental roles in natural language understanding (NLU), natural language processing (NLP), and natural language generation (NLG) tasks. Many of the inherent difficulties in these…

Computation and Language · Computer Science 2025-12-18 Lifeng Han , Gareth J. F. Jones , Alan F. Smeaton

Attention mechanism is one of the most successful techniques in deep learning based Natural Language Processing (NLP). The transformer network architecture is completely based on attention mechanisms, and it outperforms sequence-to-sequence…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-30 Sevinj Yolchuyeva , Géza Németh , Bálint Gyires-Tóth

Despite the development of pre-trained language models (PLMs) significantly raise the performances of various Chinese natural language processing (NLP) tasks, the vocabulary for these Chinese PLMs remain to be the one provided by Google…

Computation and Language · Computer Science 2020-11-18 Wei Zhu

Prepositions are frequently occurring polysemous words. Disambiguation of prepositions is crucial in tasks like semantic role labelling, question answering, text entailment, and noun compound paraphrasing. In this paper, we propose a novel…

Computation and Language · Computer Science 2021-11-30 Siddhesh Pawar , Shyam Thombre , Anirudh Mittal , Girishkumar Ponkiya , Pushpak Bhattacharyya

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was performed on…

Computation and Language · Computer Science 2025-09-30 Bassam Matar , Mohamed Fayed , Ayman Khalafallah

Encoder-only Transformers have advanced along three axes -- architecture, data, and systems -- yielding Pareto gains in accuracy, speed, and memory efficiency. Yet these improvements have not fully transferred to Chinese, where tokenization…

Computation and Language · Computer Science 2025-10-15 Zeyu Zhao , Ningtao Wang , Xing Fu , Yu Cheng

Lightweight, real-time text-to-speech systems are crucial for accessibility. However, the most efficient TTS models often rely on lightweight phonemizers that struggle with context-dependent challenges. In contrast, more advanced…

Sound · Computer Science 2025-12-15 Mahta Fetrat , Donya Navabi , Zahra Dehghanian , Morteza Abolghasemi , Hamid R. Rabiee

Recently, pre-trained language representation flourishes as the mainstay of the natural language understanding community, e.g., BERT. These pre-trained language representations can create state-of-the-art results on a wide range of…

Machine Learning · Computer Science 2019-12-24 Fu-Ming Guo , Sijia Liu , Finlay S. Mungall , Xue Lin , Yanzhi Wang

Masked Language Modeling (MLM) is widely used to pretrain language models. The standard random masking strategy in MLM causes the pre-trained language models (PLMs) to be biased toward high-frequency tokens. Representation learning of rare…

Computation and Language · Computer Science 2023-05-25 Linhan Zhang , Qian Chen , Wen Wang , Chong Deng , Xin Cao , Kongzhang Hao , Yuxin Jiang , Wei Wang

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon…

This paper proposes an additive phoneme-aware margin softmax (APM-Softmax) loss to train the multi-task learning network with phonetic information for language recognition. In additive margin softmax (AM-Softmax) loss, the margin is set as…

Sound · Computer Science 2021-06-25 Zheng Li , Yan Liu , Lin Li , Qingyang Hong