中文
相关论文

相关论文: Fine-Tashkeel: Finetuning Byte-Level Models for Ac…

200 篇论文

Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In this paper, we…

计算与语言 · 计算机科学 2024-01-11 Abderrahman Skiredj , Ismail Berrada

We propose a novel multitask learning method for diacritization which trains a model to both diacritize and translate. Our method addresses data sparsity by exploiting large, readily available bitext corpora. Furthermore, translation…

计算与语言 · 计算机科学 2021-09-30 Brian Thompson , Ali Alshehri

Diacritization of Arabic text is both an interesting and a challenging problem at the same time with various applications ranging from speech synthesis to helping students learning the Arabic language. Like many other tasks or problems in…

计算与语言 · 计算机科学 2019-05-07 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

In this work, we present several deep learning models for the automatic diacritization of Arabic text. Our models are built using two main approaches, viz. Feed-Forward Neural Network (FFNN) and Recurrent Neural Network (RNN), with several…

计算与语言 · 计算机科学 2019-11-12 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

Tashkeel, or Arabic Text Diacritization (ATD), greatly enhances the comprehension of Arabic text by removing ambiguity and minimizing the risk of misinterpretations caused by its absence. It plays a crucial role in improving Arabic text…

计算与语言 · 计算机科学 2024-07-16 Faris Alasmary , Orjuwan Zaafarani , Ahmad Ghannam

Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Standard Arabic (MSA)…

计算与语言 · 计算机科学 2019-06-03 Ahmed Abdelali , Mohammed Attia , Younes Samih , Kareem Darwish , Hamdy Mubarak

We propose a novel architecture for labelling character sequences that achieves state-of-the-art results on the Tashkeela Arabic diacritization benchmark. The core is a two-level recurrence hierarchy that operates on the word and character…

计算与语言 · 计算机科学 2020-11-03 Badr AlKhamissi , Muhammad N. ElNokrashy , Mohamed Gabr

This paper presents a methodology for inserting phrases in Arabic poems to conform to a specific rhythm using ByT5, a byte-level multilingual transformer-based model. Our work discusses a rule-based grapheme-to-beat transformation tailored…

计算与语言 · 计算机科学 2025-12-09 Mohamad Elzohbi , Richard Zhao

Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic TTS training on…

计算与语言 · 计算机科学 2026-03-03 Ahmed Musleh , Yifan Zhang , Kareem Darwish

In many languages like Arabic, diacritics are used to specify pronunciations as well as meanings. Such diacritics are often omitted in written text, increasing the number of possible pronunciations and meanings for a word. This results in a…

计算与语言 · 计算机科学 2020-06-09 Sawsan Alqahtani , Ajay Mishra , Mona Diab

Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like speech synthesis. While most of the previous works focused on…

计算与语言 · 计算机科学 2023-08-01 Parnia Bahar , Mattia Di Gangi , Nick Rossenbach , Mohammad Zeineldeen

Diacritics (short vowels) are typically omitted when writing Arabic text, and readers have to reintroduce them to correctly pronounce words. There are two types of Arabic diacritics: the first are core-word diacritics (CW), which specify…

计算与语言 · 计算机科学 2020-02-05 Kareem Darwish , Ahmed Abdelali , Hamdy Mubarak , Mohamed Eldesouki

Advances in English language representation enabled a more sample-efficient pre-training task by Efficiently Learning an Encoder that Classifies Token Replacements Accurately (ELECTRA). Which, instead of training a model to recover masked…

计算与语言 · 计算机科学 2021-03-09 Wissam Antoun , Fady Baly , Hazem Hajj

A sufficient amount of annotated data is usually required to fine-tune pre-trained language models for downstream tasks. Unfortunately, attaining labeled data can be costly, especially for multiple language varieties and dialects. We…

计算与语言 · 计算机科学 2021-02-04 Muhammad Khalifa , Muhammad Abdul-Mageed , Khaled Shaalan

We observe a recent behaviour on social media, in which users intentionally remove consonantal dots from Arabic letters, in order to bypass content-classification algorithms. Content classification is typically done by fine-tuning…

计算与语言 · 计算机科学 2021-11-19 Aviad Rom , Kfir Bar

We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires producing fully diacritized Arabic text from speech audio and undiacritized transcripts, with…

计算与语言 · 计算机科学 2026-05-26 Meshal Alamr , Hassan Alqaeri , Abdullah Aldahlawi

This study uses a character level neural machine translation approach trained on a long short-term memory-based bi-directional recurrent neural network architecture for diacritization of Medieval Arabic. The results improve from the online…

计算与语言 · 计算机科学 2020-10-13 Khalid Alnajjar , Mika Hämäläinen , Niko Partanen , Jack Rueter

Tokenization is a critical preprocessing step for large language models (LLMs), directly impacting training efficiency and downstream performance. General-purpose tokenizers trained predominantly on English and Latin-script languages…

计算与语言 · 计算机科学 2025-12-23 Mark Kashirskiy , Artiom Lipinski , Ilya Makarov

State-of-the-art performance on language understanding tasks is now achieved with increasingly large networks; the current record holder has billions of parameters. Given a language model pre-trained on massive unlabeled text corpora, only…

计算与语言 · 计算机科学 2020-04-30 Evani Radiya-Dixit , Xin Wang

Large Language Models (LLMs) have significantly advanced the field of natural language processing, enhancing capabilities in both language understanding and generation across diverse domains. However, developing LLMs for Arabic presents…

‹ 上一页 1 2 3 10 下一页 ›