中文
相关论文

相关论文: Morphological Synthesizer for Ge'ez Language: Addr…

200 篇论文

Arabic is a Semitic language which is widely spoken with many dialects. Given the success of pre-trained language models, many transformer models trained on Arabic and its dialects have surfaced. While these models have been compared with…

计算与语言 · 计算机科学 2022-11-18 Ahmed Abdelali , Nadir Durrani , Fahim Dalvi , Hassan Sajjad

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent resources: a corpus…

计算与语言 · 计算机科学 2022-07-12 Elisa Gugliotta , Marco Dinarelli

Named Entity Recognition (NER) is a fundamental NLP task, commonly formulated as classification over a sequence of tokens. Morphologically-Rich Languages (MRLs) pose a challenge to this basic formulation, as the boundaries of Named Entities…

计算与语言 · 计算机科学 2021-09-14 Dan Bareket , Reut Tsarfaty

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource and morphologically rich languages are still not well understood due to limited annotated resources and the absence of standardized…

计算与语言 · 计算机科学 2026-04-01 Hailay Kidu Teklehaymanot , Gebrearegawi Gebremariam , Wolfgang Nejdl

Proponents of the Distributed Morphology framework have posited the existence of two levels of morphological word formation: a lower one, leading to loose input-output semantic relationships; and an upper one, leading to tight input-output…

计算与语言 · 计算机科学 2022-08-12 Ido Benbaji , Omri Doron , Adèle Hénot-Mortier

Pre-trained language models (PLMs) have shown remarkable successes in acquiring a wide range of linguistic knowledge, relying solely on self-supervised training on text streams. Nevertheless, the effectiveness of this language-agnostic…

计算与语言 · 计算机科学 2023-11-02 Eylon Gueta , Omer Goldman , Reut Tsarfaty

A morphological analyzer, which is a significant component of many natural language processing applications especially for morphologically rich languages, divides an input word into all its composing morphemes and identifies their…

计算与语言 · 计算机科学 2022-03-04 Morteza Naserzade , Aso Mahmudi , Hadi Veisi , Hawre Hosseini , Mohammad MohammadAmini

This memoir explores two fundamental aspects of Natural Language Processing (NLP): the creation of linguistic resources and the evaluation of NLP system performance. Over the past decade, my work has focused on developing a morpheme-based…

计算与语言 · 计算机科学 2026-02-16 Jungyeul Park

Text generation systems are ubiquitous in natural language processing applications. However, evaluation of these systems remains a challenge, especially in multilingual settings. In this paper, we propose L'AMBRE -- a metric to evaluate the…

In the last few years, pre-trained neural architectures have provided impressive improvements across several NLP tasks. Still, generative language models are available mainly for English. We develop GePpeTto, the first generative language…

计算与语言 · 计算机科学 2020-04-30 Lorenzo De Mattei , Michele Cafagna , Felice Dell'Orletta , Malvina Nissim , Marco Guerini

Grapheme-to-phoneme (G2P) conversion is critical in speech processing, particularly for applications like speech synthesis. G2P systems must possess linguistic understanding and contextual awareness of languages with polyphone words and…

计算与语言 · 计算机科学 2024-09-16 Mahta Fetrat Qharabagh , Zahra Dehghanian , Hamid R. Rabiee

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people. The Amharic writing system, Abugida, has 282 syllables, 15…

计算机视觉与模式识别 · 计算机科学 2022-04-12 Wondimu Dikubab , Dingkang Liang , Minghui Liao , Xiang Bai

Morphological parsing is the task of decomposing words into morphemes, the smallest units of meaning in a language, and labelling their grammatical roles. It is a particularly challenging task for agglutinative languages, such as the Nguni…

计算与语言 · 计算机科学 2025-05-20 Cael Marquard , Simbarashe Mawere , Francois Meyer

As large language models (LLMs) become increasingly central to Arabic NLP applications, evaluating their understanding of regional dialects and cultural nuances is essential, particularly in linguistically diverse settings like Saudi…

计算与语言 · 计算机科学 2025-12-23 Renad Al-Monef , Hassan Alhuzali , Nora Alturayeif , Ashwag Alasmari

Tokenization is a critical part of modern NLP pipelines. However, contemporary tokenizers for Large Language Models are based on statistical analysis of text corpora, without much consideration to the linguistic features. I propose a…

计算与语言 · 计算机科学 2024-02-06 Haris Jabbar

Efficiently supporting sound gradual typing in a language with structural types is challenging. To date, the Grift compiler is the only close-to-the-metal implementation of gradual typing in this setting, exploiting coercions for runtime…

编程语言 · 计算机科学 2025-12-30 José Luis Romero , Cristóbal Isla , Matías Toro , Éric Tanter

With over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world. Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems…

计算与语言 · 计算机科学 2025-10-03 Jesujoba O. Alabi , Michael A. Hedderich , David Ifeoluwa Adelani , Dietrich Klakow

Named Entity Recognition is an information extraction task that serves as a preprocessing step for other natural language processing tasks, such as machine translation, information retrieval, and question answering. Named entity recognition…

计算与语言 · 计算机科学 2022-07-05 Ebrahim Chekol Jibril , A. Cüneyd Tantğ

We present Afro-MNIST, a set of synthetic MNIST-style datasets for four orthographies used in Afro-Asiatic and Niger-Congo languages: Ge`ez (Ethiopic), Vai, Osmanya, and N'Ko. These datasets serve as "drop-in" replacements for MNIST. We…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Daniel J Wu , Andrew C Yang , Vinay U Prabhu

Pre-trained language models such as BERT have been successful at tackling many natural language processing tasks. However, the unsupervised sub-word tokenization methods commonly used in these models (e.g., byte-pair encoding - BPE) are…

计算与语言 · 计算机科学 2023-04-26 Antoine Nzeyimana , Andre Niyongabo Rubungo