English
Related papers

Related papers: A UD Treebank for Bohairic Coptic

200 papers

Foundational Hebrew NLP tasks such as segmentation, tagging and parsing, have relied to date on various versions of the Hebrew Treebank (HTB, Sima'an et al. 2001). However, the data in HTB, a single-source newswire corpus, is now over 30…

Computation and Language · Computer Science 2022-10-19 Amir Zeldes , Nick Howell , Noam Ordan , Yifat Ben Moshe

Yor\`ub\'a is a widely spoken West African language with a writing system rich in orthographic and tonal diacritics. They provide morphological information, are crucial for lexical disambiguation, pronunciation and are vital for any…

Computation and Language · Computer Science 2020-03-25 Iroro Orife , David I. Adelani , Timi Fasubaa , Victor Williamson , Wuraola Fisayo Oyewusi , Olamilekan Wahab , Kola Tubosun

This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey…

Computation and Language · Computer Science 2026-03-31 Sercan Karakaş

Scholarship on underresourced languages bring with them a variety of challenges which make access to the full spectrum of source materials and their evaluation difficult. For Coptic in particular, large scale analyses and any kind of…

Computation and Language · Computer Science 2023-06-22 Caroline T. Schroeder , Amir Zeldes

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

Computation and Language · Computer Science 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

We develop machine translation and speech synthesis systems to complement the efforts of revitalizing Judeo-Spanish, the exiled language of Sephardic Jews, which survived for centuries, but now faces the threat of extinction in the digital…

Computation and Language · Computer Science 2022-06-01 Alp Öktem , Rodolfo Zevallos , Yasmin Moslem , Güneş Öztürk , Karen Şarhon

Developing effective educational technologies for low-resource agglutinative languages like Uyghur is often hindered by the mismatch between existing annotation frameworks and specific grammatical structures. To address this challenge, this…

Computation and Language · Computer Science 2026-01-21 Jiaxin Zuo , Yiquan Wang , Yuan Pan , Xiadiya Yibulayin

This paper presents UD-NewsCrawl, the largest Tagalog treebank to date, containing 15.6k trees manually annotated according to the Universal Dependencies framework. We detail our treebank development process, including data collection,…

Computation and Language · Computer Science 2025-05-28 Angelina A. Aquino , Lester James V. Miranda , Elsie Marie T. Or

In this study, we aim to offer linguistically motivated solutions to resolve the issues of the lack of representation of null morphemes, highly productive derivational processes, and syncretic morphemes of Turkish in the BOUN Treebank…

Despite the success of the Universal Dependencies (UD) project exemplified by its impressive language breadth, there is still a lack in `within-language breadth': most treebanks focus on standard languages. Even for German, the language…

Computation and Language · Computer Science 2024-03-18 Verena Blaschke , Barbara Kovačić , Siyao Peng , Hinrich Schütze , Barbara Plank

In this paper we present a sample treebank for Old English based on the UD Cairo sentences, collected and annotated as part of a classroom curriculum in Historical Linguistics. To collect the data, a sample of 20 sentences illustrating a…

Computation and Language · Computer Science 2025-06-13 Lauren Levine , Junghyun Min , Amir Zeldes

Natural language processing for the Turkic language family, spoken by over 200 million people across Eurasia, remains fragmented, with most languages lacking unified tooling and resources. We present TurkicNLP, an open-source Python library…

Computation and Language · Computer Science 2026-05-25 Sherzod Hakimov

Amharic is one of the official languages of the Federal Democratic Republic of Ethiopia. It is one of the languages that use an Ethiopic script which is derived from Gee'z, ancient and currently a liturgical language. Amharic is also one of…

Computer Vision and Pattern Recognition · Computer Science 2022-02-28 Mesay Samuel Gondere , Lars Schmidt-Thieme , Durga Prasad Sharma , Abiot Sinamo Boltena

Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. While substantial progress has been made in English, the task…

Computation and Language · Computer Science 2026-01-13 Kareem Elozeiri , Mervat Abassy , Preslav Nakov , Yuxia Wang

Meroitic is the still undeciphered language of the ancient civilization of Kush. Over the years, various techniques for decipherment such as finding a bilingual text or cognates from modern or other ancient languages in the Sudan and…

Computation and Language · Computer Science 2009-08-24 Reginald D. Smith

Machine translation tools do not yet exist for the Yup'ik language, a polysynthetic language spoken by around 8,000 people who live primarily in Southwest Alaska. We compiled a parallel text corpus for Yup'ik and English and developed a…

Computation and Language · Computer Science 2020-09-10 Christopher Liu , Laura Dominé , Kevin Chavez , Richard Socher

Kurdish, an Indo-European language spoken by over 30 million speakers, is considered a dialect continuum and known for its diversity in language varieties. Previous studies addressing language and speech technology for Kurdish handle it in…

Computation and Language · Computer Science 2024-03-05 Sina Ahmadi , Daban Q. Jaff , Md Mahfuz Ibn Alam , Antonios Anastasopoulos

The Tajik language, written in Cyrillic script, remains severely under-resourced in terms of publicly available natural language processing (NLP) toolkits, hindering both linguistic research and applied development. This paper introduces…

Computation and Language · Computer Science 2026-05-29 Mullosharaf K. Arabov

The Kyrgyz language, as a low-resource language, requires significant effort to create high-quality syntactic corpora. This study proposes an approach to simplify the development process of a syntactic corpus for Kyrgyz. We present a tool…

Computation and Language · Computer Science 2024-12-18 Anton Alekseev , Alina Tillabaeva , Gulnara Dzh. Kabaeva , Sergey I. Nikolenko

There has been an increasing interest in learning cross-lingual word embeddings to transfer knowledge obtained from a resource-rich language, such as English, to lower-resource languages for which annotated data is scarce, such as Turkish,…

Computation and Language · Computer Science 2020-05-19 Elmurod Kuriyozov , Yerai Doval , Carlos Gómez-Rodríguez
‹ Prev 1 2 3 10 Next ›