中文
相关论文

相关论文: A UD Treebank for Bohairic Coptic

200 篇论文

Natural language processing, as a data analytics related technology, is used widely in many research areas such as artificial intelligence, human language processing, and translation. At present, due to explosive growth of data, there are…

计算与语言 · 计算机科学 2016-08-17 Emre Erturk , Hong Shi

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard…

计算与语言 · 计算机科学 2025-06-02 Bashar Alhafni , Sarah Al-Towaity , Ziyad Fawzy , Fatema Nassar , Fadhl Eryani , Houda Bouamor , Nizar Habash

Hausa language belongs to the Afroasiatic phylum, and with more first-language speakers than any other sub-Saharan African language. With a majority of its speakers residing in the Northern and Southern areas of Nigeria and the Republic of…

计算与语言 · 计算机科学 2021-02-18 Isa Inuwa-Dutse

Treebanks are valuable linguistic resources that include the syntactic structure of a language sentence in addition to POS-tags and morphological features. They are mainly utilized in modeling statistical parsers. Although the statistical…

计算与语言 · 计算机科学 2020-07-14 Dana Halabi , Ebaa Fayyoumi , Arafat Awajan

We release S\={a}mayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited…

We present the first linguistically annotated treebank of Ashokan Prakrit, an early Middle Indo-Aryan dialect continuum attested through Emperor Ashoka Maurya's 3rd century BCE rock and pillar edicts. For annotation, we used the…

计算与语言 · 计算机科学 2021-12-14 Adam Farris , Aryaman Arora

This paper presents the phonological, morphological, and syntactic distinctions between formal and informal Persian, showing that these two variants have fundamental differences that cannot be attributed solely to pronunciation…

计算与语言 · 计算机科学 2022-01-12 Roya Kabiri , Simin Karimi , Mihai Surdeanu

Yor\`ub\'a an African language with roughly 47 million speakers encompasses a continuum with several dialects. Recent efforts to develop NLP technologies for African languages have focused on their standard dialects, resulting in…

This paper explores the possibility of improving the performance of specialized parsers for pre-modern Slavic by training them on data from different related varieties. Because of their linguistic heterogeneity, pre-modern Slavic varieties…

计算与语言 · 计算机科学 2020-11-13 Nilo Pedrazzini

Low-resource machine translation requires methods that differ from those used for high-resource languages. This paper proposes a novel in-context learning approach to support low-resource machine translation of the Coptic language to…

计算与语言 · 计算机科学 2026-05-28 Abhishek Purushothama , Emma Thronson , Alexia Guo , Amir Zeldes

We present the second ever evaluated Arabic dialect-to-dialect machine translation effort, and the first to leverage external resources beyond a small parallel corpus. The subject has not previously received serious attention due to lack of…

计算与语言 · 计算机科学 2017-12-19 Alexander Erdmann , Nizar Habash , Dima Taji , Houda Bouamor

St. Lawrence Island Yupik (ISO 639-3: ess) is an endangered polysynthetic language in the Inuit-Yupik language family indigenous to Alaska and Chukotka. This work presents a step-by-step pipeline for the digitization of written texts, and…

计算与语言 · 计算机科学 2021-01-27 Lane Schwartz , Emily Chen , Hyunji Hayley Park , Edward Jahn , Sylvia L. R. Schreiner

One of the major challenges that under-represented and endangered language communities face in language technology is the lack or paucity of language data. This is also the case of the Southern varieties of the Kurdish and Laki languages…

计算与语言 · 计算机科学 2023-04-05 Sina Ahmadi , Zahra Azin , Sara Belelli , Antonios Anastasopoulos

Despite having a large number of speakers, the Kurdish language is among the less-resourced languages. In this work we highlight the challenges and problems in providing the required tools and techniques for processing texts written in…

信息检索 · 计算机科学 2012-12-04 Kyumars Sheykh Esmaili

The main source of information regarding ancient Mesopotamian history and culture are clay cuneiform tablets. Despite being an invaluable resource, many tablets are fragmented leading to missing information. Currently these missing parts…

计算与语言 · 计算机科学 2022-06-08 Ethan Fetaya , Yonatan Lifshitz , Elad Aaron , Shai Gordin

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

计算与语言 · 计算机科学 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

Machine translation between Arabic and Hebrew has so far been limited by a lack of parallel corpora, despite the political and cultural importance of this language pair. Previous work relied on manually-crafted grammars or pivoting via…

计算与语言 · 计算机科学 2016-09-27 Yonatan Belinkov , James Glass

Amharic is the official language of the Federal Democratic Republic of Ethiopia. There are lots of historic Amharic and Ethiopic handwritten documents addressing various relevant issues including governance, science, religious, social…

计算机视觉与模式识别 · 计算机科学 2022-10-04 Mesay Samuel Gondere , Lars Schmidt-Thieme , Abiot Sinamo Boltena , Hadi Samer Jomaa

This paper presents the first systematic study of strategies for translating Coptic into French. Our comprehensive pipeline systematically evaluates: pivot versus direct translation, the impact of pre-training, the benefits of multi-version…

计算与语言 · 计算机科学 2026-05-14 Nasma Chaoui , Richard Khoury

This paper presents the first publicly available treebank of Odia, a morphologically rich low resource Indian language. The treebank contains approx. 1082 tokens (100 sentences) in Odia selected from "Samantar", the largest available…