中文
相关论文

相关论文: A UD Treebank for Bohairic Coptic

200 篇论文

The present project endeavors to enrich the linguistic resources available for Italian by constructing a Universal Dependencies treebank for the KIParla corpus (Mauri et al., 2019, Ballar\`e et al., 2020), an existing and well known…

计算与语言 · 计算机科学 2024-10-08 Ludovica Pannitto

CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold-standard Universal…

The success of several architectures to learn semantic representations from unannotated text and the availability of these kind of texts in online multilingual resources such as Wikipedia has facilitated the massive and automatic creation…

计算与语言 · 计算机科学 2020-03-31 Jesujoba O. Alabi , Kwabena Amponsah-Kaakyire , David I. Adelani , Cristina España-Bonet

We present state-of-the-art results on morphosyntactic tagging across different varieties of Arabic using fine-tuned pre-trained transformer language models. Our models consistently outperform existing systems in Modern Standard Arabic and…

计算与语言 · 计算机科学 2022-03-22 Go Inoue , Salam Khalifa , Nizar Habash

This paper presents a comprehensive survey of corpora and lexical resources available for Turkish. We review a broad range of resources, focusing on the ones that are publicly available. In addition to providing information about the…

计算与语言 · 计算机科学 2023-02-28 Çağrı Çöltekin , A. Seza Doğruöz , Özlem Çetinoğlu

We present models which complete missing text given transliterations of ancient Mesopotamian documents, originally written on cuneiform clay tablets (2500 BCE - 100 CE). Due to the tablets' deterioration, scholars often rely on contextual…

计算与语言 · 计算机科学 2021-10-26 Koren Lazar , Benny Saret , Asaf Yehudai , Wayne Horowitz , Nathan Wasserman , Gabriel Stanovsky

We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in the Hebrew language, consisting of a large variety of…

Real-time text-to-speech (TTS) for Modern Hebrew is challenging due to the language's orthographic complexity. Existing solutions ignore crucial phonetic features such as stress that remain underspecified even when vowel marks are added. To…

计算与语言 · 计算机科学 2025-10-13 Yakov Kolani , Maxim Melichov , Cobi Calev , Morris Alper

This study uses a character level neural machine translation approach trained on a long short-term memory-based bi-directional recurrent neural network architecture for diacritization of Medieval Arabic. The results improve from the online…

计算与语言 · 计算机科学 2020-10-13 Khalid Alnajjar , Mika Hämäläinen , Niko Partanen , Jack Rueter

The widespread absence of diacritical marks in Arabic text poses a significant challenge for Arabic natural language processing (NLP). This paper explores instances of naturally occurring diacritics, referred to as "diacritics in the wild,"…

计算与语言 · 计算机科学 2024-06-11 Salman Elgamal , Ossama Obeid , Tameem Kabbani , Go Inoue , Nizar Habash

Machine translation is the task of translating texts from one language to another using computers. It has been one of the major tasks in natural language processing and computational linguistics and has been motivating to facilitate human…

计算与语言 · 计算机科学 2020-10-14 Sina Ahmadi , Mariam Masoud

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Despite increasing interest in Syriac studies and growing digital availability of Syriac texts, there is currently no up-to-date infrastructure for discovering, identifying, classifying, and referencing works of Syriac literature. The…

数字图书馆 · 计算机科学 2021-07-01 Nathan P. Gibson , David A. Michelson , Daniel L. Schwartz

Communication plays a vital role in human interaction. Studying language is a worthwhile task and more recently has become quantitative in nature with developments of fields like quantitative comparative linguistics and lexicostatistics.…

应用统计 · 统计学 2024-05-13 Garett Ordway , Vic Patrangenaru

The Linguistic Data Consortium (LDC) has developed hundreds of data corpora for natural language processing (NLP) research. Among these are a number of annotated treebank corpora for Arabic. Typically, these corpora consist of a single…

计算与语言 · 计算机科学 2013-09-24 Mona Diab , Nizar Habash , Owen Rambow , Ryan Roth

This research stems from the urgency to automate the thematic grouping of hadith in line with the growing digitalization of Islamic texts. Based on a literature review, the unsupervised learning approach with the Apriori algorithm has…

In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful…

计算与语言 · 计算机科学 2020-05-27 Tae Hwan Oh , Ji Yoon Han , Hyonsu Choe , Seokwon Park , Han He , Jinho D. Choi , Na-Rae Han , Jena D. Hwang , Hansaem Kim

The Arabic language is a complex language; it is different from Western languages especially at the morphological and spelling variations. Indeed, the performance of information retrieval systems in the Arabic language is still a problem.…

信息检索 · 计算机科学 2012-04-06 Abd El Salam Al Hajjar , Anis Ismail , Mohammad Hajjar , Mazen El-Sayed

This paper presents and discusses the first Universal Dependencies treebank for the Apurin\~a language. The treebank contains 76 fully annotated sentences, applies 14 parts-of-speech, as well as seven augmented or new features - some of…

Low-resource languages serve as invaluable repositories of human history, preserving cultural and intellectual diversity. Despite their significance, they remain largely absent from modern natural language processing systems. While progress…

计算与语言 · 计算机科学 2026-03-17 Offiong Bassey Edet , Mbuotidem Sunday Awak , Emmanuel Oyo-Ita , Benjamin Okon Nyong , Ita Etim Bassey