中文
相关论文

相关论文: Experiments in Cuneiform Language Identification

200 篇论文

Identification of the languages written using cuneiform symbols is a difficult task due to the lack of resources and the problem of tokenization. The Cuneiform Language Identification task in VarDial 2019 addresses the problem of…

计算与语言 · 计算机科学 2020-09-24 Ehsan Doostmohammadi , Minoo Nassajian

This article introduces a corpus of cuneiform texts from which the dataset for the use of the Cuneiform Language Identification (CLI) 2019 shared task was derived as well as some preliminary language identification experiments conducted…

计算与语言 · 计算机科学 2019-03-14 Tommi Jauhiainen , Heidi Jauhiainen , Tero Alstola , Krister Lindén

The work in this paper describes the training and evaluation of machine learning (ML) techniques for the classification of cuneiform signs. There is a lot of variability in cuneiform signs, depending on where they come from, for what and by…

机器学习 · 计算机科学 2025-07-21 Eli Verwimp , Gustav Ryberg Smidt , Hendrik Hameeuw , Katrien De Graef

Cuneiform is the earliest known system of writing, first developed for the Sumerian language of southern Mesopotamia in the second half of the 4th millennium BC. Cuneiform signs are obtained by impressing a stylus on fresh clay tablets. For…

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc.,…

计算与语言 · 计算机科学 2024-06-27 Milind Agarwal , Joshua Otten , Antonios Anastasopoulos

Cuneiform writing, an old art style, allows us to see into the past. Aside from Egyptian hieroglyphs, the cuneiform script is one of the oldest writing systems. Many historians place Hebrew's origins in antiquity. For example, we used the…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Elaf A. Saeed , Ammar D. Jasim , Munther A. Abdul Malik

Despite the recent advancements of attention-based deep learning architectures across a majority of Natural Language Processing tasks, their application remains limited in a low-resource setting because of a lack of pre-trained models for…

计算与语言 · 计算机科学 2021-06-01 Rachit Bansal , Himanshu Choudhary , Ravneet Punia , Niko Schenk , Jacob L Dahl , Émilie Pagé-Perron

This paper presents a thoroughly automated method for identifying and interpreting cuneiform characters via advanced deep-learning algorithms. Five distinct deep-learning models were trained on a comprehensive dataset of cuneiform…

计算与语言 · 计算机科学 2025-05-09 Shahad Elshehaby , Alavikunhu Panthakkan , Hussain Al-Ahmad , Mina Al-Saad

Discriminating between closely-related language varieties is considered a challenging and important task. This paper describes our submission to the DSL 2016 shared-task, which included two sub-tasks: one on discriminating similar languages…

计算与语言 · 计算机科学 2016-09-27 Yonatan Belinkov , James Glass

This paper presents an ensemble system combining the output of multiple SVM classifiers to native language identification (NLI). The system was submitted to the NLI Shared Task 2017 fusion track which featured students essays and spoken…

计算与语言 · 计算机科学 2017-07-25 Marcos Zampieri , Alina Maria Ciobanu , Liviu P. Dinu

Dialect Identification is a crucial task for localizing various Large Language Models. This paper outlines our approach to the VarDial 2023 shared task. Here we have to identify three or two dialects from three languages each which results…

计算与语言 · 计算机科学 2023-03-29 Ankit Vaidya , Aditya Kane

This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing…

计算与语言 · 计算机科学 2021-03-10 Tommi Jauhiainen , Tharindu Ranasinghe , Marcos Zampieri

The cuneiform script constitutes one of the earliest systems of writing and is realized by wedge-shaped marks on clay tablets. A tremendous number of cuneiform tablets have already been discovered and are incrementally digitalized and made…

计算机视觉与模式识别 · 计算机科学 2018-03-12 Nils M. Kriege , Matthias Fey , Denis Fisseler , Petra Mutzel , Frank Weichert

Hawrami, a dialect of Kurdish, is classified as an endangered language as it suffers from the scarcity of data and the gradual loss of its speakers. Natural Language Processing projects can be used to partially compensate for data…

计算与语言 · 计算机科学 2024-09-26 Aram Khaksar , Hossein Hassani

Language Identification (LID) is an important component of many multilingual natural language processing pipelines, where it facilitates corpus curation, training data analysis, and cross-lingual evaluation of large language models. Despite…

计算与语言 · 计算机科学 2026-02-20 Clara Meister , Ahmetcan Yavuz , Pietro Lesci , Tiago Pimentel

Sumerian transliteration is a conventional system for representing a scholar's interpretation of a tablet in the Latin script. Thanks to visionary digital Assyriology projects such as ETCSL, CDLI, and Oracc, a large number of Sumerian…

计算与语言 · 计算机科学 2026-02-26 Cole Simmons , Richard Diehl Martinez , Dan Jurafsky

In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate…

计算与语言 · 计算机科学 2018-08-15 Liviu P. Dinu , Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi

Segmentation is a fundamental step for most Natural Language Processing tasks. The Kurdish language is a multi-dialect, under-resourced language which is written in different scripts. The lack of various segmented corpora is one of the…

计算与语言 · 计算机科学 2020-05-01 Roshna Omer Abdulrahman , Hossein Hassani

The Perso-Arabic scripts are a family of scripts that are widely adopted and used by various linguistic communities around the globe. Identifying various languages using such scripts is crucial to language technologies and challenging in…

计算与语言 · 计算机科学 2023-04-05 Sina Ahmadi , Milind Agarwal , Antonios Anastasopoulos

Twenty-five hundred years ago, the paperwork of the Achaemenid Empire was recorded on clay tablets. In 1933, archaeologists from the University of Chicago's Oriental Institute (OI) found tens of thousands of these tablets and fragments…

计算机视觉与模式识别 · 计算机科学 2025-02-04 Edward C. Williams , Grace Su , Sandra R. Schloen , Miller C. Prosser , Susanne Paulus , Sanjay Krishnan
‹ 上一页 1 2 3 10 下一页 ›