中文
相关论文

相关论文: Towards Transliteration between Sindhi Scripts fro…

200 篇论文

Cross-lingual transfer has become an effective way of transferring knowledge between languages. In this paper, we explore an often overlooked aspect in this domain: the influence of the source language of a language model on language…

计算与语言 · 计算机科学 2024-05-08 Wondimagegnhue Tsegaye Tufa , Ilia Markov , Piek Vossen

The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such…

计算与语言 · 计算机科学 2025-11-19 Adrian Benton , Alexander Gutkin , Christo Kirov , Brian Roark

In this work, we develop language models for the Sanskrit language, namely Bidirectional Encoder Representations from Transformers (BERT) and its variants: A Lite BERT (ALBERT), and Robustly Optimized BERT (RoBERTa) using Devanagari…

计算与语言 · 计算机科学 2023-04-05 Kartik Bhatnagar , Sampath Lonka , Jammi Kunal , Mahabala Rao M G

The OpenITI team has achieved Optical Character Recognition (OCR) accuracy rates for classical Arabic-script texts in the high nineties. These numbers are based on our tests of seven different Arabic-script texts of varying quality and…

计算机视觉与模式识别 · 计算机科学 2017-03-29 Maxim Romanov , Matthew Thomas Miller , Sarah Bowen Savant , Benjamin Kiessling

Generating code-switched text is a problem of growing interest, especially given the scarcity of corpora containing large volumes of real code-switched text. In this work, we adapt a state-of-the-art neural machine translation model to…

计算与语言 · 计算机科学 2021-07-15 Ishan Tarunesh , Syamantak Kumar , Preethi Jyothi

Language Identification is a very important part of several text processing pipelines. Extensive research has been done in this field. This paper proposes a procedure for automatic language identification of poems for poem analysis task,…

计算与语言 · 计算机科学 2021-01-01 Priyankit Acharya , Aditya Ku. Pathak , Rakesh Ch. Balabantaray , Anil Ku. Singh

We present a novel technique to remove spurious ambiguity from transition systems for dependency parsing. Our technique chooses a canonical sequence of transition operations (computation) for a given dependency tree. Our technique can be…

计算与语言 · 计算机科学 2012-06-29 Shay B. Cohen , Carlos Gómez-Rodríguez , Giorgio Satta

Sandhi means to join two or more words to coin new word. Sandhi literally means `putting together' or combining (of sounds), It denotes all combinatory sound-changes effected (spontaneously) for ease of pronunciation. Sandhi-vicheda…

计算与语言 · 计算机科学 2009-09-15 Priyanka Gupta , Vishal Goyal

Machine Translation for Indian languages is an emerging research area. Transliteration is one such module that we design while designing a translation system. Transliteration means mapping of source language text into the target language.…

计算与语言 · 计算机科学 2013-07-15 Juhi Ameta , Nisheeth Joshi , Iti Mathur

Searching for words in Sanskrit E-text is a problem that is accompanied by complexities introduced by features of Sanskrit such as euphonic conjunctions or sandhis. A word could occur in an E-text in a transformed form owing to the…

计算与语言 · 计算机科学 2014-09-16 S. V. Kasmir Raja , V. Rajitha , Meenakshi Lakshmanan

In large societies like India there is a huge demand to convert one human language into another. Lots of work has been done in this area. Many transfer based MTS have developed for English to other languages, as MANTRA CDAC Pune, MATRA CDAC…

计算与语言 · 计算机科学 2015-07-09 Akanksha Gehlot , Vaishali Sharma , Shashi Pal Singh , Ajai Kumar

The problem of converting images of text into plain text is a widely researched topic in both academia and industry. Arabic handwritten Text Recognation (AHTR) poses additional challenges due to diverse handwriting styles and limited…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Alhossien Waly , Bassant Tarek , Ali Feteha , Rewan Yehia , Gasser Amr , Ahmed Fares

Comprehensively searching for words in Sanskrit E-text is a non-trivial problem because words could change their forms in different contexts. One such context is sandhi or euphonic conjunctions, which cause a word to change owing to the…

计算与语言 · 计算机科学 2019-08-17 S. V. Kasmir Raja , V. Rajitha , Meenakshi Lakshmanan

Most human languages use scripts other than the Latin alphabet. Search users in these languages often formulate their information needs in a transliterated -- usually Latinized -- form for ease of typing. For example, Greek speakers might…

信息检索 · 计算机科学 2025-05-14 Andreas Chari , Iadh Ounis , Sean MacAvaney

Multilingual pre-trained models (mPLMs) have shown impressive performance on cross-lingual transfer tasks. However, the transfer performance is often hindered when a low-resource target language is written in a different script than the…

计算与语言 · 计算机科学 2024-10-10 Orgest Xhelili , Yihong Liu , Hinrich Schütze

An automated approach to text readability assessment is essential to a language and can be a powerful tool for improving the understandability of texts written and published in that language. However, the Persian language, which is spoken…

计算与语言 · 计算机科学 2020-04-23 Hamid Mohammadi , Seyed Hossein Khasteh

Farsi, also known as Persian, is the official language of Iran and Tajikistan and one of the two main languages spoken in Afghanistan. Farsi enjoys a unified Arabic script as its writing system. In this paper we briefly introduce the…

计算与语言 · 计算机科学 2014-04-21 Behrang QasemiZadeh , Saeed Rahimi , Mehdi Safaee Ghalati

Extracting Handwritten text is one of the most important components of digitizing information and making it available for large scale setting. Handwriting Optical Character Reader (OCR) is a research problem in computer vision and natural…

计算机视觉与模式识别 · 计算机科学 2022-06-10 Mohammad Daniyal Shaiq , Musa Dildar Ahmed Cheema , Ali Kamal

The paper presents a two stage classification approach for handwritten devanagari characters The first stage is using structural properties like shirorekha, spine in character and second stage exploits some intersection features of…

计算机视觉与模式识别 · 计算机科学 2010-07-01 Sandhya Arora , Debotosh Bhattacharjee , Mita Nasipuri , Latesh Malik

Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like speech synthesis. While most of the previous works focused on…

计算与语言 · 计算机科学 2023-08-01 Parnia Bahar , Mattia Di Gangi , Nick Rossenbach , Mohammad Zeineldeen