English
Related papers

Related papers: Processing South Asian Languages Written in the La…

200 papers

Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improving cross-lingual transfer (XLT) in multilingual language models (mLMs). Most prior work, however, focused on setups that favor romanization…

Computation and Language · Computer Science 2026-01-12 Benedikt Ebing , Lennart Keller , Goran Glavaš

In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time.…

Computation and Language · Computer Science 2026-02-13 Mariia Fedorova , Andrey Kutuzov , Khonzoda Umarova

Diacritic characters can be considered as a unique set of characters providing us with adequate and significant clue in identifying a given language with considerably high accuracy. Diacritics, though associated with phonetics often serve…

Computer Vision and Pattern Recognition · Computer Science 2021-12-07 Shubham Vatsal , Nikhil Arora , Gopi Ramena , Sukumar Moharana , Dhruval Jain , Naresh Purre , Rachit S Munjal

Communication plays a vital role in human interaction. Studying language is a worthwhile task and more recently has become quantitative in nature with developments of fields like quantitative comparative linguistics and lexicostatistics.…

Applications · Statistics 2024-05-13 Garett Ordway , Vic Patrangenaru

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…

Handwritten text recognition and optical character recognition solutions show excellent results with processing data of modern era, but efficiency drops with Latin documents of medieval times. This paper presents a deep learning method to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Maksym Voloshchuk , Bohdana Zarembovska , Mykola Kozlenko

Data linkage is increasingly used in health research and policy making and is relied on for understanding health inequalities. However, linked data is only as useful as the underlying data quality, and differential linkage rates may induce…

Computation and Language · Computer Science 2026-01-13 Joseph Lam , Mario Cortina-Borja , Robert Aldridge , Ruth Blackburn , Katie Harron

Distributed representation of words has improved the performance for many natural language tasks. In many methods, however, only one meaning is considered for one label of a word, and multiple meanings of polysemous words depending on the…

Computation and Language · Computer Science 2020-06-01 Yusuke Takimoto , Yosuke Fukuchi , Shoya Matsumori , Michita Imai

In this paper we present a bottom up procedure for segmentation of text lines written or printed in the Latin script. The proposed method uses a combination of image morphology, feature extraction and Gaussian mixture model to perform this…

Computer Vision and Pattern Recognition · Computer Science 2017-10-10 Himanshu Jain , Archana Praveen Kumar

In this paper, we propose a Seed-Augment-Train/Transfer (SAT) framework that contains a synthetic seed image dataset generation procedure for languages with different numeral systems using freely available open font file datasets. This seed…

Computer Vision and Pattern Recognition · Computer Science 2019-05-22 Vinay Uday Prabhu , Sanghyun Han , Dian Ang Yap , Mihail Douhaniaris , Preethi Seshadri , John Whaley

Determining the readability of a text is the first step to its simplification. In this paper, we present a readability analysis tool capable of analyzing text written in the Bengali language to provide in-depth information on its…

Computation and Language · Computer Science 2020-12-15 Susmoy Chakraborty , Mir Tafseer Nayeem , Wasi Uddin Ahmad

Most human languages use scripts other than the Latin alphabet. Search users in these languages often formulate their information needs in a transliterated -- usually Latinized -- form for ease of typing. For example, Greek speakers might…

Information Retrieval · Computer Science 2025-05-14 Andreas Chari , Iadh Ounis , Sean MacAvaney

SciLaD is a novel, large-scale dataset of scientific language constructed entirely using open-source frameworks and publicly available data sources. It comprises a curated English split containing over 10 million scientific publications and…

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural language processing…

Computation and Language · Computer Science 2020-09-29 Maha J. Althobaiti

We present our approach to automatically designing and implementing keyboard layouts on mobile devices for typing low-resource languages written in the Latin script. For many speakers, one of the barriers in accessing and creating text…

Computation and Language · Computer Science 2019-01-21 Theresa Breiner , Chieu Nguyen , Daan van Esch , Jeremy O'Brien

As the fourth largest language family in the world, the Dravidian languages have become a research hotspot in natural language processing (NLP). Although the Dravidian languages contain a large number of languages, there are relatively few…

Computation and Language · Computer Science 2021-12-06 Xiaotian Lin , Nankai Lin , Kanoksak Wattanachote , Shengyi Jiang , Lianxi Wang

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

Computation and Language · Computer Science 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

We review the recent literature (January 2022- October 2024) in South Asian languages on text-based language processing, multimodal models, and speech processing, and provide a spotlight analysis focused on 21 low-resource South Asian…

Computation and Language · Computer Science 2025-01-03 Pranav Gupta

Large Language Models (LLMs) exhibit strong multilingual performance despite being predominantly trained on English-centric corpora. This raises a fundamental question: How do LLMs achieve such multilingual capabilities? Focusing on…

Computation and Language · Computer Science 2025-12-23 Alan Saji , Jaavid Aktar Husain , Thanmay Jayakumar , Raj Dabre , Anoop Kunchukuttan , Ratish Puduppully
‹ Prev 1 3 4 5 6 7 10 Next ›