English
Related papers

Related papers: Survey on Publicly Available Sinhala Natural Langu…

200 papers

In this paper, we examine the research conducted in the field of Nepali Automatic Speech Recognition (ASR). The primary objective of this survey is to conduct a comprehensive review of the works on Nepali Automatic Speech Recognition…

Sound · Computer Science 2024-02-06 Rupak Raj Ghimire , Bal Krishna Bal , Prakash Poudyal

Tamil, a Dravidian language of South Asia, is a highly diglossic language with two very different registers in everyday use: Literary Tamil (preferred in writing and formal communication) and Spoken Tamil (confined to speech and informal…

Computation and Language · Computer Science 2023-11-15 Kabilan Prasanna , Aryaman Arora

Large Language Models (LLMs) have shown strong generalization across tasks in high-resource languages; however, their linguistic competence in low-resource and morphologically rich languages such as Tamil remains largely unexplored.…

Computation and Language · Computer Science 2025-11-18 Jeyarajalingam Varsha , Menan Velayuthan , Sumirtha Karunakaran , Rasan Nivethiga , Kengatharaiyer Sarveswaran

SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annotated carefully based on the written date, after filtering…

Computation and Language · Computer Science 2026-05-19 Nevidu Jayatilleke , Nisansa de Silva

This thesis argues that the currently widely used Natural Language Processing algorithms possibly have various limitations related to the properties of the texts they handle and produce. With the wide adoption of these tools in rapid…

Computation and Language · Computer Science 2024-09-17 Josef Jon

Slang is a predominant form of informal language making flexible and extended use of words that is notoriously hard for natural language processing systems to interpret. Existing approaches to slang interpretation tend to rely on context…

Computation and Language · Computer Science 2022-05-03 Zhewei Sun , Richard Zemel , Yang Xu

Natural language processing (NLP) has largely focused on modelling standardized languages. More recently, attention has increasingly shifted to local, non-standardized languages and dialects. However, the relevant speaker populations' needs…

Computation and Language · Computer Science 2024-06-10 Verena Blaschke , Christoph Purschke , Hinrich Schütze , Barbara Plank

Stereotype repositories are critical to assess generative AI model safety, but currently lack adequate global coverage. It is imperative to prioritize targeted expansion, strategically addressing existing deficits, over merely increasing…

Computation and Language · Computer Science 2026-02-27 Aishwarya Verma , Laud Ammah , Olivia Nercy Ndlovu Lucas , Andrew Zaldivar , Vinodkumar Prabhakaran , Sunipa Dev

In the last half-decade, the field of natural language processing (NLP) has undergone two major transitions: the switch to neural networks as the primary modeling paradigm and the homogenization of the training regime (pre-train, then…

Computation and Language · Computer Science 2021-10-19 Artur Kulmizev , Joakim Nivre

Developing Information Retrieval (IR) tools and techniques in African languages suffers from the dual problems of a lack of algorithms and very small test data collections. This affects the creation of practical IR systems and limits the…

Information Retrieval · Computer Science 2018-06-14 Hussein Suleman

The term natural language refers to any system of symbolic communication (spoken, signed or written) without intentional human planning and design. This distinguishes natural languages such as Arabic and Japanese from artificially…

Multilingual Large Language Models are capable of using powerful Large Language Models to handle and respond to queries in multiple languages, which achieves remarkable success in multilingual natural language processing tasks. Despite…

Computation and Language · Computer Science 2024-04-09 Libo Qin , Qiguang Chen , Yuhang Zhou , Zhi Chen , Yinghui Li , Lizi Liao , Min Li , Wanxiang Che , Philip S. Yu

S\'ami, an indigenous language group comprising multiple languages, faces digital marginalization due to the limited availability of data and sophisticated language models designed for its linguistic intricacies. This work focuses on…

Computation and Language · Computer Science 2024-05-10 Ronny Paul , Himanshu Buckchash , Shantipriya Parida , Dilip K. Prasad

Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs,…

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

Measuring the semantic similarity between two sentences (or Semantic Textual Similarity - STS) is fundamental in many NLP applications. Despite the remarkable results in supervised settings with adequate labeling, little attention has been…

Computation and Language · Computer Science 2018-10-31 Xin Tang , Shanbo Cheng , Loc Do , Zhiyu Min , Feng Ji , Heng Yu , Ji Zhang , Haiqin Chen

This work explores the utilization of Romanized Sinhala social media data to identify individuals at risk of depression. A machine learning-based framework is presented for the automatic screening of depression symptoms by analyzing…

Computation and Language · Computer Science 2024-04-01 Jayathi Hewapathirana , Deshan Sumanathilaka

SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the 5th to the 20th century CE in terms of written dates. The…

Computation and Language · Computer Science 2026-03-12 Nevidu Jayatilleke , Nisansa de Silva , Uthpala Nimanthi , Gagani Kulathilaka , Azra Safrullah , Johan Sofalas

The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effectively excluding indigenous languages of the Global South from the AI revolution. The Tharu language, an Indo-Aryan vernacular spoken by…

Computation and Language · Computer Science 2026-03-19 Prajwal Panth , Agniva Maiti

Since 2022 we have been exploring application areas and technologies in which Artificial Intelligence (AI) and modern Natural Language Processing (NLP), such as Large Language Models (LLMs), can be employed to foster the usage and…