English
Related papers

Related papers: Jambu: A historical linguistic database for South …

200 papers

Language models (LMs) have excelled in various broad domains. However, to ensure their safe and effective integration into real-world educational settings, they must demonstrate proficiency in specific, granular areas of knowledge. Existing…

Computation and Language · Computer Science 2025-05-27 Sagi Shaier , George Arthur Baker , Chiranthan Sridhar , Lawrence E Hunter , Katharina von der Wense

Existing datasets for relation classification and extraction often exhibit limitations such as restricted relation types and domain-specific biases. This work presents a generic framework to generate well-structured sentences from given…

Information Retrieval · Computer Science 2024-12-31 Mansi , Pranshu Pandya , Mahek Bhavesh Vora , Soumya Bharadwaj , Ashish Anand

The Houma Alliance Book is one of the national treasures of the Museum in Shanxi Museum Town in China. It has great historical significance in researching ancient history. To date, the research on the Houma Alliance Book has been staying in…

Computer Vision and Pattern Recognition · Computer Science 2022-08-03 Xiaoyu Yuan , Zhibo Zhang , Yabo Sun , Zekai Xue , Xiuyan Shao , Xiaohua Huang

Indigenous African languages are categorized as under-served in Natural Language Processing. They therefore experience poor digital inclusivity and information access. The processing challenge with such languages has been how to use machine…

Computation and Language · Computer Science 2025-01-17 Barack Wanjawa , Lilian Wanzare , Florence Indede , Owen McOnyango , Edward Ombui , Lawrence Muchemi

We release Samas\=amayik, a novel, meticulously curated, large-scale Hindi-Sanskrit corpus, comprising 92,196 parallel sentences. Unlike most data available in Sanskrit, which focuses on classical era text and poetry, this corpus aggregates…

Computation and Language · Computer Science 2026-03-26 N J Karthika , Keerthana Suryanarayanan , Jahanvi Purohit , Ganesh Ramakrishnan , Jitin Singla , Anil Kumar Gourishetty

Although the Indonesian language is spoken by almost 200 million people and the 10th most spoken language in the world, it is under-represented in NLP research. Previous work on Indonesian has been hampered by a lack of annotated datasets,…

Computation and Language · Computer Science 2020-11-03 Fajri Koto , Afshin Rahimi , Jey Han Lau , Timothy Baldwin

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

Computation and Language · Computer Science 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Sentiment Analysis (SA) is an action research area in the digital age. With rapid and constant growth of online social media sites and services, and the increasing amount of textual data such as - statuses, comments, reviews etc. available…

Computation and Language · Computer Science 2016-11-28 A. Hassan , M. R. Amin , N. Mohammed , A. K. A. Azad

This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects Lisan corpora. Lisan features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The…

Computation and Language · Computer Science 2022-12-20 Mustafa Jarrar , Fadi A Zaraket , Tymaa Hammouda , Daanish Masood Alavi , Martin Waahlisch

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

Computation and Language · Computer Science 2026-05-18 Nuwan I. Senaratna

If today some African languages like Swahili have enough resources to develop high-performing Natural Language Processing (NLP) systems, many other languages spoken on the continent are still lacking such support. For these languages, still…

Computation and Language · Computer Science 2024-12-19 Naira Abdou Mohamed , Zakarya Erraji , Abdessalam Bahafid , Imade Benelallam

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approximately seven million people. This performance disparity…

Computation and Language · Computer Science 2026-01-06 Haq Nawaz Malik

Bangla, the seventh most widely spoken language worldwide with 300 million native speakers, faces digital under-representation due to limited resources and lack of annotated datasets. Stemming, a critical preprocessing step in language…

Computation and Language · Computer Science 2025-08-22 Abhijit Paul , Mashiat Amin Farin , Sharif Md. Abdullah , Ahmedul Kabir , Zarif Masud , Shebuti Rayana

We present BabyBabelLM, a multilingual collection of datasets modeling the language a person observes from birth until they acquire a native language. We curate developmentally plausible pretraining data aiming to cover the equivalent of…

The rapid growth of machine translation (MT) systems has necessitated comprehensive studies to meta-evaluate evaluation metrics being used, which enables a better selection of metrics that best reflect MT quality. Unfortunately, most of the…

Computation and Language · Computer Science 2023-07-04 Ananya B. Sai , Vignesh Nagarajan , Tanay Dixit , Raj Dabre , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

Accessing and gaining insight into the Rigveda poses a non-trivial challenge due to its extremely ancient Sanskrit language, poetic structure, and large volume of text. By using NLP techniques, this study identified topics and semantic…

Computation and Language · Computer Science 2025-03-25 Venkatesh Bollineni , Igor Crk , Eren Gultepe

Sentiment analysis is one of the most widely studied applications in NLP, but most work focuses on languages with large amounts of data. We introduce the first large-scale human-annotated Twitter sentiment dataset for the four most widely…

The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent…

Computation and Language · Computer Science 2025-07-15 Ce Li , Xiaofan Liu , Zhiyan Song , Ce Chi , Chen Zhao , Jingjing Yang , Zhendong Wang , Kexin Yang , Boshen Shi , Xing Wang , Chao Deng , Junlan Feng

We introduce SCRum-9, the largest multilingual Stance Classification dataset for Rumour analysis in 9 languages, containing 7,516 tweets from X. SCRum-9 goes beyond existing stance classification datasets by covering more languages, linking…

Computation and Language · Computer Science 2025-11-18 Yue Li , Jake Vasilakes , Zhixue Zhao , Carolina Scarton
‹ Prev 1 8 9 10 Next ›