English
Related papers

Related papers: Mukhyansh: A Headline Generation Dataset for Indic…

200 papers

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an…

Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those…

Computation and Language · Computer Science 2025-06-25 N J Karthika , Maharaj Brahma , Rohit Saluja , Ganesh Ramakrishnan , Maunendra Sankar Desarkar

Neural Machine Translation (NMT) models are typically trained on datasets with limited exposure to Scientific, Technical and Educational domains. Translation models thus, in general, struggle with tasks that involve scientific understanding…

Computation and Language · Computer Science 2024-12-13 Advait Joglekar , Srinivasan Umesh

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

Current advancements in Natural Language Processing (NLP) have largely favored resource-rich languages, leaving a significant gap in high-quality datasets for low-resource languages like Hindi. This scarcity is particularly evident in text…

Computation and Language · Computer Science 2026-01-06 Praveenkumar Katwe , RakeshChandra Balabantaray , Kaliprasad Vittala

Natural language generation (NLG) benchmarks provide an important avenue to measure progress and develop better NLG systems. Unfortunately, the lack of publicly available NLG benchmarks for low-resource languages poses a challenging barrier…

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…

Bengali is an underrepresented language in NLP research. However, it remains a challenge due to its unique linguistic structure and computational constraints. In this work, we systematically investigate the challenges that hinder Bengali…

Computation and Language · Computer Science 2025-08-01 Shimanto Bhowmik , Tawsif Tashwar Dipto , Md Sazzad Islam , Sheryl Hsu , Tahsin Reasat

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low-resource language settings where the…

The development of robust language models for low-resource languages is frequently bottlenecked by the scarcity of high-quality, coherent, and domain-appropriate training corpora. In this paper, we introduce the Multilingual TinyStories…

Computation and Language · Computer Science 2026-03-17 Deepon Halder , Angira Mukherjee

The performance of a text-to-speech (TTS) synthesis model depends on various factors, of which the quality of the training data is of utmost importance. Millions of data are collected around the globe for various languages, but resources…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-21 Sujitha Sathiyamoorthy , N Mohana , Anusha Prakash , Hema A Murthy

In the era dominated by information overload and its facilitation with Large Language Models (LLMs), the prevalence of misinformation poses a significant threat to public discourse and societal well-being. A critical concern at present…

Computation and Language · Computer Science 2024-11-05 Cem Üyük , Danica Rovó , Shaghayegh Kolli , Rabia Varol , Georg Groh , Daryna Dementieva

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their…

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly important: despite Hindi being the third most spoken language…

Computation and Language · Computer Science 2026-01-16 Tarun Sharma , Manikandan Ravikiran , Sourava Kumar Behera , Pramit Bhattacharya , Arnab Bhattacharya , Rohit Saluja

Text generation is a highly active area of research in the computational linguistic community. The evaluation of the generated text is a challenging task and multiple theories and metrics have been proposed over the years. Unfortunately,…

Computation and Language · Computer Science 2021-07-09 Vivek Srivastava , Mayank Singh

The multi-sentential long sequence textual data unfolds several interesting research directions pertaining to natural language processing and generation. Though we observe several high-quality long-sequence datasets for English and other…

Computation and Language · Computer Science 2023-02-24 Rahul Gupta , Vivek Srivastava , Mayank Singh

A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating…

Computation and Language · Computer Science 2022-12-16 Tahir Javed , Kaushal Santosh Bhogale , Abhigyan Raman , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

The advancements in the Large Language Model (LLM) have helped in solving several problems related to language processing. Most of the researches have focused on the English language only, because of its popularity and abundance on the…

Computation and Language · Computer Science 2024-12-31 Sanjay Chouhan , Shubha Brata Nath , Aparajita Dutta

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to…

In this paper, we report the results of the TeamNRC's participation in the BHASHA-Task 1 Grammatical Error Correction shared task https://github.com/BHASHA-Workshop/IndicGEC2025/ for 5 Indian languages. Our approach, focusing on…

Computation and Language · Computer Science 2025-11-20 Sowmya Vajjala