English
Related papers

Related papers: Mukhyansh: A Headline Generation Dataset for Indic…

200 papers

The advancement in technology and accessibility of internet to each individual is revolutionizing the real time information. The liberty to express your thoughts without passing through any credibility check is leading to dissemination of…

Computation and Language · Computer Science 2021-02-24 Shivangi Singhal , Rajiv Ratn Shah , Ponnurangam Kumaraguru

One of the major problems writers and poets face is the writer's block. It is a condition in which an author loses the ability to produce new work or experiences a creative slowdown. The problem is more difficult in the context of poetry…

Computation and Language · Computer Science 2021-08-02 Shakeeb A. M. Mukhtar , Pushkar S. Joglekar

Assessing the capabilities and limitations of large language models (LLMs) has garnered significant interest, yet the evaluation of multiple models in real-world scenarios remains rare. Multilingual evaluation often relies on translated…

Computation and Language · Computer Science 2025-12-02 Varun Gumma , Ananditha Raghunath , Mohit Jain , Sunayana Sitaram

The development of robust transliteration techniques to enhance the effectiveness of transforming Romanized scripts into native scripts is crucial for Natural Language Processing tasks, including sentiment analysis, speech recognition,…

Computation and Language · Computer Science 2025-12-01 Kanchon Gharami , Quazi Sarwar Muhtaseem , Deepti Gupta , Lavanya Elluri , Shafika Showkat Moni

Transliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs. However, few training and evaluation sets are publicly available. We introduce Aksharantar,…

Computation and Language · Computer Science 2023-10-27 Yash Madhani , Sushane Parthan , Priyanka Bedekar , Gokul NC , Ruchi Khapra , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

The recent surge of complex attention-based deep learning architectures has led to extraordinary results in various downstream NLP tasks in the English language. However, such research for resource-constrained and morphologically rich…

Computation and Language · Computer Science 2021-02-23 Atharva Kulkarni , Amey Hengle , Rutuja Udyawar

Large Language Models (LLMs) perform well on unseen tasks in English, but their abilities in non English languages are less explored due to limited benchmarks and training data. To bridge this gap, we introduce the Indic QA Benchmark, a…

Machine Learning · Computer Science 2025-02-25 Abhishek Kumar Singh , Vishwajeet kumar , Rudra Murthy , Jaydeep Sen , Ashish Mittal , Ganesh Ramakrishnan

Current Text-to-Speech models pose a multilingual challenge, where most of the models traditionally focus on English and European languages, thereby hurting the potential to provide access to information to many more people. To address this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Jaskaran Singh , Amartya Roy Chowdhury , Raghav Prabhakar , Varshul C. W

Developing benchmark datasets for low-resource languages poses significant challenges, primarily due to the limited availability of native linguistic experts and the substantial time and cost involved in annotation. Given these challenges,…

Computation and Language · Computer Science 2025-10-28 Rahul Ranjan , Mahendra Kumar Gurve , Anuj , Nitin , Yamuna Prasad

Current evaluation benchmarks for question answering (QA) in Indic languages often rely on machine translation of existing English datasets. This approach suffers from bias and inaccuracies inherent in machine translation, leading to…

Computation and Language · Computer Science 2024-05-01 Vaishak Narayanan , Prabin Raj KP , Saifudheen Nouphal

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

Treebanks are important linguistic resources, which are structured and annotated corpora with rich linguistic annotations. These resources are used in Natural Language Processing (NLP) applications, supporting linguistic analyses, and are…

Computation and Language · Computer Science 2024-09-24 Kengatharaiyer Sarveswaran

Neural models have recently been used in text summarization including headline generation. The model can be trained using a set of document-headline pairs. However, the model does not explicitly consider topical similarities and differences…

Computation and Language · Computer Science 2016-08-23 Lei Xu , Ziyun Wang , Ayana , Zhiyuan Liu , Maosong Sun

Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in…

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Model (LLM) development. Existing tokenizers often exhibit high…

Cognates are present in multiple variants of the same text across different languages (e.g., "hund" in German and "hound" in English language mean "dog"). They pose a challenge to various Natural Language Processing (NLP) applications such…

Computation and Language · Computer Science 2021-12-20 Diptesh Kanojia , Pushpak Bhattacharyya , Malhar Kulkarni , Gholamreza Haffari

Despite the major advances in NLP, significant disparities in NLP system performance across languages still exist. Arguably, these are due to uneven resource allocation and sub-optimal incentives to work on less resourced languages. To…

Handwritten Text Recognition (HTR) is a well-established research area. In contrast, Handwritten Text Generation (HTG) is an emerging field with significant potential. This task is challenging due to the variation in individual handwriting…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Md. Rakibul Islam , Md. Kamrozzaman Bhuiyan , Safwan Muntasir , Arifur Rahman Jawad , Most. Sharmin Sultana Samu

Named Entity Recognition (NER) is a useful component in Natural Language Processing (NLP) applications. It is used in various tasks such as Machine Translation, Summarization, Information Retrieval, and Question-Answering systems. The…

The ILSUM shared task focuses on text summarization for two major Indian languages- Hindi and Gujarati, along with English. In this task, we experiment with various pretrained sequence-to-sequence models to find out the best model for each…

Computation and Language · Computer Science 2023-03-28 Ashok Urlana , Sahil Manoj Bhatt , Nirmal Surange , Manish Shrivastava