English
Related papers

Related papers: Bhasha-Abhijnaanam: Native-script and romanized La…

200 papers

This research investigates biases in text-to-image (TTI) models for the Indic languages widely spoken across India. It evaluates and compares the generative performance and cultural relevance of leading TTI models in these languages against…

Computation and Language · Computer Science 2024-08-02 Surbhi Mittal , Arnav Sudan , Mayank Vatsa , Richa Singh , Tamar Glaser , Tal Hassner

Large Language Models (LLMs) are increasingly deployed in high-stakes clinical applications in India. Speakers of Indian languages frequently communicate using romanized text rather than native scripts, yet existing research rarely…

Computation and Language · Computer Science 2026-04-01 Manurag Khullar , Utkarsh Desai , Poorva Malviya , Aman Dalmia , Zheyuan Ryan Shi

In this paper, we discuss an attempt to develop an automatic language identification system for 5 closely-related Indo-Aryan languages of India, Awadhi, Bhojpuri, Braj, Hindi and Magahi. We have compiled a comparable corpora of varying…

Computation and Language · Computer Science 2018-03-28 Ritesh Kumar , Bornini Lahiri , Deepak Alok , Atul Kr. Ojha , Mayank Jain , Abdul Basit , Yogesh Dawer

This review paper provides a comprehensive overview of large language model (LLM) research directions within Indic languages. Indic languages are those spoken in the Indian subcontinent, including India, Pakistan, Bangladesh, Sri Lanka,…

Computation and Language · Computer Science 2024-06-17 Sankalp KJ , Vinija Jain , Sreyoshi Bhaduri , Tamoghna Roy , Aman Chadha

Large language models (LLMs) demonstrated transformative capabilities in many applications that require automatically generating responses based on human instruction. However, the major challenge for building LLMs, particularly in Indic…

Computation and Language · Computer Science 2024-07-16 Shantipriya Parida , Shakshi Panwar , Kusum Lata , Sanskruti Mishra , Sambit Sekhar

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their…

Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrained generation of…

Computation and Language · Computer Science 2026-02-26 Thanmay Jayakumar , Mohammed Safi Ur Rahman Khan , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan

We present Vakyansh, an end to end toolkit for Speech Recognition in Indic languages. India is home to almost 121 languages and around 125 crore speakers. Yet most of the languages are low resource in terms of data and pretrained models.…

Computation and Language · Computer Science 2022-06-16 Harveen Singh Chadha , Anirudh Gupta , Priyanshi Shah , Neeraj Chhimwal , Ankur Dhuriya , Rishabh Gaur , Vivek Raghavan

The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-agnostic, limiting their applicability to India-centric…

Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a LID model which achieves a…

Computation and Language · Computer Science 2023-08-31 Laurie Burchell , Alexandra Birch , Nikolay Bogoychev , Kenneth Heafield

Language Identification (LID) is a core task in multilingual NLP, yet current systems often overfit to clean, monolingual data. This work introduces DIVERS-BENCH, a comprehensive evaluation of state-of-the-art LID models across diverse…

Computation and Language · Computer Science 2025-09-23 Jessica Ojo , Zina Kamel , David Ifeoluwa Adelani

Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Anik De , Abhirama Subramanyam Penamakuri , Rajeev Yadav , Aditya Rathore , Harshiv Shah , Devesh Sharma , Sagar Agarwal , Pravin Kumar , Anand Mishra

Script identification and text recognition are some of the major domains in the application of Artificial Intelligence. In this era of digitalization, the use of digital note-taking has become a common practice. Still, conventional methods…

Artificial Intelligence · Computer Science 2023-08-14 Sidhantha Poddar , Rohan Gupta

Evaluating Large Language Models (LLMs) in low-resource and linguistically diverse languages remains a significant challenge in NLP, particularly for languages using non-Latin scripts like those spoken in India. Existing benchmarks…

Computation and Language · Computer Science 2025-02-05 Sshubam Verma , Mohammed Safi Ur Rahman Khan , Vishwajeet Kumar , Rudra Murthy , Jaydeep Sen

India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the…

Computation and Language · Computer Science 2025-02-24 Utkarsh P

Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems extend this coverage under certain circumstances, but for most…

Computation and Language · Computer Science 2026-02-10 Rasul Dent , Pedro Ortiz Suarez , Thibault Clérice , Benoît Sagot

As large language models (LLMs) see increasing adoption across the globe, it is imperative for LLMs to be representative of the linguistic diversity of the world. India is a linguistically diverse country of 1.4 Billion people. To…

Computation and Language · Computer Science 2024-08-09 Harman Singh , Nitish Gupta , Shikhar Bharadwaj , Dinesh Tewari , Partha Talukdar

Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid…

Computation and Language · Computer Science 2026-02-24 Mariia Fedorova , Nikolay Arefyev , Maja Buljan , Jindřich Helcl , Stephan Oepen , Egil Rønningstad , Yves Scherrer

Language Identification in textual documents is the process of automatically detecting the language contained in a document based on its content. The present Language Identification techniques presume that a document contains text in one of…

Computation and Language · Computer Science 2021-06-30 Mohd Zeeshan Ansari , Tanvir Ahmad , Noaima Bari

Multimodal research has predominantly focused on single-image reasoning, with limited exploration of multi-image scenarios. Recent models have sought to enhance multi-image understanding through large-scale pretraining on interleaved…

Computation and Language · Computer Science 2026-03-26 Shaharukh Khan , Ali Faraz , Abhinav Ravi , Mohd Nauman , Mohd Sarfraz , Akshat Patidar , Raja Kolla , Chandra Khatri , Shubham Agarwal