English
Related papers

Related papers: COSMMIC: Comment-Sensitive Multimodal Multilingual…

200 papers

Tokenization plays a pivotal role in multilingual NLP. However, existing tokenizers are often skewed towards high-resource languages, limiting their effectiveness for linguistically diverse and morphologically rich languages such as those…

Computation and Language · Computer Science 2025-06-25 N J Karthika , Maharaj Brahma , Rohit Saluja , Ganesh Ramakrishnan , Maunendra Sankar Desarkar

The rapid growth of machine translation (MT) systems has necessitated comprehensive studies to meta-evaluate evaluation metrics being used, which enables a better selection of metrics that best reflect MT quality. Unfortunately, most of the…

Computation and Language · Computer Science 2023-07-04 Ananya B. Sai , Vignesh Nagarajan , Tanay Dixit , Raj Dabre , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

Large Language Models (LLMs) have made significant progress in incorporating Indic languages within multilingual models. However, it is crucial to quantitatively assess whether these languages perform comparably to globally dominant ones,…

Computation and Language · Computer Science 2024-10-31 Pritika Rohera , Chaitrali Ginimav , Akanksha Salunke , Gayatri Sawant , Raviraj Joshi

We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families).…

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems and reinforced biases…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Umair Hassan

Multi-modal information retrieval (MMIR) is a rapidly evolving field, where significant progress, particularly in image-text pairing, has been made through advanced representation learning and cross-modality alignment research. However,…

In the context of pretraining of Large Language Models (LLMs), synthetic data has emerged as an alternative for generating high-quality pretraining data at scale. This is particularly beneficial in low-resource language settings where the…

While model architecture and training objectives are well-studied, tokenization, particularly in multilingual contexts, remains a relatively neglected aspect of Large Language Model (LLM) development. Existing tokenizers often exhibit high…

In the healthcare domain, summarizing medical questions posed by patients is critical for improving doctor-patient interactions and medical decision-making. Although medical data has grown in complexity and quantity, the current body of…

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic diversity, and…

Computation and Language · Computer Science 2025-01-29 Sankalp KJ , Ashutosh Kumar , Laxmaan Balaji , Nikunj Kotecha , Vinija Jain , Aman Chadha , Sreyoshi Bhaduri

Summarizing Indian legal court judgments is a complex task not only due to the intricate language and unstructured nature of the legal texts, but also since a large section of the Indian population does not understand the complex English in…

Computation and Language · Computer Science 2026-02-10 Debtanu Datta , Rajdeep Mukherjee , Adrijit Goswami , Saptarshi Ghosh

We present a new publicly available dataset with the goal of advancing multi-modality learning by offering vision and language data within the same context. This is achieved by obtaining data from a social media website with posts…

Computation and Language · Computer Science 2020-06-16 Bofan Xue , David Chan , John Canny

Most existing medical dialogue systems operate in a single-turn question--answering paradigm or rely on template-based datasets, limiting conversational realism and multilingual applicability. We introduce IndicMedDialog, a parallel…

Computation and Language · Computer Science 2026-05-14 Shubham Kumar Nigam , Suparnojit Sarkar , Piyush Patel

We introduce DRISHTIKON, a first-of-its-kind multimodal and multilingual benchmark centered exclusively on Indian culture, designed to evaluate the cultural understanding of generative AI systems. Unlike existing benchmarks with a generic…

Computation and Language · Computer Science 2025-09-24 Arijit Maji , Raghvendra Kumar , Akash Ghosh , Anushka , Nemil Shah , Abhilekh Borah , Vanshika Shah , Nishant Mishra , Sriparna Saha

The goal of news image captioning is to generate captions by integrating news article content with corresponding images, highlighting the relationship between textual context and visual elements. The majority of research on news image…

Computation and Language · Computer Science 2026-03-12 Yuji Chen , Alistair Plum , Hansi Hettiarachchi , Diptesh Kanojia , Saroj Basnet , Marcos Zampieri , Tharindu Ranasinghe

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

Computation and Language · Computer Science 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

Automated image captioning using the content from the image is very appealing when done by harnessing the capability of computer vision and natural language processing. Extensive research has been done in the field with a major focus on the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Wasim Akram Khan , Anil Kumar Vuppala

In this paper, we report the results of the TeamNRC's participation in the BHASHA-Task 1 Grammatical Error Correction shared task https://github.com/BHASHA-Workshop/IndicGEC2025/ for 5 Indian languages. Our approach, focusing on…

Computation and Language · Computer Science 2025-11-20 Sowmya Vajjala

Understanding the sentiment of a comment from a video or an image is an essential task in many applications. Sentiment analysis of a text can be useful for various decision-making processes. One such application is to analyse the popular…

Computation and Language · Computer Science 2021-06-09 Bharathi Raja Chakravarthi , Vigneshwaran Muralidaran , Ruba Priyadharshini , John P. McCrae

Small Language Models (SLMs) offer efficient alternatives to LLMs for specific domains. The 2023 TinyStories study developed an English dataset that allows SLMs with 1 to 10 million parameters to produce coherent outputs. Our research…