English
Related papers

Related papers: Complexity counts: global and local perspectives o…

200 papers

Tamil language has an agglutinative, diglossic, alpha-syllabary structure which provides a significant combinatorial explosion of morphological forms all of which are effectively used in Tamil prose, poetry from antiquity to the modern age…

Computation and Language · Computer Science 2019-09-24 Muthiah Annamalai , T. Shrinivasan

The recent advances in deep-learning have led to the development of highly sophisticated systems with an unquenchable appetite for data. On the other hand, building good deep-learning models for low-resource languages remains a challenging…

Computation and Language · Computer Science 2024-02-20 Maithili Sabane , Onkar Litake , Aman Chadha

In this paper, we discuss an attempt to develop an automatic language identification system for 5 closely-related Indo-Aryan languages of India, Awadhi, Bhojpuri, Braj, Hindi and Magahi. We have compiled a comparable corpora of varying…

Computation and Language · Computer Science 2018-03-28 Ritesh Kumar , Bornini Lahiri , Deepak Alok , Atul Kr. Ojha , Mayank Jain , Abdul Basit , Yogesh Dawer

India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP applications for…

Computation and Language · Computer Science 2024-12-25 Rasika Ransing , Mohammed Amaan Dhamaskar , Ayush Rajpurohit , Amey Dhoke , Sanket Dalvi

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi and Magahi. This…

Computation and Language · Computer Science 2018-04-17 Priya Rani , Atul Kr. Ojha , Girish Nath Jha

The difficulties involved in spelling error detection and correction in a language have been investigated in this work through the conceptualization of SpellNet - the weighted network of words, where edges indicate orthographic proximity…

Physics and Society · Physics 2007-05-23 Monojit Choudhury , Markose Thomas , Animesh Mukherjee , Anupam Basu , Niloy Ganguly

Much recent work has shown how cross-linguistic variation is constrained by competing pressures from efficient communication. However, little attention has been paid to the role of the systematicity of forms (regularity), a key property of…

Computation and Language · Computer Science 2026-02-03 Ponrawee Prasertsom , Andrea Silvi , Jennifer Culbertson , Moa Johansson , Devdatt Dubhashi , Kenny Smith

While language competition models of diachronic language shift are increasingly sophisticated, drawing on sociolinguistic components like variable language prestige, distance from language centers and intermediate bilingual transitionary…

Computation and Language · Computer Science 2016-01-12 Rana D. Parshad , Vineeta Chand , Neha Sinha , Nitu Kumari

Large Language Models (LLMs) have emerged as powerful general-purpose reasoning systems, yet their development remains dominated by English-centric data, architectures, and optimization paradigms. This exclusionary design results in…

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their…

Transliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs. However, few training and evaluation sets are publicly available. We introduce Aksharantar,…

Computation and Language · Computer Science 2023-10-27 Yash Madhani , Sushane Parthan , Priyanka Bedekar , Gokul NC , Ruchi Khapra , Anoop Kunchukuttan , Pratyush Kumar , Mitesh M. Khapra

Speech systems are sensitive to accent variations. This is especially challenging in the Indian context, with an abundance of languages but a dearth of linguistic studies characterising pronunciation variations. The growing number of L2…

Computation and Language · Computer Science 2022-12-20 Shelly Jain , Priyanshi Pal , Anil Vuppala , Prasanta Ghosh , Chiranjeevi Yarra

Code-mixing, the blending of linguistic elements from distinct languages to form meaningful sentences, is common in multilingual settings, yielding hybrid languages like Hinglish and Minglish. Marathi, India's third most spoken language,…

Many NP-complete problems take integers as part of their input instances. These input integers are generally binarized, that is, provided in the form of the "binary" numeral representation, and the lengths of such binary forms are used as a…

Computational Complexity · Computer Science 2023-12-08 Tomoyuki Yamakami

The origin of the numerals that we inherited from the arabo-Islamic civilization remained one enigma. The hypothesis of the Indian origin remained, with controversies, without serious rival. It was the dominant hypothesis since more of one…

History and Overview · Mathematics 2007-07-24 Ahmed Boucenna

Recent methods in speech and language technology pretrain very LARGE models which are fine-tuned for specific tasks. However, the benefits of such LARGE models are often limited to a few resource rich languages of the world. In this work,…

Analysis of Indian English (IE) pronunciation variabilities are useful in building systems for Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) synthesis in the Indian context. Typically, these pronunciation variabilities have…

Computation and Language · Computer Science 2022-12-12 Priyanshi Pal , Shelly Jain , Anil Vuppala , Chiranjeevi Yarra , Prasanta Ghosh

Despite the extensive amount of scholarly work done on Indian mathematics in the last 200 years, the conditions under which it originated and evolved is still not clear. Often, one reads the ancient texts with the present concepts and…

History and Overview · Mathematics 2025-01-10 Jaidev Dasgupta

Indian languages are inflectional and agglutinative and typically follow clause-free word order. The structure of sentences across most major Indian languages are similar when their dependency parse trees are considered. While some…

Computation and Language · Computer Science 2025-01-08 N J Karthika , Adyasha Patra , Nagasai Saketh Naidu , Arnab Bhattacharya , Ganesh Ramakrishnan , Chaitali Dangarikar

Transformer-based models have revolutionized the field of natural language processing. To understand why they perform so well and to assess their reliability, several studies have focused on questions such as: Which linguistic properties…

Computation and Language · Computer Science 2025-11-04 Akhilesh Aravapalli , Mounika Marreddy , Radhika Mamidi , Manish Gupta , Subba Reddy Oota