English
Related papers

Related papers: Cross Script Hindi English NER Corpus from Wikiped…

200 papers

In this paper, we discuss the development of a multilingual annotated corpus of misogyny and aggression in Indian English, Hindi, and Indian Bangla as part of a project on studying and automatically identifying misogyny and communalism on…

Computation and Language · Computer Science 2020-03-18 Shiladitya Bhattacharya , Siddharth Singh , Ritesh Kumar , Akanksha Bansal , Akash Bhagat , Yogesh Dawer , Bornini Lahiri , Atul Kr. Ojha

In this paper, we conduct one of the very first studies for cross-corpora performance evaluation in the spoken language identification (LID) problem. Cross-corpora evaluation was not explored much in LID research, especially for the Indian…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-13 Spandan Dey , Goutam Saha , Md Sahidullah

Multilingual encoder-based language models are widely adopted for code-mixed analysis tasks, yet we know surprisingly little about how they represent code-mixed inputs internally - or whether those representations meaningfully connect to…

Computation and Language · Computer Science 2026-03-23 Debajyoti Mazumder , Divyansh Pathak , Prashant Kodali , Jasabanta Patro

Cross-lingual text classification aims at training a classifier on the source language and transferring the knowledge to target languages, which is very useful for low-resource languages. Recent multilingual pretrained language models…

Computation and Language · Computer Science 2021-05-25 Ziyun Wang , Xuan Liu , Peiji Yang , Shixing Liu , Zhisheng Wang

With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English and their native…

Computation and Language · Computer Science 2025-05-01 Manish Pandey , Nageshwar Prasad Yadav , Mokshada Adduru , Sawan Rai

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

Multilingual speakers often switch between languages to express themselves on social communication platforms. Sometimes, the original script of the language is preserved, while using a common script for all the languages is quite popular as…

Computation and Language · Computer Science 2018-03-19 Soumil Mandal , Dipankar Das

Toxic content is one of the most critical issues for social media platforms today. India alone had 518 million social media users in 2020. In order to provide a good experience to content creators and their audience, it is crucial to flag…

Computation and Language · Computer Science 2022-01-04 Manan Jhaveri , Devanshu Ramaiya , Harveen Singh Chadha

Codeswitching has become one of the most common occurrences across multilingual speakers of the world, especially in countries like India which encompasses around 23 official languages with the number of bilingual speakers being around 300…

Computation and Language · Computer Science 2022-01-03 Dhruval Jain , Arun D Prabhu , Shubham Vatsal , Gopi Ramena , Naresh Purre

Language Identification in textual documents is the process of automatically detecting the language contained in a document based on its content. The present Language Identification techniques presume that a document contains text in one of…

Computation and Language · Computer Science 2021-06-30 Mohd Zeeshan Ansari , Tanvir Ahmad , Noaima Bari

Named entity recognition (NER) is the process of recognising and classifying important information (entities) in text. Proper nouns, such as a person's name, an organization's name, or a location's name, are examples of entities. The NER is…

Computation and Language · Computer Science 2023-02-28 Onkar Litake , Maithili Sabane , Parth Patil , Aparna Ranade , Raviraj Joshi

One of the components of natural language processing that has received a lot of investigation recently is semantic textual similarity. In computational linguistics and natural language processing, assessing the semantic similarity of words,…

Computation and Language · Computer Science 2024-09-06 Mohammad Abdous , Poorya Piroozfar , Behrouz Minaei Bidgoli

The increasing diversity of languages used on the web introduces a new level of complexity to Information Retrieval (IR) systems. We can no longer assume that textual content is written in one language or even the same language family. In…

Computation and Language · Computer Science 2014-10-15 Rami Al-Rfou , Vivek Kulkarni , Bryan Perozzi , Steven Skiena

Code-switching is a phenomenon of mixing grammatical structures of two or more languages under varied social constraints. The code-switching data differ so radically from the benchmark corpora used in NLP community that the application of…

Computation and Language · Computer Science 2018-04-25 Irshad Ahmad Bhat , Riyaz Ahmad Bhat , Manish Shrivastava , Dipti Misra Sharma

The Marathi language is one of the prominent languages used in India. It is predominantly spoken by the people of Maharashtra. Over the past decade, the usage of language on online platforms has tremendously increased. However, research on…

Computation and Language · Computer Science 2022-01-12 Atharva Kulkarni , Meet Mandhane , Manali Likhitkar , Gayatri Kshirsagar , Jayashree Jagdale , Raviraj Joshi

Social Media platforms have been seeing adoption and growth in their usage over time. This growth has been further accelerated with the lockdown in the past year when people's interaction, conversation, and expression were limited…

Computation and Language · Computer Science 2022-04-06 Ekagra Ranjan , Naman Poddar

This article presents the application of the Universal Named Entity framework to generate automatically annotated corpora. By using a workflow that extracts Wikipedia data and meta-data and DBpedia information, we generated an English…

Computation and Language · Computer Science 2022-12-15 Diego Alves , Gaurish Thakkar , Marko Tadić

Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for processing such…

Computation and Language · Computer Science 2026-01-21 Dwip Dalal , Vivek Srivastava , Mayank Singh

Cross-lingual document classification aims at training a document classifier on resources in one language and transferring it to a different language without any additional resources. Several approaches have been proposed in the literature…

Computation and Language · Computer Science 2018-05-28 Holger Schwenk , Xian Li

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

Computation and Language · Computer Science 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha