English
Related papers

Related papers: My Boli: Code-mixed Marathi-English Corpora, Pretr…

200 papers

The rapid expansion in the usage of social media networking sites leads to a huge amount of unprocessed user generated data which can be used for text mining. Author profiling is the problem of automatically determining profiling aspects…

Computation and Language · Computer Science 2018-06-15 Ankush Khandelwal , Sahil Swami , Syed Sarfaraz Akhtar , Manish Shrivastava

Digital platforms have an ever-expanding user base, and act as a hub for communication, business, and connectivity. However, this has also allowed for the spread of hate speech and misogyny. Artificial intelligence models have emerged as an…

Artificial Intelligence · Computer Science 2026-01-14 Sargam Yadav , Abhishek Kaushik , Kevin Mc Daid

This paper describes the system submitted to Dravidian-Codemix-HASOC2021: Hate Speech and Offensive Language Identification in Dravidian Languages (Tamil-English and Malayalam-English). This task aims to identify offensive content in…

Computation and Language · Computer Science 2021-12-08 Sean Benhur , Kanchana Sivanraju

In recent times, we have seen an increased use of text chat for communication on social networks and smartphones. This particularly involves the use of Hindi-English code-mixed text which contains words which are not recognized in English…

Computation and Language · Computer Science 2021-11-16 Divyansh Singh

Social media has effectively become the prime hub of communication and digital marketing. As these platforms enable the free manifestation of thoughts and facts in text, images and video, there is an extensive need to screen them to protect…

The ubiquity of offensive content on social media is a growing cause for concern among companies and government organizations. Recently, transformer-based models such as BERT, XLNET, and XLM-R have achieved state-of-the-art performance in…

Computation and Language · Computer Science 2023-12-07 Tharindu Ranasinghe , Marcos Zampieri

Sentiment Analysis for Indian Languages (SAIL)-Code Mixed tools contest aimed at identifying the sentence level sentiment polarity of the code-mixed dataset of Indian languages pairs (Hi-En, Ben-Hi-En). Hi-En dataset is henceforth referred…

Computation and Language · Computer Science 2018-08-13 Pruthwik Mishra , Prathyusha Danda , Pranav Dhakras

Due to the wide adoption of social media platforms like Facebook, Twitter, etc., there is an emerging need of detecting online posts that can go against the community acceptance standards. The hostility detection task has been well explored…

Computation and Language · Computer Science 2021-01-14 Arkadipta De , Venkatesh E , Kaushal Kumar Maurya , Maunendra Sankar Desarkar

The zero-shot cross-lingual ability of models pretrained on multilingual and even monolingual corpora has spurred many hypotheses to explain this intriguing empirical result. However, due to the costs of pretraining, most research uses…

Computation and Language · Computer Science 2022-09-28 Hugo Abonizio , Leandro Rodrigues de Souza , Roberto Lotufo , Rodrigo Nogueira

The problems of online hate speech and cyberbullying have significantly worsened since the increase in popularity of social media platforms such as YouTube and Twitter (X). Natural Language Processing (NLP) techniques have proven to provide…

Computation and Language · Computer Science 2024-03-18 Sargam Yadav , Abhishek Kaushik , Kevin McDaid

The present paper introduces new sentiment data, MaCMS, for Magahi-Hindi-English (MHE) code-mixed language, where Magahi is a less-resourced minority language. This dataset is the first Magahi-Hindi-English code-mixed dataset for sentiment…

Computation and Language · Computer Science 2024-03-25 Priya Rani , Gaurav Negi , Theodorus Fransen , John P. McCrae

Embedding models are pivotal in industrial information retrieval systems like search and advertising. However, existing pretrained models often exhibit fixed architectures and embedding dimensionalities, posing significant challenges when…

Computation and Language · Computer Science 2026-05-20 Yaoxiang Wang , Simiao Zuo , Qingguo Hu , Yucheng Ding , Yeyun Gong , Jian Jiao , Jinsong Su

Natural Language Understanding (NLU) for low-resource languages remains a major challenge in NLP due to the scarcity of high-quality data and language-specific models. Maithili, despite being spoken by millions, lacks adequate computational…

Computation and Language · Computer Science 2026-02-03 Sumit Yadav , Raju Kumar Yadav , Utsav Maskey , Gautam Siddharth Kashyap , Ganesh Gautam , Usman Naseem

Multilinguality is a core capability for modern foundation models, yet training high-quality multilingual models remains challenging due to uneven data availability across languages. A further challenge is the performance interference that…

Offensive Language detection in social media platforms has been an active field of research over the past years. In non-native English spoken countries, social media users mostly use a code-mixed form of text in their posts/comments. This…

Computation and Language · Computer Science 2022-12-13 Charangan Vasantharajan , Uthayasanker Thayasivam

Online social media platforms are central to everyday communication and information seeking. While these platforms serve positive purposes, they also provide fertile ground for the spread of hate speech, offensive language, and bullying…

Computation and Language · Computer Science 2025-10-03 Md Arid Hasan , Firoj Alam , Md Fahad Hossain , Usman Naseem , Syed Ishtiaque Ahmed

We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remains underrepresented…

Computation and Language · Computer Science 2026-05-27 Param Thakkar , Anushka Yadav , Michael Tiemann , Abhi Mehta , Akshita Bhasin , Shrinivas Khedkar

Hate speech detection across contemporary social media presents unique challenges due to linguistic diversity and the informal nature of online discourse. These challenges are further amplified in settings involving code-mixing,…

Computation and Language · Computer Science 2025-06-17 Daman Deep Singh , Ramanuj Bhattacharjee , Abhijnan Chakraborty

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature. To pretrain BanglaBERT, we collect 27.5 GB of Bangla…

Computation and Language · Computer Science 2022-05-11 Abhik Bhattacharjee , Tahmid Hasan , Wasi Uddin Ahmad , Kazi Samin , Md Saiful Islam , Anindya Iqbal , M. Sohel Rahman , Rifat Shahriyar

Code-mixing is the phenomenon of using more than one language in a sentence. It is a very frequently observed pattern of communication on social media platforms. Flexibility to use multiple languages in one text message might help to…

Computation and Language · Computer Science 2020-04-21 Vivek Srivastava , Mayank Singh