English
Related papers

Related papers: My Boli: Code-mixed Marathi-English Corpora, Pretr…

200 papers

The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fundamental text classification task across several languages…

Computation and Language · Computer Science 2024-12-11 Sadia Alam , Md Farhan Ishmam , Navid Hasin Alvee , Md Shahnewaz Siddique , Md Azam Hossain , Abu Raihan Mostofa Kamal

Hate speech detection in low-resource languages like Telugu is a growing challenge in NLP. This study investigates transformer-based models, including TeluguHateBERT, HateBERT, DeBERTa, Muril, IndicBERT, Roberta, and Hindi-Abusive-MuRIL,…

Computation and Language · Computer Science 2025-02-18 Santhosh Kakarla , Gautama Shastry Bulusu Venkata

As the interaction over the web has increased, incidents of aggression and related events like trolling, cyberbullying, flaming, hate speech, etc. too have increased manifold across the globe. While most of these behaviour like bullying or…

Computation and Language · Computer Science 2018-03-28 Ritesh Kumar , Aishwarya N. Reganti , Akshit Bhatia , Tushar Maheshwari

Multilingual Machine Comprehension (MMC) is a Question-Answering (QA) sub-task that involves quoting the answer for a question from a given snippet, where the question and the snippet can be in different languages. Recently released…

Computation and Language · Computer Science 2020-06-03 Somil Gupta , Nilesh Khade

Due to the sheer volume of online hate, the AI and NLP communities have started building models to detect such hateful content. Recently, multilingual hate is a major emerging challenge for automated detection where code-mixing or more than…

Computation and Language · Computer Science 2022-05-12 Mithun Das , Punyajoy Saha , Binny Mathew , Animesh Mukherjee

The multilingual Sentence-BERT (SBERT) models map different languages to common representation space and are useful for cross-language similarity and mining tasks. We propose a simple yet effective approach to convert vanilla multilingual…

Computation and Language · Computer Science 2023-04-25 Samruddhi Deode , Janhavi Gadre , Aditi Kajale , Ananya Joshi , Raviraj Joshi

Numerous machine learning (ML) and deep learning (DL)-based approaches have been proposed to utilize textual data from social media for anti-social behavior analysis like cyberbullying, fake news detection, and identification of hate speech…

Computation and Language · Computer Science 2022-12-22 Md. Rezaul Karim , Sumon Kanti Dey , Tanhim Islam , Md. Shajalal , Bharathi Raja Chakravarthi

In the recent past, social media platforms have helped people in connecting and communicating to a wider audience. But this has also led to a drastic increase in cyberbullying. It is essential to detect and curb hate speech to keep the…

Computation and Language · Computer Science 2021-11-02 Ravindra Nayak , Raviraj Joshi

The multi-sentential long sequence textual data unfolds several interesting research directions pertaining to natural language processing and generation. Though we observe several high-quality long-sequence datasets for English and other…

Computation and Language · Computer Science 2023-02-24 Rahul Gupta , Vivek Srivastava , Mayank Singh

This paper describes our multiclass classification system developed as part of the LTEDI@RANLP-2023 shared task. We used a BERT-based language model to detect homophobic and transphobic content in social media comments across five language…

Computation and Language · Computer Science 2023-08-28 Sidney G. -J. Wong , Matthew Durward , Benjamin Adams , Jonathan Dunn

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several works have been conducted on building datasets and performing downstream NLP tasks on code-mixed data. Although it is not…

Computation and Language · Computer Science 2023-11-28 Dhiman Goswami , Md Nishat Raihan , Antara Mahmud , Antonios Anastasopoulos , Marcos Zampieri

This paper reports an increment to the state-of-the-art in hate speech detection for English-Hindi code-mixed tweets. We compare three typical deep learning models using domain-specific embeddings. On experimenting with a benchmark dataset…

Computation and Language · Computer Science 2018-11-14 Satyajit Kamble , Aditya Joshi

In this paper, we introduce HateBERT, a re-trained BERT model for abusive language detection in English. The model was trained on RAL-E, a large-scale dataset of Reddit comments in English from communities banned for being offensive,…

Computation and Language · Computer Science 2021-02-05 Tommaso Caselli , Valerio Basile , Jelena Mitrović , Michael Granitzer

In the last few years, emotion detection in social-media text has become a popular problem due to its wide ranging application in better understanding the consumers, in psychology, in aiding human interaction with computers, designing smart…

Computation and Language · Computer Science 2021-03-02 Anshul Wadhawan , Akshita Aggarwal

Pre-training large neural language models, such as BERT, has led to impressive gains on many natural language processing (NLP) tasks. Although this method has proven to be effective for many domains, it might not always provide desirable…

Computation and Language · Computer Science 2022-12-13 Omkar Gokhale , Aditya Kane , Shantanu Patankar , Tanmay Chavan , Raviraj Joshi

Code-mixing, the blending of linguistic elements from distinct languages to form meaningful sentences, is common in multilingual settings, yielding hybrid languages like Hinglish and Minglish. Marathi, India's third most spoken language,…

Exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices but also enables people to express anti-social behaviour like online harassment,…

Computation and Language · Computer Science 2020-04-21 Md. Rezaul Karim , Bharathi Raja Chakravarthi , John P. McCrae , Michael Cochez

Sentiment analysis is the most basic NLP task to determine the polarity of text data. There has been a significant amount of work in the area of multilingual text as well. Still hate and offensive speech detection faces a challenge due to…

Computation and Language · Computer Science 2021-11-02 Abhishek Velankar , Hrushikesh Patil , Amol Gore , Shubham Salunke , Raviraj Joshi

Social networking platforms provide a conduit to disseminate our ideas, views and thoughts and proliferate information. This has led to the amalgamation of English with natively spoken languages. Prevalence of Hindi-English code-mixed data…

Computation and Language · Computer Science 2021-05-12 Ananya Srivastava , Mohammed Hasan , Bhargav Yagnik , Rahee Walambe , Ketan Kotecha

In our increasingly interconnected digital world, social media platforms have emerged as powerful channels for the dissemination of hate speech and offensive content. This work delves into the domain of hate speech detection, placing…

Computation and Language · Computer Science 2023-10-04 Ananya Joshi , Raviraj Joshi