English
Related papers

Related papers: L3Cube-MahaSent-MD: A Multi-domain Marathi Sentime…

200 papers

Sentiment analysis is one of the most fundamental tasks in Natural Language Processing. Popular languages like English, Arabic, Russian, Mandarin, and also Indian languages such as Hindi, Bengali, Tamil have seen a significant amount of…

Computation and Language · Computer Science 2021-06-29 Atharva Kulkarni , Meet Mandhane , Manali Likhitkar , Gayatri Kshirsagar , Raviraj Joshi

Social media platforms are used by a large number of people prominently to express their thoughts and opinions. However, these platforms have contributed to a substantial amount of hateful and abusive content as well. Therefore, it is…

Computation and Language · Computer Science 2022-05-24 Abhishek Velankar , Hrushikesh Patil , Amol Gore , Shubham Salunke , Raviraj Joshi

Emotion recognition in low-resource languages like Marathi remains challenging due to limited annotated data. We present L3Cube-MahaEmotions, a high-quality Marathi emotion recognition dataset with 11 fine-grained emotion labels. The…

Computation and Language · Computer Science 2025-09-01 Nidhi Kowtal , Raviraj Joshi

Despite being the third most popular language in India, the Marathi language lacks useful NLP resources. Moreover, popular NLP libraries do not have support for the Marathi language. With L3Cube-MahaNLP, we aim to build resources and a…

Computation and Language · Computer Science 2022-06-01 Raviraj Joshi

The availability of text or topic classification datasets in the low-resource Marathi language is limited, typically consisting of fewer than 4 target labels, with some achieving nearly perfect accuracy. In this work, we introduce…

Computation and Language · Computer Science 2024-04-30 Saloni Mittal , Vidula Magdum , Omkar Dhekane , Sharayu Hiwarkhedkar , Raviraj Joshi

This work introduces the L3Cube-MahaSocialNER dataset, the first and largest social media dataset specifically designed for Named Entity Recognition (NER) in the Marathi language. The dataset comprises 18,000 manually labeled sentences…

Computation and Language · Computer Science 2024-02-15 Harsh Chaudhari , Anuja Patil , Dhanashree Lavekar , Pranav Khairnar , Raviraj Joshi

We present L3Cube-MahaCorpus a Marathi monolingual data set scraped from different internet sources. We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We further present, MahaBERT, MahaAlBERT, and…

Computation and Language · Computer Science 2022-03-15 Raviraj Joshi

Sentiment analysis plays a crucial role in understanding the sentiment expressed in text data. While sentiment analysis research has been extensively conducted in English and other Western languages, there exists a significant gap in…

Computation and Language · Computer Science 2023-10-03 Aabha Pingle , Aditya Vyawahare , Isha Joshi , Rahul Tangsali , Geetanjali Kale , Raviraj Joshi

The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fundamental text classification task across several languages…

Computation and Language · Computer Science 2024-12-11 Sadia Alam , Md Farhan Ishmam , Navid Hasin Alvee , Md Shahnewaz Siddique , Md Azam Hossain , Abu Raihan Mostofa Kamal

We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along with MahaSBERT-STS-v2, a fine-tuned Sentence-BERT model optimized for regression-based similarity scoring. The MahaSTS dataset consists of…

Computation and Language · Computer Science 2025-09-01 Aishwarya Mirashi , Ananya Joshi , Raviraj Joshi

We present the MahaSUM dataset, a large-scale collection of diverse news articles in Marathi, designed to facilitate the training and evaluation of models for abstractive summarization tasks in Indic languages. The dataset, containing 25k…

Computation and Language · Computer Science 2024-10-15 Pranita Deshmukh , Nikita Kulkarni , Sanhita Kulkarni , Kareena Manghani , Raviraj Joshi

The research on code-mixed data is limited due to the unavailability of dedicated code-mixed datasets and pre-trained language models. In this work, we focus on the low-resource Indian language Marathi which lacks any prior work in…

Computation and Language · Computer Science 2023-07-21 Tanmay Chavan , Omkar Gokhale , Aditya Kane , Shantanu Patankar , Raviraj Joshi

The analysis of consumer sentiment, as expressed through reviews, can provide a wealth of insight regarding the quality of a product. While the study of sentiment analysis has been widely explored in many popular languages, relatively less…

Computation and Language · Computer Science 2023-06-09 Mohsinul Kabir , Obayed Bin Mahfuz , Syed Rifat Raiyan , Hasan Mahmud , Md Kamrul Hasan

Multi-aspect sentiment analysis of Bangla e-commerce reviews remains challenging due to limited annotated datasets, morphological complexity, code-mixing phenomena, and domain shift issues, affecting 300 million Bangla-speaking users.…

Machine Learning · Computer Science 2025-12-01 Ariful Islam , Md Rifat Hossen , Tanvir Mahmud

Human communication is inherently multimodal and asynchronous. Analyzing human emotions and sentiment is an emerging field of artificial intelligence. We are witnessing an increasing amount of multimodal content in local languages on social…

Paraphrases are a vital tool to assist language understanding tasks such as question answering, style transfer, semantic parsing, and data augmentation tasks. Indic languages are complex in natural language processing (NLP) due to their…

Computation and Language · Computer Science 2025-08-26 Suramya Jadhav , Abhay Shanbhag , Amogh Thakurdesai , Ridhima Sinare , Ananya Joshi , Raviraj Joshi

Semantic evaluation in low-resource languages remains a major challenge in NLP. While sentence transformers have shown strong performance in high-resource settings, their effectiveness in Indic languages is underexplored due to a lack of…

Computation and Language · Computer Science 2025-09-03 Nishant Tanksale , Tanmay Kokate , Darshan Gohad , Sarvadnyaa Barate , Raviraj Joshi

Sentiment analysis for the Bengali language has attracted increasing research interest in recent years. However, progress remains constrained by the scarcity of large-scale and diverse annotated datasets. Although several Bengali sentiment…

Computation and Language · Computer Science 2026-01-29 Akif Islam , Sujan Kumar Roy , Md. Ekramul Hamid

We present BhashaSetu, a linguistically enriched English--Marathi parallel dataset addressing persistent data limitations in low-resource neural machine translation (NMT). Marathi, spoken by over 95 million people, remains underrepresented…

Computation and Language · Computer Science 2026-05-27 Param Thakkar , Anushka Yadav , Michael Tiemann , Abhi Mehta , Akshita Bhasin , Shrinivas Khedkar

The rapid expansion of the digital world has propelled sentiment analysis into a critical tool across diverse sectors such as marketing, politics, customer service, and healthcare. While there have been significant advancements in sentiment…

Computation and Language · Computer Science 2024-04-08 Md. Arid Hasan , Shudipta Das , Afiyat Anjum , Firoj Alam , Anika Anjum , Avijit Sarker , Sheak Rashed Haider Noori
‹ Prev 1 2 3 10 Next ›