中文
相关论文

相关论文: IIITG-ADBU@HASOC-Dravidian-CodeMix-FIRE2020: Offen…

200 篇论文

Social media often acts as breeding grounds for different forms of offensive content. For low resource languages like Tamil, the situation is more complex due to the poor performance of multilingual or language-specific models and lack of…

计算与语言 · 计算机科学 2021-02-22 Debjoy Saha , Naman Paharia , Debajit Chakraborty , Punyajoy Saha , Animesh Mukherjee

In this paper we present our submission for the EACL 2021-Shared Task on Offensive Language Identification in Dravidian languages. Our final system is an ensemble of mBERT and XLM-RoBERTa models which leverage task-adaptive pre-training of…

计算与语言 · 计算机科学 2021-03-15 Sai Muralidhar Jayanthi , Akshat Gupta

With the popularity of social media, communications through blogs, Facebook, Twitter, and other plat-forms have increased. Initially, English was the only medium of communication. Fortunately, now we can communicate in any language. It has…

计算与语言 · 计算机科学 2020-10-20 Sara Renjit , Sumam Mary Idicula

This paper presents the system descriptions submitted at the FIRE Shared Task 2021 on Urdu's Abusive and Threatening Language Detection Task. This challenge aims at automatically identifying abusive and threatening tweets written in Urdu.…

计算与语言 · 计算机科学 2022-04-08 Muhammad Humayoun

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

In social-media platforms such as Twitter, Facebook, and Reddit, people prefer to use code-mixed language such as Spanish-English, Hindi-English to express their opinions. In this paper, we describe different models we used, using the…

计算与语言 · 计算机科学 2020-10-13 Abhishek Singh , Surya Pratap Singh Parmar

This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing…

计算与语言 · 计算机科学 2021-03-10 Tommi Jauhiainen , Tharindu Ranasinghe , Marcos Zampieri

In the recent past, social media platforms have helped people in connecting and communicating to a wider audience. But this has also led to a drastic increase in cyberbullying. It is essential to detect and curb hate speech to keep the…

计算与语言 · 计算机科学 2021-11-02 Ravindra Nayak , Raviraj Joshi

In recent years, several systems have been developed to regulate the spread of negativity and eliminate aggressive, offensive or abusive contents from the online platforms. Nevertheless, a limited number of researches carried out to…

计算与语言 · 计算机科学 2021-03-02 Eftekhar Hossain , Omar Sharif , Mohammed Moshiul Hoque

This paper describes the participation of LIMSI UPV team in SemEval-2020 Task 9: Sentiment Analysis for Code-Mixed Social Media Text. The proposed approach competed in SentiMix Hindi-English subtask, that addresses the problem of predicting…

计算与语言 · 计算机科学 2020-09-01 Somnath Banerjee , Sahar Ghannay , Sophie Rosset , Anne Vilnat , Paolo Rosso

In an era of rapidly evolving internet technology, the surge in multimodal content, including videos, has expanded the horizons of online communication. However, the detection of toxic content in this diverse landscape, particularly in…

人工智能 · 计算机科学 2024-07-16 Krishanu Maity , A. S. Poornash , Sriparna Saha , Pushpak Bhattacharyya

Social media platforms often act as breeding grounds for various forms of trolling or malicious content targeting users or communities. One way of trolling users is by creating memes, which in most cases unites an image with a short piece…

多媒体 · 计算机科学 2022-04-28 Mithun Das , Somnath Banerjee , Animesh Mukherjee

Hateful content on social media increasingly appears as multimodal memes that combine images and text to convey harmful narratives. In low-resource languages such as Bengali, automated detection remains challenging due to limited annotated…

计算与语言 · 计算机科学 2026-02-24 Raihan Tanvir , Md. Golam Rabiul Alam

The widespread of offensive content online has become a reason for great concern in recent years, motivating researchers to develop robust systems capable of identifying such content automatically. With the goal of carrying out a fair…

计算与语言 · 计算机科学 2022-11-21 Tharindu Ranasinghe , Kai North , Damith Premasiri , Marcos Zampieri

Social media has penetrated into multilingual societies, however most of them use English to be a preferred language for communication. So it looks natural for them to mix their cultural language with English during conversations resulting…

计算与语言 · 计算机科学 2020-10-16 Shubhanker Banerjee , Arun Jayapal , Sajeetha Thavareesan

In the current era of the internet, where social media platforms are easily accessible for everyone, people often have to deal with threats, identity attacks, hate, and bullying due to their association with a cast, creed, gender, religion,…

计算与语言 · 计算机科学 2021-12-21 Zaki Mustafa Farooqi , Sreyan Ghosh , Rajiv Ratn Shah

Sarcasm detection is a significant challenge in sentiment analysis, particularly due to its nature of conveying opinions where the intended meaning deviates from the literal expression. This challenge is heightened in social media contexts…

计算与语言 · 计算机科学 2025-03-14 Aniket Deroy , Subhankar Maity

This paper presents our systems and results for the Hope Speech Detection in Code-Mixed Tulu Language shared task at the Sixth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages (DravidianLangTech-2026). We…

计算与语言 · 计算机科学 2026-05-12 Andrew Li , Sidney Wong

Social media platforms serve as accessible outlets for individuals to express their thoughts and experiences, resulting in an influx of user-generated data spanning all age groups. While these platforms enable free expression, they also…

计算与语言 · 计算机科学 2023-12-12 Nikhil Narayan , Mrutyunjay Biswal , Pramod Goyal , Abhranta Panigrahi

Using code-mixed data in natural language processing (NLP) research currently gets a lot of attention. Language identification of social media code-mixed text has been an interesting problem of study in recent years due to the advancement…