中文
相关论文

相关论文: Developing a Multilingual Annotated Corpus of Miso…

200 篇论文

Sarcasm detection and humor classification are inherently subtle problems, primarily due to their dependence on the contextual and non-verbal information. Furthermore, existing studies in these two topics are usually constrained in…

计算与语言 · 计算机科学 2021-06-01 Manjot Bedi , Shivani Kumar , Md Shad Akhtar , Tanmoy Chakraborty

Memes are the new-age conveyance mechanism for humor on social media sites. Memes often include an image and some text. Memes can be used to promote disinformation or hatred, thus it is crucial to investigate in details. We introduce…

We present the first English corpus study on abusive language towards three conversational AI systems gathered "in the wild": an open-domain social bot, a rule-based chatbot, and a task-based system. To account for the complexity of the…

计算与语言 · 计算机科学 2021-09-21 Amanda Cercas Curry , Gavin Abercrombie , Verena Rieser

Social media platforms are used by a large number of people prominently to express their thoughts and opinions. However, these platforms have contributed to a substantial amount of hateful and abusive content as well. Therefore, it is…

计算与语言 · 计算机科学 2022-05-24 Abhishek Velankar , Hrushikesh Patil , Amol Gore , Shubham Salunke , Raviraj Joshi

The rise of emergence of social media platforms has fundamentally altered how people communicate, and among the results of these developments is an increase in online use of abusive content. Therefore, automatically detecting this content…

计算与语言 · 计算机科学 2023-02-20 Khouloud Mnassri , Praboda Rajapaksha , Reza Farahbakhsh , Noel Crespi

The detection of offensive, hateful and profane language has become a critical challenge since many users in social networks are exposed to cyberbullying activities on a daily basis. In this paper, we present an analysis of combining…

计算与语言 · 计算机科学 2021-12-10 Sherzod Hakimov , Ralph Ewerth

Identifying adverse and hostile content on the web and more particularly, on social media, has become a problem of paramount interest in recent years. With their ever increasing popularity, fine-tuning of pretrained Transformer-based…

计算与语言 · 计算机科学 2021-01-12 Tathagata Raha , Sayar Ghosh Roy , Ujwal Narayan , Zubair Abid , Vasudeva Varma

Current multimodal toxicity benchmarks typically use a single binary hatefulness label. This coarse approach conflates two fundamentally different characteristics of expression: tone and content. Drawing on communication science theory, we…

计算与语言 · 计算机科学 2026-03-25 Nils A. Herrmann , Tobias Eder , Jingyi He , Georg Groh

Due to the sheer volume of online hate, the AI and NLP communities have started building models to detect such hateful content. Recently, multilingual hate is a major emerging challenge for automated detection where code-mixing or more than…

计算与语言 · 计算机科学 2022-05-12 Mithun Das , Punyajoy Saha , Binny Mathew , Animesh Mukherjee

This work focuses on two subtasks related to hate speech detection and target identification in Devanagari-scripted languages, specifically Hindi, Marathi, Nepali, Bhojpuri, and Sanskrit. Subtask B involves detecting hate speech in online…

计算与语言 · 计算机科学 2024-12-31 Siddhant Gupta , Siddh Singhal , Azmine Toushik Wasi

The datasets most widely used for abusive language detection contain lists of messages, usually tweets, that have been manually judged as abusive or not by one or more annotators, with the annotation performed at message level. In this…

计算与语言 · 计算机科学 2021-03-30 Stefano Menini , Alessio Palmero Aprosio , Sara Tonelli

The digital age has expanded social media and online forums, allowing free expression for nearly 45% of the global population. Yet, it has also fueled online harassment, bullying, and harmful behaviors like hate speech and toxic comments…

计算与语言 · 计算机科学 2026-03-12 Vuong M. Ngo , Cach N. Dang , Kien V. Nguyen , Mark Roantree

Social media has become a bedrock for people to voice their opinions worldwide. Due to the greater sense of freedom with the anonymity feature, it is possible to disregard social etiquette online and attack others without facing severe…

计算与语言 · 计算机科学 2021-10-19 Arushi Sharma , Anubha Kabra , Minni Jain

This paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India. Our corpus provides a training and testing ground for four language families, 14 languages, and the largest to date…

计算与语言 · 计算机科学 2023-10-23 Ashok Urlana , Pinzhen Chen , Zheng Zhao , Shay B. Cohen , Manish Shrivastava , Barry Haddow

Sentiment analysis is the most basic NLP task to determine the polarity of text data. There has been a significant amount of work in the area of multilingual text as well. Still hate and offensive speech detection faces a challenge due to…

计算与语言 · 计算机科学 2021-11-02 Abhishek Velankar , Hrushikesh Patil , Amol Gore , Shubham Salunke , Raviraj Joshi

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

计算与语言 · 计算机科学 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha

The interest in offensive content identification in social media has grown substantially in recent years. Previous work has dealt mostly with post level annotations. However, identifying offensive spans is useful in many ways. To help…

计算与语言 · 计算机科学 2021-04-20 Tharindu Ranasinghe , Marcos Zampieri

The automatic identification of offensive language such as hate speech is important to keep discussions civil in online communities. Identifying hate speech in multimodal content is a particularly challenging task because offensiveness can…

As open-ended human-chatbot interaction becomes commonplace, sensitive content detection gains importance. In this work, we propose a two stage semi-supervised approach to bootstrap large-scale data for automatic sensitive language…

计算与语言 · 计算机科学 2018-12-03 Chandra Khatri , Behnam Hedayatnia , Rahul Goel , Anushree Venkatesh , Raefer Gabriel , Arindam Mandal

While human annotations play a crucial role in language technologies, annotator subjectivity has long been overlooked in data collection. Recent studies that have critically examined this issue are often situated in the Western context, and…

计算与语言 · 计算机科学 2024-04-18 Aida Mostafazadeh Davani , Mark Díaz , Dylan Baker , Vinodkumar Prabhakaran