中文
相关论文

相关论文: Methods for Detoxification of Texts for the Russia…

200 篇论文

In this paper, we investigate the issue of hate speech by presenting a novel task of translating hate speech into non-hate speech text while preserving its meaning. As a case study, we use Spanish texts. We provide a dataset and several…

计算与语言 · 计算机科学 2023-06-05 Yevhen Kostiuk , Atnafu Lambebo Tonja , Grigori Sidorov , Olga Kolesnikova

Transformer language models (LMs) are fundamental to NLP research methodologies and applications in various languages. However, developing such models specifically for the Russian language has received little attention. This paper…

One of the stratagems used to deceive spam filters is to substitute vocables with synonyms or similar words that turn the message unrecognisable by the detection algorithms. In this paper we investigate whether the recent development of…

计算与语言 · 计算机科学 2021-07-16 Sergio Rojas-Galeano

In this paper we propose a word-wise intonation model for Russian language and show how it can be generalized for other languages. The proposed model is suitable for automatic data markup and its extended application to text-to-speech…

计算与语言 · 计算机科学 2024-10-01 Tomilov A. A. , Gromova A. Y. , Svischev A. N

This study aims to develop an efficient and accurate model for detecting malicious comments, addressing the increasingly severe issue of false and harmful content on social media platforms. We propose a deep learning model that combines…

计算与语言 · 计算机科学 2025-03-17 Zhou Fang , Hanlu Zhang , Jacky He , Zhen Qi , Hongye Zheng

This paper presents one of the top-performing solutions to the UNLP 2025 Shared Task on Detecting Manipulation in Social Media. The task focuses on detecting and classifying rhetorical and stylistic manipulation techniques used to influence…

计算与语言 · 计算机科学 2025-06-02 Kateryna Akhynko , Oleksandr Kosovan , Mykola Trokhymovych

Large Language Models (LLMs) have become integral to Software Engineering (SE), increasingly used in development workflows. However, their widespread adoption raises concerns about the presence and propagation of toxic language - harmful or…

机器学习 · 计算机科学 2026-01-21 Hao Zhuo , Yicheng Yang , Kewen Peng

Abstract: In this paper we present an approach to develop a text-classification model which would be able to identify populist content in text. The developed BERT-based model is largely successful in identifying populist content in text and…

计算与语言 · 计算机科学 2021-06-11 Jogilė Ulinskaitė , Lukas Pukelis

Automatic detection of toxic language plays an essential role in protecting social media users, especially minority groups, from verbal abuse. However, biases toward some attributes, including gender, race, and dialect, exist in most…

计算与语言 · 计算机科学 2021-06-15 Yung-Sung Chuang , Mingye Gao , Hongyin Luo , James Glass , Hung-yi Lee , Yun-Nung Chen , Shang-Wen Li

In this paper, we present a system for information extraction from scientific texts in the Russian language. The system performs several tasks in an end-to-end manner: term recognition, extraction of relations between terms, and term…

计算与语言 · 计算机科学 2021-09-15 Elena Bruches , Anastasia Mezentseva , Tatiana Batura

Fine-tuning pre-trained language models like BERT has become an effective way in NLP and yields state-of-the-art results on many downstream tasks. Recent studies on adapting BERT to new tasks mainly focus on modifying the model structure,…

计算与语言 · 计算机科学 2020-02-25 Yige Xu , Xipeng Qiu , Ligao Zhou , Xuanjing Huang

When trained on large, unfiltered crawls from the internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: they often generate racist, sexist, violent or otherwise toxic language. As…

计算与语言 · 计算机科学 2021-09-10 Timo Schick , Sahana Udupa , Hinrich Schütze

Topic models are typically represented by top-$m$ word lists for human interpretation. The corpus is often pre-processed with lemmatization (or stemming) so that those representations are not undermined by a proliferation of words with…

计算与语言 · 计算机科学 2019-05-13 Chandler May , Ryan Cotterell , Benjamin Van Durme

Risk perception is subjective, and youth's understanding of toxic content differs from that of adults. Although previous research has conducted extensive studies on toxicity detection in social media, the investigation of youth's unique…

计算与语言 · 计算机科学 2025-08-05 Yaqiong Li , Peng Zhang , Lin Wang , Hansu Gu , Siyuan Qiao , Ning Gu , Tun Lu

We present UniDetox, a universally applicable method designed to mitigate toxicity across various large language models (LLMs). Previous detoxification methods are typically model-specific, addressing only individual models or model…

计算与语言 · 计算机科学 2025-04-30 Huimin Lu , Masaru Isonuma , Junichiro Mori , Ichiro Sakata

Large language models have many beneficial applications, but can they also be used to attack content-filtering algorithms in social media platforms? We investigate the challenge of generating adversarial examples to test the robustness of…

计算与语言 · 计算机科学 2025-09-04 Piotr Przybyła , Euan McGill , Horacio Saggion

Despite regulations imposed by nations and social media platforms, such as recent EU regulations targeting digital violence, abusive content persists as a significant challenge. Existing approaches primarily rely on binary solutions, such…

In this paper, we introduce an advanced Russian general language understanding evaluation benchmark -- RussianGLUE. Recent advances in the field of universal language models and transformers require the development of a methodology for…

A significant challenge in automating hate speech detection on social media is distinguishing hate speech from regular and offensive language. These identify an essential category of content that web filters seek to remove. Only automated…

计算与语言 · 计算机科学 2024-11-12 Faria Naznin , Md Touhidur Rahman , Shahran Rahman Alve

An ever-increasing amount of social media content requires advanced AI-based computer programs capable of extracting useful information. Specifically, the extraction of health-related content from social media is useful for the development…

人工智能 · 计算机科学 2023-10-31 Pervaiz Iqbal Khan , Muhammad Nabeel Asim , Andreas Dengel , Sheraz Ahmed