中文
相关论文

相关论文: Language Models for Adult Service Website Text Ana…

200 篇论文

This project tackles the pressing issue of human trafficking in online C2C marketplaces through advanced Natural Language Processing (NLP) techniques. We introduce a novel methodology for generating pseudo-labeled datasets with minimal…

机器学习 · 计算机科学 2023-11-23 Alejandro Rodriguez Perez , Pablo Rivas

Data augmentation, the artificial creation of training data for machine learning by transformations, is a widely studied research field across machine learning disciplines. While it is useful for increasing a model's generalization…

计算与语言 · 计算机科学 2022-09-09 Markus Bayer , Marc-André Kaufhold , Christian Reuter

Textual data used to train large language models (LLMs) exhibits multifaceted bias manifestations encompassing harmful language and skewed demographic distributions. Regulations such as the European AI Act require identifying and mitigating…

Sentiment analysis (SA) has been, and is still, a thriving research area. However, the task of Arabic sentiment analysis (ASA) is still underrepresented in the body of research. This study offers the first in-depth and in-breadth analysis…

计算与语言 · 计算机科学 2024-03-05 Latifah Almurqren , Ryan Hodgson , Alexandra Cristea

There are more than 7,000 languages around the world, and current Large Language Models (LLMs) only support hundreds of languages. Dictionary-based prompting methods can enhance translation on them, but most methods use all the available…

计算与语言 · 计算机科学 2026-05-20 Hongyuan Lu , Zixuan Li , Zefan Zhang , Wai Lam

Annotating abusive language is expensive, logistically complex and creates a risk of psychological harm. However, most machine learning research has prioritized maximizing effectiveness (i.e., F1 or accuracy score) rather than data…

计算与语言 · 计算机科学 2022-09-22 Hannah Rose Kirk , Bertie Vidgen , Scott A. Hale

It is well known that textual data on the internet and other digital platforms contain significant levels of bias and stereotypes. Although many such texts contain stereotypes and biases that inherently exist in natural language for reasons…

计算与语言 · 计算机科学 2022-01-24 Ewoenam Kwaku Tokpo , Toon Calders

Abusive behaviors are common on online social networks. The increasing frequency of antisocial behaviors forces the hosts of online platforms to find new solutions to address this problem. Automating the moderation process has thus received…

社会与信息网络 · 计算机科学 2021-01-21 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

Public figures receive a disproportionate amount of abuse on social media, impacting their active participation in public life. Automated systems can identify abuse at scale but labelling training data is expensive, complex and potentially…

Code-switching (CSW) text generation has been receiving increasing attention as a solution to address data scarcity. In light of this growing interest, we need more comprehensive studies comparing different augmentation approaches. In this…

计算与语言 · 计算机科学 2023-10-25 Injy Hamed , Nizar Habash , Ngoc Thang Vu

Online sexism has become an increasing concern in social media platforms as it has affected the healthy development of the Internet and can have negative effects in society. While research in the sexism detection domain is growing, most of…

计算与语言 · 计算机科学 2021-08-09 Aiqi Jiang , Xiaohan Yang , Yang Liu , Arkaitz Zubiaga

Word embeddings are representations of individual words of a text document in a vector space and they are often use- ful for performing natural language pro- cessing tasks. Current state of the art al- gorithms for learning word embeddings…

计算与语言 · 计算机科学 2018-05-15 Prathusha Kameswara Sarma , Bill Sethares

Human trafficking (HT) is a pervasive global issue affecting vulnerable individuals, violating their fundamental human rights. Investigations reveal that a significant number of HT cases are associated with online advertisements (ads),…

计算与语言 · 计算机科学 2023-10-10 Vageesh Saxena , Benjamin Bashpole , Gijs Van Dijck , Gerasimos Spanakis

Recent work has shown that distributional word vector spaces often encode human biases like sexism or racism. In this work, we conduct an extensive analysis of biases in Arabic word embeddings by applying a range of recently introduced bias…

计算与语言 · 计算机科学 2020-11-04 Anne Lauscher , Rafik Takieddin , Simone Paolo Ponzetto , Goran Glavaš

Tweets are specific text data when compared to general text. Although sentiment analysis over tweets has become very popular in the last decade for English, it is still difficult to find huge annotated corpora for non-English languages. The…

计算与语言 · 计算机科学 2020-10-08 Valentin Barriere , Alexandra Balahur

Human trafficking is among the most challenging law enforcement problems which demands persistent fight against from all over the globe. In this study, we leverage readily available data from the website "Backpage"-- used for classified…

机器学习 · 计算机科学 2016-08-03 Hamidreza Alvari , Paulo Shakarian , J. E. Kelly Snyder

The pervasiveness of the Internet and social media have enabled the rapid and anonymous spread of Hate Speech content on microblogging platforms such as Twitter. Current EU and US legislation against hateful language, in conjunction with…

计算与语言 · 计算机科学 2021-02-10 Chrysoula Themeli , George Giannakopoulos , Nikiforos Pittaras

The increasing size and complexity of pre-trained language models have demonstrated superior performance in many applications, but they usually require large training datasets to be adequately trained. Insufficient training sets could…

计算与语言 · 计算机科学 2025-02-03 Yaping Chai , Haoran Xie , Joe S. Qin

Speech AI Technologies are largely trained on publicly available datasets or by the massive web-crawling of speech. In both cases, data acquisition focuses on minimizing collection effort, without necessarily taking the data subjects'…

计算机与社会 · 计算机科学 2023-05-04 Orestis Papakyriakopoulos , Alice Xiang

Cross-lingual document representations enable language understanding in multilingual contexts and allow transfer learning from high-resource to low-resource languages at the document level. Recently large pre-trained language models such as…

计算与语言 · 计算机科学 2021-06-08 Hongyu Gong , Vishrav Chaudhary , Yuqing Tang , Francisco Guzmán