中文
相关论文

相关论文: Is preprocessing of text really worth your time fo…

200 篇论文

Toxic comment detection on social media has proven to be essential for content moderation. This paper compares a wide set of different models on a highly skewed multi-label hate speech dataset. We consider inference time and several metrics…

计算与语言 · 计算机科学 2023-01-27 Corentin Duchene , Henri Jamet , Pierre Guillaume , Reda Dehak

Text classification approaches have usually required task-specific model architectures and huge labeled datasets. Recently, thanks to the rise of text-based transfer learning techniques, it is possible to pre-train a language model in an…

计算与语言 · 计算机科学 2019-06-10 Enkhbold Bataa , Joshua Wu

The censorship of toxic comments is often left to the judgment of imperfect models. Perspective API, a creation of Google technology incubator Jigsaw, is perhaps the most widely used toxicity classifier in industry; the model is employed by…

计算机与社会 · 计算机科学 2020-07-01 Elizabeth Reichert , Helen Qiu , Jasmine Bayrooti

Social media are becoming an increasingly important source of information about the public mood regarding issues such as elections, Brexit, stock market, etc. In this paper we focus on sentiment classification of Twitter data. Construction…

计算与语言 · 计算机科学 2021-08-31 Igor Mozetič , Luis Torgo , Vitor Cerqueira , Jasmina Smailović

Considering the large amount of content created online by the minute, slang-aware automatic tools are critically needed to promote social good, and assist policymakers and moderators in restricting the spread of offensive language, abuse,…

计算与语言 · 计算机科学 2023-02-02 Aravinda Kolla , Filip Ilievski , Hông-Ân Sandlin , Alain Mermoud

We use over 350,000 Yelp reviews on 5,000 restaurants to perform an ablation study on text preprocessing techniques. We also compare the effectiveness of several machine learning and deep learning models on predicting user sentiment…

计算与语言 · 计算机科学 2020-04-30 Siqi Liu

Large language models pretrained on extensive web corpora demonstrate remarkable performance across a wide range of downstream tasks. However, a growing concern is data contamination, where evaluation datasets may be contained in the…

计算与语言 · 计算机科学 2024-07-12 Medha Palavalli , Amanda Bertsch , Matthew R. Gormley

Progress in natural language generation research has been shaped by the ever-growing size of language models. While large language models pre-trained on web data can generate human-sounding text, they also reproduce social biases and…

计算与语言 · 计算机科学 2023-06-06 Celine Wald , Lukas Pfahler

Research into the classification of Image with Text (IWT) troll memes has recently become popular. Since the online community utilizes the refuge of memes to express themselves, there is an abundance of data in the form of memes. These…

社会与信息网络 · 计算机科学 2022-05-11 Shardul Suryawanshi , Bharathi Raja Chakravarthi , Mihael Arcan , Suzanne Little , Paul Buitelaar

Data contamination in model evaluation has become increasingly prevalent with the growing popularity of large language models. It allows models to "cheat" via memorisation instead of displaying true capabilities. Therefore, contamination…

计算与语言 · 计算机科学 2024-01-30 Yucheng Li , Frank Guerin , Chenghua Lin

Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are wrongly annotated as safe ones, the model fine-tuned on…

机器学习 · 计算机科学 2024-03-26 Zhaowei Zhu , Jialu Wang , Hao Cheng , Yang Liu

Language models (LMs) are pretrained to imitate internet text, including content that would violate human preferences if generated by an LM: falsehoods, offensive comments, personally identifiable information, low-quality or buggy code, and…

The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving…

机器学习 · 计算机科学 2025-07-08 Jing Yu , Yibo Zhao , Jiapeng Zhu , Wenming Shao , Bo Pang , Zhao Zhang , Xiang Li

The amount of text generated daily on social media is gigantic and analyzing this text is useful for many purposes. To understand what lies beneath a huge amount of text, we need dependable and effective computing techniques from…

信息检索 · 计算机科学 2025-08-04 Ngozichukwuka Onah , Nadine Steinmetz , Hani Al-Sayeh , Kai-Uwe Sattler

Probing or fine-tuning (large-scale) pre-trained models results in state-of-the-art performance for many NLP tasks and, more recently, even for computer vision tasks when combined with image data. Unfortunately, these approaches also entail…

计算机视觉与模式识别 · 计算机科学 2021-10-11 Patrick Schramowski , Kristian Kersting

Pretrained neural language models (LMs) are prone to generating racist, sexist, or otherwise toxic language which hinders their safe deployment. We investigate the extent to which pretrained LMs can be prompted to generate toxic language,…

计算与语言 · 计算机科学 2020-09-29 Samuel Gehman , Suchin Gururangan , Maarten Sap , Yejin Choi , Noah A. Smith

Large language models' (LLMs) abilities are drawn from their pretraining data, and model development begins with data curation. However, decisions around what data is retained or removed during this initial stage are under-scrutinized. In…

计算与语言 · 计算机科学 2024-06-24 Li Lucy , Suchin Gururangan , Luca Soldaini , Emma Strubell , David Bamman , Lauren F. Klein , Jesse Dodge

Text classification is one of the most widely studied tasks in natural language processing. Motivated by the principle of compositionality, large multilayer neural network models have been employed for this task in an attempt to effectively…

计算与语言 · 计算机科学 2018-08-07 Devendra Singh Sachan , Manzil Zaheer , Ruslan Salakhutdinov

The emergent capabilities of large language models (LLMs) have prompted interest in using them as surrogates for human subjects in opinion surveys. However, prior evaluations of LLM-based opinion simulation have relied heavily on costly,…

计算机与社会 · 计算机科学 2025-11-17 Terrence Neumann , Maria De-Arteaga , Sina Fazelpour

User posts whose perceived toxicity depends on the conversational context are rare in current toxicity detection datasets. Hence, toxicity detectors trained on existing datasets will also tend to disregard context, making the detection of…

计算与语言 · 计算机科学 2021-11-22 Alexandros Xenos , John Pavlopoulos , Ion Androutsopoulos , Lucas Dixon , Jeffrey Sorensen , Leo Laugier