中文
相关论文

相关论文: Enhancing Data Quality through Simple De-duplicati…

200 篇论文

In NLP, models are usually evaluated by reporting single-number performance scores on a number of readily available benchmarks, without much deeper analysis. Here, we argue that - especially given the well-known fact that benchmarks often…

计算与语言 · 计算机科学 2022-10-05 Daniel Simig , Tianlu Wang , Verna Dankers , Peter Henderson , Khuyagbaatar Batsuren , Dieuwke Hupkes , Mona Diab

The prevalence of social media presents a growing opportunity to collect and analyse examples of English varieties. Whilst usage of these varieties was - and, in many cases, still is - used only in spoken contexts or hard-to-access private…

计算与语言 · 计算机科学 2024-01-23 Nhi Pham , Lachlan Pham , Adam L. Meyers

When trained on large, unfiltered crawls from the internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: they often generate racist, sexist, violent or otherwise toxic language. As…

计算与语言 · 计算机科学 2021-09-10 Timo Schick , Sahana Udupa , Hinrich Schütze

Clinical NLP tasks such as mental health assessment from text, must take social constraints into account - the performance maximization must be constrained by the utmost importance of guaranteeing privacy of user data. Consumer protection…

计算与语言 · 计算机科学 2022-09-07 Ramit Sawhney , Atula Tejaswi Neerkaje , Ivan Habernal , Lucie Flek

Social media platforms enable the rapid dissemination and consumption of information. However, users instantly consume such content regardless of the reliability of the shared data. Consequently, the latter crowdsourcing model is exposed to…

This study addresses the critical challenge of detecting DeepFake tweets by leveraging advanced natural language processing (NLP) techniques to distinguish between genuine and AI-generated texts. Given the increasing prevalence of…

计算与语言 · 计算机科学 2024-06-25 Adam Frej , Adrian Kaminski , Piotr Marciniak , Szymon Szmajdzinski , Soveatin Kuntur , Anna Wroblewska

Crowdsourcing has been the prevalent paradigm for creating natural language understanding datasets in recent years. A common crowdsourcing practice is to recruit a small number of high-quality workers, and have them massively generate…

计算与语言 · 计算机科学 2019-08-29 Mor Geva , Yoav Goldberg , Jonathan Berant

Malicious accounts spreading misinformation has led to widespread false and misleading narratives in recent times, especially during the COVID-19 pandemic, and social media platforms struggle to eliminate these contents rapidly. This is…

社会与信息网络 · 计算机科学 2022-02-28 Karishma Sharma , Emilio Ferrara , Yan Liu

This paper aims to understand the impacts of various data combinations (e.g., web text, Wikipedia, GitHub, books) on the pretraining of large language models using SlimPajama. SlimPajama is a rigorously deduplicated, multi-source dataset,…

With the recent advances in natural language processing (NLP), a vast number of applications have emerged across various use cases. Among the plethora of NLP applications, many academic researchers are motivated to do work that has a…

计算与语言 · 计算机科学 2023-10-24 Fernando Gonzalez , Zhijing Jin , Bernhard Schölkopf , Tom Hope , Mrinmaya Sachan , Rada Mihalcea

Recent work in natural language processing (NLP) has yielded appealing results from scaling model parameters and training data; however, using only scale to improve performance means that resource consumption also grows. Such resources…

User generated text on social media often suffers from a lot of undesired characteristics including hatespeech, abusive language, insults etc. that are targeted to attack or abuse a specific group of people. Often such text is written…

计算与语言 · 计算机科学 2019-10-03 Sravan Babu Bodapati , Spandana Gella , Kasturi Bhattacharjee , Yaser Al-Onaizan

The difficulty of anonymizing text data hinders the development and deployment of NLP in high-stakes domains that involve private data, such as healthcare and social services. Poorly anonymized sensitive data cannot be easily shared with…

计算与语言 · 计算机科学 2024-10-14 Krithika Ramesh , Nupoor Gandhi , Pulkit Madaan , Lisa Bauer , Charith Peris , Anjalie Field

Self-supervised learning techniques are celebrating immense success in natural language processing (NLP) by enabling models to learn from broad language data at unprecedented scales. Here, we aim to leverage the success of these techniques…

神经元与认知 · 定量生物学 2023-01-18 Armin W. Thomas , Christopher Ré , Russell A. Poldrack

Deceptive news posts shared in online communities can be detected with NLP models, and much recent research has focused on the development of such models. In this work, we use characteristics of online communities and authors -- the context…

计算机与社会 · 计算机科学 2021-04-29 Galen Weld , Ellyn Ayton , Tim Althoff , Maria Glenski

Online social media platforms increasingly rely on Natural Language Processing (NLP) techniques to detect abusive content at scale in order to mitigate the harms it causes to their users. However, these techniques suffer from various…

计算与语言 · 计算机科学 2021-10-01 Sayan Ghosh , Dylan Baker , David Jurgens , Vinodkumar Prabhakaran

In machine learning, temporal shifts occur when there are differences between training and test splits in terms of time. For streaming data such as news or social media, models are commonly trained on a fixed corpus from a certain period of…

计算与语言 · 计算机科学 2024-05-24 Asahi Ushio , Jose Camacho-Collados

Social media data has become a vital resource for studying mental health, offering real-time insights into thoughts, emotions, and behaviors that traditional methods often miss. Progress in this area has been facilitated by benchmark…

计算与语言 · 计算机科学 2025-11-27 Saad Mankarious , Ayah Zirikly , Daniel Wiechmann , Elma Kerz , Edward Kempa , Yu Qiao

The collection and examination of social media has become a useful mechanism for studying the mental activity and behavior tendencies of users. Through the analysis of collected Twitter data, models were developed for classifying…

社会与信息网络 · 计算机科学 2020-03-26 Joseph Tassone , Peizhi Yan , Mackenzie Simpson , Chetan Mendhe , Vijay Mago , Salimur Choudhury

In the era of increasingly sophisticated natural language processing (NLP) systems, large language models (LLMs) have demonstrated remarkable potential for diverse applications, including tasks requiring nuanced textual understanding and…

计算与语言 · 计算机科学 2025-05-16 Poli Apollinaire Nemkova , Solomon Ubani , Mark V. Albert