中文
相关论文

相关论文: Single-Pass, Adaptive Natural Language Filtering: …

200 篇论文

Automatic abusive language detection is a difficult but important task for online social media. Our research explores a two-step approach of performing classification on abusive language and then classifying into specific types and compares…

计算与语言 · 计算机科学 2017-06-06 Ji Ho Park , Pascale Fung

Platforms are increasingly relying on algorithms to curate the content within users' social media feeds. However, the growing prominence of proprietary, algorithmically curated feeds has concealed what factors influence the presentation of…

人机交互 · 计算机科学 2025-08-11 Jackie Chan , Fred Choi , Koustuv Saha , Eshwar Chandrasekharan

The profusion of user generated content caused by the rise of social media platforms has enabled a surge in research relating to fields such as information retrieval, recommender systems, data mining and machine learning. However, the lack…

社会与信息网络 · 计算机科学 2018-01-23 Nuno Moniz , Luís Torgo

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

计算与语言 · 计算机科学 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

Identifying argument components from unstructured texts and predicting the relationships expressed among them are two primary steps of argument mining. The intrinsic complexity of these tasks demands powerful learning models. While…

计算与语言 · 计算机科学 2022-03-25 Subhabrata Dutta , Jeevesh Juneja , Dipankar Das , Tanmoy Chakraborty

This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new Multi-Topic…

计算与语言 · 计算机科学 2022-11-09 Yao Dou , Chao Jiang , Wei Xu

With the growth of fake news and disinformation, the NLP community has been working to assist humans in fact-checking. However, most academic research has focused on model accuracy without paying attention to resource efficiency, which is…

计算机与社会 · 计算机科学 2021-09-03 Mykola Trokhymovych , Diego Saez-Trumper

After the launch of ChatGPT v.4 there has been a global vivid discussion on the ability of this artificial intelligence powered platform and some other similar ones for the automatic production of all kinds of texts, including scientific…

计算与语言 · 计算机科学 2024-04-16 Javier J. Sanchez-Medina

On social media platforms, hateful and offensive language negatively impact the mental well-being of users and the participation of people from diverse backgrounds. Automatic methods to detect offensive language have largely relied on…

计算与语言 · 计算机科学 2022-01-26 Rishav Hada , Sohi Sudhir , Pushkar Mishra , Helen Yannakoudakis , Saif M. Mohammad , Ekaterina Shutova

Building open-domain dialogue systems capable of rich human-like conversational ability is one of the fundamental challenges in language generation. However, even with recent advancements in the field, existing open-domain generative models…

计算与语言 · 计算机科学 2022-06-14 Ritvik Choudhary , Daisuke Kawahara

We present an approach for selecting objectively informative and subjectively helpful annotations to social media posts. We draw on data from on an online environment where contributors annotate misinformation and simultaneously rate the…

社会与信息网络 · 计算机科学 2022-10-31 Stefan Wojcik , Sophie Hilgard , Nick Judd , Delia Mocanu , Stephen Ragain , M. B. Fallin Hunzaker , Keith Coleman , Jay Baxter

Automatic quality evaluation of Web information is a task with many fields of applications and of great relevance, especially in critical domains like the medical one. We move from the intuition that the quality of content of medical Web…

信息检索 · 计算机科学 2016-03-08 Vittoria Cozza , Marinella Petrocchi , Angelo Spognardi

We introduce an unsupervised approach to efficiently discover the underlying features in a data set via crowdsourcing. Our queries ask crowd members to articulate a feature common to two out of three displayed examples. In addition we also…

机器学习 · 统计学 2015-04-02 James Y. Zou , Kamalika Chaudhuri , Adam Tauman Kalai

Sampling is a common strategy for generating text from probabilistic models, yet standard ancestral sampling often results in text that is incoherent or ungrammatical. To alleviate this issue, various modifications to a model's sampling…

计算与语言 · 计算机科学 2024-01-08 Clara Meister , Tiago Pimentel , Luca Malagutti , Ethan G. Wilcox , Ryan Cotterell

Text normalization is an essential task in the processing and analysis of social media that is dominated with informal writing. It aims to map informal words to their intended standard forms. Previously proposed text normalization…

计算与语言 · 计算机科学 2017-12-29 Salman Ahmad Ansari , Usman Zafar , Asim Karim

This study introduces ValueScope, a framework leveraging language models to quantify social norms and values within online communities, grounded in social science perspectives on normative structures. We employ ValueScope to dissect and…

Recent works have demonstrated success in controlling sentence attributes ($e.g.$, sentiment) and structure ($e.g.$, syntactic structure) based on the diffusion language model. A key component that drives theimpressive performance for…

计算与语言 · 计算机科学 2024-03-26 Shujian Zhang , Lemeng Wu , Chengyue Gong , Xingchao Liu

Twitter, a popular social media outlet, has evolved into a vast source of linguistic data, rich with opinion, sentiment, and discussion. Due to the increasing popularity of Twitter, its perceived potential for exerting social influence has…

Social media feed ranking algorithms fail when they too narrowly focus on engagement as their objective. The literature has asserted a wide variety of values that these algorithms should account for as well -- ranging from well-being to…

人机交互 · 计算机科学 2025-05-19 Akaash Kolluri , Renn Su , Farnaz Jahanbakhsh , Dora Zhao , Tiziano Piccardi , Michael S. Bernstein

Data augmentation promises to alleviate data scarcity. This is most important in cases where the initial data is in short supply. This is, for existing methods, also where augmenting is the most difficult, as learning the full data…

计算与语言 · 计算机科学 2020-03-24 Guillaume Raille , Sandra Djambazovska , Claudiu Musat