中文
相关论文

相关论文: Reddit is all you need: Authorship profiling for R…

200 篇论文

Being around for decades, the problem of Authorship Attribution is still very much in focus currently. Some of the more recent instruments used are the pre-trained language models, the most prevalent being BERT. Here we used such a model to…

人工智能 · 计算机科学 2023-01-31 Sanda-Maria Avram

Determining the author of a text is a difficult task. Here we compare multiple AI techniques for classifying literary texts written by multiple authors by taking into account a limited number of speech parts (prepositions, adverbs, and…

人工智能 · 计算机科学 2023-01-25 Sanda Maria Avram , Mihai Oltean

Authorship style transfer involves altering text to match the style of a target author whilst preserving the original meaning. Existing unsupervised approaches like STRAP have largely focused on style transfer to target authors with many…

计算与语言 · 计算机科学 2024-11-05 Ajay Patel , Nicholas Andrews , Chris Callison-Burch

This study addresses the problem of authorship attribution for Romanian texts using the ROST corpus, a standard benchmark in the field. We systematically evaluate six machine learning techniques: Support Vector Machine (SVM), Logistic…

计算与语言 · 计算机科学 2025-06-30 Dana Lupsa , Sanda-Maria Avram , Radu Lupsa

We introduce PoPreRo, the first dataset for Popularity Prediction of Romanian posts collected from Reddit. The PoPreRo dataset includes a varied compilation of post samples from five distinct subreddits of Romania, totaling 28,107 data…

计算与语言 · 计算机科学 2024-11-26 Ana-Cristina Rogoz , Maria Ilinca Nechita , Radu Tudor Ionescu

Developing natural language processing (NLP) systems for social media analysis remains an important topic in artificial intelligence research. This article introduces RoBERTweet, the first Transformer architecture trained on Romanian…

计算与语言 · 计算机科学 2023-06-13 Iulian-Marius Tăiatu , Andrei-Marius Avram , Dumitru-Clementin Cercel , Florin Pop

Part of speech tagging is a fundamental NLP task often regarded as solved for high-resource languages such as English. Current state-of-the-art models have achieved high accuracy, especially on the news domain. However, when these models…

计算与语言 · 计算机科学 2020-04-30 Shabnam Behzad , Amir Zeldes

User identification has been a major field of research in privacy and security topics. Users might utilize multiple Online Social Networks (OSNs) to access a variety of text, videos, and links, and connect to their friends. Identifying user…

社会与信息网络 · 计算机科学 2024-06-05 Yasamin Kowsari

Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content.…

计算与语言 · 计算机科学 2023-02-21 Meng Ye , Karan Sikka , Katherine Atwell , Sabit Hassan , Ajay Divakaran , Malihe Alikhani

Understanding the sociodemographic composition of online platforms is essential for accurately interpreting digital behavior and its societal implications. Yet, current methods often lack the transparency and reliability required, risking…

社会与信息网络 · 计算机科学 2025-11-04 Federico Cinus , Corrado Monti , Paolo Bajardi , Gianmarco De Francisci Morales

The irreplaceable key to the triumph of Question & Answer (Q&A) platforms is their users providing high-quality answers to the challenging questions posted across various topics of interest. From more than a decade, the expert finding…

计算机与社会 · 计算机科学 2023-06-28 Sofia Strukova , José A. Ruipérez-Valiente , Félix Gómez Mármol

Authorship attribution refers to the task of automatically determining the author based on a given sample of text. It is a problem with a long history and has a wide range of application. Building author profiles using language models is…

计算与语言 · 计算机科学 2016-02-25 Zhenhao Ge , Yufang Sun

In pseudonymous online fora like Reddit, the benefits of self-disclosure are often apparent to users (e.g., I can vent about my in-laws to understanding strangers), but the privacy risks are more abstract (e.g., will my partner be able to…

人机交互 · 计算机科学 2024-12-20 Isadora Krsek , Anubha Kabra , Yao Dou , Tarek Naous , Laura A. Dabbish , Alan Ritter , Wei Xu , Sauvik Das

Authorship obfuscation techniques hold the promise of helping people protect their privacy in online communications by automatically rewriting text to hide the identity of the original author. However, obfuscation has been evaluated in…

计算与语言 · 计算机科学 2024-05-17 Calvin Bao , Marine Carpuat

Author profiling is the task of inferring characteristics about individuals by analyzing content they share. Supervised machine learning still dominates automatic systems that perform this task, despite the popularity of prompting large…

计算与语言 · 计算机科学 2025-05-29 Jan Hofmann , Cornelia Sindermann , Roman Klinger

Advice forums are a crowdsourced way to reinforce cultural norms and moral behavior. Sites like Reddit contain massive amounts of natural language human interaction, with rules and norms unique to each individual subreddit community. To…

人机交互 · 计算机科学 2021-09-21 Emily Cannon , Bianca Crouse , Souvick Ghosh , Nicholas Rihn , Kristen Chua

In online communities, recent studies have strongly improved our knowledge about the different types or profiles of contributors, from casual to very involved ones, through focused people. However they do so by using very complex…

人机交互 · 计算机科学 2018-03-28 Shubham Krishna , Romain Billot , Nicolas Jullien

Authorship identification is a process in which the author of a text is identified. Most known literary texts can easily be attributed to a certain author because they are, for example, signed. Yet sometimes we find unfinished pieces of…

计算与语言 · 计算机科学 2019-12-24 Rahul Radhakrishnan Iyer , Carolyn Penstein Rose

Recently, research on mental health conditions using public online data, including Reddit, has surged in NLP and health research but has not reported user characteristics, which are important to judge generalisability of findings. This…

计算与语言 · 计算机科学 2021-04-26 Glorianna Jagfeld , Fiona Lobban , Paul Rayson , Steven H. Jones

Massive amounts of contributed content -- including traditional literature, blogs, music, videos, reviews and tweets -- are available on the Internet today, with authors numbering in many millions. Textual information, such as product or…

数字图书馆 · 计算机科学 2014-05-21 Mishari Almishari , Ekin Oguz , Gene Tsudik
‹ 上一页 1 2 3 10 下一页 ›