中文
相关论文

相关论文: Reddit is all you need: Authorship profiling for R…

200 篇论文

Authorship Attribution is the task of creating an appropriate characterization of text that captures the authors' writing style to identify the original author of a given piece of text. With increased anonymity on the internet, this task…

计算与语言 · 计算机科学 2024-03-11 Aisha Khatun , Anisur Rahman , Md Saiful Islam , Hemayet Ahmed Chowdhury , Ayesha Tasnim

Recently, powerful Large Language Models (LLMs) have become easily accessible to hundreds of millions of users world-wide. However, their strong capabilities and vast world knowledge do not come without associated privacy risks. In this…

机器学习 · 计算机科学 2024-11-05 Hanna Yukhymenko , Robin Staab , Mark Vero , Martin Vechev

Authorship attribution aims to identify the origin or author of a document. Traditional approaches have heavily relied on manual features and fail to capture long-range correlations, limiting their effectiveness. Recent advancements…

计算与语言 · 计算机科学 2024-10-30 Zhengmian Hu , Tong Zheng , Heng Huang

Mental health remains a significant challenge of public health worldwide. With increasing popularity of online platforms, many use the platforms to share their mental health conditions, express their feelings, and seek help from the…

计算与语言 · 计算机科学 2022-06-03 Sajad Sotudeh , Nazli Goharian , Zachary Young

With the increasing availability of online scholarly databases, publication records can be easily extracted and analysed. Researchers can promptly keep abreast of others' scientific production and, in principle, can select new collaborators…

社会与信息网络 · 计算机科学 2020-09-30 Xiancheng Li , Luca Verginer , Massimo Riccaboni , Pietro Panzarasa

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

Inference of online social network users' attributes and interests has been an active research topic. Accurate identification of users' attributes and interests is crucial for improving the performance of personalization and recommender…

社会与信息网络 · 计算机科学 2015-04-21 Quanzeng You , Sumit Bhatia , Jiebo Luo

With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods…

计算与语言 · 计算机科学 2021-01-19 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

Assigning qualified, unbiased and interested reviewers to paper submissions is vital for maintaining the integrity and quality of the academic publishing system and providing valuable reviews to authors. However, matching thousands of…

信息检索 · 计算机科学 2022-11-09 Omer Anjum , Alok Kamatar , Toby Liang , Jinjun Xiong , Wen-mei Hwu

This paper refers to the syntactic analysis of phrases in Romanian, as an important process of natural language processing. We will suggest a real-time solution, based on the idea of using some words or groups of words that indicate…

计算与语言 · 计算机科学 2013-01-10 Bogdan Patrut

Having a quality annotated corpus is essential especially for applied research. Despite the recent focus of Web science community on researching about cyberbullying, the community dose not still have standard benchmarks. In this paper, we…

We present a novel corpus for personality prediction in Italian, containing a larger number of authors and a different genre compared to previously available resources. The corpus is built exploiting Distant Supervision, assigning…

计算与语言 · 计算机科学 2020-11-12 Elisa Bassignana , Malvina Nissim , Viviana Patti

Large text corpora, such as Reddit posts, have become an increasingly prevalent site of qualitative inquiry. However, most large text corpora are intractable for qualitative researchers. Instead, teams rely on statistical subsampling to…

In this paper, we propose a deep, globally normalized topic model that incorporates structural relationships connecting documents in socially generated corpora, such as online forums. Our model (1) captures discursive interactions along…

机器学习 · 计算机科学 2020-05-11 Nikita Srivatsan , Zachary Wojtowicz , Taylor Berg-Kirkpatrick

Convolutional neural networks (CNNs) have demonstrated superior capability for extracting information from raw signals in computer vision. Recently, character-level and multi-channel CNNs have exhibited excellent performance for sentence…

计算与语言 · 计算机科学 2016-09-22 Sebastian Ruder , Parsa Ghaffari , John G. Breslin

Many methods have been used to recognize author personality traits from text, typically combining linguistic feature engineering with shallow learning models, e.g. linear regression or Support Vector Machines. This work uses…

计算与语言 · 计算机科学 2016-10-17 Fei Liu , Julien Perez , Scott Nowson

Languages shared by people differ in different regions based on their accents, pronunciation and word usages. In this era sharing of language takes place mainly through social media and blogs. Every second swing of such a micro posts exist…

计算与语言 · 计算机科学 2018-04-13 Barathi Ganesh HB

Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. This technique has multiple uses in tasks related to…

The primary goal of a news headline is to summarize an event in as few words as possible. Depending on the media outlet, a headline can serve as a means to objectively deliver a summary or improve its visibility. For the latter, specific…

This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new Multi-Topic…

计算与语言 · 计算机科学 2022-11-09 Yao Dou , Chao Jiang , Wei Xu