中文
相关论文

相关论文: Forensic Authorship Analysis of Microblogging Text…

200 篇论文

Concepts and methods of complex networks can be used to analyse texts at their different complexity levels. Examples of natural language processing (NLP) tasks studied via topological analysis of networks are keyword identification,…

计算与语言 · 计算机科学 2017-02-07 Vanessa Queiroz Marinho , Graeme Hirst , Diego Raphael Amancio

The advent of instruction-tuned language models that convincingly mimic human writing poses a significant risk of abuse. However, such abuse may be counteracted with the ability to detect whether a piece of text was composed by a language…

计算与语言 · 计算机科学 2024-05-09 Rafael Rivera Soto , Kailin Koch , Aleem Khan , Barry Chen , Marcus Bishop , Nicholas Andrews

With the increasing use of the Internet and mobile devices, social networks are becoming the most used media to communicate citizens' ideas and thoughts. This information is very useful to identify communities with common ideas based on…

社会与信息网络 · 计算机科学 2019-04-22 Vargas-Calderón Vladimir , Camargo Jorge

Mining social media messages for health and drug related information has received significant interest in pharmacovigilance research. Social media sites (e.g., Twitter), have been used for monitoring drug abuse, adverse reactions of drug…

计算与语言 · 计算机科学 2018-05-17 Debanjan Mahata , Jasper Friedrichs , Hitkul , Rajiv Ratn Shah

The authors of fake news often use facts from verified news sources and mix them with misinformation to create confusion and provoke unrest among the readers. The spread of fake news can thereby have serious implications on our society.…

计算与语言 · 计算机科学 2020-09-30 Inna Vogel , Meghana Meghana

In November 2017, Twitter doubled the maximum allowed tweet length from 140 to 280 characters, a drastic switch on one of the world's most influential social media platforms. In the first long-term study of how the new length limit was…

社会与信息网络 · 计算机科学 2020-09-17 Kristina Gligorić , Ashton Anderson , Robert West

140 characters seems like too small a space for any meaningful information to be exchanged, but Twitter users have found creative ways to get the most out of each Tweet by using different communication tools. This paper looks into how 73…

计算机与社会 · 计算机科学 2012-02-20 Kristen Lovejoy , Richard Waters , Gregory D. Saxton

Twitter has become one of the main sources of news for many people. As real-world events and emergencies unfold, Twitter is abuzz with hundreds of thousands of stories about the events. Some of these stories are harmless, while others could…

社会与信息网络 · 计算机科学 2016-06-21 Soroush Vosoughi , Deb Roy

In this paper, we introduce an authorship attribution method called Authorial Language Models (ALMs) that involves identifying the most likely author of a questioned document based on the perplexity of the questioned document calculated for…

计算与语言 · 计算机科学 2024-02-14 Weihang Huang , Akira Murakami , Jack Grieve

The web-based microblogging system Twitter is a very popular altmetrics source for measuring the broader impact of science. In this case study, we demonstrate how problematic the use of Twitter data for research evaluation can be, even…

数字图书馆 · 计算机科学 2019-04-16 Lutz Bornmann , Robin Haunschild

Authorship identification has proven unsettlingly effective in inferring the identity of the author of an unsigned document, even when sensitive personal information has been carefully omitted. In the digital era, individuals leave a…

计算与语言 · 计算机科学 2023-10-04 Haining Wang

Statistical methods have been widely employed in many practical natural language processing applications. More specifically, complex networks concepts and methods from dynamical systems theory have been successfully applied to recognize…

计算与语言 · 计算机科学 2015-03-04 Diego R. Amancio

Authorship identification tasks, which rely heavily on linguistic styles, have always been an important part of Natural Language Understanding (NLU) research. While other tasks based on linguistic style understanding benefit from deep…

计算与语言 · 计算机科学 2020-10-01 Weicheng Ma , Ruibo Liu , Lili Wang , Soroush Vosoughi

Social media has become a very popular source of information. With this popularity comes an interest in systems that can classify the information produced. This study tries to create such a system detecting irony in Twitter users. Recent…

计算与语言 · 计算机科学 2023-11-09 Tibor L. R. Krols , Marie Mortensen , Ninell Oldenburg

The task of written language identification involves typically the detection of the languages present in a sample of text. Moreover, a sequence of text may not belong to a single inherent language but also may be mixture of text written in…

计算与语言 · 计算机科学 2020-07-14 Mohd Zeeshan Ansari , Tanvir Ahmad , Ana Fatima

Reliance on anonymity in social media has increased its popularity on these platforms among all ages. The availability of public Wi-Fi networks has facilitated a vast variety of online content, including social media applications. Although…

计算与语言 · 计算机科学 2025-01-15 Hiba Fallatah , Ching Suen , Olga Ormandjieva

Text from social media provides a set of challenges that can cause traditional NLP approaches to fail. Informal language, spelling errors, abbreviations, and special characters are all commonplace in these posts, leading to a prohibitively…

机器学习 · 计算机科学 2016-05-18 Bhuwan Dhingra , Zhong Zhou , Dylan Fitzpatrick , Michael Muehl , William W. Cohen

Authorship attribution aims to identify the origin or author of a document. Traditional approaches have heavily relied on manual features and fail to capture long-range correlations, limiting their effectiveness. Recent advancements…

计算与语言 · 计算机科学 2024-10-30 Zhengmian Hu , Tong Zheng , Heng Huang

Handwritten document analysis is an area of forensic science, with the goal of establishing authorship of documents through examination of inherent characteristics. Law enforcement agencies use standard protocols based on manual processing…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Eleonora Breci , Luca Guarnera , Sebastiano Battiato

Understanding the semantic of a collection of texts is a challenging task. Topic models are probabilistic models that aims at extracting "topics" from a corpus of documents. This task is particularly difficult when the corpus is composed of…

信息检索 · 计算机科学 2022-03-22 Hugo Schnoering