English
Related papers

Related papers: CEAI: CCM based Email Authorship Identification Mo…

200 papers

Enterprise relational databases increasingly contain vast amounts of non-semantic data - IP addresses, product identifiers, encoded keys, and timestamps - that challenge traditional semantic analysis. This paper introduces a novel…

Machine Learning · Computer Science 2025-11-12 Veera V S Bhargav Nunna , Shinae Kang , Zheyuan Zhou , Virginia Wang , Sucharitha Boinapally , Michael Foley

Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. This technique has multiple uses in tasks related to…

Robust analysis of coauthorship networks is based on high quality data. However, ground-truth data are usually unavailable. Empirical data suffer several types of errors, a typical one of which is called merging error, identifying different…

Physics and Society · Physics 2018-12-27 Zheng Xie

Characters are the smallest unit of text that can extract stylometric signals to determine the author of a text. In this paper, we investigate the effectiveness of character-level signals in Authorship Attribution of Bangla Literature and…

Computation and Language · Computer Science 2020-11-06 Aisha Khatun , Anisur Rahman , Md. Saiful Islam , Marium-E-Jannat

Authorship attribution (AA) is the task of identifying the most likely author of a query document from a predefined set of candidate authors. We introduce a two-stage retrieve-and-rerank framework that finetunes LLMs for cross-genre AA.…

Computation and Language · Computer Science 2025-10-21 Shantanu Agarwal , Joel Barry , Steven Fincke , Scott Miller

In this paper, the process of converting the Enron email dataset (the version cited in the preprint) to thousands of features per email for a selected set of 2400 labelled emails is explained and evaluated. The final features are tailored…

Information Retrieval · Computer Science 2022-05-13 Farshad Barahimi

Artificial intelligence systems increasingly generate text intended to provide social and emotional support. Understanding how users perceive empathic qualities in such content is therefore critical. We examined differences in perceived…

Computers and Society · Computer Science 2026-02-20 Jonas Festor , Ivo Snels , Bennett Kleinberg

Translating nuanced, textually-defined authorial writing styles into compelling visual representations presents a novel challenge in generative AI. This paper introduces a pipeline that leverages Author Writing Sheets (AWS) - structured…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Sagar Gandhi , Vishal Gandhi

How can we distinguish whether a peer review was written by a human or generated by an AI model? We argue that, in this setting, authorship should not be attributed solely from the textual features of a review, but also from the ideas,…

Computation and Language · Computer Science 2026-05-22 André V. Duarte , Brian Tufts , Aditya Oke , Fei Fang , Arlindo L. Oliveira , Lei Li

We present an approach to email filtering based on the suffix tree data structure. A method for the scoring of emails using the suffix tree is developed and a number of scoring and score normalisation functions are tested. Our results show…

Artificial Intelligence · Computer Science 2007-05-23 Rajesh M. Pampapathi , Boris Mirkin , Mark Levene

Authorship identification tasks, which rely heavily on linguistic styles, have always been an important part of Natural Language Understanding (NLU) research. While other tasks based on linguistic style understanding benefit from deep…

Computation and Language · Computer Science 2020-10-01 Weicheng Ma , Ruibo Liu , Lili Wang , Soroush Vosoughi

Computing with words (CWW) has emerged as a powerful tool for processing the linguistic information, especially the one generated by human beings. Various CWW approaches have emerged since the inception of CWW, such as perceptual computing,…

Artificial Intelligence · Computer Science 2020-05-01 Prashant K Gupta , Pranab K. Muhuri

The use of copyrighted books for training AI has sparked lawsuits from authors concerned about AI generating derivative content. Yet whether these models can produce high-quality literary text emulating authors' voices remains unclear. We…

Computation and Language · Computer Science 2026-03-18 Tuhin Chakrabarty , Jane C. Ginsburg , Paramveer Dhillon

Traditionally, authorship attribution (AA) tasks relied on statistical data analysis and classification based on stylistic features extracted from texts. In recent years, pre-trained language models (PLMs) have attracted significant…

Computation and Language · Computer Science 2025-04-14 Taisei Kanda , Mingzhe Jin , Wataru Zaitsu

To train algorithms for supervised author name disambiguation, many studies have relied on hand-labeled truth data that are very laborious to generate. This paper shows that labeled training data can be automatically generated using…

Digital Libraries · Computer Science 2021-02-08 Jinseok Kim , Jinmo Kim , Jason Owen-Smith

Using noisy crowdsourced labels from multiple annotators, a deep learning-based end-to-end (E2E) system aims to learn the label correction mechanism and the neural classifier simultaneously. To this end, many E2E systems concatenate the…

Machine Learning · Computer Science 2023-06-07 Shahana Ibrahim , Tri Nguyen , Xiao Fu

LLM-assisted writing has seen rapid adoption in interpersonal communication, yet current systems often fail to capture the subtle tones essential for effectiveness. Email writing exemplifies this challenge: effective messages require…

Human-Computer Interaction · Computer Science 2026-02-20 Rui Yao , Qiuyuan Ren , Felicia Fang-Yi Tan , Chen Yang , Xiaoyu Zhang , Shengdong Zhao

AA is the process of attributing an unidentified document to its true author from a predefined group of known candidates, each possessing multiple samples. The nature of AA necessitates accommodating emerging new authors, as each individual…

Information Retrieval · Computer Science 2024-08-20 Mostafa Rahgouy , Hamed Babaei Giglou , Mehnaz Tabassum , Dongji Feng , Amit Das , Taher Rahgooy , Gerry Dozier , Cheryl D. Seals

Authorial clustering involves the grouping of documents written by the same author or team of authors without any prior positive examples of an author's writing style or thematic preferences. For authorial clustering on shorter texts…

Computation and Language · Computer Science 2020-12-01 Rafi Trad , Myra Spiliopoulou

The problem of detecting phishing emails through machine learning techniques has been discussed extensively in the literature. Conventional and state-of-the-art machine learning algorithms have demonstrated the possibility of building…

Cryptography and Security · Computer Science 2021-01-01 Luis Felipe Gutiérrez , Faranak Abri , Miriam Armstrong , Akbar Siami Namin , Keith S. Jones