English
Related papers

Related papers: Datasets and Models for Authorship Attribution on …

200 papers

This study aims to evaluate the accuracy of authorship attributions in scientific publications, focusing on the fairness and precision of individual contributions within academic works. The study analyzes 81,823 publications from the…

Digital Libraries · Computer Science 2025-04-25 Abdelghani Maddi , Jaime A. Teixeira da Silva

In this work we develop a series of techniques to quantify the presence of bias and censorship in newspapers. These algorithms are tested analyzing the occurrence of keywords `killed' and `suicide' (`morti', `suicidio' in Italian) and their…

Physics and Society · Physics 2020-07-29 M. Casolino

Humans are naturally endowed with the ability to write in a particular style. They can, for instance, re-phrase a formal letter in an informal way, convey a literal message with the use of figures of speech or edit a novel by mimicking the…

Computation and Language · Computer Science 2022-09-30 Enrica Troiano , Aswathy Velutharambath , Roman Klinger

Translating from languages without productive grammatical gender like English into gender-marked languages is a well-known difficulty for machines. This difficulty is also due to the fact that the training data on which models are built…

Computation and Language · Computer Science 2020-06-11 Luisa Bentivogli , Beatrice Savoldi , Matteo Negri , Mattia Antonino Di Gangi , Roldano Cattoni , Marco Turchi

We describe a technique for attributing parts of a written text to a set of unknown authors. Nothing is assumed to be known a priori about the writing styles of potential authors. We use multiple independent clusterings of an input text to…

Computation and Language · Computer Science 2015-03-27 David Fifield , Torbjørn Follan , Emil Lunde

As human-AI collaboration becomes increasingly prevalent in educational contexts, understanding and measuring the extent and nature of such interactions pose significant challenges. This research investigates the use of authorship…

Computation and Language · Computer Science 2025-09-09 Eduardo Araujo Oliveira , Madhavi Mohoni , Sonsoles López-Pernas , Mohammed Saqr

Authorship identification has proven unsettlingly effective in inferring the identity of the author of an unsigned document, even when sensitive personal information has been carefully omitted. In the digital era, individuals leave a…

Computation and Language · Computer Science 2023-10-04 Haining Wang

We are addressing two fundamental problems in authorship verification (AV): Topic variability and miscalibration. Variations in the topic of two disputed texts are a major cause of error for most AV systems. In addition, it is observed that…

Computation and Language · Computer Science 2021-06-22 Benedikt Boenninghoff , Dorothea Kolossa , Robert M. Nickel

Information access research (and development) sometimes makes use of gender, whether to report on the demographics of participants in a user study, as inputs to personalized results or recommendations, or to make systems gender-fair,…

Information Retrieval · Computer Science 2023-01-18 Christine Pinney , Amifa Raj , Alex Hanna , Michael D. Ekstrand

Large language models (LLMs) can convincingly imitate human writing styles, yet it remains unclear how much stylistic information is encoded in embeddings from any language model and retained after LLM rewriting. We investigate these…

Computation and Language · Computer Science 2026-05-12 Benjamin Icard , Lila Sainero , Alice Breton , Evangelia Zve , Jean-Gabriel Ganascia

Documents as short as a single sentence may inadvertently reveal sensitive information about their authors, including e.g. their gender or ethnicity. Style transfer is an effective way of transforming texts in order to remove any…

Computation and Language · Computer Science 2021-09-21 David Ifeoluwa Adelani , Miaoran Zhang , Xiaoyu Shen , Ali Davody , Thomas Kleinbauer , Dietrich Klakow

Spoken language datasets are vital for advancing linguistic research, Natural Language Processing, and speech technology. However, resources dedicated to Italian, a linguistically rich and diverse Romance language, remain underexplored…

Computation and Language · Computer Science 2025-03-13 Marco Giordano , Claudia Rinaldi

Authorship attribution aims to identify the author of a text based on the stylometric analysis. Authorship obfuscation, on the other hand, aims to protect against authorship attribution by modifying a text's style. In this paper, we…

Computation and Language · Computer Science 2020-05-05 Asad Mahmood , Zubair Shafiq , Padmini Srinivasan

Authorship verification (AV), the task of determining whether a questioned text was written by a specific individual, is a critical part of forensic linguistics. While manual authorial impersonation by perpetrators has long been a…

Computation and Language · Computer Science 2026-04-01 Baoyi Zeng , Andrea Nini

Retrieving indexed documents, not by their topical content but their writing style opens the door for a number of applications in information retrieval (IR). One application is to retrieve textual content of a certain author X, where the…

Information Retrieval · Computer Science 2019-01-03 Oren Halvani , Christian Winter , Lukas Graner

Writing style is a combination of consistent decisions associated with a specific author at different levels of language production, including lexical, syntactic, and structural. In this paper, we introduce a style-aware neural model to…

Computation and Language · Computer Science 2019-09-16 Fereshteh Jafariakinabad , Kien A. Hua

Critical scholarship has elevated the problem of gender bias in data sets used to train virtual assistants (VAs). Most work has focused on explicit biases in language, especially against women, girls, femme-identifying people, and…

Computation and Language · Computer Science 2023-04-26 Katie Seaborn , Shruti Chandra , Thibault Fabre

Most of privacy protection studies for textual data focus on removing explicit sensitive identifiers. However, personal writing style, as a strong indicator of the authorship, is often neglected. Recent studies, such as SynTF, have shown…

Cryptography and Security · Computer Science 2021-05-14 Haohan Bo , Steven H. H. Ding , Benjamin C. M. Fung , Farkhund Iqbal

In practice, training language models for individual authors is often expensive because of limited data resources. In such cases, Neural Network Language Models (NNLMs), generally outperform the traditional non-parametric N-gram models.…

Computation and Language · Computer Science 2016-02-18 Zhenhao Ge , Yufang Sun , Mark J. T. Smith

Identity documents automatic reading and verification is an appealing technology for nowadays service industry, since this task is still mostly performed manually, leading to waste of economic and time resources. In this paper the prototype…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Filippo Attivissimo , Nicola Giaquinto , Marco Scarpetta , Maurizio Spadavecchia
‹ Prev 1 4 5 6 7 8 10 Next ›