English
Related papers

Related papers: Critical biblical studies via word frequency analy…

200 papers

Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. This technique has multiple uses in tasks related to…

Disparities in authorship and citations across gender can have substantial adverse consequences not just on the disadvantaged genders, but also on the field of study as a whole. Measuring gender gaps is a crucial step towards addressing…

Digital Libraries · Computer Science 2020-09-07 Saif M. Mohammad

Author co-citation studies employ factor analysis to reduce high-dimensional co-citation matrices to low-dimensional and possibly interpretable factors, but these studies do not use any information from the text bodies of publications. We…

Digital Libraries · Computer Science 2013-05-08 Arnim Bleier , Andreas Strotmann

In this article we propose a novel method to estimate the frequency distribution of linguistic variables while controlling for statistical non-independence due to shared ancestry. Unlike previous approaches, our technique uses all available…

Populations and Evolution · Quantitative Biology 2021-03-22 Gerhard Jäger , Johannes Wahle

Text Categorization is traditionally done by using the term frequency and inverse document frequency.This type of method is not very good because, some words which are not so important may appear in the document .The term frequency of…

Information Retrieval · Computer Science 2016-11-25 Srikanth Bethu , G Charless Babu , J Vinoda , E Priyadarshini , M Raghavendra rao

Several complex systems are characterized by presenting intricate characteristics taking place at several scales of time and space. These multiscale characterizations are used in various applications, including better understanding…

Computation and Language · Computer Science 2023-05-12 Bárbara C. e Souza , Filipi N. Silva , Henrique F. de Arruda , Giovana D. da Silva , Luciano da F. Costa , Diego R. Amancio

Understanding the complexity of human language requires an appropriate analysis of the statistical distribution of words in texts. We consider the information retrieval problem of detecting and ranking the relevant words of a text by means…

Computation and Language · Computer Science 2008-06-07 Juan P. Herrera , Pedro A. Pury

Cultural products are a source to acquire individual values and behaviours. Therefore, the differences in the content of the magazines aimed specifically at women or men are a means to create and reproduce gender stereotypes. In this study,…

Computation and Language · Computer Science 2022-03-17 Diego Kozlowski , Gabriela Lozano , Carla M. Felcher , Fernando Gonzalez , Edgar Altszyler

The sequence of documents produced by any given author varies in style and content, but some documents are more typical or representative of the source than others. We quantify the extent to which a given short text is characteristic of a…

Computation and Language · Computer Science 2019-09-10 Charuta Pethe , Steven Skiena

A common task in computational text analyses is to quantify how two corpora differ according to a measurement like word frequency, sentiment, or information content. However, collapsing the texts' rich stories into a single number is often…

Measures of textual similarity and divergence are increasingly used to study cultural change. But which measures align, in practice, with social evidence about change? We apply three different representations of text (topic models, document…

Computation and Language · Computer Science 2024-11-25 Sarah Griebel , Becca Cohen , Lucian Li , Jaihyun Park , Jiayu Liu , Jana Perkins , Ted Underwood

This chapter argues for more informed target metrics for the statistical processing of stylistic variation in text collections. Much as operationalised relevance proved a useful goal to strive for in information retrieval, research in…

Computation and Language · Computer Science 2022-05-11 Jussi Karlgren

The paper explores stylometry as a method to distinguish between texts created by Large Language Models (LLMs) and humans, addressing issues of model attribution, intellectual property, and ethical AI use. Stylometry has been used…

Computation and Language · Computer Science 2025-07-25 Karol Przystalski , Jan K. Argasiński , Iwona Grabska-Gradzińska , Jeremi K. Ochab

The advent of instruction-tuned language models that convincingly mimic human writing poses a significant risk of abuse. However, such abuse may be counteracted with the ability to detect whether a piece of text was composed by a language…

Computation and Language · Computer Science 2024-05-09 Rafael Rivera Soto , Kailin Koch , Aleem Khan , Barry Chen , Marcus Bishop , Nicholas Andrews

Research has continued to shed light on the extent and significance of gender disparity in social, cultural and economic spheres. More recently, computational tools from the Natural Language Processing (NLP) literature have been proposed…

Computers and Society · Computer Science 2022-04-13 Akarsh Nagaraj , Mayank Kejriwal

Writing and reading are dynamic processes. As an author composes a text, a sequence of words is produced. This sequence is one that, the author hopes, causes a revisitation of certain thoughts and ideas in others. These processes of…

Computation and Language · Computer Science 2018-03-21 Rick Dale , Nicholas D. Duran , Moreno Coco

Authorship verification is the task of determining if two distinct writing samples share the same author and is typically concerned with the attribution of written text. In this paper, we explore the attribution of transcribed speech, which…

Computation and Language · Computer Science 2025-05-19 Cristina Aggazzotti , Nicholas Andrews , Elizabeth Allyn Smith

Recent advances in text mining and natural language processing technology have enabled researchers to detect an authors identity or demographic characteristics, such as age and gender, in several text genres by automatically analysing the…

Cryptography and Security · Computer Science 2022-11-30 Claudia Peersman , Matthew Edwards , Emma Williams , Awais Rashid

The problem of online threats and abuse could potentially be mitigated with a computational approach, where sources of abuse are better understood or identified through author profiling. However, abusive language constitutes a specific…

Computation and Language · Computer Science 2020-09-04 Isabelle van der Vegt , Bennett Kleinberg , Paul Gill

Inspired by the authorship controversy of Dream of the Red Chamber and the application of machine learning in the study of literary stylometry, we develop a rigorous new method for the mathematical analysis of authorship by testing for a…

Machine Learning · Computer Science 2014-12-22 Xianfeng Hu , Yang Wang , Qiang Wu