English
Related papers

Related papers: Alberti's letter counts

200 papers

We present PhiloBERTA, a cross-lingual transformer model that measures semantic relationships between ancient Greek and Latin lexicons. Through analysis of selected term pairs from classical texts, we use contextual embeddings and angular…

Computation and Language · Computer Science 2025-08-26 Rumi Allbert , Makai L. Allbert

The extent to which men and women use language differently has been questioned previously. Finding clear and consistent gender differences in language is not conclusive in general, and the research is heavily influenced by the context and…

Computation and Language · Computer Science 2022-11-14 Md Zobaer Hossain , Ahnaf Mozib Samin

This study analyzes Walt Whitman's stylistic changes in his phenomenal work Leaves of Grass from a computational perspective and relates findings to standard literary criticism on Whitman. The corpus consists of all 7 editions of Leaves of…

Computation and Language · Computer Science 2021-11-11 Jieyan Zhu

Large language models (LLMs) are increasingly impacting human society, particularly in textual information. Based on more than 30,000 papers and 1,000 presentations from machine learning conferences, we examined and compared the words used…

Computation and Language · Computer Science 2024-10-23 Mingmeng Geng , Caixi Chen , Yanru Wu , Dongping Chen , Yao Wan , Pan Zhou

Prior studies have shown that distinguishing text generated by Large Language Models (LLMs) from human-written one is highly challenging for humans, and often no better than random guessing. To verify the generalizability of this finding…

Taylor's law describes the fluctuation characteristics underlying a system in which the variance of an event within a time span grows by a power law with respect to the mean. Although Taylor's law has been applied in many natural and social…

Computation and Language · Computer Science 2018-06-08 Tatsuru Kobayashi , Kumiko Tanaka-Ishii

In this paper we quantify the consistency of word usage in written texts represented by complex networks, where words were taken as nodes, by measuring the degree of preservation of the node neighborhood.} Words were considered highly…

Physics and Society · Physics 2013-02-19 Diego R. Amancio , Osvaldo N. Oliveira , Luciano da F. Costa

The use of programming languages can wax and wane across the decades. We examine the split-apply- combine pattern that is common in statistical computing, and consider how its invocation or implementation in languages like MATLAB and APL…

Programming Languages · Computer Science 2018-08-16 Jiahao Chen

The notion of number line was formed in XX c. We consider the generation of this conception in works by M. Stiefel (1544), Galilei (1633), Euler (1748), Lambert (1766), Bolzano (1830-1834), Meray (1869-1872), Cantor (1872), Dedekind (1872),…

History and Overview · Mathematics 2017-10-17 Galina Sinkevich

We present a mechanistic analysis of literary style in GPT-2, identifying individual neurons that discriminate between exemplary prose and rigid AI-generated text. Using Herman Melville's Bartleby, the Scrivener as a corpus, we extract…

Computation and Language · Computer Science 2025-10-22 Tsogt-Ochir Enkhbayar

There are different ways of measuring diversity in complex systems. In particular, in language, lexical diversity is characterized in terms of the type-token ratio and the word entropy. We here investigate both diversity metrics in six…

Computation and Language · Computer Science 2025-07-16 Pablo Rosillo-Rodes , Maxi San Miguel , David Sanchez

We prove that, for any pure morphic word $w$, if the frequencies of all letters in $w$ exist, then the frequencies of all factors in $w$ exist as well. This result answers a question of Saari in his doctoral thesis.

Combinatorics · Mathematics 2024-05-30 Shuo Li

Zipf's law of abbreviation, namely the tendency of more frequent words to be shorter, has been viewed as a manifestation of compression, i.e. the minimization of the length of forms -- a universal principle of natural communication.…

Computation and Language · Computer Science 2026-03-31 Sonia Petrini , Antoni Casas-i-Muñoz , Jordi Cluet-i-Martinell , Mengxue Wang , Christian Bentz , Ramon Ferrer-i-Cancho

Large language models (LLMs) are capable of writing grammatical text that follows instructions, answers questions, and solves problems. As they have advanced, it has become difficult to distinguish their output from human-written text.…

Computation and Language · Computer Science 2025-08-25 Alex Reinhart , Ben Markey , Michael Laudenbach , Kachatad Pantusen , Ronald Yurko , Gordon Weinberg , David West Brown

Recent work has demonstrated that language models can be trained to identify the author of much shorter literary passages than has been thought feasible for traditional stylometry. We replicate these results for authorship and extend them…

Computation and Language · Computer Science 2025-02-07 Rebecca M. M. Hicke , David Mimno

The problem of comparing two bodies of text and searching for words that differ in their usage between them arises often in digital humanities and computational social science. This is commonly approached by training word embeddings on each…

Computation and Language · Computer Science 2021-12-30 Hila Gonen , Ganesh Jawahar , Djamé Seddah , Yoav Goldberg

This study examines phonetic texture in Persian poetry as a literary-historical phenomenon rather than a by-product of meter or a feature used only for classification. The analysis draws on a large corpus of 1,116,306 mesras from 31,988…

Computation and Language · Computer Science 2026-04-02 Kourosh Shahnazari , Seyed Moein Ayyoubzadeh , Mohammadali Keshtparvar

Zipf's law of abbreviation, the tendency of more frequent words to be shorter, is one of the most solid candidates for a linguistic universal, in the sense that it has the potential for being exceptionless or with a number of exceptions…

Computation and Language · Computer Science 2023-10-13 Sonia Petrini , Antoni Casas-i-Muñoz , Jordi Cluet-i-Martinell , Mengxue Wang , Chris Bentz , Ramon Ferrer-i-Cancho

This study investigates the potential influence of Herman Melville reading on his own writings through computational semantic similarity analysis. Using documented records of books known to have been owned or read by Melville, we compare…

Computation and Language · Computer Science 2026-03-17 Nudrat Habib , Elisa Barney Smith , Steven Olsen Smith

This article contains ideas and their elaboration for quantifiers, which appeared after checking in practice the experimental language of the formal knowledge representation YAFOLL [1]: - looking at for_all and exists quantifiers as…

Logic in Computer Science · Computer Science 2019-08-30 Alex Shkotin