English
Related papers

Related papers: Co-Occurrence Patterns in the Voynich Manuscript

200 papers

In the past several decades, many authorship attribution studies have used computational methods to determine the authors of disputed texts. Disputed authorship is a common problem in Classics, since little information about ancient…

Computation and Language · Computer Science 2017-11-07 Anjalie Field

Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched…

Computation and Language · Computer Science 2020-07-24 Sunayana Sitaram , Khyathi Raghavi Chandu , Sai Krishna Rallabandi , Alan W Black

Unlike static documents, version-controlled documents are edited by one or more authors over a certain period of time. Examples include large scale computer code, papers authored by a team of scientists, and online discussion boards. Such…

Human-Computer Interaction · Computer Science 2012-05-16 Seungyeon Kim , Joshua V. Dillon , Guy Lebanon

A universal cycle, or u-cycle, for a given set of words is a circular word that contains each word from the set exactly once as a contiguous subword. The celebrated de Bruijn sequences are a particular case of such a u-cycle, where a set in…

Combinatorics · Mathematics 2019-08-06 Herman Z. Q. Chen , Sergey Kitaev , Brian Y. Sun

A double occurrence word $w$ over a finite alphabet $\Sigma$ is a word in which each alphabet letter appears exactly twice. Such words arise naturally in the study of topology, graph theory, and combinatorics. Recently, double occurrence…

Combinatorics · Mathematics 2012-05-01 Jonathan Burns , Tilahun Muche

We present an algorithm that takes an unannotated corpus as its input, and returns a ranked list of probable morphologically related pairs as its output. The algorithm tries to discover morphologically related pairs by looking for pairs…

Computation and Language · Computer Science 2007-05-23 Marco Baroni , Johannes Matiasek , Harald Trost

Word embeddings are a popular way to improve downstream performances in contemporary language modeling. However, the underlying geometric structure of the embedding space is not well understood. We present a series of explorations using…

Computation and Language · Computer Science 2020-09-17 Hongwei , Zhou , Oskar Elek , Pranav Anand , Angus G. Forbes

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres,…

Computation and Language · Computer Science 2017-05-25 Jeremy Ferrero , Laurent Besacier , Didier Schwab , Frederic Agnes

Dynamical systems on the interval were widely studied because they are among the simplest systems and nevertheless they turn out to have complex dynamics. Many works on chaos were inspired by the behaviour of interval maps. However these…

Dynamical Systems · Mathematics 2018-04-13 Sylvie Ruette

This paper deals with the task of practical and open source Handwritten Text Recognition (HTR) on German medieval manuscripts. We report on our efforts to construct mixed recognition models which can be applied out-of-the-box without any…

Computer Vision and Pattern Recognition · Computer Science 2022-01-20 Christian Reul , Stefan Tomasek , Florian Langhanki , Uwe Springmann

Certain upper triangular matrices, termed as Parikh matrices, are often used in the combinatorial study of words. Given a word, the Parikh matrix of that word elegantly computes the number of occurrences of certain predefined subwords in…

Combinatorics · Mathematics 2018-08-14 Adrian Atanasiu , Ghajendran Poovanandran , Wen Chean Teh

_Uncertainty expressions_ such as "probably" or "highly unlikely" are pervasive in human language. While prior work has established that there is population-level agreement in terms of how humans quantitatively interpret these expressions,…

Computation and Language · Computer Science 2024-11-08 Catarina G Belem , Markelle Kelly , Mark Steyvers , Sameer Singh , Padhraic Smyth

Understanding how ideas relate to each other is a fundamental question in many domains, ranging from intellectual history to public communication. Because ideas are naturally embedded in texts, we propose the first framework to…

Social and Information Networks · Computer Science 2017-07-18 Chenhao Tan , Dallas Card , Noah A. Smith

We present the new empirical parameter $f_c$, the most probable usage frequency of a word in a language, computed via the distribution of documents over frequency $x$ of the word. This parameter allows for filtering the core lexicon of a…

Disordered Systems and Neural Networks · Physics 2007-05-23 Dmitri Volchenkov , Philippe Blanchard , Serge Sharoff

Statistical analysis of repeat misprints in scientific citations leads to the conclusion that about 80% of scientific citations are copied from the lists of references used in othe papers. Based on this finding a mathematical theory of…

Statistics Theory · Mathematics 2007-06-13 M. V. Simkin , V. P. Roychowdhury

This article focuses on the transcription of medieval manuscripts. Whereas problems of transcription have long interested medievalists, few workable options in the era of printed editions were available besides normalisation. The automation…

Digital Libraries · Computer Science 2024-08-07 Estelle Guéville , David Joseph Wrisley

Several years ago, one of us, having noticed that inexperienced scientists tend to make largely the same mistakes while writing their first papers, was compelled to write a one-page note summarizing some dos and don'ts intended to help take…

Physics Education · Physics 2016-07-12 Dmitry Budker , Derek F. Jackson Kimball

Word matches are often used in sequence comparison methods, either as a measure of sequence similarity or in the first search steps of algorithms such as BLAST or BLAT. The D2 statistic is the number of matches of words of k letters between…

Quantitative Methods · Quantitative Biology 2009-09-09 Sylvain Foret , Susan R. Wilson , Conrad J. Burden

The absence of standardized spelling conventions and the organic evolution of human language present an inherent linguistic challenge within historical documents, a longstanding concern for scholars in the humanities. Addressing this issue,…

Computation and Language · Computer Science 2025-07-01 Miguel Domingo , Francisco Casacuberta

Texts exhibit considerable stylistic variation. This paper reports an experiment where a corpus of documents (N= 75 000) is analyzed using various simple stylistic metrics. A subset (n = 1000) of the corpus has been previously assessed to…

cmp-lg · Computer Science 2008-02-03 Jussi Karlgren