English
Related papers

Related papers: Classifying the typefaces of the Gutenberg 42-line…

200 papers

This paper presents a comparison of classification methods for linguistic typology for the purpose of expanding an extensive, but sparse language resource: the World Atlas of Language Structures (WALS) (Dryer and Haspelmath, 2013). We…

Computation and Language · Computer Science 2016-04-28 Reed Coke , Ben King , Dragomir Radev

ABCDE is a technique for evaluating clusterings of very large populations of items. Given two clusterings, namely a Baseline clustering and an Experiment clustering, ABCDE can characterize their differences with impact and quality metrics,…

Information Retrieval · Computer Science 2024-09-23 Stephan van Staden

I introduce a generic method for inference about a scalar parameter in research designs with a finite number of heterogeneous clusters where only a single cluster received treatment. This situation is commonplace in…

Econometrics · Economics 2020-10-09 Andreas Hagemann

The analysis of pseudo-colour diagrams, the so-called chromosome maps, of Galactic globular clusters (GCs) permits to classify them into type I and type II clusters. Type II GCs are characterized by an above-the-average complexity of their…

Astrophysics of Galaxies · Physics 2020-04-15 Matteo Simioni , Antonio Aparicio , Giampaolo Piotto

Graph classification aims to categorize graphs based on their structural and attribute features, with applications in diverse fields such as social network analysis and bioinformatics. Among the methods proposed to solve this task, those…

Machine Learning · Computer Science 2025-07-23 Lucas Potin , Rosa Figueiredo , Vincent Labatut , Christine Largeron

We apply ideas from the cluster method to q-count the permutations of a multiset according to the number of occurrences of certain generalized patterns, as defined by Babson and Steingrimsson. In particular, we consider those patterns with…

Combinatorics · Mathematics 2009-06-01 Andrew M. Baxter

In the prose style transfer task a system, provided with text input and a target prose style, produces output which preserves the meaning of the input text but alters the style. These systems require parallel data for evaluation of results…

Computation and Language · Computer Science 2021-09-01 Keith Carlson , Allen Riddell , Daniel Rockmore

Perfectly clustering words are one of many possible generalizations of Christoffel words. In this article, we propose a factorization of a perfectly clustering word on a $n$ letters alphabet into a product of $n-1$ palindromes with a letter…

Combinatorics · Mathematics 2024-07-30 Mélodie Lapointe , Christophe Reutenauer

Multiple-choice questions with item-writing flaws can negatively impact student learning and skew analytics. These flaws are often present in student-generated questions, making it difficult to assess their quality and suitability for…

Computation and Language · Computer Science 2023-07-18 Steven Moore , Huy A. Nguyen , Tianying Chen , John Stamper

We present the clustering of galaxy clusters as a useful addition to the common set of cosmological observables. The clustering of clusters probes the large-scale structure of the Universe, extending galaxy clustering analysis to the…

Cosmology and Nongalactic Astrophysics · Physics 2014-02-03 Annalisa Mana , Tommaso Giannantonio , Jochen Weller , Ben Hoyle , Gert Huetsi , Barbara Sartoris

Despite being a paradigm of quantitative linguistics, Zipf's law for words suffers from three main problems: its formulation is ambiguous, its validity has not been tested rigorously from a statistical point of view, and it has not been…

Applications · Statistics 2016-02-17 Isabel Moreno-Sánchez , Francesc Font-Clos , Álvaro Corral

The field of scientometrics has shown the power of citation-based clusters for literature analysis, yet this technique has barely been used for information retrieval tasks. This work evaluates the performance of citation based-clusters for…

Digital Libraries · Computer Science 2023-10-06 Juan Pablo Bascur , Suzan Verberne , Nees Jan van Eck , Ludo Waltman

In this paper we study a problem within Dempster-Shafer theory where 2**n - 1 pieces of evidence are clustered by a neural structure into n clusters. The clustering is done by minimizing a metaconflict function. Previously we developed a…

Artificial Intelligence · Computer Science 2007-05-23 Johan Schubert

In cancer research, clustering techniques are widely used for exploratory analyses and dimensionality reduction, playing a critical role in the identification of novel cancer subtypes, often with direct implications for patient management.…

Clustering has many important applications in computer science, but real-world datasets often contain outliers. Moreover, the presence of outliers can make the clustering problems to be much more challenging. To reduce the complexities,…

Data Structures and Algorithms · Computer Science 2020-05-04 Hu Ding , Jiawei Huang , Haikuo Yu

The Gibbs Mixing Paradox is a conceptual touchstone for understanding mixtures in statistical mechanics. While debates over the theoretical subtleties of particle distinguishability continue to this day, we seek to extend the discussion in…

Statistical Mechanics · Physics 2018-07-31 Cato Sandford , Daniel Seeto , Alexander Y. Grosberg

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

There are many different relatedness measures, based for instance on citation relations or textual similarity, that can be used to cluster scientific publications. We propose a principled methodology for evaluating the accuracy of…

Digital Libraries · Computer Science 2019-08-15 Ludo Waltman , Kevin W. Boyack , Giovanni Colavizza , Nees Jan van Eck

The evaluation of clustering algorithms can involve running them on a variety of benchmark problems, and comparing their outputs to the reference, ground-truth groupings provided by experts. Unfortunately, many research papers and graduate…

Machine Learning · Computer Science 2023-10-27 Marek Gagolewski

We present empirical data on misprints in citations to twelve high-profile papers. The great majority of misprints are identical to misprints in articles that earlier cited the same paper. The distribution of the numbers of misprint…

Physics and Society · Physics 2011-09-13 M. V. Simkin , V. P. Roychowdhury