English
Related papers

Related papers: On the Entropy of Written Spanish

200 papers

Word embeddings are real-valued word representations able to capture lexical semantics and trained on natural language corpora. Models proposing these representations have gained popularity in the recent years, but the issue of the most…

Computation and Language · Computer Science 2018-01-30 Amir Bakarov

This work describes the language resources and models developed for automatic simplification of Spanish texts in three domains: Finance, Medicine and History studies. We created several corpora in each domain, annotation and simplification…

Computation and Language · Computer Science 2024-10-01 Antonio Moreno-Sandoval , Leonardo Campillos-Llanos , Ana García-Serrano

Grammar compression represents a string as a context free grammar. Achieving compression requires encoding such grammar as a binary string; there are a few commonly used encodings. We bound the size of practically used encodings for several…

Data Structures and Algorithms · Computer Science 2020-05-21 Michał Gańczorz

What happens when a new social convention replaces an old one? While the possible forces favoring norm change - such as institutions or committed activists - have been identified since a long time, little is known about how a population…

Physics and Society · Physics 2018-09-13 Roberta Amato , Lucas Lacasa , Albert Díaz-Guilera , Andrea Baronchelli

We propose a compression-based version of the empirical entropy of a finite string over a finite alphabet. Whereas previously one considers the naked entropy of (possibly higher order) Markov processes, we consider the sum of the…

Information Theory · Computer Science 2011-04-05 Paul M. B. Vitányi

The code base of software projects evolves essentially through inserting and removing information to and from the source code. We can measure this evolution via the elements of information - tokens, words, nodes - of the respective…

Software Engineering · Computer Science 2025-06-10 Adriano Torres , Sebastian Baltes , Christoph Treude , Markus Wagner

During a conversation, our brain is responsible for combining information obtained from multiple senses in order to improve our ability to understand the message we are perceiving. Different studies have shown the importance of presenting…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle with known concepts presented in novel contexts. Although…

Computation and Language · Computer Science 2025-05-28 Sondre Wold , Lucas Georges Gabriel Charpentier , Étienne Simon

In this paper I propose a new way of measuring linguistic productivity that objectively assesses the ability of an affix to be used to coin new complex words and, unlike other popular measures, is not directly dependent upon token…

Computation and Language · Computer Science 2023-08-25 Sergei Monakhov

The number of senses of a given word, or polysemy, is a very subjective notion, which varies widely across annotators and resources. We propose a novel method to estimate polysemy, based on simple geometry in the contextual embedding space.…

Computation and Language · Computer Science 2023-05-03 Christos Xypolopoulos , Antoine J. -P. Tixier , Michalis Vazirgiannis

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data to be (pre-)trained and there is a lack of corpora in…

Computation and Language · Computer Science 2022-07-04 Asier Gutiérrez-Fandiño , David Pérez-Fernández , Jordi Armengol-Estapé , David Griol , Zoraida Callejas

Classification is a machine learning method used in many practical applications: text mining, handwritten character recognition, face recognition, pattern classification, scene labeling, computer vision, natural langage processing. A…

Machine Learning · Computer Science 2025-11-05 Doulaye Dembélé

The World Wide Web has grown so big, in such an anarchic fashion, that it is difficult to describe. One of the evident intrinsic characteristics of the World Wide Web is its multilinguality. Here, we present a technique for estimating the…

Computation and Language · Computer Science 2021-08-23 Gregory Grefenstette , Julien Nioche

The paper explores how the human natural language structure can be seen as a product of evolution of inter-personal communication code, targeting maximisation of such culture-agnostic and cross-lingual metrics such as anti-entropy,…

Computation and Language · Computer Science 2023-06-13 Anton Kolonin

In this paper we derived a variational principle for the specific entropy on the context of symbolic dynamics of compact metric space alphabets and use this result to obtain the uniqueness of the equilibrium states associated to a Walters…

Dynamical Systems · Mathematics 2018-06-11 D. Aguiar , L. Cioletti , R. Ruviaro

The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglicisms and a baseline…

Computation and Language · Computer Science 2020-04-08 Elena Álvarez-Mellado

We present a theoretical and empirical investigation of the statistical behaviour of the words in a text produced by human language. To this aim, we analyse the word distribution of various texts of Italian language selected from a specific…

Neurons and Cognition · Quantitative Biology 2025-04-15 Diederik Aerts , Jonito Aerts Arguëlles , Lester Beltran , Massimiliano Sassoli de Bianchi , Sandro Sozzo

Word complexity is defined in a number of different ways. Psycholinguistic, morphological and lexical proxies are often used. Human ratings are also used. The problem here is that these proxies do not measure complexity directly, and human…

Computation and Language · Computer Science 2024-08-06 Michael Dalvean

To transcribe speech, automatic speech recognition systems use statistical methods, particularly hidden Markov model and N-gram models. Although these techniques perform well and lead to efficient systems, they approach their maximum…

Human-Computer Interaction · Computer Science 2016-08-16 Stéphane Huet , Pascale Sébillot , Guillaume Gravier

Natural languages exhibit striking regularities in their statistical structure, including notably the emergence of Zipf's and Heaps' laws. Despite this, it remains broadly unclear how these properties relate to the modern tokenisation…

Computation and Language · Computer Science 2026-01-08 David S. Berman , Alexander G. Stapleton
‹ Prev 1 8 9 10 Next ›