English
Related papers

Related papers: The distribution of syntactic dependency distances

200 papers

We assign binary and ternary error-correcting codes to the data of syntactic structures of world languages and we study the distribution of code points in the space of code parameters. We show that, while most codes populate the lower…

Computation and Language · Computer Science 2016-10-04 Kevin Shu , Matilde Marcolli

The structure of all the permutations of a sequence can be represented as a permutohedron, a graph where vertices are permutations and two vertices are linked if a swap of adjacent elements in the permutation of one of the vertices produces…

Computation and Language · Computer Science 2026-05-14 Ramon Ferrer-i-Cancho

The recent dramatic increase in online data availability has allowed researchers to explore human culture with unprecedented detail, such as the growth and diversification of language. In particular, it provides statistical tools to explore…

In a physical system, changing parameters such as temperature can induce a phase transition: an abrupt change from one state of matter to another. Analogous phenomena have recently been observed in large language models. Typically, the task…

Machine Learning · Computer Science 2024-05-28 Julian Arnold , Flemming Holtorf , Frank Schäfer , Niels Lörch

We study a discrete-time duplication-deletion random graph model and analyse its asymptotic degree distribution. The random graphs consists of disjoint cliques. In each time step either a new vertex is brought in with probability $0<p<1$…

Probability · Mathematics 2017-02-24 Erik Thörnblad

According to the small-world concept, the entire world is connected through short chains of acquaintances. In popular imagination this is captured in the phrase six degrees of separation, implying that any two individuals are, at most, six…

Social and Information Networks · Computer Science 2010-08-20 Prerana Laddha

Sentence formation is a highly structured, history-dependent, and sample-space reducing (SSR) process. While the first word in a sentence can be chosen from the entire vocabulary, typically, the freedom of choosing subsequent words gets…

Computation and Language · Computer Science 2018-12-31 Rudolf Hanel , Stefan Thurner

This work lists and describes the main recent strategies for building fixed-length, dense and distributed representations for words, based on the distributional hypothesis. These representations are now commonly called word embeddings and,…

Computation and Language · Computer Science 2023-05-03 Felipe Almeida , Geraldo Xexéo

A dynamic model of a society is studied where each person is an uncorrelated and non-interacting random walker. A dynamical random graph represents the acquaintance network of the society whose nodes are the individuals and links are the…

Statistical Mechanics · Physics 2009-11-10 S. S. Manna

We explore the link between the extent to which syntactic relations are preserved in translation and the ease of correctly constructing a parse tree in a zero-shot setting. While previous work suggests such a relation, it tends to focus on…

Computation and Language · Computer Science 2021-10-12 Ofir Arviv , Dmitry Nikolaev , Taelin Karidi , Omri Abend

We study the Susceptible-Infected-Recovered model in complex networks, considering that not all individuals in the population interact in the same way between them. This heterogeneity between contacts is modeled by a continuous disorder. In…

Physics and Society · Physics 2013-01-01 C. Buono , C. Lagorio , P. A. Macri , L. A. Braunstein

Statistical studies of languages have focused on the rank-frequency distribution of words. Instead, we introduce here a measure of how word ranks change in time and call this distribution \emph{rank diversity}. We calculate this diversity…

Computation and Language · Computer Science 2015-05-15 Germinal Cocho , Jorge Flores , Carlos Gershenson , Carlos Pineda , Sergio Sánchez

In this article we present a multivariate model for determining the different syntactic, semantic, and form (surface-structure) processes underlying the comprehension of simple phrases. This model is applied to EEG signals recorded during a…

Computation and Language · Computer Science 2018-04-17 Sabine Ploux , Viviane Déprez

This paper presents a new approach for measuring semantic similarity/distance between words and concepts. It combines a lexical taxonomy structure with corpus statistical information so that the semantic distance between nodes in the…

cmp-lg · Computer Science 2008-02-03 Jay J. Jiang , David W. Conrath

Distributed representations of sentences have been developed recently to represent their meaning as real-valued vectors. However, it is not clear how much information such representations retain about the polarity of sentences. To study…

Computation and Language · Computer Science 2017-09-07 Edoardo Maria Ponti , Ivan Vulić , Anna Korhonen

This thesis presents a broad-coverage probabilistic top-down parser, and its application to the problem of language modeling for speech recognition. The parser builds fully connected derivations incrementally, in a single pass from…

Computation and Language · Computer Science 2007-05-23 Brian Roark

Natural languages are full of rules and exceptions. One of the most famous quantitative rules is Zipf's law which states that the frequency of occurrence of a word is approximately inversely proportional to its rank. Though this `law' of…

Computation and Language · Computer Science 2015-05-27 Jake Ryland Williams , James P. Bagrow , Christopher M. Danforth , Peter Sheridan Dodds

We study the fractal structure of language, aiming to provide a precise formalism for quantifying properties that may have been previously suspected but not formally shown. We establish that language is: (1) self-similar, exhibiting…

Computation and Language · Computer Science 2024-05-24 Ibrahim Alabdulmohsin , Vinh Q. Tran , Mostafa Dehghani

Several studies on sentence processing suggest that the mental lexicon keeps track of the mutual expectations between words. Current DSMs, however, represent context words as separate features, thereby loosing important information for word…

Computation and Language · Computer Science 2016-10-06 Emmanuele Chersoni , Enrico Santus , Alessandro Lenci , Philippe Blache , Chu-Ren Huang

Most of the syntax-based metrics obtain the similarity by comparing the sub-structures extracted from the trees of hypothesis and reference. These sub-structures are defined by human and can't express all the information in the trees…

Computation and Language · Computer Science 2016-11-07 Hui Yu , Xiaofeng Wu , Wenbin Jiang , Qun Liu , ShouXun Lin