Related papers: Invariance of generalized wordlength patterns
Compositional generalization allows efficient learning and human-like inductive biases. Since most research investigating compositional generalization in NLP is done on English, important questions remain underexplored. Do the necessary…
This essay proposes an interpretive analogy between large language models (LLMs) and quasicrystals, systems that exhibit global coherence without periodic repetition, generated through local constraints. While LLMs are typically evaluated…
We introduce "representative generation," extending the theoretical framework for generation proposed by Kleinberg et al. (2024) and formalized by Li et al. (2024), to additionally address diversity and bias concerns in generative models.…
In the application of autoregressive models the order of the model is often estimated using either a sequence of likelihood ratio tests, a likelihood based information criterion, or a residual based test. The properties of such procedures…
In language learning in the limit, the most common type of hypothesis is to give an enumerator for a language. This so-called $W$-index allows for naming arbitrary computably enumerable languages, with the drawback that even the membership…
While the use of cluster features became ubiquitous in core NLP tasks, most cluster features in NLP are based on distributional similarity. We propose a new type of clustering criteria, specific to the task of part-of-speech tagging.…
Designing models that are both expressive and preserve known invariances of tasks is an increasingly hard problem. Existing solutions tradeoff invariance for computational or memory resources. In this work, we show how to leverage…
Let $L(G)$ denote the space of integer-valued length functions on a countable group $G$ endowed with the topology of pointwise convergence. Assuming that $G$ does not satisfy any non-trivial mixed identity, we prove that a generic (in the…
We consider continuous-time models with a large panel of moment conditions, where the structural parameter depends on a set of characteristics, whose effects are of interest. The leading example is the linear factor model in financial…
The width $\wid(G,W)$ of the verbal subgroup $v(G,W)$ of a group $G$ defined by a collection of group words $W$ is the smallest number $m$ in $\mathbb N \cup {+\infty}$ such that every element of $v(G,W)$ is can be represented as the…
The Johnson-Lindenstrauss (JL) theorem states that a set of points in high-dimensional space can be embedded into a lower-dimensional space while approximately preserving pairwise distances with high probability Johnson and Lindenstrauss…
The point-line geometry known as a \textit{partial quadrangle} (introduced by Cameron in 1975) has the property that for every point/line non-incident pair $(P,\ell)$, there is at most one line through $P$ concurrent with $\ell$. So in…
We define matrix groups $FG_n(P)$ for each natural number $n$ and finite set of primes $P$, such that every rational-valued upper triangular matrix group is a (possibly distorted) subgroup. Brofferio and Schapira [Brofferio2011poisson],…
About forty years ago it was realized by several researchers that the essential features of certain objects of Probability theory, notably Gaussian processes and limit theorems, may be better understood if they are considered in settings…
This paper defines unification based ID/LP grammars based on typed feature structures as nonterminals and proposes a variant of Earley's algorithm to decide whether a given input sentence is a member of the language generated by a…
Word segmentation is a low-level NLP task that is non-trivial for a considerable number of languages. In this paper, we present a sequence tagging framework and apply it to word segmentation for a wide range of languages with different…
Recently, a new characterization of Lyndon words that are also perfectly clustering was proposed by Lapointe and Reutenauer (2024). A word over a ternary alphabet {a,b,c} is called perfectly clustering Lyndon if and only if it is the…
We show that the algorithm for computing an element of the Clarke generalized Jacobian of a max-type function proposed by Zheng-da Huang and Guo-chun Ma can be extended to a much wider class of functions representable as a difference of…
We introduce a stochastic process with Wishart marginals: the generalised Wishart process (GWP). It is a collection of positive semi-definite random matrices indexed by any arbitrary dependent variable. We use it to model dynamic (e.g. time…
There has been widespread use of causal inference methods for the rigorous analysis of observational studies and to identify policy evaluations. In this article, we consider a class of generalized coarsened procedures for confounding. At a…