Related papers: Enumerating Lambda Terms by Weighted Length of The…
Denote by $\lambda(n)$ Liouville's function concerning the parity of the number of prime divisors of $n$. Using a theorem of Allouche, Mend\`es France, and Peyri\`ere and many classical results from the theory of the distribution of prime…
Safety is a syntactic condition of higher-order grammars that constrains occurrences of variables in the production rules according to their type-theoretic order. In this paper, we introduce the safe lambda calculus, which is obtained by…
Large language models (LLMs) with billions of parameters excel at predicting the next token in a sequence. Recent work computes non-vacuous compression-based generalization bounds for LLMs, but these bounds are vacuous for large models at…
This is the first part of a series of four articles. In this work, we are interested in weighted norm estimates. We put the emphasis on two results of different nature: one is based on a good-$\lambda$ inequality with two-parameters and the…
Let $\lambda$ and $\mu$ denote the Liouville and M\"obius functions respectively. Hildebrand showed that all eight possible sign patterns for $(\lambda(n), \lambda(n+1), \lambda(n+2))$ occur infinitely often. By using the recent result of…
Word embeddings trained on large corpora have shown to encode high levels of unfair discriminatory gender, racial, religious and ethnic biases. In contrast, human-written dictionaries describe the meanings of words in a concise, objective…
We show that, in an alphabet of $n$ symbols, the number of words of length $n$ whose number of different symbols is away from $(1-1/e)n$, which is the value expected by the Poisson distribution, has exponential decay in $n$. We use…
In this work we provide alternative formulations of the concepts of lambda theory and extensional theory without introducing the notion of substitution and the sets of all, free and bound variables occurring in a term. We also clarify the…
Extending the lambda-calculus with a construct for sharing, such as let expressions, enables a special representation of terms: iterated applications are decomposed by introducing sharing points in between any two of them, reducing to the…
We determine the complexity of counting models of bounded size of specifications expressed in Linear-time Temporal Logic. Counting word models is #P-complete, if the bound is given in unary, and as hard as counting accepting runs of…
This paper addresses the problem of mapping natural language sentences to lambda-calculus encodings of their meaning. We describe a learning algorithm that takes as input a training set of sentences labeled with expressions in the lambda…
The ability (and inability) of large language models (LLMs) to perform arithmetic tasks has been the subject of much theoretical and practical debate. We show that LLMs are frequently able to correctly and confidently predict the first…
We study properties of a sequence $\Lambda$ obtained by a randomselection of integers $n$, where $n\in\Lambda$ with probability $\varpi_{n}$, independently of the other choices. We distinguish two cases : if…
It is common to model inductive datatypes as least fixed points of functors. We show that within the Cedille type theory we can relax functoriality constraints and generically derive an induction principle for Mendler-style lambda-encoded…
We introduce a call-by-name lambda-calculus $\lambda Jn$ with generalized applications which is equipped with distant reduction. This allows to unblock $\beta$-redexes without resorting to the standard permutative conversions of generalized…
Binary time series data are very common in many applications, and are typically modelled independently via a Bernoulli process with a single probability of success. However, the probability of a success can be dependent on the outcome…
We investigate the properties of a Block Decomposition Method (BDM), which extends the power of a Coding Theorem Method (CTM) that approximates local estimations of algorithmic complexity based upon Solomonoff-Levin's theory of algorithmic…
If $\gcd(r,t)=1$, then a theorem of Alladi offers the M\"obius sum identity $$-\sum_{\substack{ n \geq 2 \\ p_{\rm{min}}(n) \equiv r \pmod{t}}} \mu(n)n^{-1}= \frac{1}{\varphi(t)}. $$ Here $p_{\rm{min}}(n)$ is the smallest prime divisor of…
Word embeddings are commonly used as a starting point in many NLP models to achieve state-of-the-art performances. However, with a large vocabulary and many dimensions, these floating-point representations are expensive both in terms of…
For a word $\pi$ and integer $i$, we define $L^i(\pi)$ to be the length of the longest subsequence of the form $i(i+1)\cdots j$, and we let $L(\pi):=\max_i L^i(\pi)$. In this paper we estimate the expected values of $L^1(\pi)$ and $L(\pi)$…