Related papers: Multiplicative Turing Ensembles, Pareto's Law, and…
Infinite time Turing machines (ITTMs) have been introduced by Hamkins and Lewis in their seminal article arXiv:math/9808093. The strength of the model comes from a limit rule which allows the ITTM to compute through ordinal stages. This…
In this paper, we propose a novel maximum causal Tsallis entropy (MCTE) framework for imitation learning which can efficiently learn a sparse multi-modal policy distribution from demonstrations. We provide the full mathematical analysis of…
The distribution of income and wealth in developed economies exhibits a robust two-class structure: an exponential (Boltzmann--Gibbs) bulk covering $\sim\!97\%$ of the population, and a power-law (Pareto) tail in the upper $\sim\!3\%$. We…
Using the Generalized Maximium Entropy Principle based on the nonextensive q entropy a new family of random matrix ensembles is generated. This family unifies previous extensions of Random Matrix Theory and gives rise to an orthogonal…
We study, from the viewpoint of metrical number theory and (infinite) ergodic theory, the probabilistic laws governing the occurrence of prime numbers as digits in continued fraction expansions of real numbers.
Topological entanglement entropy (TEE), the sub-leading term in the entanglement entropy of topological order, is the direct evidence of the long-range entanglement. While effective in characterizing topological orders on closed manifolds,…
Probabilistic graphical models (PGMs) are widely used to discover latent structure in data, but their success hinges on selecting an appropriate model design. In practice, model specification is difficult and often requires iterative…
The optimal mixing evolutionary algorithms (OMEAs) have recently drawn much attention for their robustness, small size of required population, and efficiency in terms of number of function evaluations (NFE). In this paper, the performances…
We describe $k$-MLE, a fast and efficient local search algorithm for learning finite statistical mixtures of exponential families such as Gaussian mixture models. Mixture models are traditionally learned using the expectation-maximization…
``Extended Ensemble Monte Carlo''is a generic term that indicates a set of algorithms which are now popular in a variety of fields in physics and statistical information processing. Exchange Monte Carlo (Metropolis-Coupled Chain, Parallel…
This paper discusses scalability of standard genetic programming (GP) and the probabilistic incremental program evolution (PIPE). To investigate the need for both effective mixing and linkage learning, two test problems are considered:…
One method to generate random permutations involves using Gaussian elimination with partial pivoting (GEPP) on a random matrix $A$ and storing the permutation matrix factor $P$ from the resulting GEPP factorization $PA=LU$. We are…
A celebrated and deep result of Green and Tao states that the primes contain arbitrarily long arithmetic progressions. In this note I provide a straightforward argument demonstrating that the primes get arbitrarily close to arbitrarily long…
The study of automorphisms of computable and other structures connects computability theory with classical group theory. Among the noncomputable countable structures, computably enumerable structures are one of the most important objects of…
Identifying the generating mechanism of a network is challenging as, more often than not, only snapshots are available, but not the full evolution. One candidate for the generating mechanism is preferential attachment which, in its simplest…
The use of machine learning algorithms in finance, medicine, and criminal justice can deeply impact human lives. As a consequence, research into interpretable machine learning has rapidly grown in an attempt to better control and fix…
This work aims at providing new bounds for the diversity multiplexing gain trade-off of a general class of division algebra based lattice codes. In the low multiplexing gain regime, some bounds were previously obtained from the high…
We develop a theory of generalization and scaling for Mixture-of-Experts (MoE) Transformers that cleanly separates \emph{active} per-input capacity from routing combinatorics. By conditioning on fixed routing patterns and union-bounding…
We study the statistical properties of piecewise expanding maps in the general setting of metric measure spaces. We provide sufficient conditions for exponential mixing of such systems with explicit estimates on the constants. We also…
Large language models, such as OpenAI's ChatGPT, have demonstrated exceptional language understanding capabilities in various NLP tasks. Sparsely activated mixture-of-experts (MoE) has emerged as a promising solution for scaling models…