English
Related papers

Related papers: Asymptotic Log-loss of Prequential Maximum Likelih…

200 papers

We consider the problem of tracking an unknown time varying parameter that characterizes the probabilistic evolution of a sequence of independent observations. To this aim, we propose a stochastic gradient descent-based recursive scheme in…

Statistics Theory · Mathematics 2023-03-01 Alberto Lanconelli , Christopher S. A. Lauria

We model the development of the linear complexity of multisequences by a stochastic infinite state machine, the Battery-Discharge-Model, BDM. The states s in S of the BDM have asymptotic probabilities or mass Pr(s)=1/(P(q,M) q^K(s)), where…

Information Theory · Computer Science 2007-07-13 Michael Vielhaber , Monica del Pilar Canales

English words and the outputs of many other natural processes are well-known to follow a Zipf distribution. Yet this thoroughly-established property has never been shown to help compress or predict these important processes. We show that…

Information Theory · Computer Science 2015-05-04 Moein Falahatgar , Ashkan Jafarpour , Alon Orlitsky , Venkatadheeraj Pichapati , Ananda Theertha Suresh

Can autoregressive large language models (LLMs) learn consistent probability distributions when trained on sequences in different token orders? We prove formally that for any well-defined probability distribution, sequence perplexity is…

Computation and Language · Computer Science 2025-05-14 Xiaoliang Luo , Xinyi Xu , Michael Ramscar , Bradley C. Love

A positive linear recurrence sequence is of the form $H_{n+1} = c_1 H_n + \cdots + c_L H_{n+1-L}$ with each $c_i \ge 0$ and $c_1 c_L > 0$, with appropriately chosen initial conditions. There is a notion of a legal decomposition (roughly,…

Number Theory · Mathematics 2016-07-19 Steven J. Miller , Dawn Nelson , Zhao Pan , Huanzhong Xu

This paper deals with the problem of universal lossless coding on a countable infinite alphabet. It focuses on some classes of sources defined by an envelope condition on the marginal distribution, namely exponentially decreasing envelope…

Information Theory · Computer Science 2011-07-07 Dominique Bontemps

We study numerically the distributions of the length $L$ of the longest increasing subsequence (LIS) for the two cases of random permutations and of one-dimensional random walks. Using sophisticated large-deviation algorithms, we are able…

Disordered Systems and Neural Networks · Physics 2019-04-05 Jörn Börjes , Hendrik Schawe , Alexander K. Hartmann

Let M_n denote the number of sites in the largest cluster in critical site percolation on the triangular lattice inside a box side length n. We give lower and upper bounds on the probability that M_n / E(M_n) > x of the form exp(- C…

Probability · Mathematics 2014-04-09 Demeter Kiss

We show that for unconstrained Deep Linear Discriminant Analysis (LDA) classifiers, maximum-likelihood training admits pathological solutions in which class means drift together, covariances collapse, and the learned representation becomes…

Machine Learning · Statistics 2026-01-06 Maxat Tezekbayev , Rustem Takhanov , Arman Bolatov , Zhenisbek Assylbekov

Motivated by average-case trace reconstruction and coding for portable DNA-based storage systems, we initiate the study of \emph{coded trace reconstruction}, the design and analysis of high-rate efficiently encodable codes that can be…

Information Theory · Computer Science 2019-09-11 Mahdi Cheraghchi , Ryan Gabrys , Olgica Milenkovic , João Ribeiro

The article studies the almost surely asymptotics of extreme values $\bar{\xi}_n = \max_{1\leq i \leq n} \xi_i$, where $ \xi , \xi_1 , \xi_2 , \ldots$ are discrete identically distributed random variables. One of the main results on this…

Probability · Mathematics 2025-03-27 Kateryna Akbash , Ivan Matsak

Relational models for contingency tables are generalizations of log-linear models, allowing effects associated with arbitrary subsets of cells in a possibly incomplete table, and not necessarily containing the overall effect. In this…

Methodology · Statistics 2015-05-01 Anna Klimova , Tamás Rudas

We give an independent proof of the Krasikov-Litsyn bound d/n<~(1-5^{-1/4})/2 on doubly-even self-dual binary codes. The technique used (a refinement of the Mallows-Odlyzko-Sloane approach) extends easily to other families of self-dual…

Combinatorics · Mathematics 2007-05-23 Eric M. Rains

Let (X_n,Y_n) be i.i.d. random vectors. Let W(x) be the partial sum of Y_n just before that of X_n exceeds x>0. Motivated by stochastic models for neural activity, uniform convergence of the form $\sup_{c\in I}|a(c,x)\operatorname…

Probability · Mathematics 2009-09-29 Zhiyi Chi

Since the classical work of Berlekamp, McEliece and van Tilborg, it is well known that the problem of exact maximum-likelihood (ML) decoding of general linear codes is NP-hard. In this paper, we show that exact ML decoding of a classs of…

Information Theory · Computer Science 2016-11-17 Weiyu Xu , Babak Hassibi

We investigate the ratio $\rho_{n,L}$ of prefix codes to all uniquely decodable codes over an $n$-letter alphabet and with length distribution $L$. For any integers $n\geq 2$ and $m\geq 1$, we construct a lower bound and an upper bound for…

Combinatorics · Mathematics 2018-04-17 Adam Woryna

We consider the maximum coding rate achievable by uniformly-random codes for the deletion channel. We prove an upper bound that's within 0.1 of the best known lower bounds for all values of the deletion probability $d,$ and much closer for…

Information Theory · Computer Science 2022-10-17 Berivan Isik , Francisco Pernice , Tsachy Weissman

Learning the structure of Bayesian networks and causal relationships from observations is a common goal in several areas of science and technology. We show that the prequential minimum description length principle (MDL) can be used to…

Machine Learning · Computer Science 2021-07-13 Jorg Bornschein , Silvia Chiappa , Alan Malek , Rosemary Nan Ke

This article studies exponential families $\mathcal{E}$ on finite sets such that the information divergence $D(P\|\mathcal{E})$ of an arbitrary probability distribution from $\mathcal{E}$ is bounded by some constant $D>0$. A particular…

Statistics Theory · Mathematics 2014-06-18 Johannes Rauh

We employ a parameter-free distribution estimation framework where estimators are random distributions and utilize the Kullback-Leibler (KL) divergence as a loss function. Wu and Vos [J. Statist. Plann. Inference 142 (2012) 1525-1536] show…

Statistics Theory · Mathematics 2015-09-21 Paul Vos , Qiang Wu