Related papers: One-parameter counterexamples to the refined Bessi…
In this paper a new family of minimum divergence estimators based on the Bregman divergence is proposed, where the defining convex function has an exponential nature. These estimators avoid the necessity of using an intermediate kernel…
In this work, we investigate the expressiveness of the "conditional mutual information" (CMI) framework of Steinke and Zakynthinou (2020) and the prospect of using it to provide a unified framework for proving generalization bounds in the…
We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably…
We describe a novel algorithm for random sampling of freely reduced words equal to the identity in a finitely presented group. The algorithm is based on Metropolis Monte Carlo sampling. The algorithm samples from a stretched Boltzmann…
Black box variational inference (BBVI) with reparameterization gradients triggered the exploration of divergence measures other than the Kullback-Leibler (KL) divergence, such as alpha divergences. In this paper, we view BBVI with…
We present a theoretical and empirical investigation of the statistical behaviour of the words in a text produced by human language. To this aim, we analyse the word distribution of various texts of Italian language selected from a specific…
This paper proposes a simple, novel, and fully-Bayesian approach for causal inference in partially linear models with high-dimensional control variables. Off-the-shelf machine learning methods can introduce biases in the causal parameter…
The BMV conjecture states that for \(n\times n\) Hermitian matrices \(A\) and \(B\) the function \(f_{A,B}(t)=\tr e^{tA+B}\) is exponentially convex. Recently the BMV conjecture was proved by Herbert Stahl. The proof of Herbert Stahl is…
The Burrows-Wheeler-Transform (BWT), a reversible string transformation, is one of the fundamental components of many current data structures in string processing. It is central in data compression, as well as in efficient query algorithms…
It is shown that at least 50% of the probability mass of a sum of independent Rademacher random variables is within one standard deviation from its mean. This lower bound is sharp, it is much better than for instance the bound that can be…
Blundell, Buesing, Davies, Veli\v{c}kovi\'c, and Williamson (BBDVW) introduced the notion of a hypercube decomposition of an interval in Bruhat order. They conjectured a recursive formula in terms of this structure which, if shown for all…
The Bounded Height Conjecture of Bombieri, Masser, and Zannier states that for any sufficiently generic algebraic subvariety of a semiabelian $\overline{\mathbb{Q}}$-variety $G$ there is an upper bound on the Weil height of the points…
This paper studies large sample properties of a Bayesian approach to inference about slope parameters $\gamma$ in linear regression models with a structural break. In contrast to the conventional approach to inference about $\gamma$ that…
An iterative randomness extraction algorithm which generalized the Von Neumann's extraction algorithm is detailed, analyzed and implemented in standard C++. Given a sequence of independently and identically distributed biased Bernoulli…
The theory of regular variation, in its Karamata and Bojani\'c-Karamata/de Haan forms, is long established and makes essential use of homomorphisms. Both forms are subsumed within the recent theory of Beurling regular variation, developed…
A Bernstein-von Mises theorem is derived for general semiparametric functionals. The result is applied to a variety of semiparametric problems in i.i.d. and non-i.i.d. situations. In particular, new tools are developed to handle…
In recent years, word embeddings have been widely used to measure biases in texts. Even if they have proven to be effective in detecting a wide variety of biases, metrics based on word embeddings lack transparency and interpretability. We…
The designation ``Bernstein-von Mises theorem'' is apparently due to Lucien Le Cam. Roughly, the assertion of this theorem states that the posterior distribution of a parameter, conditioned on a large sample, is approximately normal,…
We show that a normal matrix $A$ with coefficient in $\mathbb C[[X]]$, $X=(X_1, \ldots, X_n)$, can be diagonalized, provided the discriminant $\Delta_A $ of its characteristic polynomial is a monomial times a unit. The proof is an…
The normalized maximum likelihood (NML) is a recent penalized likelihood that has properties that justify defining the amount of discrimination information (DI) in the data supporting an alternative hypothesis over a null hypothesis as the…