Related papers: On the Discrepancy of Jittered Sampling
We present a systematic study of the nested sampling algorithm based on the example of the Potts model. This model, which exhibits a first order phase transition for $q>4$, exemplifies a generic numerical challenge in statistical physics:…
Using a Lubachevsky-Stillinger-like growth algorithm combined with biased SWAP Monte Carlo and transient degrees of freedom, we generate ultradense disordered jammed ellipse packings. For all aspect ratios $\alpha$, these packings exhibit…
In this paper, we develop a class of samplers for the diffusion model using the operator-splitting technique. The linear drift term and the nonlinear score-driven drift of the probability flow ordinary differential equation are split and…
Let $\# K$ be a number of integer lattice points contained in a set $K$. In this paper we prove that for each $d\in {\mathbb N}$ there exists a constant $C(d)$ depending on $d$ only, such that for any origin-symmetric convex body $K \subset…
Given an integer $d \geq 2$, $s \in (0,1]$, and $t \in [0,2(d-1)]$, suppose a set $X$ in $\mathbb{R}^d$ has the following property: there is a collection of lines of packing dimension $t$ such that every line from the collection intersects…
We consider the set M_n of all n-truncated power moment sequences of probability measures on [0,1]. We endow this set with the uniform probability. Picking randomly a point in M_n, we show that the upper canonical measure associated with…
Given a sequence of $N$ positive real numbers $\{a_1,a_2,..., a_N \}$, the number partitioning problem consists of partitioning them into two sets such that the absolute value of the difference of the sums of $a_j$ over the two sets is…
Monotonic surfaces spanning finite regions of $Z^d$ arise in many contexts, including DNA-based self-assembly, card-shuffling and lozenge tilings. One method that has been used to uniformly generate these surfaces is a Markov chain that…
We make progress on a conjecture made by [DM], which states that the $d$-dimensional frames of $m$-dimensional boxes resulting from a fragmentation process satisfy Benford's law for all $1 \leq d \leq m$. We provide a sufficient condition…
We study the extreme $L_p$ discrepancy of infinite sequences in the $d$-dimensional unit cube, which uses arbitrary sub-intervals of the unit cube as test sets. This is in contrast to the classical star $L_p$ discrepancy, which uses…
A common approach to statistical learning with big-data is to randomly split it among $m$ machines and learn the parameter of interest by averaging the $m$ individual estimates. In this paper, focusing on empirical risk minimization, or…
For any natural number $d$ and positive number $\varepsilon$, we present a point set in the $d$-dimensional unit cube $[0,1]^d$ that intersects every axis-aligned box of volume greater than $\varepsilon$. These point sets are very easy to…
We propose Distributionally Balanced Designs (DBD), a new class of probability sampling designs that target representativeness at the level of the full auxiliary distribution rather than selected moments. In disciplines such as ecology,…
We carry out the asymptotic analysis of repulsive ensembles of N particles which are discrete analogues of continuous 1d log-gases or beta-ensembles of random matrix theory. The ensembles that we study have several groups of particles which…
In the first part of the series papers, we set out to answer the following question: given specific restrictions on a set of samplers, what kind of signal can be uniquely represented by the corresponding samples attained, as the foundation…
We prove that the the discrepancy of arithmetic progressions in the $d$-dimensional grid $\{1, \dots, N\}^d$ is within a constant factor depending only on $d$ of $N^{\frac{d}{2d+2}}$. This extends the case $d=1$, which is a celebrated…
The nested distance builds on the Wasserstein distance to quantify the difference of stochastic processes, including also the information modelled by filtrations. The Sinkhorn divergence is a relaxation of the Wasserstein distance, which…
Existing two-sample testing techniques, particularly those based on choosing a kernel for the Maximum Mean Discrepancy (MMD), often assume equal sample sizes from the two distributions. Applying these methods in practice can require…
In the matter of selection of sample time points for the estimation of the power spectral density of a continuous time stationary stochastic process, irregular sampling schemes such as Poisson sampling are often preferred over regular…
The maximum mean discrepancy (MMD) is a recently proposed test statistic for two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD calculation, in this…