Related papers: Two-sample KS test with approxQuantile in Apache S…
Complex data are often represented as a graph, which in turn can often be viewed as a realisation of a random graph, such as an inhomogeneous random graph model (IRG). For general fast goodness-of-fit tests in high dimensions, kernelised…
We discuss the possibility of sampling exponential moments of the canonical phase from the s-parametrized phase space functions. We show that the sampling kernels exist and are well-behaved for any s>-1, whereas for s=-1 the kernels diverge…
We derive inequalities for $n$ spin-1/2 systems under the assumption that the hidden-variable theoretical joint probability distribution for any pair of commuting observables is equal to the quantum mechanical one. Fine showed that this…
We propose a Similarity-Based Stratified Splitting (SBSS) technique, which uses both the output and input space information to split the data. The splits are generated using similarity functions among samples to place similar samples in…
In this article we present very intuitive, easy to follow, yet mathematically rigorous, approach to the so called data fitting process. Rather than minimizing the distance between measured and simulated data points, we prefer to find such…
A smooth test to simultaneously compare $K$ copulas, where $K \geq 2$ is proposed. The $K$ observed populations can be paired, and the test statistic is constructed based on the differences between moment sequences, called copula…
This paper proposes using a method named Double Score Matching (DSM) to do mass-imputation and presents an application to make inferences with a nonprobability sample. DSM is a $k$-Nearest Neighbors algorithm that uses two balance scores…
In statistics, assuming samples are independent is reasonable. However, this property can fail to hold for the features, a distinction that has led to several lines of work aiming to remove the latter assumption of independence present in…
This paper deals with two-sample Kolmogorov-Smirnov test and its biasedness. This test is not unbiased in general in case of different sample sizes. We found out most biased distribution for some values of significance level $\alpha$.…
In large-scale unconstrained optimization algorithms such as limited memory BFGS (LBFGS), a common subproblem is a line search minimizing the loss function along a descent direction. Commonly used line searches iteratively find an…
Stochastic gradients have been widely integrated into Langevin-based methods to improve their scalability and efficiency in solving large-scale sampling problems. However, the proximal sampler, which exhibits much faster convergence than…
We propose a simple quantum-key-distribution (QKD) scheme for practical single photon sources (SPSs), which works even with a moderate suppression of the second-order correlation $g^{(2)}$ of the source. The scheme utilizes a passive…
Consider a fixed universe of $N=2^n$ elements and the uniform distribution over elements of some subset of size $K$. Given samples from this distribution, the task of complement sampling is to provide a sample from the complementary subset.…
In this paper we study the problem of testing of constrained samplers over high-dimensional distributions with $(\varepsilon,\eta,\delta)$ guarantees. Samplers are increasingly used in a wide range of safety-critical ML applications, and…
Within the realm of early fault-tolerant quantum computing (EFTQC), quantum Krylov subspace diagonalization (QKSD) has emerged as a promising quantum algorithm for the approximate Hamiltonian diagonalization via projection onto the quantum…
Data selection is essential for training deep learning models. An effective data sampler assigns proper sampling probability for training data and helps the model converge to a good local minimum with high performance. Previous studies in…
An Automated Sliced Gibbs framework is proposed for fully automated Markov chain Monte Carlo sampling from arbitrary finite dimensional probability kernels. The method targets unnormalized, non-smooth, heavy tailed, and highly multimodal…
Synthetic data generation has become a key ingredient for training machine learning procedures, addressing tasks such as data augmentation, analysing privacy-sensitive data, or visualising representative samples. Assessing the quality of…
An all-sky sample of 1227 visual binaries based on Washington Double Star catalogue is constructed to infer the IMF, mass ratio, and projected distance distribution with a dedicated population synthesis model. Parallaxes from Gaia DR2 and…
With increasing point of interest (POI) datasets available with fine-grained spatial and temporal attributes, space-time Ripley's K function has been regarded as a powerful approach to analyze spatiotemporal point process. However,…