English
Related papers

Related papers: Improving discrepancy by moving a few points

200 papers

By a profound result of Heinrich, Novak, Wasilkowski, and Wo{\'z}niakowski the inverse of the star-discrepancy $n^*(s,\ve)$ satisfies the upper bound $n^*(s,\ve) \leq c_{\mathrm{abs}} s \ve^{-2}$. This is equivalent to the fact that for any…

Numerical Analysis · Mathematics 2012-11-07 Christoph Aistleitner , Markus Hofer

We introduce a methodology, labelled Non-Parametric Isolate-Detect (NPID), for the consistent estimation of the number and locations of multiple change-points in a non-parametric setting. The method can handle general distributional changes…

Statistics Theory · Mathematics 2025-05-01 Andreas Anastasiou , Piotr Fryzlewicz

We study metric learning from preference comparisons under the ideal point model, in which a user prefers an item over another if it is closer to their latent ideal item. These items are embedded into $\mathbb{R}^d$ equipped with an unknown…

Machine Learning · Computer Science 2024-07-15 Zhi Wang , Geelon So , Ramya Korlakai Vinayak

We investigate the discrepancy principle for choosing smoothing parameters for kernel density estimation. The method is based on the distance between the empirical and estimated distribution functions. We prove some new positive and…

Statistics Theory · Mathematics 2015-03-19 Thoralf Mildenberger

The paper gives a wide range, uniform, local approximation of symmetric binomial distribution. The result clearly shows how one has to modify the the classical de Moivre--Laplace normal approximation in order to give an estimate at the tail…

Probability · Mathematics 2025-06-25 Tamás Szabados

This work briefly explores the possibility of approximating spatial distance (alternatively, similarity) between data points using the Isolation Forest method envisioned for outlier detection. The logic is similar to that of isolation: the…

Machine Learning · Statistics 2019-11-26 David Cortes

A natural way of handling imbalanced data is to attempt to equalise the class frequencies and train the classifier of choice on balanced data. For two-class imbalanced problems, the classification success is typically measured by the…

Computer Vision and Pattern Recognition · Computer Science 2018-04-20 Ludmila I. Kuncheva , Álvar Arnaiz-González , José-Francisco Díez-Pastor , Iain A. D. Gunn

We present an extension of the Kolmogorov-Smirnov (KS) two-sample test, which can be more sensitive to differences in the tails. Our test statistic is an integral probability metric (IPM) defined over a higher-order total variation ball,…

Machine Learning · Statistics 2019-03-26 Veeranjaneyulu Sadhanala , Yu-Xiang Wang , Aaditya Ramdas , Ryan J. Tibshirani

In this expository note we present simple proofs of the lower bound of Ramsey numbers (Erd\"os theorem), and of the estimation of discrepancy. Neither statements nor proofs require any knowledge beyond high-school curriculum (except a minor…

Combinatorics · Mathematics 2026-01-06 A. Buchaev , A. Skopenkov

We suggest efficient and provable methods to compute an approximation for imbalanced point clustering, that is, fitting $k$-centers to a set of points in $\mathbb{R}^d$, for any $d,k\geq 1$. To this end, we utilize \emph{coresets}, which,…

Machine Learning · Computer Science 2025-03-13 David Denisov , Dan Feldman , Shlomi Dolev , Michael Segal

The Kolmogorov distances between a symmetric hypergeometric law with standard deviation $\sigma$ and its usual normal approximations are computed and shown to be less than $1/(\sqrt{8\pi}\,\sigma)$, with the order $1/\sigma$ and the…

Probability · Mathematics 2014-05-01 Lutz Mattner , Jona Schulz

Let $p_1,p_2,p_3$ be three non-collinear points in the plane, and let $P$ be a set of $n$ other points in the plane. We show that the number of distinct distances between $p_1,p_2,p_3$ and the points of $P$ is $\Omega(n^{6/11})$, improving…

Combinatorics · Mathematics 2019-02-20 Micha Sharir , Jozsef Solymosi

This paper proposes a method for OOD detection. Questioning the premise of previous studies that ID and OOD samples are separated distinctly, we consider samples lying in the intermediate of the two and use them for training a network. We…

Computer Vision and Pattern Recognition · Computer Science 2021-01-08 Engkarat Techapanurak , Anh-Chuong Dang , Takayuki Okatani

Systematics contaminate observables, leading to distribution shifts relative to theoretically simulated signals-posing a major challenge for using pre-trained models to label such observables. Since systematics are often poorly understood…

Instrumentation and Methods for Astrophysics · Physics 2025-11-18 Sultan Hassan , Sambatra Andrianomena , Benjamin D. Wandelt

A fundamental notion of distance between train and test distributions from the field of domain adaptation is discrepancy distance. While in general hard to compute, here we provide the first set of provably efficient algorithms for testing…

Data Structures and Algorithms · Computer Science 2024-06-14 Gautam Chandrasekaran , Adam R. Klivans , Vasilis Kontonis , Konstantinos Stavropoulos , Arsen Vasilyan

We define two minimum distance estimators for dependent data by minimizing some approximated Maximum Mean Discrepancy distances between the true empirical distribution of observations and their assumed (parametric) model distribution. When…

Methodology · Statistics 2026-01-19 Pierre Alquier , Jean-David Fermanian , Benjamin Poignard

Consider throwing $n$ balls at random into $m$ urns, each ball landing in urn $i$ with probability $p_i$. Let $S$ be the resulting number of singletons, i.e., urns containing just one ball. We give an error bound for the Kolmogorov distance…

Probability · Mathematics 2009-01-23 Mathew D. Penrose

Rerandomization is a strategy of increasing efficiency as compared to complete randomization. The idea with rerandomization is that of removing allocations with imbalance in the observed covariates and then randomizing within the set of…

Methodology · Statistics 2019-11-07 Junni L. Zhang , Per Johansson

The maximum mean discrepancy (MMD) is a recently proposed test statistic for two-sample test. Its quadratic time complexity, however, greatly hampers its availability to large-scale applications. To accelerate the MMD calculation, in this…

Artificial Intelligence · Computer Science 2015-06-19 Ji Zhao , Deyu Meng

We improve constants in the Rademacher-Menchov inequality.

Probability · Mathematics 2007-05-23 Witold Bednorz