English
Related papers

Related papers: Regularizing random points by deleting a few

200 papers

For every positive integer N and every $\alpha\in [0,1)$, let $B(N, \alpha)$ denote the probabilistic model in which a random set $A\subset \{1,\dots,N\}$ is constructed by choosing independently every element of $\{1,\dots,N\}$ with…

Number Theory · Mathematics 2020-05-15 Daniele Mastrostefano

We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…

Machine Learning · Computer Science 2024-12-03 Maryam Aliakbarpour , Piotr Indyk , Ronitt Rubinfeld , Sandeep Silwal

Many data-fitting applications require the solution of an optimization problem involving a sum of large number of functions of high dimensional parameter. Here, we consider the problem of minimizing a sum of $n$ functions over a convex…

Optimization and Control · Mathematics 2016-02-29 Farbod Roosta-Khorasani , Michael W. Mahoney

The small sample universal hypothesis testing problem is investigated in this paper, in which the number of samples $n$ is smaller than the number of possible outcomes $m$. The goal of this work is to find an appropriate criterion to…

Statistics Theory · Mathematics 2014-12-30 Dayu Huang , Sean Meyn

Consider a sequence of $n$ independent random variables with a common continuous distribution $F$, and consider the task of choosing an increasing subsequence where the observations are revealed sequentially and where an observation must be…

Probability · Mathematics 2016-08-02 Alessandro Arlotto , Vinh V. Nguyen , J. Michael Steele

A central problem in discrepancy theory is the challenge of evenly distributing points $\left\{x_1, \dots, x_n \right\}$ in $[0,1]^d$. Suppose a set is so regular that for some $\varepsilon> 0$ and all $y \in [0,1]^d$ the sub-region $[0,y]…

Combinatorics · Mathematics 2023-01-31 Stefan Steinerberger

The task of the binary classification problem is to determine which of two distributions has generated a length-$n$ test sequence. The two distributions are unknown; two training sequences of length $N$, one from each distribution, are…

Information Theory · Computer Science 2016-04-18 Dayu Huang , Sean Meyn

We obtain estimation error rates for estimators obtained by aggregation of regularized median-of-means tests, following a construction of Le Cam. The results hold with exponentially large probability -- as in the gaussian framework with…

Statistics Theory · Mathematics 2017-07-19 Lecué Guillaume , Lerasle Matthieu

Let $A_n$ be an $n$ by $n$ random matrix whose entries are independent real random variables with mean zero, variance one and with subexponential tail. We show that the logarithm of $|\det A_n|$ satisfies a central limit theorem. More…

Probability · Mathematics 2014-01-14 Hoi H. Nguyen , Van Vu

Despite empirical risk minimization (ERM) is widely applied in the machine learning community, its performance is limited on data with spurious correlation or subpopulation that is introduced by hidden attributes. Existing literature…

Machine Learning · Computer Science 2024-12-18 Hongyu Shen , Zhizhen Zhao

Given two disjoint sets $W_1$ and $W_2$ of points in the plane, the Optimal Discretization problem asks for the minimum size of a family of horizontal and vertical lines that separate $W_1$ from $W_2$, that is, in every region into which…

Data Structures and Algorithms · Computer Science 2026-03-16 Stefan Kratsch , Tomáš Masařík , Irene Muzi , Marcin Pilipczuk , Manuel Sorge

Statistical inference is often simplified by sample-splitting. This simplification comes at the cost of the introduction of randomness not native to the data. We propose a simple procedure for sequentially aggregating statistics constructed…

Econometrics · Economics 2024-11-18 David M. Ritzwoller , Joseph P. Romano

This paper provides new error bounds on "consistent" reconstruction methods for signals observed from quantized random projections. Those signal estimation techniques guarantee a perfect matching between the available quantized data and a…

Information Theory · Computer Science 2016-04-21 Laurent Jacques

The well-known trace reconstruction problem is the problem of inferring an unknown source string $x \in \{0,1\}^n$ from independent "traces", i.e. copies of $x$ that have been corrupted by a $\delta$-deletion channel which independently…

Data Structures and Algorithms · Computer Science 2022-11-08 Xi Chen , Anindya De , Chin Ho Lee , Rocco A. Servedio , Sandip Sinha

Let A(n) be a sequence of i.i.d. topical (i.e. isotone and additively homogeneous) operators. Let $x(n,x_0)$ be defined by $x(0,x_0)=x_0$ and $x(n,x_0)=A(n)x(n-1,x_0)$. This can modelize a wide range of systems including, task graphs, train…

Probability · Mathematics 2007-05-23 Glenn Merlet

This paper describes a new method for reducing the error in a classifier. It uses an error correction update that includes the very simple rule of either adding or subtracting the error adjustment, based on whether the variable value is…

Artificial Intelligence · Computer Science 2018-03-02 Kieran Greer

We study the space requirements of a sorting algorithm where only items that at the end will be adjacent are kept together. This is equivalent to the following combinatorial problem: Consider a string of fixed length n that starts as a…

Probability · Mathematics 2007-05-23 Svante Janson

Recovery procedures in various application in Data Science are based on \emph{stable point separation}. In its simplest form, stable point separation implies that if $f$ is "far away" from $0$, and one is given a random sample…

Probability · Mathematics 2019-04-19 Shahar Mendelson , Grigoris Paouris

Many randomized approximation algorithms operate by giving a procedure for simulating a random variable $X$ which has mean $\mu$ equal to the target answer, and a relative standard deviation bounded above by a known constant $c$. Examples…

Computation · Statistics 2019-08-16 Mark Huber

Our goal is to develop a general strategy to decompose a random variable $X$ into multiple independent random variables, without sacrificing any information about unknown parameters. A recent paper showed that for some well-known natural…

Methodology · Statistics 2025-12-23 Ameer Dharamshi , Anna Neufeld , Keshav Motwani , Lucy L. Gao , Daniela Witten , Jacob Bien
‹ Prev 1 4 5 6 7 8 10 Next ›