English
Related papers

Related papers: Telling Two Distributions Apart: a Tight Character…

200 papers

We give a general unified method that can be used for $L_1$ {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for…

Data Structures and Algorithms · Computer Science 2015-08-25 Ilias Diakonikolas , Daniel M. Kane , Vladimir Nikishkin

We consider two multi-armed bandit problems with $n$ arms: (i) given an $\epsilon > 0$, identify an arm with mean that is within $\epsilon$ of the largest mean and (ii) given a threshold $\mu_0$ and integer $k$, identify $k$ arms with means…

Machine Learning · Statistics 2019-06-18 Julian Katz-Samuels , Kevin Jamieson

A characterization of the exponential distribution based on equidistribution conditions for maxima of random samples with consecutive sizes n-1 and n for an arbitrary and fixed n>2 is proved. This solves an open problem stated recently in…

Probability · Mathematics 2015-02-24 Santanu Chakraborty , George P. Yanev

Large-scale datasets are increasingly being used to inform decision making. While this effort aims to ground policy in real-world evidence, challenges have arisen as selection bias and other forms of distribution shifts often plague…

Methodology · Statistics 2023-11-07 Santiago Cortes-Gomez , Mateo Dulce , Carlos Patino , Bryan Wilder

Two-sample hypothesis testing-determining whether two sets of data are drawn from the same distribution-is a fundamental problem in statistics and machine learning with broad scientific applications. In the context of nonparametric testing,…

Machine Learning · Statistics 2026-04-21 Antoine Chatalic , Marco Letizia , Nicolas Schreuder , Lorenzo Rosasco

With the rise of increasingly powerful and user-facing NLP systems, there is growing interest in assessing whether they have a good representation of uncertainty by evaluating the quality of their predictive distribution over outcomes. We…

Computation and Language · Computer Science 2024-02-27 Joris Baan , Raquel Fernández , Barbara Plank , Wilker Aziz

The L_2-discrepancy measures the irregularity of the distribution of a finite point set. In this note we prove lower bounds for the L_2 discrepancy of arbitrary N-point sets. Our main focus is on the two-dimensional case. Asymptotic upper…

Numerical Analysis · Mathematics 2014-02-19 Aicke Hinrichs , Lev Markhasin

Understanding how different classes are distributed in an unlabeled data set is an important challenge for the calibration of probabilistic classifiers and uncertainty quantification. Approaches like adjusted classify and count, black-box…

Machine Learning · Statistics 2024-06-19 Albert Ziegler , Paweł Czyż

The performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two…

Machine Learning · Computer Science 2025-09-24 Varun Babbar , Zhicheng Guo , Cynthia Rudin

Black-box optimization is often encountered for decision-making in complex systems management, where the knowledge of system is limited. Under these circumstances, it is essential to balance the utilization of new information with…

Computation · Statistics 2025-01-15 Teng Lian , Jian-Qiang Hu , Yuhang Wu , Zeyu Zheng

The recent success of generative adversarial networks and variational learning suggests training a classifier network may work well in addressing the classical two-sample problem. Network-based tests have the computational advantage that…

Machine Learning · Statistics 2022-06-01 Xiuyuan Cheng , Alexander Cloninger

We consider the problem of estimating how well a model class is capable of fitting a distribution of labeled data. We show that it is often possible to accurately estimate this "learnability" even when given an amount of data that is too…

Machine Learning · Computer Science 2019-03-26 Weihao Kong , Gregory Valiant

We propose the discrepancy-based generalization theories for unsupervised domain adaptation. Previous theories introduced distribution discrepancies defined as the supremum over complete hypothesis space. The hypothesis space may contain…

Machine Learning · Computer Science 2020-08-17 Yuchen Zhang , Mingsheng Long , Jianmin Wang , Michael I. Jordan

By focusing on the X-matrix part of a density matrix of two qubits we provide an algebraic lower bound for the concurrence. The lower bound is generalized for cases beyond two qubits and can serve as a sufficient condition for…

Quantum Physics · Physics 2012-04-19 S. M. Hashemi rafsanjani , S. Agarwal

Two-sample tests evaluate whether two samples are realizations of the same distribution (the null hypothesis) or two different distributions (the alternative hypothesis). We consider a new setting for this problem where sample features are…

Machine Learning · Computer Science 2022-07-20 Weizhi Li , Gautam Dasarathy , Karthikeyan Natesan Ramamurthy , Visar Berisha

This paper explores a theory of generalization for learning problems on product distributions, complementing the existing learning theories in the sense that it does not rely on any complexity measures of the hypothesis classes. The main…

Computer Science and Game Theory · Computer Science 2020-07-28 Chenghao Guo , Zhiyi Huang , Zhihao Gavin Tang , Xinzhi Zhang

The vast majority of real world classification problems are imbalanced, meaning there are far fewer data from the class of interest (the positive class) than from other classes. We propose two machine learning algorithms to handle highly…

Machine Learning · Statistics 2014-06-10 Siong Thye Goh , Cynthia Rudin

The majority of traditional classification ru les minimizing the expected probability of error (0-1 loss) are inappropriate if the class probability distributions are ill-defined or impossible to estimate. We argue that in such cases class…

Machine Learning · Statistics 2018-08-14 Robert P. W. Duin , Elzbieta Pekalska

Given a network of fixed size $n$ and an initial distribution of data, we derive sufficient connectivity conditions on a sequence of time-varying digraphs for (a) data collection and (b) data dissemination, within at most $(n-1)$…

Systems and Control · Computer Science 2016-05-03 Kevin Topley

Understanding the shape of a distribution of data is of interest to people in a great variety of fields, as it may affect the types of algorithms used for that data. We study one such problem in the framework of distribution property…

Machine Learning · Computer Science 2022-12-06 Maryam Aliakbarpour , Amartya Shankha Biswas , Kavya Ravichandran , Ronitt Rubinfeld
‹ Prev 1 3 4 5 6 7 10 Next ›