中文
相关论文

相关论文: Convergence Rates for Empirical Estimation of Bina…

200 篇论文

In the context of supervised learning, meta learning uses features, metadata and other information to learn about the difficulty, behavior, or composition of the problem. Using this knowledge can be useful to contextualize classifier…

信息论 · 计算机科学 2020-04-28 Salimeh Yasaei Sekeh , Brandon Oselio , Alfred O. Hero

Henze-Penrose divergence is a non-parametric divergence measure that can be used to estimate a bound on the Bayes error in a binary classification problem. In this paper, we show that a cross-match statistic based on optimal weighted…

信息论 · 计算机科学 2018-02-14 Salimeh Yasaei Sekeh , Brandon Oselio , Alfred O. Hero

Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal…

机器学习 · 统计学 2017-11-01 Morteza Noshad Iranzad , Alfred O. Hero

The Bayes Error Rate (BER) is the fundamental limit on the achievable generalizable classification accuracy of any machine learning model due to inherent uncertainty within the data. BER estimators offer insight into the difficulty of any…

机器学习 · 计算机科学 2025-09-24 Lesley Wheat , Martin v. Mohrenschildt , Saeid Habibi

When randomized ensembles such as bagging or random forests are used for binary classification, the prediction error of the ensemble tends to decrease and stabilize as the number of classifiers increases. However, the precise relationship…

概率论 · 数学 2019-05-01 Miles E. Lopes

We unify f-divergences, Bregman divergences, surrogate loss bounds (regret bounds), proper scoring rules, matching losses, cost curves, ROC-curves and information. We do this by systematically studying integral and variational…

机器学习 · 统计学 2009-01-06 Mark D. Reid , Robert C. Williamson

A central problem in Binary Hypothesis Testing (BHT) is to determine the optimal tradeoff between the Type I error (referred to as false alarm) and Type II (referred to as miss) error. In this context, the exponential rate of convergence of…

信息论 · 计算机科学 2021-11-29 Sebastian Espinosa , Jorge F. Silva , Pablo Piantanida

Estimating the ratio of two probability densities from finitely many observations of the densities is a central problem in machine learning and statistics with applications in two-sample testing, divergence estimation, generative modeling,…

机器学习 · 计算机科学 2024-03-12 Werner Zellinger , Stefan Kindermann , Sergei V. Pereverzyev

In statistical learning theory, determining the sample complexity of realizable binary classification for VC classes was a long-standing open problem. The results of Simon and Hanneke established sharp upper bounds in this setting. However,…

机器学习 · 计算机科学 2023-04-19 Ishaq Aden-Ali , Yeshwanth Cherapanamjeri , Abhishek Shetty , Nikita Zhivotovskiy

The Neyman-Pearson region of a simple binary hypothesis testing is the set of points whose coordinates represent the false positive rate and false negative rate of some test. The lower boundary of this region is given by the Neyman-Pearson…

统计理论 · 数学 2025-05-15 Andrew Mullhaupt , Cheng Peng

Information divergence functions play a critical role in statistics and information theory. In this paper we show that a non-parametric f-divergence measure can be used to provide improved bounds on the minimum binary classification…

信息论 · 计算机科学 2015-02-11 Visar Berisha , Alan Wisler , Alfred O. Hero , Andreas Spanias

The existing upper and lower bounds between entropy and error are mostly derived through an inequality means without linking to joint distributions. In fact, from either theoretical or application viewpoint, there exists a need to achieve a…

信息论 · 计算机科学 2013-03-06 Bao-Gang Hu , Hong-Jie Xing

In statistical classification/multiple hypothesis testing and machine learning, a model distribution estimated from the training data is usually applied to replace the unknown true distribution in the Bayes decision rule, which introduces a…

信息论 · 计算机科学 2024-09-24 Zijian Yang , Vahe Eminyan , Ralf Schlüter , Hermann Ney

In statistical classification and machine learning, classification error is an important performance measure, which is minimized by the Bayes decision rule. In practice, the unknown true distribution is usually replaced with a model…

机器学习 · 计算机科学 2025-01-28 Zijian Yang , Vahe Eminyan , Ralf Schlüter , Hermann Ney

Binary density ratio estimation (DRE), the problem of estimating the ratio $p_1/p_2$ given their empirical samples, provides the foundation for many state-of-the-art machine learning algorithms such as contrastive representation learning…

机器学习 · 计算机科学 2021-12-08 Lantao Yu , Yujia Jin , Stefano Ermon

Recent work has focused on the problem of nonparametric estimation of information divergence functionals. Many existing approaches are restrictive in their assumptions on the density support set or require difficult calculations at the…

信息论 · 计算机科学 2021-07-30 Kevin R. Moon , Kumar Sricharan , Kristjan Greenewald , Alfred O. Hero

Machine learning models have traditionally been developed under the assumption that the training and test distributions match exactly. However, recent success in few-shot learning and related problems are encouraging signs that these models…

机器学习 · 统计学 2020-10-15 James Lucas , Mengye Ren , Irene Kameni , Toniann Pitassi , Richard Zemel

Although kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to derive efficient…

机器学习 · 统计学 2022-06-03 Shingo Yashima , Atsushi Nitanda , Taiji Suzuki

We study a hypothesis testing problem in which data is compressed distributively and sent to a detector that seeks to decide between two possible distributions for the data. The aim is to characterize all achievable encoding rates and…

信息论 · 计算机科学 2011-02-01 Md. Saifur Rahman , Aaron B. Wagner

The task of the binary classification problem is to determine which of two distributions has generated a length-$n$ test sequence. The two distributions are unknown; two training sequences of length $N$, one from each distribution, are…

信息论 · 计算机科学 2016-04-18 Dayu Huang , Sean Meyn
‹ 上一页 1 2 3 10 下一页 ›