English
Related papers

Related papers: When is the majority-vote classifier beneficial?

200 papers

In this paper, we propose a diversity-aware ensemble learning based algorithm, referred to as DAMVI, to deal with imbalanced binary classification tasks. Specifically, after learning base classifiers, the algorithm i) increases the weights…

Machine Learning · Computer Science 2020-04-17 Anil Goyal , Jihed Khiari

Learning from imbalanced data is one of the most significant challenges in real-world classification tasks. In such cases, neural networks performance is substantially impaired due to preference towards the majority class. Existing…

Machine Learning · Computer Science 2022-11-13 Bronislav Yasinnik , Moshe Salhov , Ofir Lindenbaum , Amir Averbuch

Weakly-supervised text classification aims to induce text classifiers from only a few user-provided seed words. The vast majority of previous work assumes high-quality seed words are given. However, the expert-annotated seed words are…

Computation and Language · Computer Science 2021-04-21 Yiping Jin , Akshay Bhatia , Dittaya Wanvarie

Electing a single committee of a small size is a classical and well-understood voting situation. Being interested in a sequence of committees, we introduce and study two time-dependent multistage models based on simple Plurality voting.…

Computational Complexity · Computer Science 2024-01-22 Robert Bredereck , Till Fluschnik , Andrzej Kaczmarczyk

We consider the online one-class collaborative filtering (CF) problem that consists of recommending items to users over time in an online fashion based on positive ratings only. This problem arises when users respond only occasionally to a…

Machine Learning · Computer Science 2017-06-02 Reinhard Heckel , Kannan Ramchandran

A common approach in positive-unlabeled learning is to train a classification model between labeled and unlabeled data. This strategy is in fact known to give an optimal classifier under mild conditions; however, it results in biased…

Machine Learning · Statistics 2017-02-03 Shantanu Jain , Martha White , Predrag Radivojac

In many large multiple testing problems the hypotheses are divided into families. Given the data, families with evidence for true discoveries are selected, and hypotheses within them are tested. Neither controlling the error-rate in each…

Statistics Theory · Mathematics 2011-06-21 Yoav Benjamini , Marina Bogomolov

Training and evaluation of fair classifiers is a challenging problem. This is partly due to the fact that most fairness metrics of interest depend on both the sensitive attribute information and label information of the data points. In many…

Machine Learning · Computer Science 2021-02-18 Pranjal Awasthi , Alex Beutel , Matthaeus Kleindessner , Jamie Morgenstern , Xuezhi Wang

Effective patent value assessment provides decision support for patent transection and promotes the practical application of patent technology. The limitations of previous research on patent value assessment were analyzed in this work, and…

Machine Learning · Computer Science 2020-01-24 Yihui Qiu , Chiyu Zhang

We investigate a stochastic counterpart of majority votes over finite ensembles of classifiers, and study its generalization properties. While our approach holds for arbitrary distributions, we instantiate it with Dirichlet distributions:…

Machine Learning · Computer Science 2021-10-20 Valentina Zantedeschi , Paul Viallard , Emilie Morvant , Rémi Emonet , Amaury Habrard , Pascal Germain , Benjamin Guedj

We revisit the fundamental question of simple-versus-simple hypothesis testing with an eye towards computational complexity, as the statistically optimal likelihood ratio test is often computationally intractable in high-dimensional…

Statistics Theory · Mathematics 2025-05-05 Ankur Moitra , Alexander S. Wein

For a bucket test with a single criterion for success and a fixed number of samples or testing period, requiring a $p$-value less than a specified value of $\alpha$ for the success criterion produces statistical confidence at level $1 -…

Methodology · Statistics 2024-08-05 Eric Bax , Arundhyoti Sarkar , Alex Shtoff

If part of a population is hidden but two or more sources are available that each cover parts of this population, dual- or multiple-system(s) estimation can be applied to estimate this population. For this it is common to use the log-linear…

Methodology · Statistics 2023-11-06 Daan B. Zult , Peter G. M. van der Heijden , Bart F. M. Bakker

Good large sample performance is typically a minimum requirement of any model selection criterion. This article focuses on the consistency property of the Bayes factor, a commonly used model comparison tool, which has experienced a recent…

Statistics Theory · Mathematics 2016-07-04 Siddhartha Chib , Todd A. Kuffner

A statistical analysis of optimal universal cloning shows that it is possible to identify an ideal (but non-positive) copying process that faithfully maps all properties of the original Hilbert space onto two separate quantum systems. The…

Quantum Physics · Physics 2012-07-18 Holger F. Hofmann

Wisdom of the crowd revealed a striking fact that the majority answer from a crowd is often more accurate than any individual expert. We observed the same story in machine learning--ensemble methods leverage this idea to combine multiple…

Machine Learning · Computer Science 2019-10-01 Tianyi Luo , Yang Liu

We introduce a very general method for high-dimensional classification, based on careful combination of the results of applying an arbitrary base classifier to random projections of the feature vectors into a lower-dimensional space. In one…

Methodology · Statistics 2017-06-06 Timothy I. Cannings , Richard J. Samworth

Class imbalance poses a significant challenge to supervised classification, particularly in critical domains like medical diagnostics and anomaly detection where minority class instances are rare. While numerous studies have explored…

Machine Learning · Computer Science 2025-09-10 Ali Nawaz , Amir Ahmad , Shehroz S. Khan

Classifier chains have recently been proposed as an appealing method for tackling the multi-label classification task. In addition to several empirical studies showing its state-of-the-art performance, especially when being used in its…

Machine Learning · Computer Science 2019-06-10 Robin Senge , Juan José del Coz , Eyke Hüllermeier

We compare the notions "Decisiveness" and "Success" for certain weighted voting systems and various underlying voting measures. In particular, we compute the success rate for the Shapley-Shubik meassure and, more generally, for Common…

General Mathematics · Mathematics 2017-06-27 Werner Kirsch