English
Related papers

Related papers: Data organization limits the predictability of bin…

200 papers

Variance in predictions across different trained models is a significant, under-explored source of error in fair binary classification. In practice, the variance on some data examples is so large that decisions can be effectively arbitrary.…

We investigate a data-driven quasiconcave maximization problem where information about the objective function is limited to a finite sample of data points. We begin by defining an ambiguity set for admissible objective functions based on…

Optimization and Control · Mathematics 2026-04-07 Jian Wu , William B. Haskell , Wenjie Huang , Huifu Xu

We address the problem of general supervised learning when data can only be accessed through an (indefinite) similarity function between data points. Existing work on learning with indefinite kernels has concentrated solely on…

Machine Learning · Computer Science 2012-10-23 Purushottam Kar , Prateek Jain

We consider the problem of optimal decision referrals in human-automation teams performing binary classification tasks. The automation, which includes a pre-trained classifier, observes data for a batch of independent tasks, analyzes them,…

Systems and Control · Electrical Eng. & Systems 2025-04-08 Kesav Kaza , Jerome Le Ny , Aditya Mahajan

We consider a class of submodular maximization problems in which decision-makers have limited access to the objective function. We explore scenarios where the decision-maker can observe only pairwise information, i.e., can evaluate the…

Data Structures and Algorithms · Computer Science 2022-02-09 Andrew Downie , Bahman Gharesifard , Stephen L. Smith

In this paper, we elaborate on the use of the Sugeno integral in the context of machine learning. More specifically, we propose a method for binary classification, in which the Sugeno integral is used as an aggregation function that…

Machine Learning · Computer Science 2020-07-08 Sadegh Abbaszadeh , Eyke Hüllermeier

Binary discrimination between well-defined signal and background datasets is a problem of fundamental importance in particle physics. With detailed event simulation and the advent of extensive deep learning tools, identification of the…

High Energy Physics - Phenomenology · Physics 2024-02-06 Andrew J. Larkoski

The data for many classification problems, such as pattern and speech recognition, follow mixture distributions. To quantify the optimum performance for classification tasks, the Shannon mutual information is a natural information-theoretic…

Signal Processing · Electrical Eng. & Systems 2022-06-22 Yijun Ding , Amit Ashok

Learning theory has traditionally followed a model-centric approach, focusing on designing optimal algorithms for a fixed natural learning task (e.g., linear classification or regression). In this paper, we adopt a complementary…

Machine Learning · Computer Science 2025-04-29 Steve Hanneke , Shay Moran , Alexander Shlimovich , Amir Yehudayoff

A central concern in classification is the vulnerability of machine learning models to adversarial attacks. Adversarial training is one of the most popular techniques for training robust classifiers, which involves minimizing an adversarial…

Machine Learning · Computer Science 2025-10-09 Natalie S. Frank

As artificial intelligence is increasingly affecting all parts of society and life, there is growing recognition that human interpretability of machine learning models is important. It is often argued that accuracy or other similar…

Machine Learning · Statistics 2018-06-27 Kush R. Varshney , Prashant Khanduri , Pranay Sharma , Shan Zhang , Pramod K. Varshney

In this paper, we address the problem of data description using a Bayesian framework. The goal of data description is to draw a boundary around objects of a certain class of interest to discriminate that class from the rest of the feature…

Machine Learning · Computer Science 2016-02-26 Alireza Ghasemi , Hamid R. Rabiee , Mohammad T. Manzuri , M. H. Rohban

With the increasing deployment of machine learning models in many socially sensitive tasks, there is a growing demand for reliable and trustworthy predictions. One way to accomplish these requirements is to allow a model to abstain from…

Machine Learning · Computer Science 2024-09-19 Andrea Pugnana , Lorenzo Perini , Jesse Davis , Salvatore Ruggieri

Model explainability is crucial for human users to be able to interpret how a proposed classifier assigns labels to data based on its feature values. We study generalized linear models constructed using sets of feature value rules, which…

Machine Learning · Statistics 2023-11-06 Sanjeeb Dash , Soumyadip Ghosh , Joao Goncalves , Mark S. Squillante

Consider a binary decision making process where a single machine learning classifier replaces a multitude of humans. We raise questions about the resulting loss of diversity in the decision making process. We study the potential benefits of…

Machine Learning · Statistics 2017-07-03 Nina Grgić-Hlača , Muhammad Bilal Zafar , Krishna P. Gummadi , Adrian Weller

As an intrinsic and fundamental property of big data, data heterogeneity exists in a variety of real-world applications, such as precision medicine, autonomous driving, financial applications, etc. For machine learning algorithms, the…

Machine Learning · Computer Science 2023-04-04 Jiashuo Liu , Jiayun Wu , Bo Li , Peng Cui

The standard approach to supervised classification involves the minimization of a log-loss as an upper bound to the classification error. While this is a tight bound early on in the optimization, it overemphasizes the influence of…

Machine Learning · Computer Science 2016-12-30 Nicolas Le Roux

Machine learning classification methods usually assume that all possible classes are sufficiently present within the training set. Due to their inherent rarities, extreme events are always under-represented and classifiers tailored for…

Methodology · Statistics 2025-06-12 Juliette Legrand , Philippe Naveau , Marco Oesting

Clustering consists of partitioning data objects into subsets called clusters according to some similarity criteria. This paper addresses a generalization called quasi-clustering that allows overlapping of clusters, and which we link to…

Artificial Intelligence · Computer Science 2020-02-13 Fred Glover , Said Hanafi , Gintaras Palubeckis

Despite the enormous success of machine learning models in various applications, most of these models lack resilience to (even small) perturbations in their input data. Hence, new methods to robustify machine learning models seem very…

Machine Learning · Computer Science 2020-10-30 Fariborz Salehi , Babak Hassibi