English
Related papers

Related papers: Statistical Verification of Linear Classifiers

200 papers

Open-generation bias benchmarks evaluate social biases in Large Language Models (LLMs) by analyzing their outputs. However, the classifiers used in analysis often have inherent biases, leading to unfair conclusions. This study examines such…

Computation and Language · Computer Science 2025-01-22 Nathaniel Demchak , Xin Guan , Zekun Wu , Ziyi Xu , Adriano Koshiyama , Emre Kazim

In distributed and federated learning, heterogeneity across data sources remains a major obstacle to effective model aggregation and convergence. We focus on feature heterogeneity and introduce energy distance as a sensitive measure for…

Machine Learning · Statistics 2025-01-28 Mengchen Fan , Baocheng Geng , Roman Shterenberg , Joseph A. Casey , Zhong Chen , Keren Li

A criterion is proposed for testing hypothesis about the nature of the error variance in the dependent variable in linear model, which separates correctly and incorrectly specified models. In the former only measurement errors determine the…

Methodology · Statistics 2019-11-19 Alexander Kukush , Igor Mandel

We define a novel, basic, unsupervised learning problem - learning the lowest density homogeneous hyperplane separator of an unknown probability distribution. This task is relevant to several problems in machine learning, such as…

Machine Learning · Computer Science 2009-01-22 Shai Ben-David , Tyler Lu , David Pal , Miroslava Sotakova

Traditional LLM alignment methods are vulnerable to heterogeneity in human preferences. Fitting a na\"ive probabilistic model to pairwise comparison data (say over prompt-completion pairs) yields an inconsistent estimate of the…

Artificial Intelligence · Computer Science 2025-10-30 Ali Aouad , Aymane El Gadarri , Vivek F. Farias

This paper proposes nonparametric two-sample tests for the direct comparison of the probabilities of a particular transition between states of a continuous time nonhomogeneous Markov process with a finite state space. The proposed tests are…

Methodology · Statistics 2020-02-24 Giorgos Bakoyannis

Classical results on the statistical complexity of linear models have commonly identified the norm of the weights $\|w\|$ as a fundamental capacity measure. Generalizations of this measure to the setting of deep networks have been varied,…

Machine Learning · Statistics 2019-10-24 Ryan Theisen , Jason M. Klusowski , Huan Wang , Nitish Shirish Keskar , Caiming Xiong , Richard Socher

We study the problem of generalized uniformity testing \cite{BC17} of a discrete probability distribution: Given samples from a probability distribution $p$ over an {\em unknown} discrete domain $\mathbf{\Omega}$, we want to distinguish,…

Data Structures and Algorithms · Computer Science 2017-09-08 Ilias Diakonikolas , Daniel M. Kane , Alistair Stewart

Algorithms are now routinely used to make consequential decisions that affect human lives. Examples include college admissions, medical interventions or law enforcement. While algorithms empower us to harness all information hidden in vast…

Machine Learning · Computer Science 2020-12-10 Bahar Taskesen , Jose Blanchet , Daniel Kuhn , Viet Anh Nguyen

In this paper, a methodology to detect inconsistencies in classification-based image steganalysis is presented. The proposed approach uses two classifiers: the usual one, trained with a set formed by cover and stego images, and a second…

Cryptography and Security · Computer Science 2019-09-27 Daniel Lerch-Hostalot , David Megías

We introduce a two-parameter ensemble of random discrete-time Markov models that simultaneously captures critical slowing down and broken detailed balance. Extending a previously studied heterogeneous Markov ensemble, we incorporate…

Disordered Systems and Neural Networks · Physics 2026-02-06 Faheem Mosam , Eric De Giuli

Laboratory test results are an important and generally high dimensional component of a patient's Electronic Health Record (EHR). We train embedding representations (via Word2Vec and GloVe) for LOINC codes of laboratory tests from the EHRs…

Machine Learning · Computer Science 2019-08-02 Lorenzo A. Rossi , Chad Shawber , Janet Munu , Finly Zachariah

Cross-study replicability is a powerful model evaluation criterion that emphasizes generalizability of predictions. When training cross-study replicable prediction models, it is critical to decide between merging and treating the studies…

Machine Learning · Statistics 2022-07-14 Cathy Shyr , Pragya Sur , Giovanni Parmigiani , Prasad Patil

This paper is to prove the asymptotic normality of a statistic for detecting the existence of heteroscedasticity for linear regression models without assuming randomness of covariates when the sample size $n$ tends to infinity and the…

Statistics Theory · Mathematics 2018-06-11 Zhidong Bai , Guangming Pan , Yanqing Yin

Testing covariance structure is of importance in many areas of statistical analysis, such as microarray analysis and signal processing. Conventional tests for finite-dimensional covariance cannot be applied to high-dimensional data in…

Statistics Theory · Mathematics 2013-10-31 Rongmao Zhang , Liang Peng , Ruodu Wang

We study the robustness of classifiers to various kinds of random noise models. In particular, we consider noise drawn uniformly from the $\ell\_p$ ball for $p \in [1, \infty]$ and Gaussian noise with an arbitrary covariance matrix. We…

Machine Learning · Computer Science 2018-06-25 Jean-Yves Franceschi , Alhussein Fawzi , Omar Fawzi

Detecting and locating changes in highly multivariate data is a major concern in several current statistical applications. In this context, the first contribution of the paper is a novel non-parametric two-sample homogeneity test for…

Statistics Theory · Mathematics 2012-02-13 Alexandre Lung-Yut-Fong , Céline Lévy-Leduc , Olivier Cappé

We give a formal verification procedure that decides whether a classifier ensemble is robust against arbitrary randomized attacks. Such attacks consist of a set of deterministic attacks and a distribution over this set. The…

Machine Learning · Computer Science 2020-07-10 Dennis Gross , Nils Jansen , Guillermo A. Pérez , Stephan Raaijmakers

Neural network models have a reputation for being black boxes. We propose to monitor the features at every layer of a model and measure how suitable they are for classification. We use linear classifiers, which we refer to as "probes",…

Machine Learning · Statistics 2018-11-26 Guillaume Alain , Yoshua Bengio

In this paper we suggest two statistical hypothesis tests for the regression function of binary classification based on conditional kernel mean embeddings. The regression function is a fundamental object in classification as it determines…

Machine Learning · Statistics 2022-06-22 Ambrus Tamás , Balázs Csanád Csáji