English
Related papers

Related papers: Testing for Outliers with Conformal p-values

200 papers

Likelihood-based methods of statistical inference provide a useful general methodology that is appealing, as a straightforward asymptotic theory can be applied for their implementation. It is important to assess the relationships between…

Statistics Theory · Mathematics 2015-03-20 Thomas J. DiCiccio , Todd A. Kuffner , G. Alastair Young , Russell Zaretzki

We are concerned with testing replicability hypotheses for many endpoints simultaneously. This constitutes a multiple test problem with composite null hypotheses. Traditional $p$-values, which are computed under least favourable parameter…

Methodology · Statistics 2020-02-26 Anh-Tuan Hoang , Thorsten Dickhaus

Verifying that a statistically significant result is scientifically meaningful is not only good scientific practice, it is a natural way to control the Type I error rate. Here we introduce a novel extension of the p-value - a…

Methodology · Statistics 2018-07-04 Jeffrey D. Blume , Lucy DAgostino McGowan , William D. Dupont , Robert A. Greevy

In this work we provide a review of basic ideas and novel developments about Conformal Prediction -- an innovative distribution-free, non-parametric forecasting method, based on minimal assumptions -- that is able to yield in a very…

Machine Learning · Computer Science 2024-02-01 Matteo Fontana , Gianluca Zeni , Simone Vantini

In this paper, we present a new classifier, which integrates significance testing results over different random subspaces to yield consensus p-values for quantifying the uncertainty of classification decision. The null hypothesis is that…

Machine Learning · Computer Science 2024-10-17 Zengyou He , Zerun Li , Junjie Dong , Xinying Liu , Mudi Jiang , Lianyu Hu

We develop a new method for generating prediction sets that combines the flexibility of conformal methods with an estimate of the conditional distribution $P_{Y \mid X}$. Existing methods, such as conformalized quantile regression and…

Machine Learning · Statistics 2024-10-10 Vincent Plassier , Alexander Fishkov , Mohsen Guizani , Maxim Panov , Eric Moulines

We build a valid p-value based on a concentration inequality for bounded random variables introduced by Pelekis, Ramon and Wang. The motivation behind this work is the calibration of predictive algorithms in a distribution-free setting. The…

Machine Learning · Statistics 2024-05-16 Joaquin Alvarez

We consider statistical hypothesis testing simultaneously over a fairly general, possibly uncountably infinite, set of null hypotheses, under the assumption that a suitable single test (and corresponding $p$-value) is known for each…

Methodology · Statistics 2014-02-10 Gilles Blanchard , Sylvain Delattre , Etienne Roquain

Outlying observations are frequently encountered across a wide spectrum of scientific domains, posing notable challenges to the generalizability of statistical models and the reproducibility of downstream analysis. They are identified…

Methodology · Statistics 2026-03-17 Dongliang Zhang , Masoud Asgharian , Martin A. Lindquist

An overview of existing nonparametric tests of extreme-value dependence is presented. Given an i.i.d.\ sample of random vectors from a continuous distribution, such tests aim at assessing whether the underlying unknown copula is of the {\em…

Methodology · Statistics 2014-10-27 Axel Bücher , Ivan Kojadinovic

We propose a post-hoc adaptive conformal anomaly detection method for monitoring time series that leverages predictions from pre-trained foundation models without requiring additional fine-tuning. Our method yields an interpretable anomaly…

Machine Learning · Computer Science 2026-04-23 Natalia Martinez Gil , Fearghal O'Donncha , Wesley M. Gifford , Nianjun Zhou , Dhaval C. Patel , Roman Vaculin

This study addresses an important gap in time series outlier detection by proposing a novel problem setting: long-term outlier prediction. Conventional methods primarily focus on immediate detection by identifying deviations from normal…

Inequalities are key tools to prove FDR control of a multiple test. The present paper studies upper and lower bounds for the FDR under various dependence structures of p-values, namely independence, reverse martingale dependence and…

Statistics Theory · Mathematics 2015-02-18 Philipp Heesen , Arnold Janssen

Introductory statistical inference texts and courses treat the point estimation, hypothesis testing, and interval estimation problems separately, with primary emphasis on large-sample approximations. Here I present an alternative approach…

Other Statistics · Statistics 2017-07-14 Ryan Martin

Many computer vision tasks involve processing large amounts of data contaminated by outliers, which need to be detected and rejected. While outlier detection methods based on robust statistics have existed for decades, only recently have…

Computer Vision and Pattern Recognition · Computer Science 2017-04-14 Chong You , Daniel P. Robinson , René Vidal

Adaptive experiments use preliminary analyses of the data to inform further course of action and are commonly used in many disciplines including medical and social sciences. Because the null hypothesis and experimental design are…

Methodology · Statistics 2026-05-26 Tobias Freidling , Qingyuan Zhao , Zijun Gao

The presence of outlying observations may adversely affect statistical testing procedures that result in unstable test statistics and unreliable inferences depending on the distortion in parameter estimates. In spite of the fact that the…

Methodology · Statistics 2021-04-19 Beste Hamiye Beyaztas , Soutir Bandyopadhyay , Abhijit Mandal

This work proposes an unsupervised learning framework for trajectory (sequence) outlier detection that combines ranking tests with user sequence models. The overall framework identifies sequence outliers at a desired false positive rate…

Machine Learning · Computer Science 2021-11-09 Mohamed A. Zahran , Leonardo Teixeira , Vinayak Rao , Bruno Ribeiro

We study the problem of Out-of-Distribution (OOD) detection, that is, detecting whether a learning algorithm's output can be trusted at inference time. While a number of tests for OOD detection have been proposed in prior work, a formal…

Machine Learning · Statistics 2023-09-19 Akshayaa Magesh , Venugopal V. Veeravalli , Anirban Roy , Susmit Jha

A popular data-driven method for choosing the bandwidth in standard kernel regression is cross-validation. Even when there are outliers in the data, robust kernel regression can be used to estimate the unknown regression curve [Robust and…

Statistics Theory · Mathematics 2007-06-13 Denis Heng-Yan Leung