Related papers: Comment on Glenn Shafer's "Testing by betting"
We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of $(X,Y)$, and additional unlabeled samples…
One often finds in the literature connections between measures of fairness and measures of feature importance employed to interpret trained classifiers. However, there seems to be no study that compares fairness measures and feature…
When testing a statistical hypothesis, is it legitimate to deliberate on the basis of initial data about whether and how to collect further data? Game-theoretic probability's fundamental principle for testing by betting says yes, provided…
Courses on the mathematics of gambling have been offered by a number of colleges and universities, and for a number of reasons. In the past 15 years, at least seven potential textbooks for such a course have been published. In this article…
A question that comes up repeatedly is how to combine the results of two experiments if all that is known is that one experiment had a n-sigma effect and another experiment had a m-sigma effect. This question is not well-posed: depending on…
In an attempt to provide an answer to the increasing criticism against p-values and to bridge the gap between statistical inference and prediction modelling, we introduce the probability of improved prediction (PIP). In general, the PIP is…
Statistical significance testing plays an important role when drawing conclusions from experimental results in NLP papers. Particularly, it is a valuable tool when one would like to establish the superiority of one algorithm over another.…
Prediction markets are useful for estimating probabilities of claims whose truth will be revealed at some fixed time -- this includes questions about the values of real-world events (i.e. statistical uncertainty), and questions about the…
In this paper, we demonstrate that a new measure of evidence we developed called the Dempster-Shafer p-value which allow for insights and interpretations which retain most of the structure of the p-value while covering for some of the…
Comment on "Quantifying the Fraction of Missing Information for Hypothesis Testing in Statistical and Genetic Studies" [arXiv:1102.2774]
Comment on "Quantifying the Fraction of Missing Information for Hypothesis Testing in Statistical and Genetic Studies" [arXiv:1102.2774]
Comment on "Harold Jeffreys's Theory of Probability Revisited" [arXiv:0804.3173]
Comment on "Harold Jeffreys's Theory of Probability Revisited" [arXiv:0804.3173]
Comment on "Harold Jeffreys's Theory of Probability Revisited" [arXiv:0804.3173]
Comment on "Harold Jeffreys's Theory of Probability Revisited" [arXiv:0804.3173]
Financial and gambling markets are ostensibly similar and hence strategies from one could potentially be applied to the other. Financial markets have been extensively studied, resulting in numerous theorems and models, while gambling…
The main purpose of the paper is the proof of a cardinal inequality for a space with points $G_\delta$, obtained with the help of a long version of the Menger game. This result improves a similar one of Scheepers and Tall.
We study the problems of sequential nonparametric two-sample and independence testing. Sequential tests process data online and allow using observed data to decide whether to stop and reject the null hypothesis or to collect more data,…
This note concerns a search for publications in which the pragmatic concept of a test as conducted in the practice of software testing is formalized, a theory about software testing based on such a formalization is presented or it is…
The notion of p-value is a fundamental concept in statistical inference and has been widely used for reporting outcomes of hypothesis tests. However, p-value is often misinterpreted, misused or miscommunicated in practice. Part of the issue…