Related papers: Comment on Glenn Shafer's "Testing by betting"
This note outlines an approach to stress testing of covariance of financial time series, in the context of financial risk management. It discusses how the geodesic distance between covariance matrices implies a notion of plausibility of…
This article aims at clarifying the language and practice of scientific experiment, mainly by hooking observability on calculability.
In this paper, we address the problem of testing exchangeability of a sequence of random variables, $X_1, X_2,\cdots$. This problem has been studied under the recently popular framework of testing by betting. But the mapping of testing…
We introduce the prediction value (PV) as a measure of players' informational importance in probabilistic TU games. The latter combine a standard TU game and a probability distribution over the set of coalitions. Player $i$'s prediction…
We introduce and analyze several variations of Penney's game aimed to find a more equitable game.
Unsupervised performance estimation, or evaluating how well models perform on unlabeled data is a difficult task. Recently, a method was proposed by Garg et al. [2022] which performs much better than previous methods. Their method relies on…
Comment on ``Tests of scaling and universality of the distributions of trade size and share volume: Evidence from three distinct markets" by Plerou and Stanley, Phys. Rev. E 76, 046109 (2007)
Combining dependent tests of significance has broad applications but the $p$-value calculation is challenging. Current moment-matching methods (e.g., Brown's approximation) for Fisher's combination test tend to significantly inflate the…
In this note, we show that classical statistical tests for randomness are language dependent.
We point out that the traditional notion of test statistic is too narrow, and we propose a natural generalization that is arguably maximal. The study is restricted to simple statistical hypotheses.
In many fields of research null hypothesis significance tests and p values are the accepted way of assessing the degree of certainty with which research results can be extrapolated beyond the sample studied. However, there are very serious…
This paper is a comment to M Wilkinson, EPL 106 (2014) 40001, arXiv:1401.4620 [physics.ao-ph,cond-mat.soft], which draws conclusion from our data that are at variance with our observations.
Besides the classical distinction of correlation and dependence, many dependence measures bear further pitfalls in their application and interpretation. The aim of this paper is to raise and recall awareness of some of these limitations by…
As a convention, p-value is often computed in frequentist hypothesis testing and compared with the nominal significance level of 0.05 to determine whether or not to reject the null hypothesis. The smaller the p-value, the more significant…
Stemming from de Finetti's work on finitely additive coherent probabilities, the paradigm of coherence has been applied to many uncertainty calculi in order to remove structural restrictions on the domain of the assessment. Three possible…
A few remarks on hep-ph/9612213 are given.
Selective inference is a subfield of statistics that enables valid inference after selection of a data-dependent question. In this paper, we introduce selectively dominant p-values, a class of p-values that allow practitioners to easily…
This note gives an informal overview of the proof in our paper "Borel Conjecture and Dual Borel Conjecture", see arXiv:1105.0823.
This paper studies hypothesis testing and parameter estimation in the context of the divide and conquer algorithm. In a unified likelihood based framework, we propose new test statistics and point estimators obtained by aggregating various…
Probability forecasts for binary events play a central role in many applications. Their quality is commonly assessed with proper scoring rules, which assign forecasts a numerical score such that a correct forecast achieves a minimal…