Related papers: Benford-Newcomb Subsequences for Fraud Detection
When some treatments are ordered according to the categories of an ordinal categorical variable (e.g., extent of side effects) in a monotone order, one might be interested in knowing wether the treatments are equally effective or not. One…
It is pointed out that the language of quotient groups and wrapped distributions allows an elementary discussion of Benford's Law, and adds arguments supporting wide-spread observability of this statistics.
The package BeyondBenford compares the goodness of fit of Benford's and Blondeau Da Silva's (BDS's) digit distributions in a dataset. The package is used to check whether the data distribution is consistent with theoretical distributions…
The likelihood principle makes strong claims about the nature of statistical evidence but is controversial. Its claims are undermined by the existence of several examples that are assumed to show that it allows, with unity probability,…
In the era of fast-paced precision medicine, observational studies play a major role in properly evaluating new treatments in clinical practice. Yet, unobserved confounding can significantly compromise causal conclusions drawn from…
We study the properties of several likelihood-based statistics commonly used in testing for the presence of a known signal under a mixture model with known background, but unknown signal fraction. Under the null hypothesis of no signal, all…
Pre-taxation analysis plays a crucial role in ensuring the fairness of public revenue collection. It can also serve as a tool to reduce the risk of tax avoidance, one of the UK government's concerns. Our report utilises pre-tax income…
We study the individual digits for the absolute value of the characteristic polynomial for the Circular $\beta$-Ensemble. We show that, in the large $N$ limit, the first digits obey Benford's Law and the further digits become uniformly…
Some properties of diagonal binomial coefficients were studied in respect to frequency of their units digits. An approach was formulated that led to use of difference tables to predict if certain units digits can appear in the values of…
We estimate the local laws of the distribution of the middle prime factor of an integer, defined according to multiplicity or not. An asymptotic estimate with effective remainder is provided for a wide range of values. In particular this…
Data analysis in HEP experiments often uses binned likelihood from data and finite Monte Carlo sample. Statistical uncertainty of Monte Carlo sample has been introduced in Frequentist Inference in some literatures, but they are not suitable…
A universal First-Letter Law (FLL) is derived and described. It predicts the percentages of first letters for words in novels. The FLL is akin to Benford's law (BL) of first digits, which predicts the percentages of first digits in a data…
The likelihood ratio is a crucial quantity for statistical inference in science that enables hypothesis testing, construction of confidence intervals, reweighting of distributions, and more. Many modern scientific applications, however,…
We formulate the problem of fake news detection using distributed fact-checkers (agents) with unknown reliability. The stream of news/statements is modeled as an independent and identically distributed binary source (to represent true and…
The Birnbaum-Saunders regression model is commonly used in reliability studies. We address the issue of performing inference in this class of models when the number of observations is small. We show that the likelihood ratio test tends to…
The generalized gamma distribution shows up in many problems related to engineering, hydrology as well as survival analysis. Earlier work has been done that estimated the deviation of the exponential and the Weibull distribution from…
This paper studies how insurers can chose which claims to investigate for fraud. Given a prediction model, typically only claims with the highest predicted propability of being fraudulent are investigated. We argue that this can lead to…
Two separate statistical tests are described and developed in order to test un-binned data sets for adherence to the power-law form. The first test employs the TP-statistic, a function defined to deviate from zero when the sample deviates…
Many proofs in discrete mathematics and theoretical computer science are based on the probabilistic method. To prove the existence of a good object, we pick a random object and show that it is bad with low probability. This method is…
Out-of-distribution (OOD) detection is critical for ensuring the reliability of deep learning systems, particularly in safety-critical applications. Likelihood-based deep generative models have historically faced criticism for their…