English
Related papers

Related papers: Investigating the agreement between methods of dif…

200 papers

Difference-in-differences is a widely-used evaluation strategy that draws causal inference from observational panel data. Its causal identification relies on the assumption of parallel trends, which is scale dependent and may be…

Applications · Statistics 2019-06-25 Peng Ding , Fan Li

A well-known metric for quantifying the similarity between two clusterings is the adjusted mutual information. Compared to mutual information, a corrective term based on random permutations of the labels is introduced, preventing two…

Machine Learning · Computer Science 2021-03-24 Denys Lazarenko , Thomas Bonald

In change-point analysis, one aims at finding the locations of abrupt distributional changes (if any) in a sequence of multivariate observations. In this article, we propose some nonparametric methods based on averages of pairwise distances…

Statistics Theory · Mathematics 2025-11-14 Spandan Ghoshal , Bilol Banerjee , Anil K. Ghosh

Recent work has sought to quantify large language model uncertainty to facilitate model control and modulate user trust. Previous works focus on measures of uncertainty that are theoretically grounded or reflect the average overt behavior…

Computation and Language · Computer Science 2025-03-18 Kyle Moore , Jesse Roberts , Daryl Watson , Pamela Wisniewski

Unmeasured confounding is a key threat to reliable causal inference based on observational studies. Motivated from two powerful natural experiment devices, the instrumental variables and difference-in-differences, we propose a new method…

Methodology · Statistics 2021-11-09 Ting Ye , Ashkan Ertefaie , James Flory , Sean Hennessy , Dylan S. Small

The reliability of large language models (LLMs) is greatly compromised by their tendency to hallucinate, underscoring the need for precise identification of knowledge gaps within LLMs. Various methods for probing such gaps exist, ranging…

Computation and Language · Computer Science 2025-06-02 Raoyuan Zhao , Abdullatif Köksal , Ali Modarressi , Michael A. Hedderich , Hinrich Schütze

Besides the classical distinction of correlation and dependence, many dependence measures bear further pitfalls in their application and interpretation. The aim of this paper is to raise and recall awareness of some of these limitations by…

Methodology · Statistics 2020-04-17 Björn Böttcher

Many predictions are probabilistic in nature; for example, a prediction could be for precipitation tomorrow, but with only a 30 percent chance. Given both the predictions and the actual outcomes, "reliability diagrams" (also known as…

Methodology · Statistics 2020-07-20 Mark Tygert

Bell inequalities rely on an assumption that the probabilities of adopting configurations of hidden variables describing a system prior to measurement are independent of the choice of measured physical property, also known as measurement…

Quantum Physics · Physics 2025-09-05 Sophia M. Walls , Ian J. Ford

The standardized mean difference (SMD) is a widely used measure of effect size, particularly common in psychology, clinical trials, and meta-analysis involving continuous outcomes. Traditionally, under the equal variance assumption, the SMD…

Methodology · Statistics 2025-06-05 Jiandong Shi , Xiaochen Zhang , Lu Lin , Hiu Yee Kwan , Tiejun Tong

New inference methods for the multivariate coefficient of variation and its reciprocal, the standardized mean, are presented. While there are various testing procedures for both parameters in the univariate case, it is less known how to do…

Methodology · Statistics 2020-03-31 Marc Ditzhaus , Łukas Smaga

Combining several independent measurements of the same physical quantity is one of the most important tasks in metrology. Small samples, biased input estimates, not always adequate reported uncertainties, and unknown error distribution make…

Data Analysis, Statistics and Probability · Physics 2026-04-22 Zinovy Malkin

Multiple raters are often needed to be used interchangeably in practice for measurement or evaluation. Assessing agreement among these multiple raters via agreement indices are necessary before their participation. While the intuitively…

Methodology · Statistics 2020-06-09 Tongrong Wang , Huiman X. Barnhart

We consider a permutation method for testing whether observations given in their natural pairing exhibit an unusual level of similarity in situations where any two observations may be similar at some unknown baseline level. Under a null…

Statistics Theory · Mathematics 2007-06-13 Larry Goldstein , Yosef Rinott

To compare different forecasting methods on demand series we require an error measure. Many error measures have been proposed, but when demand is intermittent some become inapplicable, some give counter-intuitive results, and there is no…

Methodology · Statistics 2015-01-20 S. D. Prestwich , R. Rossi , S. A. Tarim , B. Hnich

We consider a three-level meta-analysis of standardized mean differences. The standard method of estimation uses inverse-variance weights and REML/PL estimation of variance components for the random effects. We introduce new moment-based…

Methodology · Statistics 2024-11-05 Elena Kulinskaya , David C. Hoaglin

We review the alternative proposals introduced recently in the literature to update the standard formula to estimate the uncertainty on the mean of repeated measurements, and we compare their performances on synthetic examples with normal…

Data Analysis, Statistics and Probability · Physics 2022-09-13 Pascal Pernot , Jean-Paul Berthet

In recent years, there has been a strong interest in measuring the available bandwidth of network paths. Several methods and techniques have been proposed and various measurement tools have been developed and evaluated. However, there have…

Networking and Internet Architecture · Computer Science 2007-06-28 Ahmed Ait Ali , Fabien Michaut , Francis Lepage

The median absolute deviation (MAD) is a robust measure of scale that is simple to implement and easy to interpret. Motivated by this, we introduce interval estimators of the MAD to make reliable inferences for dispersion for a single…

Statistics Theory · Mathematics 2024-08-06 Chandima N. P. G. Arachchige , Luke A. Prendergast

Adjusted similarity measures, such as Cohen's kappa for inter-rater reliability and the adjusted Rand index used to compare clustering algorithms, are a vital tool for comparing discrete labellings. These measures are intended to have the…

Methodology · Statistics 2026-01-16 William L. Lippitt , Edward J. Bedrick , Nichole E. Carlson