English
Related papers

Related papers: Smaller $p$-values via indirect information

200 papers

Models are consistently treated as approximations and all procedures are consistent with this. They do not treat the model as being true. In this context $p$-values are one measure of approximation, a small $p$-value indicating a poor…

Other Statistics · Statistics 2016-11-21 Laurie Davies

The log-normal distribution is used to describe the positive data, that it has skewed distribution with small mean and large variance. This distribution has application in many sciences for example medicine, economics, biology and…

Methodology · Statistics 2015-08-10 Saba Aghadoust , Kamel Abdollahnezhad , Farhad Yaghmaei , Ali Akbar Jafari

In certain privacy-sensitive scenarios within fields such as clinical trial simulations, federated learning, and distributed learning, researchers often face the challenge of estimating correlations between variables without access to…

Methodology · Statistics 2025-08-05 Longwen Shang , Min Tsao , Xuekui Zhang

We derive inferential procedures for large sample sizes that remain valid under data-dependent significance levels (so-called "post-hoc valid inference"). Classical statistical tools require that the significance level -- the "type-I error"…

Statistics Theory · Mathematics 2026-03-10 Ben Chugg , Etienne Gauthier , Michael I. Jordan , Aaditya Ramdas , Ian Waudby-Smith

\citet{Rosenbaum83ps} introduced the notion of the propensity score and discussed its central role in causal inference with observational studies. Their paper, however, caused a fundamental incoherence with an early paper by…

Methodology · Statistics 2022-03-29 Peng Ding , Tianyu Guo

Statistical inference for extreme values of random events is difficult in practice due to low sample sizes and inaccurate models for the studied rare events. If prior knowledge for extreme values is available, Bayesian statistics can be…

Methodology · Statistics 2022-05-18 Tobias Kallehauge

Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. Specifically, PPI methods provide tighter confidence intervals by combining small amounts of human-labeled data with…

Machine Learning · Computer Science 2024-05-13 R. Alex Hofer , Joshua Maynez , Bhuwan Dhingra , Adam Fisch , Amir Globerson , William W. Cohen

In recent years, a number of results have been developed which connect information measures and estimation measures under various models, including, predominently, Gaussian and Poisson models. More recent results due to Taborda and…

Information Theory · Computer Science 2012-08-01 Dongning Guo

We propose a new simple procedure called Population-Mean-Based Aggregation (PMBA) that enables a principal to "aggregate" information about an unknown state of the world from agents without understanding the information structure among…

Theoretical Economics · Economics 2026-04-29 Yi-Chun Chen , Manuel Mueller-Frank , Mallesh M Pai

The problem of combining p-values is an old and fundamental one, and the classic assumption of independence is often violated or unverifiable in many applications. There are many well-known rules that can combine a set of arbitrarily…

Statistics Theory · Mathematics 2025-03-21 Matteo Gasparin , Ruodu Wang , Aaditya Ramdas

Conformal Prediction (CP) is a distribution-free uncertainty estimation framework that constructs prediction sets guaranteed to contain the true answer with a user-specified probability. Intuitively, the size of the prediction set encodes a…

Machine Learning · Computer Science 2025-02-18 Alvaro H. C. Correia , Fabio Valerio Massoli , Christos Louizos , Arash Behboodi

From a machine learning point of view, identifying a subset of relevant features from a real data set can be useful to improve the results achieved by classification methods and to reduce their time and space complexity. To achieve this…

Machine Learning · Computer Science 2017-05-23 Pietro Cassara , Alessandro Rozza , Mirco Nanni

This paper proposes a Bayesian method for estimating the parameters of a normal distribution when only limited summary statistics (sample mean, minimum, maximum, and sample size) are available. To estimate the parameters of a normal…

Methodology · Statistics 2024-11-21 Tomoki Matsumoto

The core objective of modelling recommender systems from implicit feedback is to maximize the positive sample score $s_p$ and minimize the negative sample score $s_n$, which can usually be summarized into two paradigms: the pointwise and…

Information Retrieval · Computer Science 2022-03-01 Jianhuan Zhuo , Qiannan Zhu , Yinliang Yue , Yuhong Zhao

When analyzing incomplete data, is it better to use multiple imputation (MI) or full information maximum likelihood (ML)? In large samples ML is clearly better, but in small samples ML's usefulness has been limited because ML commonly uses…

Methodology · Statistics 2017-03-24 Paul T. von Hippel

We consider the problem of classification with a (peer-to-peer) network of heterogeneous and partially informative agents, each receiving local data generated by an underlying true class, and equipped with a classifier that can only…

Machine Learning · Computer Science 2024-10-01 Tong Yao , Shreyas Sundaram

There is a growing need for the ability to analyse interval-valued data. However, existing descriptive frameworks to achieve this ignore the process by which interval-valued data are typically constructed; namely by the aggregation of…

Methodology · Statistics 2019-03-08 Xin Zhang , Boris Beranger , Scott A. Sisson

Link prediction is pervasively employed to uncover the missing links in the snapshots of real-world networks, which are usually obtained from kinds of sampling methods. Contrarily, in the previous literature, in order to evaluate the…

Social and Information Networks · Computer Science 2014-10-28 Jichang Zhao , Xu Feng , Li Dong , Xiao Liang , Ke Xu

In Official Statistics, interest for data integration has been increasingly growing, due to the need of extracting information from different sources. However, the effects of these procedures on the validity of the resulting statistical…

We derive a well-defined renormalized version of mutual information that allows to estimate the dependence between continuous random variables in the important case when one is deterministically dependent on the other. This is the situation…

Machine Learning · Computer Science 2021-05-26 Leopoldo Sarra , Andrea Aiello , Florian Marquardt