English
Related papers

Related papers: Comment: Quantifying the Fraction of Missing Infor…

200 papers

Comment on ``Boosting Algorithms: Regularization, Prediction and Model Fitting'' [arXiv:0804.2752]

Methodology · Statistics 2008-12-18 Trevor Hastie

We consider the problem of estimating the number of false null hypotheses among a very large number of independently tested hypotheses, focusing on the situation in which the proportion of false null hypotheses is very small. We propose a…

Statistics Theory · Mathematics 2007-06-13 Nicolai Meinshausen , John Rice

Gene expression and phenotype association can be affected by potential unmeasured confounders from multiple sources, leading to biased estimates of the associations. Since genetic variants largely explain gene expression variations, they…

Methodology · Statistics 2019-10-23 Jiarui Lu , Hongzhe Li

We describe a statistical hypothesis test for the presence of a signal. The test allows the researcher to fix the signal location and/or width a priori, or perform a search to find the signal region that maximizes the signal. The background…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Wolfgang Rolke , Angel Lopez

In cell biology, statistical analysis means testing the hypothesis that there was no effect. This weak form of hypothesis testing neglects effect size, is universally misinterpreted, and is disastrously prone to error when combined with…

Other Quantitative Biology · Quantitative Biology 2025-05-13 Josh L. Morgan

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

Machine Learning · Statistics 2015-09-22 Qiyi Lu , Xingye Qiao

Rejoinder to ``Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data'' [arXiv:0804.2958]

Methodology · Statistics 2008-12-18 Joseph D. Y. Kang , Joseph L. Schafer

Information Value (IV) is a widely used technique for feature selection prior to the modeling phase, particularly in credit scoring and related domains. However, conventional IV-based practices rely on fixed empirical thresholds, which lack…

Statistics Theory · Mathematics 2026-01-28 Helder Rojas , Cirilo Alvarez , Nilton Rojas

Determining whether a provided context contains sufficient information to answer a question is a critical challenge for building reliable question-answering systems. While simple prompting strategies have shown success on factual questions,…

Computation and Language · Computer Science 2026-03-24 Akriti Jain , Aparna Garimella

A new method based on the rejection sampling for finding statistical tests is proposed. This method is conceptually intuitive, easy to implement, and applicable for arbitrary dimension. To illustrate its potential applicability, three…

Methodology · Statistics 2026-03-11 Markku Kuismin

We provide a means of computing and estimating the asymptotic distributions of statistics based on an outer minimization of an inner maximization. Such test statistics, which arise frequently in moment models, are of special interest in…

Econometrics · Economics 2024-04-17 Isaac Loh

It is quite common in modern research, for a researcher to test many hypotheses. The statistical (frequentist) hypothesis testing framework, does not scale with the number of hypotheses in the sense that naively performing many hypothesis…

Methodology · Statistics 2013-06-26 Jonathan Rosenblatt

This paper considers the problem of defining a measure of redundant information that quantifies how much common information two or more random variables specify about a target random variable. We discussed desired properties of such a…

Information Theory · Computer Science 2023-07-19 Virgil Griffith , Tracey Ho

Recently, it is well recognized that hypothesis testing has deep relations with other topics in quantum information theory as well as in classical information theory. These relations enable us to derive precise evaluation in the…

Quantum Physics · Physics 2017-09-25 Masahito Hayashi

Relation of genome sizes to organisms complexity is still described rather equivocally. Neither the number of genes (G-value), nor the total amount of DNA (C-value) correlates consistently with phenotype complexity. Using information theory…

Genomics · Quantitative Biology 2007-05-23 Dmitri V. Parkhomchuk

In this paper we propose a Bayesian answer to testing problems when the hypotheses are not well separated. The idea of the method is to study the posterior distribution of a discrepancy measure between the parameter and the model we want to…

Statistics Theory · Mathematics 2017-06-28 Jean-Bernard Salomond

Estimation of the $\phi$-divergence between two unknown probability distributions using empirical data is a fundamental problem in information theory and statistical learning. We consider a multi-variate generalization of the data dependent…

Probability · Mathematics 2018-01-04 Fengqiao Luo , Sanjay Mehrotra

There has been much interest in the nonparametric testing of conditional independence in the econometric and statistical literature, but the simplest and potentially most useful method, based on the sample partial correlation, seems to have…

Statistics Theory · Mathematics 2020-05-27 Wicher Bergsma

We formulate nonparametric and semiparametric hypothesis testing of multivariate stationary linear time series in a unified fashion and propose new test statistics based on estimators of the spectral density matrix. The limiting…

Statistics Theory · Mathematics 2009-09-03 Yoshihiro Yajima , Yasumasa Matsuda

We answer a question of Zeilberger and Zeilberger about certain partition statistics.

Combinatorics · Mathematics 2018-11-09 Christopher Ryba