English
Related papers

Related papers: Genome-Wide Significance Levels and Weighted Hypot…

200 papers

As a common step in refining their scientific inquiry, investigators are often interested in performing some screening of a collection of given statistical hypotheses. For example, they may wish to determine whether any one of several…

Methodology · Statistics 2022-03-04 Adam Elder , Marco Carone , Peter Gilbert , Alex Luedtke

We consider testing for two-sample means of high dimensional populations by thresholding. Two tests are investigated, which are designed for better power performance when the two population mean vectors differ only in sparsely populated…

Methodology · Statistics 2014-10-13 Song Xi Chen , Jun Li , Ping-Shou Zhong

Parameter estimates for associated genetic variants, report ed in the initial discovery samples, are often grossly inflated compared to the values observed in the follow-up replication samples. This type of bias is a consequence of the…

Applications · Statistics 2011-04-15 Lizhen Xu , Radu V. Craiu , Lei Sun

Machine learning tools are commonly used in modern high energy physics (HEP) experiments. Different models, such as boosted decision trees (BDT) and artificial neural networks (ANN), are widely used in analyses and even in the software…

Data Analysis, Statistics and Probability · Physics 2016-12-21 A. Rogozhnikov

High-dimensional logistic regression is widely used in analyzing data with binary outcomes. In this paper, global testing and large-scale multiple testing for the regression coefficients are considered in both single- and two-regression…

Methodology · Statistics 2020-11-23 Rong Ma , T. Tony Cai , Hongzhe Li

Weighted sieves are used to detect numbers with at most $S$ prime factors with $S \in \mathbb{N}$ as small as possible. When one studies problems with two variables in somewhat symmetric roles (such as Chen primes, that is primes $p$ such…

Number Theory · Mathematics 2025-08-20 Kaisa Matomäki , Sebastian Zuniga Alterman

In this work, we consider hypothesis testing and anomaly detection on datasets where each observation is a weighted network. Examples of such data include brain connectivity networks from fMRI flow data, or word co-occurrence counts for…

Machine Learning · Statistics 2018-09-10 Guilherme Gomes , Vinayak Rao , Jennifer Neville

We propose new goodness-of-fit tests for the Poisson distribution. The testing procedure entails fitting a weighted Poisson distribution, which has the Poisson as a special case, to observed data. Based on sample data, we calculate an…

Methodology · Statistics 2024-02-21 Winnie Kirui , Elzanie Bothma , Marius Smuts , Anke Steyn , Jaco Visagie

Survival outcomes are common in comparative effectiveness studies and require unique handling because they are usually incompletely observed due to right-censoring. A ``once for all'' approach for causal inference with survival outcomes…

Methodology · Statistics 2021-12-21 Shuxi Zeng , Fan Li , Liangyuan Hu , Fan Li

This paper presents a hypothesis testing method given independent samples from a number of connected populations. The method is motivated by a forestry project for monitoring change in the strength of lumber. Traditional practice has been…

Statistics Theory · Mathematics 2015-05-15 Song Cai , Jiahua Chen , James V. Zidek

Equivalence tests, otherwise known as parity or similarity tests, are frequently used in ``bioequivalence studies" to establish practical equivalence rather than the usual statistical significant difference. In this article, we propose an…

Methodology · Statistics 2025-07-29 Daniel Ochieng

Weighted histogram in Monte-Carlo simulations is often used for the estimation of a probability density function. It is obtained as a result of random experiment with random events that have weights. In this paper the bin contents of…

Data Analysis, Statistics and Probability · Physics 2008-11-28 N. D. Gagunashvili

Statistical significance testing of differences in values of metrics like recall, precision and balanced F-score is a necessary part of empirical natural language processing. Unfortunately, we find in a set of experiments that many commonly…

Computation and Language · Computer Science 2007-05-23 Alexander Yeh

We present a subset selection algorithm designed to work with arbitrary model families in a practical batch setting. In such a setting, an algorithm can sample examples one at a time but, in order to limit overhead costs, is only able to…

Machine Learning · Computer Science 2023-01-31 Gui Citovsky , Giulia DeSalvo , Sanjiv Kumar , Srikumar Ramalingam , Afshin Rostamizadeh , Yunjuan Wang

Traditionally, heritability has been estimated using family-based methods such as twin studies. Advancements in molecular genomics have facilitated the development of alternative methods that utilise large samples of unrelated or related…

The energy test is a powerful binning-free, multi-dimensional and distribution-free tool that can be applied to compare a measurement to a given prediction (goodness-of-fit) or to check whether two data samples originate from the same…

Data Analysis, Statistics and Probability · Physics 2018-04-30 G. Zech

Science can be seen as a sequential process where each new study augments evidence to the existing knowledge. To have the best prospects to make an impact in this process, a new study should be designed optimally taking into account the…

Methodology · Statistics 2017-01-03 Juha Karvanen , Mikko J. Sillanpää

Motivated by the important problem of detecting association between genetic markers and binary traits in genome-wide association studies, we present a novel Bayesian model that establishes a hierarchy between markers and genes by defining…

Applications · Statistics 2016-06-22 Ian Johnston , Timothy Hancock , Hiroshi Mamitsuka , Luis Carvalho

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

Methodology · Statistics 2024-09-05 F. Richard Guo , Rajen D. Shah

Propensity score (PS) weighting methods are often used in non-randomized studies to adjust for confounding and assess treatment effects. The most popular among them, the inverse probability weighting (IPW), assigns weights that are…

Methodology · Statistics 2020-11-04 Yunji Zhou , Roland A. Matsouaka , Laine Thomas