English
Related papers

Related papers: Using Score Distributions to Compare Statistical S…

200 papers

Sentiment analysis or opinion mining has become an open research domain after proliferation of Internet and Web 2.0 social media. People express their attitudes and opinions on social media including blogs, discussion forums, tweets, etc.…

Information Retrieval · Computer Science 2013-09-17 Anuj sharma , Shubhamoy Dey

A definition for the statistical significance of a signal in an experiment is proposed by establishing a correlation between the observed p-value and the normal distribution integral probability, which is suitable for both counting…

Data Analysis, Statistics and Probability · Physics 2009-02-20 Yong-Sheng Zhu

In this work, we take a closer look at the evaluation of two families of methods for enriching information from knowledge graphs: Link Prediction and Entity Alignment. In the current experimental setting, multiple different scores are…

Machine Learning · Computer Science 2023-09-21 Max Berrendorf , Evgeniy Faerman , Laurent Vermue , Volker Tresp

Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…

Statistics Theory · Mathematics 2017-07-25 Abhik Ghosh , Ayanendranath Basu

Research often necessitates of samples, yet obtaining large enough samples is not always possible. When it is, the researcher may use one of two methods for deciding upon the required sample size: rules-of-thumb, quick yet uncertain, and…

Methodology · Statistics 2016-04-08 Jose D. Perezgonzalez

We investigate one/two-sample mean tests for high-dimensional compositional data when the number of variables is comparable with the sample size, as commonly encountered in microbiome research. Existing methods mainly focus on max-type test…

Statistics Theory · Mathematics 2024-04-15 Qianqian Jiang , Wenbo Li , Zeng Li

Given the prevalence of missing data in modern statistical research, a broad range of methods is available for any given imputation task. How does one choose the `best' imputation method in a given application? The standard approach is to…

Applications · Statistics 2022-12-01 Jeffrey Näf , Meta-Lina Spohn , Loris Michel , Nicolai Meinshausen

In this paper, we use the method of modified signed log-likelihood ratio test for the problem of testing the equality of correlation coefficients in two independent bivariate normal distributions. We compare this method with two other…

Methodology · Statistics 2016-06-01 M. R. Kazemi , A. A. Jafari

Missing data are frequently encountered in various disciplines and can be divided into three categories: missing completely at random (MCAR), missing at random (MAR) and missing not at random (MNAR). Valid statistical approaches to missing…

Methodology · Statistics 2021-05-28 Hairu Wang , Zhiping Lu , Yukun Liu

We introduce a new statistical test based on the observed spacings of ordered data. The statistic is sensitive to detect non-uniformity in random samples, or short-lived features in event time series. Under some conditions, this new test…

Methodology · Statistics 2022-10-27 Philipp Eller , Lolian Shtembari

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

Computation · Statistics 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

We consider the problem of testing significance of predictors in multivariate nonparametric quantile regression. A stochastic process is proposed, which is based on a comparison of the responses with a nonparametric quantile regression…

Methodology · Statistics 2012-06-15 Stanislav Volgushev , Melanie Birke , Holger Dette , Natalie Neumeyer

Despite the increasing use of citation-based metrics for research evaluation purposes, we do not know yet which metrics best deliver on their promise to gauge the significance of a scientific paper or a patent. We assess 17 network-based…

Social and Information Networks · Computer Science 2020-07-10 Shuqi Xu , Manuel Sebastian Mariani , Linyuan Lü , Matúš Medo

We theoretically analyze the problem of testing for $p$-hacking based on distributions of $p$-values across multiple studies. We provide general results for when such distributions have testable restrictions (are non-increasing) under the…

Econometrics · Economics 2022-05-13 Graham Elliott , Nikolay Kudrin , Kaspar Wuthrich

Fractional scoring has been proposed to avoid inconsistencies in the attribution of publications to percentile rank classes. Uncertainties and ambiguities in the evaluation of percentile ranks can be demonstrated most easily with small…

Digital Libraries · Computer Science 2013-03-25 Michael Schreiber

Several hypothesis testing methods have been proposed to validate the assumption of isotropy in spatial point patterns. A majority of these methods are characterised by an unknown distribution of the test statistic under the null hypothesis…

Methodology · Statistics 2025-04-09 Jakub J. Pypkowski , Adam M. Sykulski , James S. Martin

Recent discussions on alternative facts, fake news, and post truth politics have motivated research on creating technologies that allow people not only to access information, but also to assess the credibility of the information presented…

Information Retrieval · Computer Science 2017-08-25 Christina Lioma , Jakob Grue Simonsen , Birger Larsen

The meaning of randomization tests has become obscure in statistics education and practice over the last century. This article makes a fresh attempt at rectifying this core concept of statistics. A new term -- "quasi-randomization test" --…

Methodology · Statistics 2023-04-05 Yao Zhang , Qingyuan Zhao

Various measures have been proposed to quantify human-like social biases in word embeddings. However, bias scores based on these measures can suffer from measurement error. One indication of measurement quality is reliability, concerning…

Computation and Language · Computer Science 2021-09-13 Yupei Du , Qixiang Fang , Dong Nguyen

Intraclass correlation in bilateral data has been investigated in recent decades with various statistical methods. In practice, stratifying bilateral data by some control variables will provide more sophisticated statistical results to…

Methodology · Statistics 2023-03-24 Wanqing Tian , Chang-Xing Ma