English
Related papers

Related papers: Isochores Merit the Prefix 'Iso'

200 papers

Given a finite-valued sample $X_1,...,X_n$ we wish to test whether it was generated by a stationary ergodic process belonging to a family $H_0$, or it was generated by a stationary ergodic process outside $H_0$. We require the Type I error…

Statistics Theory · Mathematics 2014-12-30 Daniil Ryabko

Testing hypotheses of goodness-of-fit about mixture distributions on the basis of independent but not necessarily identically distributed random vectors is considered. The hypotheses are given by a specific distribution or by a family of…

Statistics Theory · Mathematics 2016-04-21 Daniel Gaigall

In benchmarking of Information Retrieval systems, the Wilcoxon signed-rank test is often treated as a safer alternative to the t-test. This belief is fueled by textbooks and recommendations that portray Wilcoxon as the proper non-parametric…

Information Retrieval · Computer Science 2026-04-29 Julián Urbano

Statistical significance testing is widely accepted as a means to assess how well a difference in effectiveness reflects an actual difference between systems, as opposed to random noise because of the selection of topics. According to…

Information Retrieval · Computer Science 2019-06-07 Julián Urbano , Harlley Lima , Alan Hanjalic

Clustering is part of unsupervised analysis methods that consist in grouping samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to…

Methodology · Statistics 2022-10-25 Benjamin Hivert , Denis Agniel , Rodolphe Thiébaut , Boris P Hejblum

This article presents a measure of semantic similarity in an IS-A taxonomy based on the notion of shared information content. Experimental evaluation against a benchmark set of human similarity judgments demonstrates that the measure…

Artificial Intelligence · Computer Science 2011-05-30 P. Resnik

Comparing large covariance matrices has important applications in modern genomics, where scientists are often interested in understanding whether relationships (e.g., dependencies or co-regulations) among a large number of genes vary…

Methodology · Statistics 2017-04-04 Jinyuan Chang , Wen Zhou , Wen-Xin Zhou , Lan Wang

Independent component (IC) models are a standard tool for representing multivariate data in statistics, signal processing, and machine learning. Despite the extensive use of IC models, much less attention has been given to goodness-of-fit…

Statistics Theory · Mathematics 2026-05-20 Mingshuo Liu , Siyao Wang , Miles E. Lopes

This paper deals with the comparison of several stationary processes with unequal sample sizes. We provide a detailed theoretical framework on the testing problem for equality of spectral densities in the bivariate case, after which the…

Statistics Theory · Mathematics 2012-07-25 Philip Preuß , Thimo Hildebrandt

Real-world data often violates the equal-variance assumption (homoscedasticity), making it essential to account for heteroscedastic noise in causal discovery. In this work, we explore heteroscedastic symmetric noise models (HSNMs), where…

Machine Learning · Computer Science 2025-04-22 Yingyu Lin , Yuxing Huang , Wenqin Liu , Haoran Deng , Ignavier Ng , Kun Zhang , Mingming Gong , Yi-An Ma , Biwei Huang

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

Statistics Theory · Mathematics 2020-03-09 Rémy Mariétan , Stephan Morgenthaler

Graph isomorphism is a problem for which there is no known polynomial-time solution. Nevertheless, assessing (dis)similarity between two or more networks is a key task in many areas, such as image recognition, biology, chemistry, computer…

Computation · Statistics 2022-06-28 Pierre Miasnikof , Alexander Y. Shestopaloff , Cristián Bravo , Yuri Lawryshyn

Hypothesis test plays a key role in uncertain statistics based on uncertain measure. This paper extends the parametric hypothesis of a single uncertain population to multiple cases, thereby addressing a broader range of scenarios. First, an…

Methodology · Statistics 2025-12-03 Fan Zhang , Zhiming Li

Growth hormone (GH) constitutes a set of closely related protein isoforms. In clinical practice, the disagreement of test results between commercially available ligand-binding assays is still an ongoing issue, and incomplete knowledge about…

Quantitative Methods · Quantitative Biology 2014-11-12 Cristian G. Arsene , Jürgen Kratzsch , André Henrion

This commentary discusses a recently proposed measure of heterogeneity of DNA sequences and compares with the measures of complexity.

adap-org · Physics 2012-05-07 Wentian Li

We propose two tests for the equality of covariance matrices between two high-dimensional populations. One test is on the whole variance--covariance matrices, and the other is on off-diagonal sub-matrices, which define the covariance…

Statistics Theory · Mathematics 2012-06-06 Jun Li , Song Xi Chen

Tens of thousands of simultaneous hypothesis tests are routinely performed in genomic studies to identify differentially expressed genes. However, due to unmeasured confounders, many standard statistical approaches may be substantially…

Methodology · Statistics 2025-03-18 Jin-Hong Du , Larry Wasserman , Kathryn Roeder

Bayesian Improved Surname Geocoding (BISG) is the most popular method for proxying race/ethnicity in voter registration files that do not contain it. This paper benchmarks BISG against a range of previously untested machine learning…

Machine Learning · Computer Science 2022-08-02 Ari Decter-Frain

An important class of two-sample multivariate homogeneity tests is based on identifying differences between the distributions of interpoint distances. While generating distances from point clouds offers a straightforward and intuitive way…

Methodology · Statistics 2024-08-21 Annika Betken , Aljosa Marjanovic , Katharina Proksch

Universal outlier hypothesis testing is studied in a sequential setting. Multiple observation sequences are collected, a small subset of which are outliers. A sequence is considered an outlier if the observations in that sequence are…

Statistics Theory · Mathematics 2014-11-27 Yun Li , Sirin Nitinawarat , Venugopal V. Veeravalli