English
Related papers

Related papers: Generalized Rescaled Polya urn and its statistical…

200 papers

Clustering task of mixed data is a challenging problem. In a probabilistic framework, the main difficulty is due to a shortage of conventional distributions for such data. In this paper, we propose to achieve the mixed data clustering with…

Methodology · Statistics 2015-10-01 Matthieu Marbac , Christophe Biernacki , Vincent Vandewalle

There is no, nor will there ever be, single best clustering algorithm. Nevertheless, we would still like to be able to distinguish between methods that work well on certain task types and those that systematically underperform. Clustering…

Machine Learning · Computer Science 2025-10-16 Marek Gagolewski

Combining individual p-values to aggregate multiple small effects has a long-standing interest in statistics, dating back to the classic Fisher's combination test. In modern large-scale data analysis, correlation and sparsity are common…

Methodology · Statistics 2018-11-30 Yaowu Liu , Jun Xie

The environmental Kuznets curve predicts an inverted U-shaped relationship between environmental pollution and economic growth. Current analyses frequently employ models which restrict nonlinearities in the data to be explained by the…

Econometrics · Economics 2021-12-14 Yicong Lin , Hanno Reuvers

Many scientific phenomena are studied using computer experiments consisting of multiple runs of a computer model while varying the input settings. Gaussian processes (GPs) are a popular tool for the analysis of computer experiments,…

Methodology · Statistics 2021-07-21 Matthias Katzfuss , Joseph Guinness , Earl Lawrence

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

Statistics Theory · Mathematics 2010-07-28 Audrey Finkler

Statistical data is often analyzed as a contingency table, sometimes with empty cells called zeros. Such sparse tables can be due to scarse observations classified in numerous categories, as for example in genetic association studies. Thus,…

Statistics Theory · Mathematics 2010-07-28 Audrey Finkler

A successful measurement of the Stochastic Gravitational Wave Background (SGWB) in Pulsar Timing Arrays (PTAs) would open up a new window through which to test the predictions of General Relativity (GR). We consider how these measurements…

Cosmology and Nongalactic Astrophysics · Physics 2025-05-29 Qiuyue Liang , Meng-Xiang Lin , Mark Trodden

A probability forecast or probabilistic classifier is reliable or calibrated if the predicted probabilities are matched by ex post observed frequencies, as examined visually in reliability diagrams. The classical binning and counting…

Methodology · Statistics 2021-08-26 Timo Dimitriadis , Tilmann Gneiting , Alexander I. Jordan

Kernel smoothing is a widely used nonparametric method in modern statistical analysis. The problem of efficiently conducting kernel smoothing for a massive dataset on a distributed system is a problem of great importance. In this work, we…

Computation · Statistics 2024-10-08 Yuan Gao , Rui Pan , Feng Li , Riquan Zhang , Hansheng Wang

We propose a new formulation of robust regression by integrating all realizations of the uncertainty set and taking an averaged approach to obtain the optimal solution for the ordinary least squares regression problem. We show that this…

Machine Learning · Computer Science 2024-10-10 Dimitris Bertsimas , Yu Ma

We propose a new copula model that can be used with replicated spatial data. Unlike the multivariate normal copula, the proposed copula is based on the assumption that a common factor exists and affects the joint dependence of all…

Applications · Statistics 2016-12-08 Pavel Krupskii , Raphael Huser , Marc G. Genton

Generalized Nested Rollout Policy Adaptation (GNRPA) is a Monte Carlo search algorithm for optimizing a sequence of choices. We propose to improve on GNRPA by avoiding too deterministic policies that find again and again the same sequence…

Artificial Intelligence · Computer Science 2024-01-22 Tristan Cazenave

We develop a nonparametric extension of the sequential generalized likelihood ratio (GLR) test and corresponding time-uniform confidence sequences for the mean of a univariate distribution. By utilizing a geometric interpretation of the GLR…

Statistics Theory · Mathematics 2021-05-17 Jaehyeok Shin , Aaditya Ramdas , Alessandro Rinaldo

We consider a stationary linear AR($p$) model with observations subject to gross errors (outliers). The autoregression parameters are unknown as well as the distribution and moments of innoovations. The distribution of outliers $\Pi$ is…

Statistics Theory · Mathematics 2020-03-19 Michael Boldin

In this paper, we introduce the Generalized Mixed Regularized Reduced Rank Regression model (GMR4), an extension of the GMR3 model designed to improve performance in high-dimensional settings. GMR3 is a regression method for a mix of…

Methodology · Statistics 2025-12-16 Lorenza Cotugno , Mark de Rooij , Roberta Siciliano

Gaussian mixture models with eigen-decomposed covariance structures make up the most popular family of mixture models for clustering and classification, i.e., the Gaussian parsimonious clustering models (GPCM). Although the GPCM family has…

Methodology · Statistics 2014-05-05 Antonio Punzo , Ryan P. Browne , Paul D. McNicholas

For random samples of size n obtained from p-variate normal distributions, we consider the classical likelihood ratio tests (LRT) for their means and covariance matrices in the high-dimensional setting. These test statistics have been…

Statistics Theory · Mathematics 2013-06-04 Tiefeng Jiang , Fan Yang

Mendelian randomization (MR) is a method of exploiting genetic variation to unbiasedly estimate a causal effect in presence of unmeasured confounding. MR is being widely used in epidemiology and other related areas of population science. In…

Applications · Statistics 2019-01-03 Qingyuan Zhao , Jingshu Wang , Gibran Hemani , Jack Bowden , Dylan S. Small

We consider the problem of robust inference under the generalized linear model (GLM) with stochastic covariates. We derive the properties of the minimum density power divergence estimator of the parameters in GLM with random design and use…

Methodology · Statistics 2020-04-06 Ayanendranath Basu , Abhik Ghosh , Abhijit Mandal , Nirian Martin , Leandro Pardo