English
Related papers

Related papers: Rank-based concordance for zero-inflated data: New…

200 papers

High-dimensional sparse matrix data frequently arise in various applications. A notable example is the weighted word-word co-occurrence count data, which summarizes the weighted frequency of word pairs appearing within the same context…

Machine Learning · Computer Science 2025-01-03 Taejoon Kim , Haiyan Wang

The Mallows model occupies a central role in parametric modelling of ranking data to learn preferences of a population of judges. Despite the wide range of metrics for rankings that can be considered in the model specification, the choice…

Methodology · Statistics 2022-09-21 Marta Crispino , Cristina Mollica , Valerio Astuti , Luca Tardella

The advent of modern data collection and processing techniques has seen the size, scale, and complexity of data grow exponentially. A seminal step in leveraging these rich datasets for downstream inference is understanding the…

Applications · Statistics 2024-07-30 Zeyi Wang , Eric Bridgeford , Shangsi Wang , Joshua T. Vogelstein , Brian Caffo

Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings when two conditions co-occur: (1) the evaluation matrix is…

Machine Learning · Computer Science 2026-05-13 Jung Min Kang

We explore how the classical concordance measures - Kendall's $\tau$, Spearman's rank correlation $\rho$, and Spearman's footrule $\phi$ - relate to Chatterjee's rank correlation $\xi$ when restricted to lower semilinear copulas. First, we…

Methodology · Statistics 2025-08-01 Sebastian Fuchs , Carsten Limbach , Fabian Schürrer

Understanding the spatial distribution of animals, during all their life phases, as well as how the distributions are influenced by environmental covariates, is a fundamental requirement for the effective management of animal populations.…

Applications · Statistics 2020-10-26 Soraia Pereira , Raquel Menezes , Maria Manuel Angélico , Tiago Marques

A common type of zero-inflated data has certain true values incorrectly replaced by zeros due to data recording conventions (rare outcomes assumed to be absent) or details of data recording equipment (e.g. artificial zeros in gene…

Spatially correlated data with an excess of zeros, usually referred to as zero-inflated spatial data, arise in many disciplines. Examples include count data, for instance, abundance (or lack thereof) of animal species and disease counts, as…

Methodology · Statistics 2024-04-23 Ben Seiyon Lee , Murali Haran

Zero-inflated models are frequently used to deal with data having many zeros. A commonly used model for over-dispersed data containing zeros is known as the zero-inflated Poisson model. However, to account for the heterogeneity of counts…

Methodology · Statistics 2025-09-04 Ali Abbas , Sajid Ali , Ismail Shah

Claim frequency data in insurance records the number of claims on insurance policies during a finite period of time. Given that insurance companies operate with multiple lines of insurance business where the claim frequencies on different…

Applications · Statistics 2022-12-05 Pengcheng Zhang , David Pitt , Xueyuan Wu

Wearable devices collect time-varying biobehavioral data, offering opportunities to investigate how behaviors influence health outcomes. However, these data often contain measurement error and excess zeros (due to nonwear, sedentary…

Methodology · Statistics 2026-02-06 Caihong Qin , Lan Xue , Ufuk Beyaztas , Roger S. Zoh , Mark Benden , Jeff Goldsmith , Carmen D. Tekwe

Matching is a widely used causal inference design that aims to approximate a randomized experiment using observational data by forming matched sets of treated and control units based on similarities in their covariates. Ideally, treated…

Methodology · Statistics 2026-04-06 Jianan Zhu , Jeffrey Zhang , Zijian Guo , Siyu Heng

Researchers are often interested in predicting outcomes, conducting clustering analysis to detect distinct subgroups of their data, or computing causal treatment effects. Pathological data distributions that exhibit skewness and…

Methodology · Statistics 2020-08-24 Arman Oganisian , Nandita Mitra , Jason Roy

The assessment of monotone dependence between random variables $X$ and $Y$ is a classical problem in statistics and a gamut of application domains. Consequently, researchers have sought measures of association that are invariant under…

Methodology · Statistics 2025-10-22 Eva-Maria Walz , Andreas Eberl , Tilmann Gneiting

This paper deals with a general class of transformation models that contains many important semiparametric regression models as special cases. It develops a self-induced smoothing for the maximum rank correlation estimator, resulting in…

Methodology · Statistics 2013-02-28 Junyi Zhang , Zhezhen Jin , Yongzhao Shao , Zhiliang Ying

Representations of measures of concordance in terms of Pearson' s correlation coefficient are studied. All transforms of random variables are characterized such that the correlation coefficient of the transformed random variables is a…

Statistics Theory · Mathematics 2023-01-10 Takaaki Koike , Marius Hofert

It is shown that the psychometric test reliability, based on any true-score model with randomly sampled items and conditionally independent errors, converges to 1 as the test length goes to infinity, assuming some fairly general regularity…

Methodology · Statistics 2025-08-29 Jules L. Ellis

In biomedical studies, paired survival data arise naturally when two event times are observed within the same subject. Existing statistical models seldom accommodate both cure fractions and complex dependence structures. In this paper, we…

Methodology · Statistics 2026-04-28 Masaki Hino , Shogo Kato , Takeshi Emura

Bayesian inference for rank-order problems is frustrated by the absence of an explicit likelihood function. This hurdle can be overcome by assuming a latent normal representation that is consistent with the ordinal information in the data:…

Methodology · Statistics 2019-05-20 Johnny van Doorn , Alexander Ly , Maarten Marsman , Eric-Jan Wagenmakers

Score matching is a vital tool for learning the distribution of data with applications across many areas including diffusion processes, energy based modelling, and graphical model estimation. Despite all these applications, little work…

Machine Learning · Statistics 2025-06-03 Josh Givens , Song Liu , Henry W J Reeve