English
Related papers

Related papers: Fitting phase--type scale mixtures to heavy--taile…

200 papers

The masses of data now available have opened up the prospect of discovering weak signals using machine-learning algorithms, with a view to predictive or interpretation tasks. As this survey of recent results attempts to show, bringing…

Statistics Theory · Mathematics 2026-05-06 Stephan Clémençon , Anne Sabourin

The distribution of data in the world (eg, internet, etc.) significantly differs from the well-curated datasets and is often over-populated with samples from common categories. The algorithms designed for well-curated datasets perform…

Machine Learning · Computer Science 2025-07-30 Harsh Rangwani

Current out-of-distribution (OOD) detection methods typically assume balanced in-distribution (ID) data, while most real-world data follow a long-tailed distribution. Previous approaches to long-tailed OOD detection often involve balancing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-19 Yina He , Lei Peng , Yongcun Zhang , Juanjuan Weng , Zhiming Luo , Shaozi Li

This paper presents an R package EMMIXcskew for the fitting of the canonical fundamental skew t-distribution (CFUST) and finite mixtures of this distribution (FM-CFUST) via maximum likelihood (ML). The CFUST distribution provides a flexible…

Computation · Statistics 2017-02-10 Sharon X. Lee , Geoffrey J. McLachlan

We study the problem of estimating the mean of a distribution in high dimensions when either the samples are adversarially corrupted or the distribution is heavy-tailed. Recent developments in robust statistics have established efficient…

Data Structures and Algorithms · Computer Science 2021-01-20 Samuel B. Hopkins , Jerry Li , Fred Zhang

In situations where both extreme and non-extreme data are of interest, modelling the whole data set accurately is important. In a univariate framework, modelling the bulk and tail of a distribution has been extensively studied before.…

Methodology · Statistics 2023-10-11 Lídia M. André , Jennifer L. Wadsworth , Adrian O'Hagan

We study the distributed stochastic optimization (DSO) problem under a heavy-tailed noise condition by utilizing a multi-agent system. Despite the extensive research on DSO algorithms used to solve DSO problems under light-tailed noise…

Optimization and Control · Mathematics 2025-09-23 Zhan Yu , Lan Liao , Deming Yuan , Daniel W. C. Ho , Ding-Xuan Zhou

Accurate channel modeling plays a pivotal role in optimizing communication systems, and fitting field measurements to stochastic models is crucial for capturing the key propagation features and to map these to achievable system…

Signal Processing · Electrical Eng. & Systems 2025-03-10 Santiago Fernández , José David Vega-Sánchez , Juan E. Galeote-Cazorla , F. Javier López-Martínez

A tail empirical process for heavy-tailed and right-censored data is introduced and its Gaussian approximation is established. In this context, a (weighted) new Hill-type estimator for positive extreme value index is proposed and its…

Statistics Theory · Mathematics 2018-02-06 Brahim Brahimi , Djamel Meraghni , Abdelhakim Necir , Louiza Soltane

Real-world networks are generally claimed to be scale-free, meaning that the degree distributions follow the classical power-law, at least asymptotically. Yet, closer observation shows that the classical power-law distribution is often…

Statistics Theory · Mathematics 2022-07-18 Swarup Chattopadhyay , Tanujit Chakraborty , Kuntal Ghosh , Asit K. das

Mixed Poisson distributions provide a flexible approach to the analysis of count data with overdispersion, zero inflation, or heavy tails. Since the Poisson mean must be nonnegative, the mixing distribution is typically assumed to have…

Probability · Mathematics 2025-08-20 F. William Townes

Recently, high-dimensional heterogeneous data have attracted a lot of attention and discussion. Under heterogeneity, semiparametric regression is a popular choice to model data in statistics. In this paper, we take advantages of expectile…

Statistics Theory · Mathematics 2019-08-20 Jun Zhao , Guan'ao Yan , Yi Zhang

The task of modeling claim severities is addressed when data is not consistent with the classical regression assumptions. This framework is common in several lines of business within insurance and reinsurance, where catastrophic losses or…

Statistics Theory · Mathematics 2022-04-01 Martin Bladt , Jorge Yslas

This paper introduces a class of copula models for spatial data, based on multivariate Pareto-mixture distributions. We explore the tail properties of these models, demonstrating their ability to capture both tail dependence and asymptotic…

Methodology · Statistics 2026-01-28 Pavel Krupskii

Recently some papers, such as Aban, Meerschaert and Panorska (2006), Nuyts (2010) and Clark (2013), have drawn attention to possible truncation in Pareto tail modelling. Sometimes natural upper bounds exist that truncate the probability…

Statistics Theory · Mathematics 2015-05-21 Jan Beirlant , Isabel Fraga Alves , Ivette Gomes

We consider a model for multivariate data with heavy-tailed marginal distributions and a Gaussian dependence structure. The different marginals in the model are allowed to have non-identical tail behavior in contrast to most popular…

Methodology · Statistics 2023-05-23 Bikramjit Das

Finite mixture modelling is a popular method in the field of clustering and is beneficial largely due to its soft cluster membership probabilities. A common method for fitting finite mixture models is to employ spectral clustering, which…

Machine Learning · Statistics 2024-03-22 Liam Welsh , Phillip Shreeves

Motivated by two case studies using primary care records from the Clinical Practice Research Datalink, we describe statistical methods that facilitate the analysis of tall data, with very large numbers of observations. Our focus is on…

Methodology · Statistics 2018-05-14 Kirsty Rhodes , Rebecca Turner , Rupert Payne , Ian White

For measuring tail risk with scarce extreme events, extreme value analysis is often invoked as the statistical tool to extrapolate to the tail of a distribution. The presence of large datasets benefits tail risk analysis by providing more…

Methodology · Statistics 2023-12-18 Liujun Chen , Deyuan Li , Chen Zhou

High-dimensional data arise routinely in modern statistics, econometrics, finance, genomics, and machine learning. While a large body of existing methodology is developed under Gaussian or light-tailed assumptions, many real data sets…

Methodology · Statistics 2026-04-16 Long Feng