English
Related papers

Related papers: A survey of a hurdle model for heavy-tailed data b…

200 papers

Normalising flows are tractable probabilistic models that leverage the power of deep learning to describe a wide parametric family of distributions, all while remaining trainable using maximum likelihood. We discuss how these methods can be…

Machine Learning · Computer Science 2020-07-14 Simon Alexanderson , Gustav Eje Henter

The g-and-k and (generalised) g-and-h distributions are flexible univariate distributions which can model highly skewed or heavy tailed data through only four parameters: location and scale, and two shape parameters influencing the skewness…

Computation · Statistics 2017-06-22 Dennis Prangle

Our goal is to develop a Bayesian model averaging technique in linear regression models that accommodates heavier tailed error densities than the normal distribution. Motivated by the use of the Huber loss function in the presence of…

Methodology · Statistics 2024-11-26 Shamriddha De , Joyee Ghosh

Time Series Forecasting (TSF) is a widely researched topic with broad applications in weather forecasting, traffic control, and stock price prediction. Extreme values in time series often significantly impact human and natural systems, but…

Machine Learning · Computer Science 2023-10-12 Jincheng Wang , Yue Gao

The generalized Laplace (GL) distribution, which falls in the larger family of generalized hyperbolic distributions, provides a versatile model to deal with a variety of applications thanks to its shape parameters. The elliptically…

Methodology · Statistics 2025-04-15 Marco Geraci

The family of location and scale mixtures of Gaussians has the ability to generate a number of flexible distributional forms. It nests as particular cases several important asymmetric distributions like the Generalised Hyperbolic…

Methodology · Statistics 2014-08-05 Darren Wraith , Florence Forbes

Distributed data naturally arise in scenarios involving multiple sources of observations, each stored at a different location. Directly pooling all the data together is often prohibited due to limited bandwidth and storage, or due to…

Methodology · Statistics 2021-07-07 Jiyu Luo , Qiang Sun , Wenxin Zhou

1. Joint species distribution models (JSDMs) have gained considerable traction among ecologists over the past decade, due to their capacity to answer a wide range of questions at both the species- and the community-level. The family of…

Methodology · Statistics 2024-03-19 Pekka Korhonen , Francis K. C. Hui , Jenni Niku , Sara Taskinen , Bert van der Veen

We benchmark the robustness of maximum likelihood based uncertainty estimation methods to outliers in training data for regression tasks. Outliers or noisy labels in training data results in degraded performances as well as incorrect…

Machine Learning · Computer Science 2022-02-09 Deebul S. Nair , Nico Hochgeschwender , Miguel A. Olivares-Mendez

Heavy-tailed random variables have been used in insurance research to model both loss frequencies and loss severities, with substantially more emphasis on the latter. In the present work, we take a step toward addressing this imbalance by…

Methodology · Statistics 2022-11-11 Jiansheng Dai , Ziheng Huang , Michael R. Powers , Jiaxin Xu

Heavy-tailed or power-law distributions are becoming increasingly common in biological literature. A wide range of biological data has been fitted to distributions with heavy tails. Many of these studies use simple fitting methods to find…

Quantitative Methods · Quantitative Biology 2007-12-06 A. James , M. J. Plank

Many real-world prediction tasks have outcome variables that have characteristic heavy-tail distributions. Examples include copies of books sold, auction prices of art pieces, demand for commodities in warehouses, etc. By learning…

Machine Learning · Computer Science 2023-10-13 Xindi Wang , Onur Varol , Tina Eliassi-Rad

Directional or Circular statistics are pertaining to the analysis and interpretation of directions or rotations. In this work, a novel probability distribution is proposed to model multidimensional sparse directional data. The Generalised…

Sound · Computer Science 2017-08-21 Nikolaos Mitianoudis

Heavy-tailed metrics are common and often critical to product evaluation in the online world. While we may have samples large enough for Central Limit Theorem to kick in, experimentation is challenging due to the wide confidence interval of…

Applications · Statistics 2019-05-23 Jason , Wang , Pauline Burke

We provide upper bounds on the end-to-end backlog and delay in a network with heavy-tailed and self-similar traffic. The analysis follows a network calculus approach where traffic is characterized by envelope functions and service is…

Networking and Internet Architecture · Computer Science 2013-12-30 Jorg Liebeherr , Almut Burchard , Florin Ciucu

Classical information-theoretic generalization bounds typically control the generalization gap through KL-based mutual information and therefore rely on boundedness or sub-Gaussian tails via the moment generating function (MGF). In many…

Machine Learning · Statistics 2026-04-14 Huiming Zhang , Binghan Li , Wan Tian , Qiang Sun

Latency to end-users and regulatory requirements push large companies to build data centers all around the world. The resulting data is "born" geographically distributed. On the other hand, many machine learning applications require a…

Machine Learning · Computer Science 2016-03-31 Ignacio Cano , Markus Weimer , Dhruv Mahajan , Carlo Curino , Giovanni Matteo Fumarola

Federated graph learning (FGL) has become an important research topic in response to the increasing scale and the distributed nature of graph-structured data in the real world. In FGL, a global graph is distributed across different clients,…

Machine Learning · Computer Science 2024-08-27 Binchi Zhang , Minnan Luo , Shangbin Feng , Ziqi Liu , Jun Zhou , Qinghua Zheng

The possibilities of the use of the coefficient of variation over a high threshold in tail modelling are discussed. The paper also considers multiple threshold tests for a generalized Pareto distribution, together with a threshold selection…

Statistics Theory · Mathematics 2015-10-02 J. Castillo , M. Padilla

Real-world data is laden with outlying values. The challenge for machine learning is that the learner typically has no prior knowledge of whether the feedback it receives (losses, gradients, etc.) will be heavy-tailed or not. In this work,…

Machine Learning · Statistics 2020-12-16 Matthew J. Holland