Related papers: Scoring Rules with Normalized Upper Order Statisti…
In this paper we investigate potential changes which may have occurred over the last two decades in the probability mass of the right tail of the wage distribution, through the analysis of the corresponding tail index. In specific, a…
Modeling heterogeneity on heavy-tailed distributions under a regression framework is challenging, and classical statistical methodologies usually place conditions on the distribution models to facilitate the learning procedure. However,…
In domains where transparency and trustworthiness are crucial, such as healthcare, rule-based systems are widely used and often preferred over black-box models for decision support systems due to their inherent interpretability. However, as…
Several tasks in information retrieval (IR) rely on assumptions regarding the distribution of some property (such as term frequency) in the data being processed. This thesis argues that such distributional assumptions can lead to incorrect…
We investigate a way of comparing and classifying tails of random variables. Our approach extends the notion of classical indices, such as exponential and moment indices, which are widely used measuring heaviness of tail functions. A…
Score-based model research in the last few years has produced state of the art generative models by employing Gaussian denoising score-matching (DSM). However, the Gaussian noise assumption has several high-dimensional limitations,…
Diffusion models have become the de facto standard for modern visual generation, including well-established frameworks such as latent diffusion and flow matching. Recently, modeling high-order dynamics has emerged as a promising frontier in…
Heavy-tailed phenomena appear across diverse domains --from wealth and firm sizes in economics to network traffic, biological systems, and physical processes-- characterized by the disproportionate influence of extreme values. These…
To draw inference on serial extremal dependence within heavy-tailed Markov chains, Drees, Segers and Warcho{\l} [Extremes (2015) 18, 369--402] proposed nonparametric estimators of the spectral tail process. The methodology can be extended…
Both parametric distribution functions appearing in extreme value theory - the generalized extreme value distribution and the generalized Pareto distribution - have log-concave densities if the extreme value index gamma is in [-1,0].…
We introduce a trimmed version of the Hill estimator for the index of a heavy-tailed distribution, which is robust to perturbations in the extreme order statistics. In the ideal Pareto setting, the estimator is essentially finite-sample…
Recently, long-tailed image classification harvests lots of research attention, since the data distribution is long-tailed in many real-world situations. Piles of algorithms are devised to address the data imbalance problem by biasing the…
Diffusion models have emerged as powerful generative frameworks with widespread applications across machine learning and artificial intelligence systems. While current research has predominantly focused on linear diffusions, these…
The most popular approach in extreme value statistics is the modelling of threshold exceedances using the asymptotically motivated generalised Pareto distribution. This approach involves the selection of a high threshold above which the…
Asymptotic normality of extreme value tail estimators received much attention in the literature, giving rise to increasingly complicated 2nd order regularity conditions. However, such conditions are really difficult to be checked for real…
We make use of the empirical process theory to approximate the adapted Hill estimator, for censored data, in terms of Gaussian processes. Then, we derive its asymptotic normality, only under the usual second-order condition of regular…
It was shown that when one disposes of a parametric information of the truncation distribution, the semiparametric estimator of the distribution function for truncated data (Wang, 1989) is more efficient than the nonparametric one. On the…
We propose GradTail, an algorithm that uses gradients to improve model performance on the fly in the face of long-tailed training data distributions. Unlike conventional long-tail classifiers which operate on converged - and possibly…
We study the generalization properties of unregularized gradient methods applied to separable linear classification -- a setting that has received considerable attention since the pioneering work of Soudry et al. (2018). We establish tight…
This study proposed an exhaustive stable/reproducible rule-mining algorithm combined to a classifier to generate both accurate and interpretable models. Our method first extracts rules (i.e., a conjunction of conditions about the values of…