Related papers: L-moments for automatic threshold selection in ext…
Domain experts often possess valuable physical insights that are overlooked in fully automated decision-making processes such as Bayesian optimisation. In this article we apply high-throughput (batch) Bayesian optimisation alongside…
Determinantal Point Processes (DPPs) are a family of probabilistic models that have a repulsive behavior, and lend themselves naturally to many tasks in machine learning where returning a diverse set of objects is important. While there are…
We use extreme value theory to estimate the probability of successive exceedances of a threshold value of a time-series of an observable on several classes of chaotic dynamical systems. The observables have either a Fr\'echet (fat-tailed)…
Data valuation and subset selection have emerged as valuable tools for application-specific selection of important training data. However, the efficiency-accuracy tradeoffs of state-of-the-art methods hinder their widespread application to…
In extreme value statistics, the peaks-over-threshold method is widely used. The method is based on the generalized Pareto distribution characterizing probabilities of exceedances over high thresholds in $\mathbb {R}^d$. We present a…
We use point processes theory to describe the asymptotic distribution of all upper order statistics for observations collected at renewal times. As a corollary, we obtain limiting theorems for corresponding extremal processes.
For extreme value estimation we propose to use a model with a Dirichlet process mixture of gamma densities in the center and generalized Pareto densities for the tails. Due to the randomness in the center and a heavy tailed density in the…
In high-dimensional classification settings, we wish to seek a balance between high power and ensuring control over a desired loss function. In many settings, the points most likely to be misclassified are those who lie near the decision…
A/B testing is ubiquitous within the machine learning and data science operations of internet companies. Generically, the idea is to perform a statistical test of the hypothesis that a new feature is better than the existing platform---for…
A coupling method is developed for univariate extreme value theory , providing an alternative to the use of the tail empirical/quantile processes. Emphasizing the Peak-over-Threshold approach that approximates the distribution above high…
A new thresholding method, based on L-statistics and called order thresholding, is proposed as a technique for improving the power when testing against high-dimensional alternatives. The new method allows great flexibility in the choice of…
Probability density estimation is a core problem of statistics and signal processing. Moment methods are an important means of density estimation, but they are generally strongly dependent on the choice of feasible functions, which severely…
We study distributional robustness in the context of Extreme Value Theory (EVT). We provide a data-driven method for estimating extreme quantiles in a manner that is robust against incorrect model assumptions underlying the application of…
When passing from the univariate to the multivariate setting, modelling extremes becomes much more intricate. In this introductory exposition, classical multivariate extreme value theory is presented from the point of view of multivariate…
We propose a new threshold selection method for the nonparametric estimation of the extremal index of stochastic processes. The so-called discrepancy method was proposed as a data-driven smoothing tool for estimation of a probability…
Max-stable processes are a popular tool for the study of environmental extremes, and the extremal skew-$t$ process is a general model that allows for a flexible extremal dependence structure. For inference on max-stable processes with…
We study online changepoint detection in the context of a linear regression model. We propose a class of heavily weighted statistics based on the CUSUM process of the regression residuals, which are specifically designed to ensure timely…
LLMs are highly sensitive to prompt phrasing, yet standard benchmarks typically report performance using a single prompt, raising concerns about the reliability of such evaluations. In this work, we argue for a stochastic method of moments…
Generating accurate extremes from an observational data set is crucial when seeking to estimate risks associated with the occurrence of future extremes which could be larger than those already observed. Applications range from the…
The thresholding of time series of activity or intensity is frequently used to define and differentiate events. This is either implicit, for example due to resolution limits, or explicit, in order to filter certain small scale physics from…