English
Related papers

Related papers: Model misspecification in peaks over threshold ana…

200 papers

Non-Gaussian outcomes are often modeled using members of the so-called exponential family. Notorious members are the Bernoulli model for binary data, leading to logistic regression, and the Poisson model for count data, leading to Poisson…

Motivated by the fundamental problem of measuring species diversity, this paper introduces the concept of a cluster structure to define an exchangeable cluster probability function that governs the joint distribution of a random count and…

Methodology · Statistics 2014-10-14 Mingyuan Zhou , Stephen G Walker

Mixture models, such as Gaussian mixture models, are widely used in machine learning to represent complex data distributions. A key challenge, especially in high-dimensional settings, is to determine the mixture order and estimate the…

Optimization and Control · Mathematics 2025-09-30 Srećko Đurašinović , Jean-Bernard Lasserre , Victor Magron

When there is no independence, abnormal observations may have a tendency to appear in clusters instead of scattered along the time frame. Identifying clusters and estimating their size are important problems arising in statistics of…

Probability · Mathematics 2020-01-08 Miguel Abadi , Ana Cristina Moreira Freitas , Jorge Milhazes Freitas

Statistical modeling of multivariate and spatial extreme events has attracted broad attention in various areas of science. Max-stable distributions and processes are the natural class of models for this purpose, and many parametric families…

Methodology · Statistics 2017-08-09 Clement Dombry , Sebastian Engelke , Marco Oesting

Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to…

Methodology · Statistics 2007-12-04 Jeffrey Picka , Mingxia Deng

The extremal index is a quantity introduced in extreme value theory to measure the presence of clusters of exceedances. In the dynamical systems framework, it provides important information about the dynamics of the underlying systems. In…

Dynamical Systems · Mathematics 2020-01-08 Th. Caby , D. Faranda , S. Vaienti , P. Yiou

Over the last three decades, case-crossover designs have found many applications in health sciences, especially in air pollution epidemiology. They are typically used, in combination with partial likelihood techniques, to define a…

Methodology · Statistics 2024-10-23 Samuel Perreault , Gracia Y. Dong , Alex Stringer , Hwashin Shin , Patrick Brown

Estimating the parameters of max-stable parametric models poses significant challenges, particularly when some parameters lie on the boundary of the parameter space. This situation arises when a subset of variables exhibits extreme values…

Methodology · Statistics 2026-04-08 Anas Mourahib , Anna Kiriliouk , Johan Segers

A non-homogeneous Poisson cluster model is studied, motivated by insurance applications. The Poisson center process which expresses arrival times of claims, triggers off cluster member processes which correspond to number or amount of…

Probability · Mathematics 2013-12-02 Muneya Matsui

Abstract In Extreme Value methodology the choice of threshold plays an important role in efficient modelling of observations exceeding the threshold. The threshold must be chosen high enough to ensure an unbiased extreme value index but…

Methodology · Statistics 2020-06-11 Andréhette Verster , Lizanne Raubenheimer

Virtually any model we use in machine learning to make predictions does not perfectly represent reality. So, most of the learning happens under model misspecification. In this work, we present a novel analysis of the generalization…

Machine Learning · Computer Science 2020-10-23 Andres R. Masegosa

Mixture models extend the toolbox of clustering methods available to the data analyst. They allow for an explicit definition of the cluster shapes and structure within a probabilistic framework and exploit estimation and inference…

Methodology · Statistics 2025-09-15 Bettina Grün

Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often…

Machine Learning · Statistics 2018-03-30 Luzie Helfmann , Johannes von Lindheim , Mattes Mollenhauer , Ralf Banisch

High-dimensional data of discrete and skewed nature is commonly encountered in high-throughput sequencing studies. Analyzing the network itself or the interplay between genes in this type of data continues to present many challenges. As…

Methodology · Statistics 2017-12-01 Anjali Silva , Steven J. Rothstein , Paul D. McNicholas , Sanjeena Subedi

Our purpose in this paper is to apply the general methodology for model selection based on T-estimators developed in Birg\'{e} [Ann. Inst. H. Poincar\'{e} Probab. Statist. 42 (2006) 273--325] to the particular situation of the estimation of…

Statistics Theory · Mathematics 2009-09-29 Lucien Birgé

Relying on the excursion set theory, we compute the number density of local extrema and crossing statistics versus the threshold for the stock market indices. Comparing the number density of excursion sets calculated numerically with the…

Statistical Finance · Quantitative Finance 2022-07-08 M. Shadmangohar , S. M. S. Movahed

The impact of an extreme climate event depends strongly on its geographical scale. Max-stable processes can be used for the statistical investigation of climate extremes and their spatial dependencies on a continuous area. Most existing…

Methodology · Statistics 2023-06-14 Justus Contzen , Thorsten Dickhaus , Gerrit Lohmann

Probabilistic clustering models (or equivalently, mixture models) are basic building blocks in countless statistical models and involve latent random variables over discrete spaces. For these models, posterior inference methods can be…

Machine Learning · Statistics 2020-06-24 Ari Pakman , Yueqi Wang , Catalin Mitelut , JinHyung Lee , Liam Paninski

Data-driven anomaly detection methods typically build a model for the normal behavior of the target system, and score each data instance with respect to this model. A threshold is invariably needed to identify data instances with high (or…

Machine Learning · Statistics 2019-10-09 Sreelekha Guggilam , S. M. Arshad Zaidi , Varun Chandola , Abani Patra