English
Related papers

Related papers: Semi-parametric modeling of excesses above high mu…

200 papers

We consider the problem of estimating the probability density function of a circular random variable observed under censoring. To this end, we introduce a projection estimator constructed via a regression approach on linear sieves. We first…

Statistics Theory · Mathematics 2025-12-09 Nicolas Conanec , Claire Lacour , Thanh Mai Pham Ngoc

Clustering is one of the most widely used procedures in the analysis of microarray data, for example with the goal of discovering cancer subtypes based on observed heterogeneity of genetic marks between different tissues. It is well-known…

Methodology · Statistics 2009-04-21 Heng Lian

We study the problem of estimating the probability density function of a circular random variable subject to censoring. To this end, we propose a fully computable quotient estimator that combines a projection estimator on linear sieves with…

Statistics Theory · Mathematics 2025-08-11 Nicolas Conanec

Directional data require specialized probability models because of the non-Euclidean and periodic nature of their domain. When a directional variable is observed jointly with linear variables, modeling their dependence adds an additional…

Methodology · Statistics 2022-12-22 Tong Zou , Hal S. Stern

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

Statistics Theory · Mathematics 2018-05-22 Takumi Saegusa

In particle physics, as in many areas of science, parameter inference relies on simulations to bridge the gap between theory and experiment. Recent developments in simulation-based inference have boosted the sensitivity of analyses;…

High Energy Physics - Phenomenology · Physics 2026-04-23 Ezequiel Alvarez , Sean Benevedes , Manuel Szewc , Jesse Thaler

The rapid development of computing power and efficient Markov Chain Monte Carlo (MCMC) simulation algorithms have revolutionized Bayesian statistics, making it a highly practical inference method in applied work. However, MCMC algorithms…

Methodology · Statistics 2018-09-21 Matias Quiroz , Mattias Villani , Robert Kohn , Minh-Ngoc Tran , Khue-Dung Dang

Recent developments in big data and analytics research have produced an abundance of large data sets that are too big to be analyzed in their entirety, due to limits on computer memory or storage capacity. To address these issues,…

Methodology · Statistics 2016-01-06 Alexey Miroshnikov , Erin M. Conlon

Computational capability often falls short when confronted with massive data, posing a common challenge in establishing a statistical model or statistical inference method dealing with big data. While subsampling techniques have been…

Methodology · Statistics 2024-10-31 Yixiao Ruan , Zan Li , Zhaohui Li , Dennis K. J. Lin , Qingpei Hu , Dan Yu

Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the…

Methodology · Statistics 2021-07-06 Jun Yu , HaiYing Wang , Mingyao Ai , Huiming Zhang

This study explores the classification error of Mixture Discriminant Analysis (MDA) in scenarios where the number of mixture components exceeds those present in the actual data distribution, a condition known as overspecification. We use a…

Machine Learning · Statistics 2025-11-03 Arman Bolatov , Alan Legg , Igor Melnykov , Amantay Nurlanuly , Maxat Tezekbayev , Zhenisbek Assylbekov

The goal of causal mediation analysis, often described within the potential outcomes framework, is to decompose the effect of an exposure on an outcome of interest along different causal pathways. Using the assumption of sequential…

Methodology · Statistics 2021-11-09 Lexi Rene , Antonio R. Linero , Elizabeth Slate

Dirichlet process (DP) mixture models provide a flexible Bayesian framework for density estimation. Unfortunately, their flexibility comes at a cost: inference in DP mixture models is computationally expensive, even when conjugate…

Machine Learning · Computer Science 2009-07-13 Hal Daumé

Both marginal and dependence features must be described when modelling the extremes of a stationary time series. There are standard approaches to marginal modelling, but long- and short-range dependence of extremes may both appear. In…

Methodology · Statistics 2016-03-17 Thomas Lugrin , Anthony C. Davison , Jonathan A. Tawn

Panel data allows for the modeling of unobserved heterogeneity, significantly raising the number of nuisance parameters and making high dimensionality a practical issue. Meanwhile, temporal and cross-sectional dependence in panel data…

Econometrics · Economics 2025-12-23 Kaicheng Chen

Block maxima methods constitute a fundamental part of the statistical toolbox in extreme value analysis. However, most of the corresponding theory is derived under the simplifying assumption that block maxima are independent observations…

Statistics Theory · Mathematics 2019-07-24 Nan Zou , Stanislav Volgushev , Axel Bücher

A fully Bayesian approach is proposed for ultrahigh-dimensional nonparametric additive models in which the number of additive components may be larger than the sample size, though ideally the true model is believed to include only a small…

Methodology · Statistics 2013-09-24 Zuofeng Shang , Ping Li

We introduce a mixture model for censored durations (C-mix), and develop maximum likelihood inference for the joint estimation of the time distributions and latent regression parameters of the model. We consider a high-dimensional setting,…

Machine Learning · Statistics 2017-11-28 Simon Bussy , Agathe Guilloux , Stéphane Gaïffas , Anne-Sophie Jannot

This work addresses the problem of segmentation in time series data with respect to a statistical parameter of interest in Bayesian models. It is common to assume that the parameters are distinct within each segment. As such, many Bayesian…

Machine Learning · Computer Science 2017-10-27 Alireza Ahrabian , Shirin Enshaeifar , Clive Cheong-Took , Payam Barnaghi

Data-driven anomaly detection methods typically build a model for the normal behavior of the target system, and score each data instance with respect to this model. A threshold is invariably needed to identify data instances with high (or…

Machine Learning · Statistics 2019-10-09 Sreelekha Guggilam , S. M. Arshad Zaidi , Varun Chandola , Abani Patra
‹ Prev 1 8 9 10 Next ›