English
Related papers

Related papers: Measuring wage inequality under right censoring

200 papers

Real-world data usually present long-tailed distributions. Training on imbalanced data tends to render neural networks perform well on head classes while much worse on tail classes. The severe sparseness of training instances for the tail…

Machine Learning · Computer Science 2021-11-10 Chaozheng Wang , Shuzheng Gao , Cuiyun Gao , Pengyun Wang , Wenjie Pei , Lujia Pan , Zenglin Xu

Empirical likelihood is a well-known nonparametric method in statistics and has been widely applied in statistical inference. The method has been employed by Lu and Peng (2002) to constructing confidence intervals for the tail index of a…

Methodology · Statistics 2019-04-19 Yizeng Li , Yongcheng Qi

Class imbalance, which is also called long-tailed distribution, is a common problem in classification tasks based on machine learning. If it happens, the minority data will be overwhelmed by the majority, which presents quite a challenge…

Machine Learning · Computer Science 2023-03-29 Jia-Chen Zhao

Causal questions are omnipresent in many scientific problems. While much progress has been made in the analysis of causal relationships between random variables, these methods are not well suited if the causal mechanisms only manifest…

Methodology · Statistics 2020-09-23 Nicola Gnecco , Nicolai Meinshausen , Jonas Peters , Sebastian Engelke

Chebyshev's inequality provides an upper bound on the tail probability of a random variable based on its mean and variance. While tight, the inequality has been criticized for only being attained by pathological distributions that abuse the…

Optimization and Control · Mathematics 2020-10-16 Ernst Roos , Ruud Brekelmans , Wouter van Eekelen , Dick den Hertog , Johan van Leeuwaarden

The vast majority of techniques to train fair models require access to the protected attribute (e.g., race, gender), either at train time or in production. However, in many important applications this protected attribute is largely…

Machine Learning · Computer Science 2023-10-04 Hadi Elzayn , Emily Black , Patrick Vossler , Nathanael Jo , Jacob Goldin , Daniel E. Ho

This paper addresses the challenge of forecasting corporate distress, a problem marked by three key statistical hurdles: (i) right censoring, (ii) high-dimensional predictors, and (iii) mixed-frequency data. To overcome these complexities,…

Econometrics · Economics 2026-02-09 Wei Miao , Jad Beyhum , Jonas Striaukas , Ingrid Van Keilegom

We introduce a new actuarial tail-shape index, the $\theta$-index, based on a probability equal level relationship between Value at Risk and Expected Shortfall. The index is defined at each tail probability level as the parameter value for…

Risk Management · Quantitative Finance 2026-01-29 Georgios I. Papayiannis , Georgios Psarrakos

Different questions related with analysis of extreme values and outliers arise frequently in practice. To exclude extremal observations and outliers is not a good decision because they contain important information about the observed…

Methodology · Statistics 2018-01-17 Pavlina K. Jordanova , Monika P. Petkova

We present an algorithm for distributed estimation of an unknown vector parameter $\boldsymbol{\theta}^\ast \in {\mathbb R}^M$ in the presence of heavy-tailed observation and communication noises. Heavy-tailed noises frequently appear,…

Information Theory · Computer Science 2026-03-24 Dragana Bajovic , Dusan Jakovetic , Soummya Kar , Manojlo Vukovic

A new measure of income inequality that captures the heavy tail behavior of the income distribution is proposed. We discuss two different approaches to find the estimators of the proposed measure. We show that these estimators are…

Methodology · Statistics 2024-08-21 Sudheesh K Kattumannil , Saparya Suresh

Quantifying tail dependence is an important issue in insurance and risk management. The prevalent tail dependence coefficient (TDC), however, is known to underestimate the degree of tail dependence and it does not capture non-exchangeable…

Statistics Theory · Mathematics 2023-02-14 Takaaki Koike , Shogo Kato , Marius Hofert

Heavy-tailed phenomena appear across diverse domains --from wealth and firm sizes in economics to network traffic, biological systems, and physical processes-- characterized by the disproportionate influence of extreme values. These…

Statistics Theory · Mathematics 2025-11-10 Hamidreza Maleki Almani

Statistics in ranked lists is important in analyzing molecular biology measurement data, such as ChIP-seq, which yields ranked lists of genomic sequences. State of the art methods study fixed motifs in ranked lists. More flexible models…

Quantitative Methods · Quantitative Biology 2013-07-31 Limor Leibovich , Zohar Yakhini

Recently some papers, such as Aban, Meerschaert and Panorska (2006), Nuyts (2010) and Clark (2013), have drawn attention to possible truncation in Pareto tail modelling. Sometimes natural upper bounds exist that truncate the probability…

Statistics Theory · Mathematics 2015-05-21 Jan Beirlant , Isabel Fraga Alves , Ivette Gomes

This paper proposes methods for Bayesian inference in time-varying parameter (TVP) quantile regression (QR) models featuring conditional heteroskedasticity. I use data augmentation schemes to render the model conditionally Gaussian and…

Econometrics · Economics 2021-10-19 Michael Pfarrhofer

We study the interplay of information and prior (mis)perceptions in a Phelps-Aigner-Cain-type model of statistical discrimination in the labor market. We decompose the effect on average pay of an increase in how informative observables are…

Theoretical Economics · Economics 2026-01-23 Matteo Escudé , Paula Onuchic , Ludvig Sinander , Quitzé Valenzuela-Stookey

In this paper, we first provide a review of different non-parametric estimators for the cumulative distribution function under left-censoring. We then propose a new estimator based on a non-parametric likelihood approach using reversed…

Statistics Theory · Mathematics 2023-07-11 N. Balakrishnan , Christian Paroissin , Magdalena Pereda Vivo

Random survival forest and survival trees are popular models in statistics and machine learning. However, there is a lack of general understanding regarding consistency, splitting rules and influence of the censoring mechanism. In this…

Statistics Theory · Mathematics 2019-02-05 Yifan Cui , Ruoqing Zhu , Mai Zhou , Michael Kosorok

In this paper, we consider survival analysis with right-censored data which is a common situation in predictive maintenance and health field. We propose a model based on the estimation of two-parameter Weibull distribution conditionally to…

Methodology · Statistics 2020-02-24 Achraf Bennis , Sandrine Mouysset , Mathieu Serrurier