English
Related papers

Related papers: Testing Most Influential Sets

200 papers

We introduce a new class of heavy-tailed distributions for which any weighted average of independent and identically distributed random variables is larger than one such random variable in (usual) stochastic order. We show that many…

Probability · Mathematics 2025-06-18 Yuyu Chen , Seva Shneer

We consider a model for multivariate data with heavy-tailed marginal distributions and a Gaussian dependence structure. The different marginals in the model are allowed to have non-identical tail behavior in contrast to most popular…

Methodology · Statistics 2023-05-23 Bikramjit Das

The masses of data now available have opened up the prospect of discovering weak signals using machine-learning algorithms, with a view to predictive or interpretation tasks. As this survey of recent results attempts to show, bringing…

Statistics Theory · Mathematics 2026-05-06 Stephan Clémençon , Anne Sabourin

While machine-learning models are flourishing and transforming many aspects of everyday life, the inability of humans to understand complex models poses difficulties for these models to be fully trusted and embraced. Thus, interpretability…

Artificial Intelligence · Computer Science 2020-06-18 Guangyi Zhang , Aristides Gionis

I report a new statistical distribution formulated to confront the infamous, long-standing, computational/modeling challenge presented by highly skewed and/or leptokurtic ("fat- or heavy-tailed") data. The distribution is straightforward,…

Statistical Finance · Quantitative Finance 2011-11-01 Lawrence R. Thorne

A variety of methods have been proposed for inference about extreme dependence for multivariate or spatially-indexed stochastic processes and time series. Most of these proceed by first transforming data to some specific extreme value…

Statistics Theory · Mathematics 2018-05-22 James E. Johndrow , Robert L. Wolpert

Adaptive experiment designs can dramatically improve statistical efficiency in randomized trials, but they also complicate statistical inference. For example, it is now well known that the sample mean is biased in adaptive trials.…

Machine Learning · Statistics 2021-02-16 Vitor Hadad , David A. Hirshberg , Ruohan Zhan , Stefan Wager , Susan Athey

Refining one's hypotheses in the light of data is a common scientific practice; however, the dependency on the data introduces selection bias and can lead to specious statistical analysis. An approach for addressing this is via conditioning…

Machine Learning · Computer Science 2020-03-03 Jen Ning Lim , Makoto Yamada , Wittawat Jitkrittum , Yoshikazu Terada , Shigeyuki Matsui , Hidetoshi Shimodaira

Linear regression is ubiquitous in statistical analysis. It is well understood that conflicting sources of information may contaminate the inference when the classical normality of errors is assumed. The contamination caused by the light…

Methodology · Statistics 2019-06-13 Philippe Gagnon , Alain Desgagné , Mylène Bédard

In this paper, we consider the problem of linear regression with heavy-tailed distributions. Different from previous studies that use the squared loss to measure the performance, we choose the absolute loss, which is capable of estimating…

Machine Learning · Computer Science 2018-10-26 Lijun Zhang , Zhi-Hua Zhou

There is an especially strong need in modern large-scale data analysis to prioritize samples for manual inspection. For example, the inspection could target important mislabeled samples or key vulnerabilities exploitable by an adversarial…

Machine Learning · Statistics 2017-05-11 Mike Wojnowicz , Ben Cruz , Xuan Zhao , Brian Wallace , Matt Wolff , Jay Luan , Caleb Crable

Understanding the influence of a training instance on a neural network model leads to improving interpretability. However, it is difficult and inefficient to evaluate the influence, which shows how a model's prediction would be changed if a…

Machine Learning · Computer Science 2021-11-22 Sosuke Kobayashi , Sho Yokoi , Jun Suzuki , Kentaro Inui

Sub-sampling is a common and often effective method to deal with the computational challenges of large datasets. However, for most statistical models, there is no well-motivated approach for drawing a non-uniform subsample. We show that the…

Machine Learning · Statistics 2017-09-07 Daniel Ting , Eric Brochu

Count data are omnipresent in many applied fields, often with overdispersion due to an excess of zeroes or extreme values. With mixtures of Poisson distributions representing an elegant and appealing modelling strategy, we focus here on the…

Statistics Theory · Mathematics 2022-04-25 Samuel Valiquette , Frédéric Mortier , Jean Peyhardi , Gwladys Toulemonde

Power-law distributions occur in many situations of scientific interest and have significant consequences for our understanding of natural and man-made phenomena. Unfortunately, the detection and characterization of power laws is…

Data Analysis, Statistics and Probability · Physics 2009-11-12 Aaron Clauset , Cosma Rohilla Shalizi , M. E. J. Newman

Many man-made and natural phenomena, including the intensity of earthquakes, population of cities and size of international wars, are believed to follow power-law distributions. The accurate identification of power-law patterns has…

Data Analysis, Statistics and Probability · Physics 2014-04-15 Yogesh Virkar , Aaron Clauset

The presence of units with extreme values in the dependent and/or independent variables (i.e., vertical outliers, leveraged data) has the potential to severely bias regression coefficients and/or standard errors. This is common with short…

Econometrics · Economics 2023-12-12 Annalivia Polselli

We study the spread of influence in a social network based on the Linear Threshold model. We derive an analytical expression for evaluating the expected size of the eventual influenced set for a given initial set, using the probability of…

Other Computer Science · Computer Science 2010-02-09 Srinivasan Venkatramanan , Anurag Kumar

Inference over tails is usually performed by fitting an appropriate limiting distribution over observations that exceed a fixed threshold. However, the choice of such threshold is critical and can affect the inferential results. Extreme…

Statistical Finance · Quantitative Finance 2019-02-26 Chiara Lattanzi , Manuele Leonelli

We consider multivariate extreme value statistics for independent but nonidentically distributed random vectors. In particular, the data may have varying tail copulas and also heteroscedastic marginal distributions. Assuming smoothly…

Statistics Theory · Mathematics 2026-04-14 John H. J. Einmahl , Chen Zhou