English
Related papers

Related papers: Testing Most Influential Sets

200 papers

Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…

Statistics Theory · Mathematics 2017-07-25 Abhik Ghosh , Ayanendranath Basu

We study the problem of heavy-tailed mean estimation in settings where the variance of the data-generating distribution does not exist. Concretely, given a sample $\mathbf{X} = \{X_i\}_{i = 1}^n$ from a distribution $\mathcal{D}$ over…

Statistics Theory · Mathematics 2020-12-10 Yeshwanth Cherapanamjeri , Nilesh Tripuraneni , Peter L. Bartlett , Michael I. Jordan

Data exhibiting heavy-tails in one or more dimensions is often studied using the framework of regular variation. In a multivariate setting this requires identifying specific forms of dependence in the data; this means identifying that the…

Statistics Theory · Mathematics 2017-02-02 Bikramjit Das , Sidney I. Resnick

In a standard classification framework a set of trustworthy learning data are employed to build a decision rule, with the final aim of classifying unlabelled units belonging to the test set. Therefore, unreliable labelled observations,…

Applications · Statistics 2019-11-20 Andrea Cappozzo , Francesca Greselin , Thomas Brendan Murphy

Improving the quality of training samples is crucial for improving the reliability and performance of ML models. In this paper, we conduct a comparative evaluation of influence-based signals for debugging training data. These signals can…

Machine Learning · Computer Science 2025-06-16 Nikolaos Myrtakis , Ioannis Tsamardinos , Vassilis Christophides

Using the proposed by us thinning approach to describe extreme matrices, we find an explicit exponentiation formula linking classical extreme laws of Fr\'echet, Gumbel and Weibull given by Fisher-Tippet-Gnedenko classification and free…

Mathematical Physics · Physics 2020-08-19 Jacek Grela , Maciej A. Nowak

The theory of influence and sharp threshold is a key tool in probability and probabilistic combinatorics, with numerous applications. One significant aspect of the theory is directed at identifying the level of generality of the product…

Probability · Mathematics 2015-04-27 Geoffrey Grimmett , Svante Janson , James Norris

We study the problem of robust influence maximization in dynamic diffusion networks. In line with recent works, we consider the scenario where the network can undergo insertion and removal of nodes and edges, in discrete time steps, and the…

Databases · Computer Science 2024-12-17 Arkaprava Saha , Bogdan Cautis , Xiaokui Xiao , Laks V. S. Lakshmanan

In the problem of composite hypothesis testing, identifying the potential uniformly most powerful (UMP) unbiased test is of great interest. Beyond typical hypothesis settings with exponential family, it is usually challenging to prove the…

Methodology · Statistics 2022-08-03 Tianyu Zhan , Jian Kang

Ensemble learning is a popular technique to improve the accuracy of machine learning models. It traditionally hinges on the rationale that aggregating multiple weak models can lead to better models with lower variance and hence higher…

Optimization and Control · Mathematics 2026-01-06 Huajie Qian , Donghao Ying , Henry Lam , Wotao Yin

There is growing interest in developing statistical estimators that achieve exponential concentration around a population target even when the data distribution has heavier than exponential tails. More recent activity has focused on…

Statistics Theory · Mathematics 2025-04-22 Jakwang Kim , Jiyoung Park , Anirban Bhattacharya

It is well-known that trimmed sample means are robust against heavy tails and data contamination. This paper analyzes the performance of trimmed means and related methods in two novel contexts. The first one consists of estimating…

Statistics Theory · Mathematics 2025-12-03 Roberto I. Oliveira , Lucas Resende

Learning generalized models from biased data is an important undertaking toward fairness in deep learning. To address this issue, recent studies attempt to identify and leverage bias-conflicting samples free from spurious correlations…

Machine Learning · Computer Science 2024-11-04 Yeonsung Jung , Jaeyun Song , June Yong Yang , Jin-Hwa Kim , Sung-Yub Kim , Eunho Yang

The study of geometric extremes, where extremal dependence properties are inferred from the deterministic limiting shapes of scaled sample clouds, provides an exciting approach to modelling the extremes of multivariate data. These shapes,…

Methodology · Statistics 2024-09-16 Callum J. R. Murphy-Barltrop , Reetam Majumder , Jordan Richards

Models for extreme values are generally derived from limit results, which are meant to be good enough approximations when applied to finite samples. Depending on the speed of convergence of the process underlying the data, these…

Statistics Theory · Mathematics 2019-02-20 Thomas Lugrin , Anthony C. Davison , Jonathan A. Tawn

Identification of input data points relevant for the classifier (i.e. serve as the support vector) has recently spurred the interest of researchers for both interpretability as well as dataset debugging. This paper presents an in-depth…

Machine Learning · Computer Science 2020-09-30 Dominique Mercier , Shoaib Ahmed Siddiqui , Andreas Dengel , Sheraz Ahmed

Fine-tuning large language models (LLMs) on chain-of-thought (CoT) data shows that a small amount of high-quality data can outperform massive datasets. Yet, what constitutes "quality" remains ill-defined. Existing reasoning methods rely on…

Machine Learning · Computer Science 2025-12-02 Prateek Humane , Paolo Cudrano , Daniel Z. Kaplan , Matteo Matteucci , Supriyo Chakraborty , Irina Rish

Despite the successes of probabilistic models based on passing noise through neural networks, recent work has identified that such methods often fail to capture tail behavior accurately, unless the tails of the base distribution are…

Machine Learning · Statistics 2023-06-16 Feynman Liang , Liam Hodgkinson , Michael W. Mahoney

Influence maximization is a crucial issue for mining the deep information of social networks, which aims to select a seed set from the network to maximize the number of influenced nodes. To evaluate the influence spread of a seed set…

Neural and Evolutionary Computing · Computer Science 2024-10-28 Chao Wang , Jiaxuan Zhao , Lingling Li , Licheng Jiao , Jing Liu , Kai Wu

This paper proposes a new Bayesian approach for analysing moment condition models in the situation where the data may be contaminated by outliers. The approach builds upon the foundations developed by Schennach (2005) who proposed the…

Methodology · Statistics 2018-01-03 Zhichao Liu , Catherine S. Forbes , Heather M. Anderson
‹ Prev 1 3 4 5 6 7 10 Next ›