English
Related papers

Related papers: Testing Most Influential Sets

200 papers

Identifying leading measurement units from a large collection is a common inference task in various domains of large-scale inference. Testing approaches, which measure evidence against a null hypothesis rather than effect magnitude, tend to…

Methodology · Statistics 2020-11-17 Nicholas C. Henderson , Michael A. Newton

The vast majority of theoretical results in machine learning and statistics assume that the available training data is a reasonably reliable reflection of the phenomena to be learned or estimated. Similarly, the majority of machine learning…

Machine Learning · Computer Science 2017-06-13 Moses Charikar , Jacob Steinhardt , Gregory Valiant

Heavy-tailed distributions naturally occur in many real life problems. Unfortunately, it is typically not possible to compute inference in closed-form in graphical models which involve such heavy-tailed distributions. In this work, we…

Machine Learning · Computer Science 2011-03-22 Danny Bickson , Carlos Guestrin

The least cost influence maximization problem aims to determine minimum cost of partial (e.g., monetary) incentives initially given to the influential spreaders on a social network, so that these early adopters exert influence toward their…

Optimization and Control · Mathematics 2022-09-28 Cheng-Lung Chen , Eduardo Pasiliao , Vladimir Boginski

Researchers often run resource-intensive randomized controlled trials (RCTs) to estimate the causal effects of interventions on outcomes of interest. Yet these outcomes are often noisy, and estimated overall effects can be small or…

Econometrics · Economics 2023-12-21 Jann Spiess , Vasilis Syrgkanis , Victor Yaneng Wang

An important aspect of multiple hypothesis testing is controlling the significance level, or the level of Type I error. When the test statistics are not independent it can be particularly challenging to deal with this problem, without…

Statistics Theory · Mathematics 2009-03-04 Sandy Clarke , Peter Hall

Organizational growth processes have consistently been shown to exhibit a fatter-than-Gaussian growth-rate distribution in a variety of settings. Long periods of relatively small changes are interrupted by sudden changes in all size scales.…

Physics and Society · Physics 2014-08-13 Hernan Mondani , Petter Holme , Fredrik Liljeros

Analysis of competing risks data plays an important role in the lifetime data analysis. Recently Feizjavadian and Hashemi (Computational Statistics and Data Analysis, vol. 82, 19-34, 2015) provided a classical inference of a competing risks…

Methodology · Statistics 2021-05-04 Debashis Samanta , Debasis Kundu

Deep neural networks have amply demonstrated their prowess but estimating the reliability of their predictions remains challenging. Deep Ensembles are widely considered as being one of the best methods for generating uncertainty estimates…

Machine Learning · Computer Science 2021-06-28 Nikita Durasov , Timur Bagautdinov , Pierre Baque , Pascal Fua

Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples,…

Computation and Language · Computer Science 2025-09-16 Shuzheng Si , Haozhe Zhao , Gang Chen , Yunshui Li , Kangyang Luo , Chuancheng Lv , Kaikai An , Fanchao Qi , Baobao Chang , Maosong Sun

Heavy-tailed or power-law distributions are becoming increasingly common in biological literature. A wide range of biological data has been fitted to distributions with heavy tails. Many of these studies use simple fitting methods to find…

Quantitative Methods · Quantitative Biology 2007-12-06 A. James , M. J. Plank

Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determine the adequacy of a dataset prior to training and evaluating a model's performance would be…

Machine Learning · Computer Science 2026-04-28 Arya Hatamian , Lionel Levine , Haniyeh Ehsani Oskouie , Majid Sarrafzadeh

We study the problem of estimating the mean of a distribution in high dimensions when either the samples are adversarially corrupted or the distribution is heavy-tailed. Recent developments in robust statistics have established efficient…

Data Structures and Algorithms · Computer Science 2021-01-20 Samuel B. Hopkins , Jerry Li , Fred Zhang

We consider a robust version of the classical Wald test statistics for testing simple and composite null hypotheses for general parametric models. These test statistics are based on the minimum density power divergence estimators instead of…

Statistics Theory · Mathematics 2016-07-04 Abhik Ghosh , Abhijit Mandal , Nirian Martin , Leandro Pardo

In this paper we are interested in studying concise representations of concepts and dependencies, i.e., implications and association rules. Such representations are based on equivalence classes and their elements, i.e., minimal generators,…

Discrete Mathematics · Computer Science 2022-11-28 Aleksey Buzmakov , Egor Dudyrev , Sergei O. Kuznetsov , Tatiana Makhalova , Amedeo Napoli

This paper concerns the development of an inferential framework for high-dimensional linear mixed effect models. These are suitable models, for instance, when we have $n$ repeated measurements for $M$ subjects. We consider a scenario where…

Methodology · Statistics 2019-12-17 Lina Lin , Mathias Drton , Ali Shojaie

We explore extreme value phenomena in spatial scale-free random graphs in a continuum setting based on a homogeneous Poisson point process in $\mathbb{R}^d$. Vertices carry i.i.d. weights $(W_x)$ and, conditionally on the vertex set and the…

Probability · Mathematics 2026-02-17 Arnaud Rousselle , Ercan Sönmez

Expectiles define the only law-invariant, coherent and elicitable risk measure apart from the expectation. The popularity of expectile-based risk measures is steadily growing and their properties have been studied for independent data, but…

Methodology · Statistics 2021-10-13 Anthony C. Davison , Simone A. Padoan , Gilles Stupfler

Confounding variables are a recurrent challenge for causal discovery and inference. In many situations, complex causal mechanisms only manifest themselves in extreme events, or take simpler forms in the extremes. Stimulated by data on…

Methodology · Statistics 2024-11-14 Olivier C. Pasche , Valérie Chavez-Demoulin , Anthony C. Davison

The problem of estimating the coefficient of bivariate tail dependence is considered here from the robustness point of view; it combines two apparently contradictory theories of robust statistics and extreme value statistics. The usual…

Applications · Statistics 2014-07-08 Abhik Ghosh