中文
相关论文

相关论文: Comparing Two Contaminated Samples

200 篇论文

The goal of two-sample tests is to assess whether two samples, $S_P \sim P^n$ and $S_Q \sim Q^m$, are drawn from the same distribution. Perhaps intriguingly, one relatively unexplored method to build two-sample tests is the use of binary…

机器学习 · 统计学 2018-03-14 David Lopez-Paz , Maxime Oquab

$P$-values that are derived from continuously distributed test statistics are typically uniformly distributed on $(0,1)$ under least favorable parameter configurations (LFCs) in the null hypothesis. Conservativeness of a $p$-value $P$…

统计方法学 · 统计学 2023-03-13 Daniel Ochieng , Anh-Tuan Hoang , Thorsten Dickhaus

A common assumption in causal inference from observational data is that there is no hidden confounding. Yet it is, in general, impossible to verify this assumption from a single dataset. Under the assumption of independent causal mechanisms…

统计方法学 · 统计学 2023-11-07 Rickard K. A. Karlsson , Jesse H. Krijthe

In scientific inference problems, the underlying statistical modeling assumptions have a crucial impact on the end results. There exist, however, only a few automatic means for validating these fundamental modelling assumptions. The…

统计方法学 · 统计学 2019-05-21 Andreas Svensson , Dave Zachariah , Petre Stoica , Thomas B. Schön

We present a sample path dependent measure of causal influence between two time series. The proposed measure is a random variable whose expected sum is the directed information. A realization of the proposed measure may be used to identify…

信息论 · 计算机科学 2018-10-15 Gabriel Schamberg , Todd P. Coleman

Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…

统计理论 · 数学 2017-07-25 Abhik Ghosh , Ayanendranath Basu

We consider the conditional randomization test as a way to account for covariate imbalance in randomized experiments. The test accounts for covariate imbalance by comparing the observed test statistic to the null distribution of the test…

We study two-sample variable selection: identifying variables that discriminate between the distributions of two sets of data vectors. Such variables help scientists understand the mechanisms behind dataset discrepancies. Although…

Quality data is a fundamental contributor to success in statistics and machine learning. If a statistical assessment or machine learning leads to decisions that create value, data contributors may want a share of that value. This paper…

计算机科学与博弈论 · 计算机科学 2019-06-28 Eric Bax

Estimating statistical models within sensor networks requires distributed algorithms, in which both data and computation are distributed across the nodes of the network. We propose a general approach for distributed learning based on…

机器学习 · 计算机科学 2012-07-03 Qiang Liu , Alexander Ihler

An unbinned statistical test on cluster-like deviations from Poisson processes for point process data is introduced, presented in the context of time variability analysis of astrophysical sources in count rate experiments. The measure of…

天体物理学 · 物理学 2007-05-23 Juergen Prahl

In this paper, we consider testing the homogeneity of risk differences in independent binomial distributions especially when data are sparse. We point out some drawback of existing tests in either controlling a nominal size or obtaining…

统计方法学 · 统计学 2018-05-31 Junyong Park , Iris Ivy Gauran

We study the data-driven selection of causal graphical models using constraint-based algorithms, which determine the existence or non-existence of edges (causal connections) in a graph based on testing a series of conditional independence…

统计方法学 · 统计学 2026-04-29 Daniel Malinsky

Inference based on the penalized density ratio model is proposed and studied. The model under consideration is specified by assuming that the log--likelihood function of two unknown densities is of some parametric form. The model has been…

统计理论 · 数学 2008-07-17 Konstantinos Fokianos

We consider the problem of assessing whether, in an individual case, there is a causal relationship between an observed exposure and a response variable. When data are available on similar individuals we may be able to estimate prospective…

统计理论 · 数学 2023-11-15 Monica Musio , Philip Dawid

We consider testing whether a set of Gaussian variables, selected from the data, is independent of the remaining variables. We assume that this set is selected via a very simple approach that is commonly used across scientific disciplines:…

统计方法学 · 统计学 2022-11-04 Arkajyoti Saha , Daniela Witten , Jacob Bien

We consider the problem of hypotheses testing with the basic simple hypothesis: observed sequence of points corresponds to stationary Poisson process with known intensity against a composite one-sided parametric alternative that this is a…

统计理论 · 数学 2007-06-13 Serguei Dachian , Yury A. Kutoyants

We investigate the sample complexity of mutual information and conditional mutual information testing. For conditional mutual information testing, given access to independent samples of a triple of random variables $(A, B, C)$ with unknown…

数据结构与算法 · 计算机科学 2025-06-05 Jan Seyfried , Sayantan Sen , Marco Tomamichel

The analysis of count data is commonly done using Poisson models. Negative binomial models are a straightforward and readily motivated generalization for the case of overdispersed data, i.e., when the observed variance is greater than…

统计方法学 · 统计学 2016-01-06 Christian Röver , Stefan Andreas , Tim Friede

We consider the problem of distributed binary hypothesis testing of two sequences that are generated by an i.i.d. doubly-binary symmetric source. Each sequence is observed by a different terminal. The two hypotheses correspond to different…

信息论 · 计算机科学 2018-01-03 Eli Haim , Yuval Kochman