中文
相关论文

相关论文: Comparing Two Contaminated Samples

200 篇论文

When evaluating causal influence from one time series to another in a multivariate dataset it is necessary to take into account the conditioning effect of the other variables. In the presence of many variables, and possibly of a reduced…

数据分析、统计与概率 · 物理学 2012-03-26 Daniele Marinazzo , Mario Pellicoro , Sebastiano Stramaglia

Randomization tests allow simple and unambiguous tests of null hypotheses, by comparing observed data to a null ensemble in which experimentally-controlled variables are randomly resampled. In behavioral and neuroscience experiments,…

统计方法学 · 统计学 2023-11-08 Kenneth D. Harris , Kevin J. Miller

A widely used method to create a continuous representation of a discrete data-set is regression analysis. When the regression model is not based on a mathematical description of the physics underlying the data, heuristic techniques play a…

统计理论 · 数学 2013-07-18 Giovanni Mana , Paolo Alberto Giuliano Albo , Simona Lago

A mixture with varying concentrations is a modification of a finite mixture model in which the mixing probabilities (concentrations of mixture components) may be different for different observations. In the paper, we assume that the…

概率论 · 数学 2015-03-19 Alexey Doronin , Rostyslav Maiboroda

Given a set of several inputs into a system (e.g., independent variables characterizing stimuli) and a set of several stochastically non-independent outputs (e.g., random variables describing different aspects of responses), how can one…

人工智能 · 计算机科学 2011-08-30 Ehtibar N. Dzhafarov , Janne V. Kujala

We study the problem of hypothesis testing between two discrete distributions, where we only have access to samples after the action of a known reversible Markov chain, playing the role of noise. We derive instance-dependent minimax rates…

统计理论 · 数学 2018-08-15 Quentin Berthet , Varun Kanade

Analysis of three-way data is becoming ever more prevalent in the literature, especially in the area of clustering and classification. Real data, including real three-way data, are often contaminated by potential outlying observations.…

We study the problems of sequential nonparametric two-sample and independence testing. Sequential tests process data online and allow using observed data to decide whether to stop and reject the null hypothesis or to collect more data,…

机器学习 · 统计学 2023-07-21 Aleksandr Podkopaev , Aaditya Ramdas

We introduce a new conservative test for quantifying the consistency of two or more datasets. The test is based on the Bayesian answer to the question, ``How much more probable is it that all my data were generated from the same model…

天体物理学 · 物理学 2008-11-26 Phil Marshall , Nutan Rajguru , Anze Slosar

This paper investigates a statistical procedure for testing the equality of two independent estimated covariance matrices when the number of potentially dependent data vectors is large and proportional to the size of the vectors, that is,…

统计理论 · 数学 2020-06-01 Rémy Mariétan , Stephan Morgenthaler

A new combinatorial-probabilistic diagnostic entropy has been introduced. It describes the pair-wise sum of probabilities of system conditions that have to be distinguished during the diagnosing process. The proposed measure describes the…

信息论 · 计算机科学 2009-09-29 Henryk Borowczyk

The paper investigates the problem of performing correlation analysis when the number of observations is very large. In such a case, it is often necessary to combine the random observations to achieve dimensionality reduction of the…

信息论 · 计算机科学 2020-10-19 Pavel Loskot

Machine-learning classifiers can be leveraged as a two-sample statistical test. Suppose each sample is assigned a different label and that a classifier can obtain a better-than-chance result discriminating them. In this case, we can infer…

机器学习 · 计算机科学 2022-12-20 Alejandro Álvarez-Ayllón , Manuel Palomo-Duarte , Juan-Manuel Dodero

We propose a method for inferring the existence of a latent common cause ('confounder') of two observed random variables. The method assumes that the two effects of the confounder are (possibly nonlinear) functions of the confounder plus…

机器学习 · 统计学 2012-05-14 Dominik Janzing , Jonas Peters , Joris Mooij , Bernhard Schoelkopf

A distributed binary hypothesis testing problem, in which multiple observers transmit their observations to a detector over noisy channels, is studied. Given its own side information, the goal of the detector is to decide between two…

信息论 · 计算机科学 2017-04-06 Sreejith Sreekumar , Deniz Gündüz

The practice of pooling several individual test statistics to form aggregate tests is common in many statistical application where individual tests may be underpowered. While selection by aggregate tests can serve to increase power, the…

统计方法学 · 统计学 2020-12-08 Ruth Heller , Amit Meir , Nilanjan Chatterjee

In this paper, we investigate the verification of dissipativity properties for polynomial systems without an explicitly identified model but directly from noise-corrupted measurements. Contrary to most data-driven approaches for nonlinear…

系统与控制 · 电气工程与系统科学 2020-11-11 Tim Martin , Frank Allgöwer

We pose a fundamental question in computational learning theory: can we efficiently test whether a training set satisfies the assumptions of a given noise model? This question has remained unaddressed despite decades of research on learning…

机器学习 · 计算机科学 2026-05-11 Surbhi Goel , Adam R. Klivans , Konstantinos Stavropoulos , Arsen Vasilyan

We develop large sample theory for merged data from multiple sources. Main statistical issues treated in this paper are (1) the same unit potentially appears in multiple datasets from overlapping data sources, (2) duplicated items are not…

统计理论 · 数学 2018-05-22 Takumi Saegusa

Causal inference is a fundamental research topic for discovering the cause-effect relationships in many disciplines. However, not all algorithms are equally well-suited for a given dataset. For instance, some approaches may only be able to…