English
Related papers

Related papers: Comparing Two Contaminated Samples

200 papers

A validated simulation model primarily requires performing an appropriate input analysis mainly by determining the behavior of real-world processes using probability distributions. In many practical cases, probability distributions of the…

Applications · Statistics 2014-03-05 Issac Shams , Saeede Ajorlou , Kai Yang

This paper proposes a new framework based on joint statistical models for evaluating risks of automated vehicles in a naturalistic driving environment. The previous studies on the Accelerated Evaluation for automated vehicles are extended…

Systems and Control · Computer Science 2017-07-18 Zhiyuan Huang , Henry Lam , Ding Zhao

We revisit the problem of condensation for independent, identically distributed random variables with a power-law tail, conditioned by the value of their sum. For large values of the sum, and for a large number of summands, a condensation…

Statistical Mechanics · Physics 2022-03-03 Claude Godrèche

Heterogeneous data from multiple populations, sub-groups, or sources is often represented as a ``mixture model'' with a single latent class influencing all of the observed covariates. Heterogeneity can be resolved at multiple levels by…

Machine Learning · Computer Science 2024-12-16 Bijan Mazaheri , Chandler Squires , Caroline Uhler

In many real-world scenarios, interested variables are often represented as discretized values due to measurement limitations. Applying Conditional Independence (CI) tests directly to such discretized data, however, can lead to incorrect…

Artificial Intelligence · Computer Science 2025-06-11 Boyang Sun , Yu Yao , Xinshuai Dong , Zongfang Liu , Tongliang Liu , Yumou Qiu , Kun Zhang

We consider the problem of determining the top-$k$ largest measurements from a dataset distributed among a network of $n$ agents with noisy communication links. We show that this scenario can be cast as a distributed convex optimization…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-12-02 Xu Zhang , Marcos Vasconcelos

In randomized experiments, treatment and control groups should be roughly the same--balanced--in their distributions of pretreatment variables. But how nearly so? Can descriptive comparisons meaningfully be paired with significance tests?…

Methodology · Statistics 2008-08-29 Ben B. Hansen , Jake Bowers

To adapt kernel two-sample and independence testing to complex structured data, aggregation of multiple kernels is frequently employed to boost testing power compared to single-kernel tests. However, we observe a phenomenon that directly…

Machine Learning · Computer Science 2025-10-14 Zhijian Zhou , Xunye Tian , Liuhua Peng , Chao Lei , Antonin Schrab , Danica J. Sutherland , Feng Liu

In many scientific studies, it is of interest to determine whether an exposure has a causal effect on an outcome. In observational studies, this is a challenging task due to the presence of confounding variables that affect both the…

Methodology · Statistics 2020-10-07 Ted Westling

A recent method for causal discovery is in many cases able to infer whether X causes Y or Y causes X for just two observed variables X and Y. It is based on the observation that there exist (non-Gaussian) joint distributions P(X,Y) for…

Information Theory · Computer Science 2009-10-12 Dominik Janzing , Bastian Steudel

We describe a statistical hypothesis test for the presence of a signal based on the likelihood ratio statistic. We derive the test for a special case of interest. We study extensions of the test to cases where there are multiple channels…

Data Analysis, Statistics and Probability · Physics 2008-07-31 Wolfgang A. Rolke , Angel M. Lopez

Predictive values are measures of the clinical accuracy of a binary diagnostic test, and depend on the sensitivity and the specificity of the test and on the disease prevalence among the population being studied. This article studies…

Other Statistics · Statistics 2024-08-14 Jose Antonio Roldan-Nofuentes

We study the problem of multiple hypothesis testing for multidimensional data when inter-correlations are present. The problem of multiple comparisons is common in many applications. When the data is multivariate and correlated, existing…

Statistics Theory · Mathematics 2015-06-02 Mahdis Azadbakhsh , Xin Gao , Hanna Jankowski

We consider the problem of testing independence in mixed-type data that combine count variables with positive, absolutely continuous variables. We first introduce two distinct classes of test statistics in the bivariate setting, designed to…

Methodology · Statistics 2025-07-29 Dana Bucalo Jelić , Marija Cuparić , Bojana Milošević

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

Machine Learning · Statistics 2015-09-22 Qiyi Lu , Xingye Qiao

We consider the problem of two-sample testing in a semi-supervised setting with abundant unlabeled covariate data. Standard two-sample tests neglect covariate information, which has the potential to significantly boost performance. However,…

Machine Learning · Statistics 2026-05-05 Gyumin Lee , Shubhanshu Shekhar , Ilmun Kim

In this note, we focus on a selection model problem: a mono-exponential model versus a bi-exponential one. This is done in the biological context of living cells, where small data are available. Classical statistics are revisited to improve…

Applications · Statistics 2011-05-31 Ph. Heinrich , J. Kahn , L. Héliot , D. Trinel

Image noise can often be accurately fitted to a Poisson-Gaussian distribution. However, estimating the distribution parameters from a noisy image only is a challenging task. Here, we study the case when paired noisy and noise-free samples…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Nicolas Bähler , Majed El Helou , Étienne Objois , Kaan Okumuş , Sabine Süsstrunk

This paper is concerned with test of the conditional independence. We first establish an equivalence between the conditional independence and the mutual independence. Based on the equivalence, we propose an index to measure the conditional…

Methodology · Statistics 2021-05-18 Zhanrui Cai , Runze Li , Yaowu Zhang

Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results. The relationship between data quality and the accuracy of results could be applied on the…

Databases · Computer Science 2021-04-27 Zhixin Qi , Hongzhi Wang , Jianzhong Li , Hong Gao