中文
相关论文

相关论文: Comparing Two Contaminated Samples

200 篇论文

Hypothesis testing is a statistical inference approach used to determine whether data supports a specific hypothesis. An important type is the two-sample test, which evaluates whether two sets of data points are from identical…

机器学习 · 计算机科学 2025-01-08 Weizhi Li , Visar Berisha , Gautam Dasarathy

The recent success of generative adversarial networks and variational learning suggests training a classifier network may work well in addressing the classical two-sample problem. Network-based tests have the computational advantage that…

机器学习 · 统计学 2022-06-01 Xiuyuan Cheng , Alexander Cloninger

We present a general framework for hypothesis testing on distributions of sets of individual examples. Sets may represent many common data sources such as groups of observations in time series, collections of words in text or a batch of…

统计方法学 · 统计学 2021-02-03 Alexis Bellot , Mihaela van der Schaar

Integrating data from multiple heterogeneous sources has become increasingly popular to achieve a large sample size and diverse study population. This paper reviews development in causal inference methods that combines multiple datasets…

统计方法学 · 统计学 2021-10-05 Xu Shi , Ziyang Pan , Wang Miao

Missing data is a common issue in many biomedical studies. Under a paired design, some subjects may have missing values in either one or both of the conditions due to loss of follow-up, insufficient biological samples, etc. Such partially…

In this paper, we study the problem of determining $k$ anomalous random variables that have different probability distributions from the rest $(n-k)$ random variables. Instead of sampling each individual random variable separately as in the…

信息论 · 计算机科学 2024-09-09 Myung Cho , Weiyu Xu , Lifeng Lai

For three natural classes of dynamic decision problems; 1. additively separable problems, 2. discounted problems, and 3. discounted problems for a fixed discount factor; we provide necessary and sufficient conditions for one sequential…

理论经济学 · 经济学 2024-05-24 Mark Whitmeyer , Cole Williams

Given a pair of multivariate time-series data of the same length and dimensions, an approach is proposed to select variables and time intervals where the two series are significantly different. In applications where one time series is an…

统计方法学 · 统计学 2024-12-11 Kensuke Mitsuzawa , Margherita Grossi , Stefano Bortoli , Motonobu Kanagawa

Data can be collected in scientific studies via a controlled experiment or passive observation. Big data is often collected in a passive way, e.g. from social media. In studies of causation great efforts are made to guard against bias and…

统计方法学 · 统计学 2018-11-21 Elena Pesce , Eva Riccomagno , Henry P. Wynn

Methods for quantifying the similarity of datasets are relevant in applications where two or more datasets, or their underlying distributions, need to be compared, ranging from two- and k-sample testing to applications in machine learning…

统计方法学 · 统计学 2026-04-15 Marieke Stolte , Jörg Rahnenführer , Andrea Bommert

We introduce a new discrepancy score between two distributions that gives an indication on their similarity. While much research has been done to determine if two samples come from exactly the same distribution, much less research…

机器学习 · 计算机科学 2012-10-16 Maayan Harel , Shie Mannor

We investigate the problem of jointly testing two hypotheses and estimating a random parameter based on data that is observed sequentially by sensors in a distributed network. In particular, we assume the data to be drawn from a Gaussian…

信号处理 · 电气工程与系统科学 2020-03-04 Dominik Reinhard , Michael Fauß , Abdelhak M. Zoubir

The paper provides a simple test for deciding, from a given causal diagram, whether two sets of variables have the same bias-reducing potential under adjustment. The test requires that one of the following two conditions holds: either (1)…

统计方法学 · 统计学 2012-03-19 Judea Pearl , Azaria Paz

We propose and study a general method for construction of consistent statistical tests on the basis of possibly indirect, corrupted, or partially available observations. The class of tests devised in the paper contains Neyman's smooth…

统计理论 · 数学 2017-09-22 Mikhail Langovoy

We propose a new nonparametric test for the supposition of independence between two continuous random variables. The test is based on the size of the longest increasing subsequence of a random permutation. We identified the independence…

统计方法学 · 统计学 2015-03-13 Jesus E. Garcia , Veronica A. Gonzalez-Lopez

Several approaches to testing the hypothesis that two histograms are drawn from the same distribution are investigated. We note that single-sample continuous distribution tests may be adapted to this two-sample grouped data situation. The…

数据分析、统计与概率 · 物理学 2008-04-03 Frank C. Porter

We consider testing statistical hypotheses about densities of signals in deconvolution models. A new approach to this problem is proposed. We constructed score tests for the deconvolution with the known noise density and efficient score…

统计理论 · 数学 2013-12-02 Mikhail Langovoy

Data taking values on discrete sample spaces are the embodiment of modern biological research. "Omics" experiments produce millions of symbolic outcomes in the form of reads (i.e., DNA sequences of a few dozens to a few hundred…

统计方法学 · 统计学 2023-07-14 Antony Pearson , Manuel E. Lladser

Economic data are often generated by stochastic processes that take place in continuous time, though observations may occur only at discrete times. For example, electricity and gas consumption take place in continuous time. Data generated…

计量经济学 · 经济学 2021-06-15 Federico A. Bugni , Joel L. Horowitz

In clinical studies with paired organs, binary outcomes often exhibit intra-subject correlation and may include a mixture of unilateral and bilateral observations. Under Donner's constant correlation model, we develop three likelihood-based…

统计方法学 · 统计学 2025-10-22 Jia Zhou , Chang-Xing Ma