English
Related papers

Related papers: Partial identification of kernel based two sample …

200 papers

The modelling of data on a spherical surface requires the consideration of directional probability distributions. To model asymmetrically distributed data on a three-dimensional sphere, Kent distributions are often used. The moment…

Machine Learning · Computer Science 2015-06-29 Parthan Kasarapu

Numerous machine learning (ML) models have been developed, including those for software engineering (SE) tasks, under the assumption that training and testing data come from the same distribution. However, training and testing distributions…

Software Engineering · Computer Science 2025-03-04 Yanfu Yan , Viet Duong , Huajie Shao , Denys Poshyvanyk

Machine learning models are vulnerable to membership inference attack, which can be used to determine whether a given sample appears in the training data. Most existing methods assume the attacker has full access to the features of the…

Machine Learning · Computer Science 2025-12-24 Xurun Wang , Guangrui Liu , Xinjie Li , Haoyu He , Lin Yao , Zhongyun Hua , Weizhe Zhang

We provide a new estimation method for conditional moment models via the martingale difference divergence (MDD).Our MDD-based estimation method is formed in the framework of a continuum of unconditional moment restrictions. Unlike the…

Econometrics · Economics 2024-04-18 Kunyang Song , Feiyu Jiang , Ke Zhu

We introduce credal two-sample testing, a new hypothesis testing framework for comparing credal sets -- convex sets of probability measures where each element captures aleatoric uncertainty and the set itself represents epistemic…

Machine Learning · Statistics 2025-03-14 Siu Lun Chau , Antonin Schrab , Arthur Gretton , Dino Sejdinovic , Krikamol Muandet

Under the null hypothesis, the marginal probability of the positive response is symmetric at any specified correlated coefficient, and the discordance probability is also symmetric to the positive response probability. The marginal…

Methodology · Statistics 2022-11-11 Guanghui Huang

Image classification models deployed in the real world may receive inputs outside the intended data distribution. For critical applications such as clinical decision making, it is important that a model can detect such out-of-distribution…

Computer Vision and Pattern Recognition · Computer Science 2021-07-07 Christoph Berger , Magdalini Paschali , Ben Glocker , Konstantinos Kamnitsas

This paper studies the high-dimensional mixed linear regression (MLR) where the output variable comes from one of the two linear regression models with an unknown mixing proportion and an unknown covariance structure of the random…

Methodology · Statistics 2020-11-10 Linjun Zhang , Rong Ma , T. Tony Cai , Hongzhe Li

Linear mixed models (LMMs) are used as an important tool in the data analysis of repeated measures and longitudinal studies. The most common form of LMMs utilize a normal distribution to model the random effects. Such assumptions can often…

Methodology · Statistics 2016-02-16 Hien D. Nguyen , Geoffrey J. McLachlan

Class imbalance in supervised classification often degrades model performance by biasing predictions toward the majority class, particularly in critical applications such as medical diagnosis and fraud detection. Traditional oversampling…

Machine Learning · Statistics 2025-09-16 Suman Cha , Hyunjoong Kim

In kernel methods, the median heuristic has been widely used as a way of setting the bandwidth of RBF kernels. While its empirical performances make it a safe choice under many circumstances, there is little theoretical understanding of why…

Statistics Theory · Mathematics 2018-10-31 Damien Garreau , Wittawat Jitkrittum , Motonobu Kanagawa

We introduce a kernel-based goodness-of-fit test for censored data, where observations may be missing in random time intervals: a common occurrence in clinical trials and industrial life-testing. The test statistic is straightforward to…

Methodology · Statistics 2018-10-11 Tamara Fernández , Arthur Gretton

In the field of machine learning, model performance is usually assessed by randomly splitting data into training and test sets. Different random splits, however, can yield markedly different performance estimates, so a genuinely good model…

One of the most well-known and simplest models for diversity maximization is the Max-Min Diversification (MMD) model, which has been extensively studied in the data mining and database literature. In this paper, we initiate the study of the…

Data Structures and Algorithms · Computer Science 2025-02-05 Iiro Kumpulainen , Florian Adriaens , Nikolaj Tatti

The problem of binary hypothesis testing between two probability measures is considered. New sharp bounds are derived for the best achievable error probability of such tests based on independent and identically distributed observations.…

Information Theory · Computer Science 2024-05-30 Valentinian Lungu , Ioannis Kontoyiannis

Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness and reliability of machine learning systems deployed in real-world applications. Recent approaches have explored the use of unlabeled data, showing…

Machine Learning · Computer Science 2025-10-09 Momin Abbas , Ali Falahati , Hossein Goli , Mohammad Mohammadi Amiri

We show that for local alternatives to uniformity which are determined by a sequence of square integrable densities the moderate deviation (MD) theorem for the corresponding Neyman-Pearson statistic does not hold in the full range for all…

Statistics Theory · Mathematics 2020-03-27 Tadeusz Inglot

A new goodness-of-fit test for normality in high-dimension (and Reproducing Kernel Hilbert Space) is proposed. It shares common ideas with the Maximum Mean Discrepancy (MMD) it outperforms both in terms of computation time and applicability…

Statistics Theory · Mathematics 2014-04-14 Jérémie Kellner , Alain Celisse

The nonparametric problem of detecting existence of an anomalous interval over a one dimensional line network is studied. Nodes corresponding to an anomalous interval (if exists) receive samples generated by a distribution q, which is…

Information Theory · Computer Science 2016-04-11 Shaofeng Zou , Yingbin Liang , H. Vincent Poor

Widely used methods for analyzing missing data can be biased in small samples. To understand these biases, we evaluate in detail the situation where a small univariate normal sample, with values missing at random, is analyzed using either…

Statistics Theory · Mathematics 2017-03-27 Paul T. von Hippel
‹ Prev 1 8 9 10 Next ›