English
Related papers

Related papers: Post Selection Inference with Incomplete Maximum M…

200 papers

This paper presents a theoretical analysis of sample selection bias correction. The sample bias correction technique commonly used in machine learning consists of reweighting the cost of an error on each training point of a biased sample to…

Machine Learning · Computer Science 2008-12-18 Corinna Cortes , Mehryar Mohri , Michael Riley , Afshin Rostamizadeh

The applications of traditional statistical feature selection methods to high-dimension, low sample-size data often struggle and encounter challenging problems, such as overfitting, curse of dimensionality, computational infeasibility, and…

Machine Learning · Statistics 2023-12-19 Kexuan Li , Fangfang Wang , Lingli Yang , Ruiqi Liu

We develop a framework for post model selection inference, via marginal screening, in linear regression. At the core of this framework is a result that characterizes the exact distribution of linear functions of the response $y$,…

Methodology · Statistics 2014-03-03 Jason D Lee , Jonathan E Taylor

Adversarial detection aims to determine whether a given sample is an adversarial one based on the discrepancy between natural and adversarial distributions. Unfortunately, estimating or comparing two data distributions is extremely…

Machine Learning · Computer Science 2023-05-26 Shuhai Zhang , Feng Liu , Jiahao Yang , Yifan Yang , Changsheng Li , Bo Han , Mingkui Tan

Approximate inference in probability models is a fundamental task in machine learning. Approximate inference provides powerful tools to Bayesian reasoning, decision making, and Bayesian deep learning. The main goal is to estimate the…

Machine Learning · Computer Science 2020-03-10 Jun Han

Feature selection is essential in the analysis of molecular systems and many other fields, but several uncertainties remain: What is the optimal number of features for a simplified, interpretable model that retains essential information?…

Machine Learning · Computer Science 2025-01-22 Romina Wild , Felix Wodaczek , Vittorio Del Tatto , Bingqing Cheng , Alessandro Laio

Maximum mean discrepancy (MMD) has been widely adopted in domain adaptation to measure the discrepancy between the source and target domain distributions. Many existing domain adaptation approaches are based on the joint MMD, which is…

Machine Learning · Computer Science 2020-04-13 Wen Zhang , Dongrui Wu

In this paper we propose a heterogeneous modeling framework which achieves individual-wise feature selection and individualized covariates' effects subgrouping simultaneously. In contrast to conventional model selection approaches, the new…

Methodology · Statistics 2019-06-11 Xiwei Tang , Fei Xue , Annie Qu

Assessing the importance of individual features in Machine Learning is critical to understand the model's decision-making process. While numerous methods exist, the lack of a definitive ground truth for comparison highlights the need for…

Machine Learning · Computer Science 2025-12-05 Eddie Conti , Álvaro Parafita , Axel Brando

The principal support vector machines method (Li et al., 2011) is a powerful tool for sufficient dimension reduction that replaces original predictors with their low-dimensional linear combinations without loss of information. However, the…

Machine Learning · Statistics 2019-12-02 Jun Jin , Chao Ying , Zhou Yu

While statistics focusses on hypothesis testing and on estimating (properties of) the true sampling distribution, in machine learning the performance of learning algorithms on future data is the primary issue. In this paper we bridge the…

Machine Learning · Computer Science 2009-12-30 Marcus Hutter

Outlying observations are frequently encountered across a wide spectrum of scientific domains, posing notable challenges to the generalizability of statistical models and the reproducibility of downstream analysis. They are identified…

Methodology · Statistics 2026-03-17 Dongliang Zhang , Masoud Asgharian , Martin A. Lindquist

Understanding and identifying different types of single molecules' diffusion that occur in a broad range of systems (including living matter) is extremely important, as it can provide information on the physical and chemical characteristics…

Quantitative Methods · Quantitative Biology 2023-03-07 Patrycja Kowalek , Hanna Loch-Olszewska , Łukasz Łaszczuk , Jarosław Opała , Janusz Szwabiński

Measuring distances in a multidimensional setting is a challenging problem, which appears in many fields of science and engineering. In this paper, to measure the distance between two multivariate distributions, we introduce a new measure…

Methodology · Statistics 2024-11-05 Gennaro Auricchio , Giovanni Brigati , Paolo Giudici , Giuseppe Toscani

Data collection is a fundamental problem in the scenario of big data, where the size of sampling sets plays a very important role, especially in the characterization of data structure. This paper considers the information collection process…

Information Theory · Computer Science 2018-01-23 Shanyun Liu , Rui She , Pingyi Fan

We propose a scalable divergence estimation method based on hashing. Consider two continuous random variables $X$ and $Y$ whose densities have bounded support. We consider a particular locality sensitive random hashing, and consider the…

Information Theory · Computer Science 2018-01-03 Morteza Noshad , Alfred O. Hero

There has been great interest recently in applying nonparametric kernel mixtures in a hierarchical manner to model multiple related data samples jointly. In such settings several data features are commonly present: (i) the related samples…

Methodology · Statistics 2017-04-18 Jacopo Soriano , Li Ma

Selecting relevant features is an important and necessary step for intelligent machines to maximize their chances of success. However, intelligent machines generally have no enough computing resources when faced with huge volume of data.…

Machine Learning · Computer Science 2025-07-04 Hexiang Bai , Deyu Li , Jiye Liang , Yanhui Zhai

Automated feature selection is important for text categorization to reduce the feature size and to speed up the learning process of classifiers. In this paper, we present a novel and efficient feature selection framework based on the…

Machine Learning · Statistics 2016-11-15 Bo Tang , Steven Kay , Haibo He

Quantifying the difference between probability distributions is crucial in machine learning. However, estimating statistical divergences from empirical samples is challenging due to unknown underlying distributions. This work proposes the…

Machine Learning · Computer Science 2024-10-25 Jhoan K. Hoyos-Osorio , Luis G. Sanchez-Giraldo
‹ Prev 1 3 4 5 6 7 10 Next ›