中文
相关论文

相关论文: Improving importance estimation in covariate shift…

200 篇论文

Calibration$\unicode{x2014}$the problem of ensuring that predicted probabilities align with observed class frequencies$\unicode{x2014}$is a basic desideratum for reliable prediction with machine learning systems. Calibration error is…

机器学习 · 统计学 2026-03-02 Eugène Berta , Sacha Braun , David Holzmüller , Francis Bach , Michael I. Jordan

Neural networks make accurate predictions but often fail to provide reliable uncertainty estimates, especially under covariate distribution shifts between training and testing. To address this problem, we propose a Bayesian framework for…

机器学习 · 统计学 2025-12-22 Yuli Slavutsky , David M. Blei

Modern foundation models are trained on diverse datasets to enhance generalization across tasks and domains A central challenge in this process is determining how to effectively mix and sample data from multiple sources This naturally leads…

机器学习 · 计算机科学 2025-06-05 Zeman Li , Yuan Deng , Peilin Zhong , Meisam Razaviyayn , Vahab Mirrokni

Probabilistic classification of unassociated Fermi-LAT sources using machine learning methods has an implicit assumption that the distributions of associated and unassociated sources are the same as a function of source parameters, which is…

高能天体物理现象 · 物理学 2024-01-04 Dmitry V. Malyshev

An important feature of kernel mean embeddings (KME) is that the rate of convergence of the empirical KME to the true distribution KME can be bounded independently of the dimension of the space, properties of the distribution and smoothness…

统计理论 · 数学 2025-04-17 Geoffrey Wolfer , Pierre Alquier

In the face of dataset shift, model calibration plays a pivotal role in ensuring the reliability of machine learning systems. Calibration error (CE) is an indicator of the alignment between the predicted probabilities and the classifier…

机器学习 · 计算机科学 2023-12-15 Teodora Popordanoska , Gorjan Radevski , Tinne Tuytelaars , Matthew B. Blaschko

Interpreting black-box machine learning models is challenging due to their strong dependence on data and inherently non-parametric nature. This paper reintroduces the concept of importance through "Marginal Variable Importance Metric"…

机器学习 · 统计学 2025-01-30 Mohammad Kaviul Anam Khan , Olli Saarela , Rafal Kustra

Large Language Models (LLMs) heavily rely on high-quality training data, making data valuation crucial for optimizing model performance, especially when working within a limited budget. In this work, we aim to offer a third-party data…

机器学习 · 计算机科学 2025-05-14 Yanzhou Pan , Huawei Lin , Yide Ran , Jiamin Chen , Xiaodong Yu , Weijie Zhao , Denghui Zhang , Zhaozhuo Xu

When analyzing data from randomized clinical trials, covariate adjustment can be used to account for chance imbalance in baseline covariates and to increase precision of the treatment effect estimate. A practical barrier to covariate…

统计方法学 · 统计学 2023-07-04 Chia-Rui Chang , Yue Song , Fan Li , Rui Wang

The ratio between the probability that two distributions $R$ and $P$ give to points $x$ are known as importance weights or propensity scores and play a fundamental role in many different fields, most notably, statistics and machine…

机器学习 · 计算机科学 2021-03-11 Parikshit Gopalan , Omer Reingold , Vatsal Sharan , Udi Wieder

Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap…

机器学习 · 统计学 2025-08-25 Hongbo Chen , Li Charlie Xia

Different features have different relevance to a particular learning problem. Some features are less relevant; while some very important. Instead of selecting the most relevant features using feature selection, an algorithm can be given…

机器学习 · 计算机科学 2011-01-26 Ridwan Al Iqbal

This article focuses on measurement error in covariates in regression analyses in which the aim is to estimate the association between one or more covariates and an outcome, adjusting for confounding. Error in covariate measurements, if…

统计方法学 · 统计学 2019-10-16 Ruth H. Keogh , Jonathan W. Bartlett

We propose a new estimator for nonparametric binary choice models that does not impose a parametric structure on either the systematic function of covariates or the distribution of the error term. A key advantage of our approach is its…

计量经济学 · 经济学 2026-01-13 Guo Yan

Importance sampling with data-driven proposal distributions is widely used in practice. A common workflow first generates an auxiliary sample of size $N$ from an approximation of the target distribution, constructs a density estimate $\hat…

统计理论 · 数学 2026-05-20 Cathrine Aeckerle-Willems , Ilja Klebanov , Simon Weissmann

In machine learning applications, distribution shifts between training and target environments can lead to significant drops in model performance. This study investigates the impact of such shifts on binary classification models within the…

机器学习 · 统计学 2024-08-20 Minji Kim , Seong Jin Lee , Bumsik Kim

Importance sampling is a well developed method in statistics. Given a random variable $X$, the problem of estimating its expected value $\mu$ is addressed. The standard approach is to use the sample mean as an estimator $\bar x$. In…

应用统计 · 统计学 2014-05-09 Georg Hofmann

Estimating frequencies of certain items among a population is a basic step in data analytics, which enables more advanced data analytics (e.g., heavy hitter identification, frequent pattern mining), client software optimization, and…

密码学与安全 · 计算机科学 2018-12-12 Jinyuan Jia , Neil Zhenqiang Gong

The discovery of causal relationships in a set of random variables is a fundamental objective of science and has also recently been argued as being an essential component towards real machine intelligence. One class of causal discovery…

机器学习 · 统计学 2024-02-01 Tim Tse , Zhitang Chen , Shengyu Zhu , Yue Liu

Selecting important features in non-linear or kernel spaces is a difficult challenge in both classification and regression problems. When many of the features are irrelevant, kernel methods such as the support vector machine and kernel…

机器学习 · 统计学 2009-06-25 Genevera I. Allen