中文
相关论文

相关论文: Gower's similarity coefficients with automatic wei…

200 篇论文

The effect that weighted summands have on each other in approximations of $S=w_1S_1+w_2S_2+\cdots+w_NS_N$ is investigated. Here, $S_i$'s are sums of integer-valued random variables, and $w_i$ denote weights, $i=1,\dots,N$. Two cases are…

概率论 · 数学 2018-06-12 Vydas Čekanavičius , Palaniappan Vellaisamy

Nearest neighbor imputation is popular for handling item nonresponse in survey sampling. In this article, we study the asymptotic properties of the nearest neighbor imputation estimator for general population parameters, including…

统计方法学 · 统计学 2017-07-05 Shu Yang , Jae Kwang Kim

Trustworthy classifiers are essential to the adoption of machine learning predictions in many real-world settings. The predicted probability of possible outcomes can inform high-stakes decision making, particularly when assessing the…

机器学习 · 计算机科学 2023-02-22 Kiri L. Wagstaff , Thomas G. Dietterich

It is well known that general variational inequalities provide us with a unified, natural, novel and simple framework to study a wide class of unrelated problems, which arise in pure and applied sciences. In this paper, we present a number…

最优化与控制 · 数学 2020-09-24 M. A. Noor , K. I. Noor , M. Th. Rassias

Understanding and developing a correlation measure that can detect general dependencies is not only imperative to statistics and machine learning, but also crucial to general scientific discovery in the big data age. In this paper, we…

机器学习 · 统计学 2024-06-27 Cencheng Shen , Carey E. Priebe , Joshua T. Vogelstein

A learned generative model often produces biased statistics relative to the underlying data distribution. A standard technique to correct this bias is importance sampling, where samples from the model are weighted by the likelihood ratio…

We propose an extensive simulation study to compare some variable selection procedures in a high-dimensional framework. Assuming that the relationship between the actives variables and the response variable is linear, the high-dimensional…

应用统计 · 统计学 2025-03-21 Perrine Lacroix , Mélina Gallopin , Marie-Laure Martin

Deep metric learning aims at learning the distance metric between pair of samples, through the deep neural networks to extract the semantic feature embeddings where similar samples are close to each other while dissimilar samples are…

计算机视觉与模式识别 · 计算机科学 2019-05-31 Haijun Liu , Jian Cheng , Wen Wang , Yanzhou Su

Statistical tasks such as density estimation and approximate Bayesian inference often involve densities with unknown normalising constants. Score-based methods, including score matching, are popular techniques as they are free of…

机器学习 · 统计学 2021-12-22 Li K. Wenliang , Heishiro Kanagawa

Attribute weighting and differential weighting, two major mechanisms for computing context-dependent similarity or dissimilarity measures are studied and compared. A dissimilarity measure based on subset size in the context is proposed and…

人工智能 · 计算机科学 2013-04-05 Yizong Cheng

Spaces with locally varying scale of measurement, like multidimensional structures with differently scaled dimensions, are pretty common in statistics and machine learning. Nevertheless, it is still understood as an open question how to…

The propensity score analysis is one of the most widely used methods for studying the causal treatment effect in observational studies. This paper studies treatment effect estimation with the method of matching weights. This method…

统计方法学 · 统计学 2011-05-17 Liang Li

Feature selection can facilitate the learning of mixtures of discrete random variables as they arise, e.g. in crowdsourcing tasks. Intuitively, not all workers are equally reliable but, if the less reliable ones could be eliminated, then…

机器学习 · 统计学 2017-11-28 Vincent Zhao , Steven W. Zucker

In theory, the probabilistic linkage method provides two distinct advantages over non-probabilistic methods, including minimal rates of linkage error and accurate measures of these rates for data users. However, implementations can fall…

统计方法学 · 统计学 2019-11-06 Abel Dasylva , Arthur Goussanou , David Ajavon , Hanan Abousaleh

Model averaging is a useful and robust method for dealing with model uncertainty in statistical analysis. Often, it is useful to consider data subset selection at the same time, in which model selection criteria are used to compare models…

统计方法学 · 统计学 2023-10-26 Ethan T. Neil , Jacob W. Sitison

One of the fundamental problems in Bayesian statistics is the approximation of the posterior distribution. Gibbs sampler and coordinate ascent variational inference are renownedly utilized approximation techniques that rely on stochastic…

统计理论 · 数学 2021-06-18 Se Yoon Lee

We propose a new method for multivariate response regression and covariance estimation when elements of the response vector are of mixed types, for example some continuous and some discrete. Our method is based on a model which assumes the…

统计方法学 · 统计学 2022-03-04 Karl Oskar Ekvall , Aaron J. Molstad

Measures of similarity (or dissimilarity) are a key ingredient to many machine learning algorithms. We introduce DID, a pairwise dissimilarity measure applicable to a wide range of data spaces, which leverages the data's internal structure…

机器学习 · 统计学 2022-03-08 Théophile Cantelobre , Carlo Ciliberto , Benjamin Guedj , Alessandro Rudi

Distance correlation is a novel class of multivariate dependence measure, taking positive values between 0 and 1, and applicable to random vectors of arbitrary dimensions, not necessarily equal. It offers several advantages over the…

统计计算 · 统计学 2024-05-06 Blanca E. Monroy-Castillo , M. A , Jácome , Ricardo Cao

Gaussian processes (GPs) are powerful models for human-in-the-loop experiments due to their flexibility and well-calibrated uncertainty. However, GPs modeling human responses typically ignore auxiliary information, including a priori domain…

机器学习 · 计算机科学 2025-03-07 Kaiwen Wu , Craig Sanders , Benjamin Letham , Phillip Guan