中文
相关论文

相关论文: Balanced $k$-nearest neighbor imputation

200 篇论文

Imputation of missing values is a strategy for handling non-responses in surveys or data loss in measurement processes, which may be more effective than ignoring them. When the variable represents a count, the literature dealing with this…

应用统计 · 统计学 2020-07-31 Gilma Hernández-Herrera , Albert Navarro , David Moriña

Missing feature values are a significant hurdle for downstream machine-learning tasks such as classification. However, imputation methods for classification might be time-consuming for high-dimensional data, and offer few theoretical…

机器学习 · 计算机科学 2025-05-15 Rahul Bordoloi , Clémence Réda , Saptarshi Bej , Olaf Wolkenhauer

The paper considers a distributed robust estimation problem over a network with directed topology involving continuous time observers. While measurements are available to the observers continuously, the nodes interact according to a…

系统与控制 · 计算机科学 2014-05-09 V. Ugrinovskii , E. Fridman

We introduce a new randomization procedure for experiments based on the cube method, which achieves near-exact covariate balance. This ensures compliance with standard balance tests and allows for balancing on many covariates, enabling more…

计量经济学 · 经济学 2025-07-21 Laurent Davezies , Guillaume Hollard , Pedro Vergara Merino

In multicenter research, individual-level data are often protected against sharing across sites. To overcome the barrier of data sharing, many distributed algorithms, which only require sharing aggregated information, have been developed.…

统计方法学 · 统计学 2021-03-25 Rui Duan , Yang Ning , Yong Chen

Many high dimensional integrals can be reduced to the problem of finding the relative measures of two sets. Often one set will be exponentially larger than the other, making it difficult to compare the sizes. A standard method of dealing…

概率论 · 数学 2011-12-19 Mark Huber , Sarah Schott

Iterative imputation, in which variables are imputed one at a time each given a model predicting from all the others, is a popular technique that can be convenient and flexible, as it replaces a potentially difficult multivariate modeling…

统计理论 · 数学 2012-04-04 Jingchen Liu , Andrew Gelman , Jennifer Hill , Yu-Sung Su

Missing data are often dealt with multiple imputation. A crucial part of the multiple imputation process is selecting sensible models to generate plausible values for incomplete data. A method based on posterior predictive checking is…

统计计算 · 统计学 2026-05-14 Mingyang Cai , Stef van Buuren , Gerko Vink

This paper proposes using a method named Double Score Matching (DSM) to do mass-imputation and presents an application to make inferences with a nonprobability sample. DSM is a $k$-Nearest Neighbors algorithm that uses two balance scores…

统计方法学 · 统计学 2021-10-19 Ali Furkan Kalay

Multiple importance sampling estimators are widely used for computing intractable constants due to its reliability and robustness. The celebrated balance heuristic estimator belongs to this class of methods and has proved very successful in…

统计计算 · 统计学 2019-09-05 Felipe J Medina-Aguayo , Richard G Everitt

Imputation methods for dealing with incomplete data typically assume that the missingness mechanism is at random (MAR). These methods can also be applied to missing not at random (MNAR) situations, where the user specifies some adjustment…

统计方法学 · 统计学 2024-04-24 Shahab Jolani , Stef van Buuren

In a multiple testing context, we consider a semiparametric mixture model with two components where one component is known and corresponds to the distribution of $p$-values under the null hypothesis and the other component $f$ is…

应用统计 · 统计学 2013-04-04 Van Hanh Nguyen , Catherine Matias

Item nonresponse is frequently encountered in practice. Ignoring missing data can lose efficiency and lead to misleading inference. Fractional imputation is a frequentist approach of imputation for handling missing data. However, the…

统计方法学 · 统计学 2018-09-18 Hejian Sang , Jae Kwang Kim

As in other estimation scenarios, likelihood based estimation in the normal mixture set-up is highly non-robust against model misspecification and presence of outliers (apart from being an ill-posed optimization problem). A robust…

统计方法学 · 统计学 2023-12-20 Soumya Chakraborty , Ayanendranath Basu , Abhik Ghosh

With the ubiquitous availability of unstructured data, growing attention is paid as how to adjust for selection bias in such non-probability samples. The majority of the robust estimators proposed by prior literature are either fully or…

统计方法学 · 统计学 2022-04-08 Ali Rafei , Michael R. Elliott , Carol A. C. Flannagan

This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represented in the data. Existing approaches require a finite, known…

机器学习 · 统计学 2025-10-16 Tianmin Xie , Yanfei Zhou , Ziyi Liang , Stefano Favaro , Matteo Sesia

There are p heterogeneous objects to be assigned to n competing agents (n > p) each with unit demand. It is required to design a Groves mechanism for this assignment problem satisfying weak budget balance, individual rationality, and…

计算机科学与博弈论 · 计算机科学 2014-01-17 Sujit Gujar , Yadati Narahari

The k-nearest-neighbor method performs classification tasks for a query sample based on the information contained in its neighborhood. Previous studies into the k-nearest-neighbor algorithm usually achieved the decision value for a class by…

机器学习 · 计算机科学 2018-12-10 Chengsheng Mao , Bin Hu , Lei Chen , Philip Moore , Xiaowei Zhang

Motivated by normalizing DNA microarray data and by predicting the interest rates, we explore nonparametric estimation of additive models with highly correlated covariates. We introduce two novel approaches for estimating the additive…

统计理论 · 数学 2010-10-05 Jiancheng Jiang , Yingying Fan , Jianqing Fan

Given a linear dynamical system, we consider the problem of constructing an approximate system using only a subset of the sensors out of the total set such that the observability Gramian of the new system is approximately equal to that of…

系统与控制 · 计算机科学 2018-11-08 Shaunak D. Bopardikar