中文
相关论文

相关论文: Balanced $k$-nearest neighbor imputation

200 篇论文

Researchers are often interested in predicting outcomes, conducting clustering analysis to detect distinct subgroups of their data, or computing causal treatment effects. Pathological data distributions that exhibit skewness and…

统计方法学 · 统计学 2020-08-24 Arman Oganisian , Nandita Mitra , Jason Roy

Distributed resource allocation is a central task in network systems such as smart grids, water distribution networks, and urban transportation systems. When solving such problems in practice it is often important to have nonasymptotic…

最优化与控制 · 数学 2021-03-30 Xuyang Wu , Sindri Magnusson , Mikael Johansson

We investigate the problem of jointly testing two hypotheses and estimating a random parameter based on data that is observed sequentially by sensors in a distributed network. In particular, we assume the data to be drawn from a Gaussian…

信号处理 · 电气工程与系统科学 2020-03-04 Dominik Reinhard , Michael Fauß , Abdelhak M. Zoubir

The task of reconstructing a matrix given a sample of observedentries is known as the matrix completion problem. It arises ina wide range of problems, including recommender systems, collaborativefiltering, dimensionality reduction, image…

统计理论 · 数学 2014-12-20 Jean Lafond , Olga Klopp , Eric Moulines , Jospeh Salmon

A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty. Minimizing overall prediction error often encourages models to prioritize informativeness over calibration,…

机器学习 · 统计学 2026-02-17 Ádám Jung , Domokos M. Kelen , András A. Benczúr

Consider a setting with multiple units (e.g., individuals, cohorts, geographic locations) and outcomes (e.g., treatments, times, items), where the goal is to learn a multivariate distribution for each unit-outcome entry, such as the…

机器学习 · 统计学 2025-10-21 Kyuseong Choi , Jacob Feitelberg , Caleb Chin , Anish Agarwal , Raaz Dwivedi

Randomized controlled trials are susceptible to imbalance on covariates predictive of the outcome. Rerandomization and deterministic treatment assignment are two proposed solutions. This paper explores the relationship between…

统计方法学 · 统计学 2023-10-03 Connor T. Jerzak , Rebecca Goldstein

Imputation of missing attribute values in medical datasets for extracting hidden knowledge from medical datasets is an interesting research topic of interest which is very challenging. One cannot eliminate missing values in medical records.…

数据库 · 计算机科学 2016-03-11 Yelipe UshaRani , P. Sammulal

Linear regression is a fundamental and popular statistical method. There are various kinds of linear regression, such as mean regression and quantile regression. In this paper, we propose a new one called distribution regression, which…

统计方法学 · 统计学 2017-12-27 Xin Chen , Xuejun Ma , Wang Zhou

This article discusses the problem of estimation of parameters in finite mixtures when the mixture components are assumed to be symmetric and to come from the same location family. We refer to these mixtures as semi-parametric because no…

统计理论 · 数学 2007-08-07 David R. Hunter , Shaoli Wang , Thomas P. Hettmansperger

Entropy estimation is of practical importance in information theory and statistical science. Many existing entropy estimators suffer from fast growing estimation bias with respect to dimensionality, rendering them unsuitable for…

信息论 · 计算机科学 2023-08-22 Ziqiao Ao , Jinglai Li

This paper addresses distributed parameter estimation in randomized one-hidden-layer neural networks. A group of agents sequentially receive measurements of an unknown parameter that is only partially observable to them. In this paper, we…

系统与控制 · 电气工程与系统科学 2020-03-23 Yinsong Wang , Shahin Shahrampour

This work unifies the analysis of various randomized methods for solving linear and nonlinear inverse problems by framing the problem in a stochastic optimization setting. By doing so, we show that many randomized methods are variants of a…

数值分析 · 数学 2023-06-21 Jonathan Wittmer , C. G. Krishnanunni , Hai V. Nguyen , Tan Bui-Thanh

A class of simultaneous equation models arise in the many domains where observed binary outcomes are themselves a consequence of the existing choices of of one of the agents in the model. These models are gaining increasing interest in the…

计量经济学 · 经济学 2025-12-30 Shakeeb Khan , Elie Tamer , Qingsong Yao

Many statistical methodologies for high-dimensional data assume the population is normal. Although a few multivariate normality tests have been proposed, to the best of our knowledge, none of them can properly control the type I error when…

统计方法学 · 统计学 2021-05-04 Hao Chen , Yin Xia

It is often of interest to assess whether a function-valued statistical parameter, such as a density function or a mean regression function, is equal to any function in a class of candidate null parameters. This can be framed as a…

统计方法学 · 统计学 2023-06-14 Aaron Hudson

We present single imputation method for missing values which borrows the idea of data depth---a measure of centrality defined for an arbitrary point of a space with respect to a probability distribution or data cloud. This consists in…

统计方法学 · 统计学 2018-08-08 Pavlo Mozharovskyi , Julie Josse , Francois Husson

Deconvolution is a statistical inverse problem to estimate the distribution of a random variable based on its noisy observations. Despite the extensive studies on the topic, deconvolution with unknown noise distribution remains as a…

统计理论 · 数学 2020-04-06 Devavrat Shah , Dogyoon Song

In some previous works, two of the authors introduced a technique to design high-order numerical methods for one-dimensional balance laws that preserve all their stationary solutions. The basis of these methods is a well-balanced…

We investigate the issue of parameter estimation with nonuniform negative sampling for imbalanced data. We first prove that, with imbalanced data, the available information about unknown parameters is only tied to the relatively small…

机器学习 · 统计学 2021-10-26 HaiYing Wang , Aonan Zhang , Chong Wang