中文
相关论文

相关论文: Missing data imputation using a truncated Gaussian…

200 篇论文

Imputation of missing data is a common application in various classification problems where the feature training matrix has missingness. A widely used solution to this imputation problem is based on the lazy learning technique, $k$-nearest…

机器学习 · 统计学 2020-02-26 Arkopal Choudhury , Michael R. Kosorok

We study semiparametric factor models in high-dimensional panels where the factor loadings consist of a nonparametric component explained by observed covariates and an idiosyncratic component capturing unobserved heterogeneity. A key…

统计方法学 · 统计学 2025-12-09 Sijie Zheng

Multivariate time-series data are used in many classification and regression predictive tasks, and recurrent models have been widely used for such tasks. Most common recurrent models assume that time-series data elements are of equal length…

机器学习 · 计算机科学 2020-09-21 Mehak Gupta , Rahmatollah Beheshti

Missing values pose a persistent challenge in modern data science. Consequently, there is an ever-growing number of publications introducing new imputation methods in various fields. The present paper attempts to take a step back and…

统计理论 · 数学 2026-01-21 Jeffrey Näf , Erwan Scornet , Julie Josse

We consider the problem of multivariate density estimation when the unknown density is assumed to follow a particular form of dimensionality reduction, a noisy independent factor analysis (IFA) model. In this model the data are generated by…

应用统计 · 统计学 2009-06-17 Umberto Amato , Anestis Antoniadis , Alexander Samarov , Alexander Tsybakov

This paper presents a new modeling strategy for joint unsupervised analysis of multiple high-throughput biological studies. As in Multi-study Factor Analysis, our goals are to identify both common factors shared across studies and…

应用统计 · 统计学 2018-06-27 Roberta De Vito , Ruggero Bellio , Lorenzo Trippa , Giovanni Parmigiani

Recent work on overfitting Bayesian mixtures of distributions offers a powerful framework for clustering multivariate data using a latent Gaussian model which resembles the factor analysis model. The flexibility provided by overfitting…

统计方法学 · 统计学 2019-08-29 Panagiotis Papastamoulis

This paper proposes an imputation procedure that uses the factors estimated from a tall block along with the re-rotated loadings estimated from a wide block to impute missing values in a panel of data. Assuming that a strong factor…

计量经济学 · 经济学 2021-08-13 Jushan Bai , Serena Ng

Metabolic flux balance analyses are a standard tool in analysing metabolic reaction rates compatible with measurements, steady-state and the metabolic reaction network stoichiometry. Flux analysis methods commonly place unrealistic…

We develop Probabilistic Targeted Factor Analysis (PTFA), a likelihood-based framework for constructing latent factors that are explicitly targeted to variables of economic interest. PTFA provides a probabilistic foundation for Partial…

计量经济学 · 经济学 2026-01-12 Miguel C. Herculano , Santiago Montoya-Blandón

We study mean estimation for a Gaussian distribution with identity covariance in $\mathbb{R}^d$ under a missing data scheme termed realizable $\epsilon$-contamination model. In this model an adversary can choose a function $r(x)$ between 0…

机器学习 · 计算机科学 2026-03-18 Ilias Diakonikolas , Daniel M. Kane , Thanasis Pittas

We propose an adaption of the multiple imputation random lasso procedure tailored to longitudinal data with unobserved fixed effects which provides robust variable selection in the presence of complex missingness, high dimensionality and…

应用统计 · 统计学 2024-12-04 Lotta Rüter , Melanie Schienle

Background. Emerging technologies now allow for mass spectrometry based profiling of up to thousands of small molecule metabolites (metabolomics) in an increasing number of biosamples. While offering great promise for revealing insight into…

A nonparametric Bayesian extension of Factor Analysis (FA) is proposed where observed data $\mathbf{Y}$ is modeled as a linear superposition, $\mathbf{G}$, of a potentially infinite number of hidden factors, $\mathbf{X}$. The Indian Buffet…

应用统计 · 统计学 2011-07-29 David Knowles , Zoubin Ghahramani

Missing data imputation, where a model is trained on observed data to estimate unobserved values, is a fundamental problem in machine learning. In this paper, we rigorously formulate imputation model learning as a mean-squared error risk…

机器学习 · 统计学 2026-05-14 Luke Shannon , Song Liu , Katarzyna Reluga

Motivation: Untargeted metabolomics comprehensively characterizes small molecules and elucidates activities of biochemical pathways within a biological sample. Despite computational advances, interpreting collected measurements and…

定量方法 · 定量生物学 2020-03-10 Ramtin Hosseini , Neda Hassanpour , Li-Ping Liu , Soha Hassoun

Missing data imputation is an important research topic in data mining. Large-scale Molecular descriptor data may contains missing values (MVs). However, some methods for downstream analyses, including some prediction tools, require a…

计算工程、金融与科学 · 计算机科学 2013-12-13 Doreswamy , Chanabasayya . M. Vastrad

Spurious measurements frequently occur in surface data from technical components. Excluding or ignoring these spurious points may lead to incorrect surface characterization if these points inherit features of the surface. Therefore, data…

统计方法学 · 统计学 2025-04-10 Arsalan Jawaid , Samuel Schmidt , Marvin Lotz , Jörg Seewig

A conventional linear model for functional data involves expressing a response variable $Y$ in terms of the explanatory function $X(t)$, via the model: $Y=a+\int_I b(t)X(t)dt+\hbox{error}$, where $a$ is a scalar, $b$ is an unknown function…

统计方法学 · 统计学 2014-07-01 Peter Hall , Giles Hooker

We propose graph-based predictable feature analysis (GPFA), a new method for unsupervised learning of predictable features from high-dimensional time series, where high predictability is understood very generically as low variance in the…

机器学习 · 计算机科学 2017-05-12 Björn Weghenkel , Asja Fischer , Laurenz Wiskott