中文
相关论文

相关论文: IlocA: An algorithm to Cluster Cells and form Impu…

200 篇论文

Multiple imputation (MI) is a popular method for dealing with missing values. However, the suitable way for applying clustering after MI remains unclear: how to pool partitions? How to assess the clustering instability when data are…

统计方法学 · 统计学 2022-05-16 Vincent Audigier , Ndèye Niang

Public policy-makers use cost-effectiveness analyses (CEA) to decide which health and social care interventions to provide. Appropriate methods have not been developed for handling missing data in complex settings, such as for CEA that use…

统计方法学 · 统计学 2012-06-27 Karla Diaz-Ordaz , Michael G. Kenward , Richard Grieve

Multiple imputation (MI) is an established technique to handle missing data in observational studies. Joint modeling (JM) and fully conditional specification (FCS) are commonly used methods for imputing multilevel clustered data. However,…

统计方法学 · 统计学 2022-09-28 Mei Dong , Aya Mitani

Many fields, such as neuroscience, are experiencing the vast proliferation of cellular data, underscoring the need for organizing and interpreting large datasets. A popular approach partitions data into manageable subsets via hierarchical…

定量方法 · 定量生物学 2024-03-07 Diek W. Wheeler , Giorgio A. Ascoli

Modern inference and learning often hinge on identifying low-dimensional structures that approximate large scale data. Subspace clustering achieves this through a union of linear subspaces. However, in contemporary applications data is…

机器学习 · 计算机科学 2018-08-03 Daniel L. Pimentel-Alarcón , Usman Mahmood

Copula-based methods provide a flexible approach to build missing data imputation models of multivariate data of mixed types. However, the choice of copula function is an open question. We consider a Bayesian nonparametric approach by using…

统计方法学 · 统计学 2019-10-15 Jiali Wang , Anton Westveld , Bronwyn Loong , Alan Welsh

The explosion in the amount of data available for analysis often necessitates a transition from batch to incremental clustering methods, which process one element at a time and typically store only a small subset of the data. In this paper,…

机器学习 · 计算机科学 2014-06-26 Margareta Ackerman , Sanjoy Dasgupta

Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population…

统计方法学 · 统计学 2015-11-04 Hélène Chaput , Guillaume Chauvet , David Haziza , Laurianne Salembier , Julie Solard

Missing data imputation forms the first critical step of many data analysis pipelines. The challenge is greatest for mixed data sets, including real, Boolean, and ordinal data, where standard techniques for imputation fail basic sanity…

统计方法学 · 统计学 2020-06-17 Yuxuan Zhao , Madeleine Udell

Clustering is a standard approach for achieving efficient and scalable performance in wireless sensor networks. Traditionally, clustering algorithms aim at generating a number of disjoint clusters that satisfy some criteria. In this paper,…

网络与互联网体系结构 · 计算机科学 2013-04-09 Moustafa Youssef , Adel Youssef , Mohamed Younis

Multivariate data are typically represented by a rectangular matrix (table) in which the rows are the objects (cases) and the columns are the variables (measurements). When there are many variables one often reduces the dimension by…

统计方法学 · 统计学 2021-01-13 Mia Hubert , Peter J. Rousseeuw , Wannes Van den Bossche

We develop a model in which interactions between nodes of a dynamic network are counted by non homogeneous Poisson processes. In a block modelling perspective, nodes belong to hidden clusters (whose number is unknown) and the intensity…

机器学习 · 统计学 2017-07-11 Marco Corneli , Pierre Latouche , Fabrice Rossi

Multiple imputation has become one of the standard methods in drawing inferences in many incomplete data applications. Applications of multiple imputation in relatively more complex settings, such as high-dimensional clustered data, require…

统计方法学 · 统计学 2025-04-08 Qiushuang Li , Recai Yucel

A first step when fitting multilevel models to continuous responses is to explore the degree of clustering in the data. Researchers fit variance-component models and then report the proportion of variation in the response that is due to…

统计方法学 · 统计学 2020-02-17 George Leckie , William Browne , Harvey Goldstein , Juan Merlo , Peter Austin

Many real-world datasets contain missing entries and mixed data types including categorical and ordered (e.g. continuous and ordinal) variables. Imputing the missing entries is necessary, since many data analysis pipelines require complete…

统计方法学 · 统计学 2022-10-14 Yuxuan Zhao , Alex Townsend , Madeleine Udell

Clinical decision support using data mining techniques offers more intelligent way to reduce the decision error in the last few years. However, clinical datasets often suffer from high missingness, which adversely impacts the quality of…

机器学习 · 计算机科学 2020-11-20 Xuetong Wu , Hadi Akbarzadeh Khorshidi , Uwe Aickelin , Zobaida Edib , Michelle Peate

Missing values frequently arise in modern biomedical studies due to various reasons, including missing tests or complex profiling technologies for different omics measurements. Missing values can complicate the application of clustering…

机器学习 · 统计学 2019-02-27 Shahin Boluki , Siamak Zamani Dadaneh , Xiaoning Qian , Edward R. Dougherty

Pairwise clustering, in general, partitions a set of items via a known similarity function. In our treatment, clustering is modeled as a transductive prediction problem. Thus rather than beginning with a known similarity function, the…

机器学习 · 计算机科学 2017-06-21 Stephen Pasteris , Fabio Vitale , Claudio Gentile , Mark Herbster

Missing data has a ubiquitous presence in real-life applications of machine learning techniques. Imputation methods are algorithms conceived for restoring missing values in the data, based on other entries in the database. The choice of the…

机器学习 · 计算机科学 2017-08-16 Unai Garciarena , Roberto Santana , Alexander Mendiburu

We consider the problem of decentralized clustering and estimation over multi-task networks, where agents infer and track different models of interest. The agents do not know beforehand which model is generating their own data. They also do…

最优化与控制 · 数学 2017-05-24 Sahar Khawatmi , Ali H. Sayed , Abdelhak M. Zoubir