中文
相关论文

相关论文: Semi-Supervised Mixture Models under the Concept o…

200 篇论文

Semi-supervised learning is being extensively applied to estimate classifiers from training data in which not all the labels of the feature vectors are available. We present gmmsslm, an R package for estimating the Bayes' classifier from…

统计计算 · 统计学 2024-04-18 Ziyang Lyu , Daniel Ahfock , Ryan Thompson , Geoffrey J. McLachlan

Semi-supervised learning (SSL) approaches have been successfully applied in a wide range of engineering and scientific fields. This paper investigates the generative model framework with a missingness mechanism for unclassified…

机器学习 · 统计学 2024-01-01 Ziyang Lyu

We investigate model based classification with partially labelled training data. In many biostatistical applications, labels are manually assigned by experts, who may leave some observations unlabelled due to class uncertainty. We analyse…

统计方法学 · 统计学 2019-04-08 Daniel Ahfock , Geoffrey J. McLachlan

Semi-supervised learning (SSL) constructs classifiers from datasets in which only a subset of observations is labelled, a situation that naturally arises because obtaining labels often requires expert judgement or costly manual effort. This…

统计计算 · 统计学 2025-12-09 Geoffrey J. McLachlan , Jinran Wu

Model-based unsupervised learning, as any learning task, stalls as soon as missing data occurs. This is even more true when the missing data are informative, or said missing not at random (MNAR). In this paper, we propose model-based…

Conditions ensuring optimal parameter estimation in the presence of missing data are well established in inference, typically relying on the Missing-at-Random (MAR) assumption. In prediction, similar principles are often assumed to apply.…

统计方法学 · 统计学 2026-03-19 Pierre Catoire , Robin Genuer , Cecile Proust-Lima

We consider the situation where the observed sample contains some observations whose class of origin is known (that is, they are classified with respect to the g underlying classes of interest), and where the remaining observations in the…

机器学习 · 统计学 2020-04-15 Geoffrey J. McLachlan , Daniel Ahfock

Missing data is an universal problem in statistics. We develop a unified framework for estimating parameters defined by general estimating equations under a missing-at-random (MAR) mechanism, based on generalized entropy calibration…

统计方法学 · 统计学 2026-03-31 Mst Moushumi Pervin , Hengfang Wang , Jae Kwang Kim

Missing data in supervised learning is well-studied, but the specific issue of missing labels during model evaluation has been overlooked. Ignoring samples with missing values, a common solution, can introduce bias, especially when data is…

机器学习 · 计算机科学 2025-04-28 Danial Dervovic , Michael Cashmore

This article considers a semi-supervised classification setting on a Gaussian mixture model, where the data is not labeled strictly as usual, but instead with uncertain labels. Our main aim is to compute the Bayes risk for this model. We…

机器学习 · 统计学 2024-03-28 Victor Leger , Romain Couillet

We study moment-based estimation with two sequentially collected variables subject to non-monotone missingness. The commonly used Missing at Random (MAR) assumption requiring all missingness mechanisms to depend on the same fully observed…

计量经济学 · 经济学 2026-05-29 Shenshen Yang

This paper provides further insight into the key concept of missing at random (MAR) in incomplete data analysis. Following the usual selection modelling approach we envisage two models with separable parameters: a model for the response of…

统计理论 · 数学 2007-06-13 Guobing Lu , John B. Copas

During the past few decades, missing-data problems have been studied extensively, with a focus on the ignorable missing case, where the missing probability depends only on observable quantities. By contrast, research into non-ignorable…

统计方法学 · 统计学 2019-08-06 Yukun Liu , Pengfei Li , Jing Qin

Conducting valid statistical analyses is challenging in the presence of missing-not-at-random (MNAR) data, where the missingness mechanism is dependent on the missing values themselves even conditioned on the observed data. Here, we…

统计方法学 · 统计学 2023-06-13 Anna Guo , Jiwei Zhao , Razieh Nabi

In the missing data literature, the Maximum Likelihood Estimator (MLE) is celebrated for its ignorability property under missing at random (MAR) data. However, its sensitivity to misspecification of the (complete) data model, even under…

统计方法学 · 统计学 2025-09-23 Badr-Eddine Chérief-Abdellatif , Jeffrey Näf

Semi-supervised learning (SSL) constructs classifiers using both labelled and unlabelled data. It leverages information from labelled samples, whose acquisition is often costly or labour-intensive, together with unlabelled data to enhance…

机器学习 · 统计学 2025-12-29 Jinran Wu , You-Gan Wang , Geoffrey J. McLachlan

Efficient estimation methods for simultaneous autoregressive (SAR) models with missing data in the response variable have been well-explored in the literature. A common practice is to introduce measurement error into SAR models to separate…

统计方法学 · 统计学 2024-10-10 Anjana Wijayawardhana , Thomas Suesse , David Gunawan

Most classification algorithms used in high energy physics fall under the category of supervised machine learning. Such methods require a training set containing both signal and background events and are prone to classification errors…

数据分析、统计与概率 · 物理学 2015-06-03 Mikael Kuusela , Tommi Vatanen , Eric Malmi , Tapani Raiko , Timo Aaltonen , Yoshikazu Nagai

In modern large-scale observational studies, data collection constraints often result in partially labeled datasets, posing challenges for reliable causal inference, especially due to potential labeling bias and relatively small size of the…

统计方法学 · 统计学 2025-04-22 Yuqian Zhang , Abhishek Chakrabortty , Jelena Bradic

A learning algorithm referred to as Maximum Margin (MM) is proposed for considering the class-imbalance data learning issue: the trained model tends to predict the majority of classes rather than the minority ones. That is, underfitting for…

机器学习 · 计算机科学 2023-03-30 Haeyong Kang , Thang Vu , Chang D. Yoo
‹ 上一页 1 2 3 10 下一页 ›