English
Related papers

Related papers: Pseudo-value regression of clustered multistate cu…

200 papers

Producing labels for unlabeled data is error-prone, making semi-supervised learning (SSL) troublesome. Often, little is known about when and why an algorithm fails to outperform a supervised baseline. Using benchmark datasets, we craft five…

Understanding treatment effect heterogeneity is vital for scientific and policy research. However, identifying and evaluating heterogeneous treatment effects pose significant challenges due to the typically unknown subgroup structure.…

Methodology · Statistics 2024-11-05 Kwangho Kim , Jisu Kim , Larry A. Wasserman , Edward H. Kennedy

Current Reinforcement Learning (RL) methods often suffer from sample-inefficiency, resulting from blind exploration strategies that neglect causal relationships among states, actions, and rewards. Although recent causal approaches aim to…

Artificial Intelligence · Computer Science 2025-02-17 Hongye Cao , Fan Feng , Tianpei Yang , Jing Huo , Yang Gao

In meta analysis, multiple hypothesis testing and many other methods, p-values are utilized as inputs and assumed to be uniformly distributed over the unit interval under the null hypotheses. If data used to generate p-values have discrete…

Methodology · Statistics 2026-02-24 Joshua Habiger , Pratyaydipta Rudra

Pseudo-labeling is a popular semi-supervised learning technique to leverage unlabeled data when labeled samples are scarce. The generation and selection of pseudo-labels heavily rely on labeled data. Existing approaches implicitly assume…

Machine Learning · Computer Science 2024-06-21 Nabeel Seedat , Nicolas Huynh , Fergus Imrie , Mihaela van der Schaar

We propose a new approach for sparse regression and marginal testing, for data with correlated features. Our procedure first clusters the features, and then chooses as the cluster prototype the most informative feature in that cluster. Then…

Methodology · Statistics 2015-03-16 Stephen Reid , Robert Tibshirani

Current status censoring or case I interval censoring takes place when subjects in a study are observed just once to check if a particular event has occurred. If the event is recurring, the data are classified as current count data; if…

Methodology · Statistics 2024-10-15 Pavithra Hariharan , P. G. Sankaran

Semi-supervised semantic segmentation (SSS) has recently gained increasing research interest as it can reduce the requirement for large-scale fully-annotated training data. The current methods often suffer from the confirmation bias from…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Zicheng Wang , Zhen Zhao , Xiaoxia Xing , Dong Xu , Xiangyu Kong , Luping Zhou

The COVID-19 pandemic has presented unprecedented challenges worldwide, with its impact varying significantly across different geographic and socioeconomic contexts. This study employs a clustering analysis to examine the diversity of…

Computational Engineering, Finance, and Science · Computer Science 2024-08-06 Morteza Maleki

Unsupervised learning has gained prominence in the big data era, offering a means to extract valuable insights from unlabeled datasets. Deep clustering has emerged as an important unsupervised category, aiming to exploit the non-linear…

Machine Learning · Computer Science 2024-02-02 Georgios Vardakas , Ioannis Papakostas , Aristidis Likas

We propose a novel semi-supervised learning (SSL) method that adopts selective training with pseudo labels. In our method, we generate hard pseudo-labels and also estimate their confidence, which represents how likely each pseudo-label is…

Machine Learning · Computer Science 2021-03-16 Masato Ishii

Local differential privacy (LDP) has been deemed as the de facto measure for privacy-preserving distributed data collection and analysis. Recently, researchers have extended LDP to the basic data type in NoSQL systems: the key-value data,…

Cryptography and Security · Computer Science 2019-07-12 Lin Sun , Jun Zhao , Xiaojun Ye , Shuo Feng , Teng Wang , Tao Bai

The restricted mean survival time (RMST) model has been garnering attention as a way to provide a clinically intuitive measure: the mean survival time. RMST models, which use methods based on pseudo time-to-event values and inverse…

Methodology · Statistics 2024-06-11 Keisuke Hanada , Masahiro Kojima

This study introduces a general semiparametric clusterwise index distribution model to analyze how latent clusters affect the covariate-response relationships. By employing sufficient dimension reduction to account for the effects of…

Methodology · Statistics 2025-09-30 Jen-Chieh Teng , Chin-Tsang Chiang

In many public health problems, an important goal is to identify the effect of some treatment/intervention on the risk of failure for the whole population. A marginal proportional hazards regression model is often used to analyze such an…

Statistics Theory · Mathematics 2007-06-13 Donglin Zeng

The restricted mean survival time is a clinically easy-to-interpret measure that does not require any assumption of proportional hazards. We focus on two ways to directly model the survival time and adjust the covariates. One is to…

Methodology · Statistics 2022-11-03 Keisuke Hanada , Junji Moriya , Masahiro Kojima

Conformal prediction is a popular technique for constructing prediction intervals with distribution-free coverage guarantees. The coverage is marginal, meaning it only holds on average over the entire population but not necessarily for any…

Methodology · Statistics 2026-05-28 Yao Zhang , Emmanuel J. Candès

Current status data abounds in the field of epidemiology and public health, where the only observable data for a subject is the random inspection time and the event status at inspection. Motivated by such a current status data from a…

Methodology · Statistics 2020-04-24 Tong Wang , Kejun He , Wei Ma , Dipankar Bandyopadhyay , Samiran Sinha

Sufficient dimension reduction (SDR) is a popular class of regression methods which aim to find a small number of linear combinations of covariates that capture all the information of the responses i.e., a central subspace. The majority of…

Methodology · Statistics 2024-10-15 Linh H. Nghiem , F. K. C. Hui

Given the potential difficulties in obtaining large quantities of labelled data, many works have explored the use of deep semi-supervised learning, which uses both labelled and unlabelled data to train a neural network architecture. The…

Machine Learning · Computer Science 2021-09-02 Philip Sellars , Angelica Aviles-Rivero , Carola Bibiane Schönlieb