English
Related papers

Related papers: Pseudo-value regression of clustered multistate cu…

200 papers

Counterfactual reasoning is an important paradigm applicable in many fields, such as healthcare, economics, and education. In this work, we propose a novel method to address the issue of \textit{selection bias}. We learn two groups of…

Machine Learning · Computer Science 2019-12-20 Zichen Zhang , Qingfeng Lan , Lei Ding , Yue Wang , Negar Hassanpour , Russell Greiner

This report is designed to clarify a few points about the article "Semiparametric modeling of grouped current duration data with preferential reporting" by McLain, Sundaram, Thoma and Louis in Statistics in Medicine (McLain et al., 2014,…

Applications · Statistics 2018-01-03 Alexander C. McLain , Rajeshwari Sundaram , Marie Thoma , Germaine M. Buck Louis

Estimation of temporal counterfactual outcomes from observed history is crucial for decision-making in many domains such as healthcare and e-commerce, particularly when randomized controlled trials (RCTs) suffer from high cost or…

Machine Learning · Computer Science 2024-02-13 Chuizheng Meng , Yihe Dong , Sercan Ö. Arık , Yan Liu , Tomas Pfister

Entrepreneurial regimes are topic, receiving ever more research attention. Existing studies on entrepreneurial regimes mainly use common methods from multivariate analysis and some type of institutional related analysis. In our analysis,…

Methodology · Statistics 2020-07-28 Andrej Srakar , Marilena Vecco

The case-cohort design is a commonly used cost-effective sampling strategy for large cohort studies, where some covariates are expensive to measure or obtain. In this paper, we consider regression analysis under a case-cohort study with…

Methodology · Statistics 2023-10-24 Qingning Zhou , Kin Yau Wong

We propose a new approach for scaling prior to cluster analysis based on the concept of pooled variance. Unlike available scaling procedures such as the standard deviation and the range, our proposed scale avoids dampening the beneficial…

Methodology · Statistics 2020-07-28 Jakob Raymaekers , Ruben H. Zamar

Comparative meta-analyses of groups of subjects by integrating multiple observational studies rely on estimated propensity scores (PSs) to mitigate covariate imbalances. However, PS estimation grapples with the theoretical and practical…

Methodology · Statistics 2024-05-09 Subharup Guha , Yi Li

Through solving pretext tasks, self-supervised learning (SSL) leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream task. A common pretext task consists in pretraining a SSL…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-14 Salah Zaiem , Titouan Parcollet , Slim Essid

Selecting the top-$m$ variables with the $m$ largest population parameters from a larger set of candidates is a fundamental problem in statistics. In this paper, we propose a novel methodology called Sequential Correct Screening (SCS),…

Methodology · Statistics 2025-08-21 Masaki Toyoda , Yoshimasa Uematsu

Text clustering methods were traditionally incorporated into multi-document summarization (MDS) as a means for coping with considerable information repetition. Particularly, clusters were leveraged to indicate information saliency as well…

Computation and Language · Computer Science 2022-05-23 Ori Ernst , Avi Caciularu , Ori Shapira , Ramakanth Pasunuru , Mohit Bansal , Jacob Goldberger , Ido Dagan

This article investigates factor-augmented sparse MIDAS (Mixed Data Sampling) regressions for high-dimensional time series data, which may be observed at different frequencies. Our novel approach integrates sparse and dense dimensionality…

Econometrics · Economics 2025-10-17 Jad Beyhum , Jonas Striaukas

Local variable selection aims to test for the effect of covariates on an outcome within specific regions. We outline a challenge that arises in the presence of non-linear effects and model misspecification. Specifically, for common…

Methodology · Statistics 2024-08-02 David Rossell , Arnold Kisuk Kseung , Ignacio Saez , Michele Guindani

Purpose: The primary goal of this study is to explore the application of evaluation metrics to different clustering algorithms using the data provided from the Canadian Longitudinal Study (CLSA), focusing on cognitive features. The…

Machine Learning · Computer Science 2025-05-19 ChenNingZhi Sheng

Deep learning is pushing the state-of-the-art in many computer vision applications. However, it relies on large annotated data repositories, and capturing the unconstrained nature of the real-world data is yet to be solved. Semi-supervised…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Mamshad Nayeem Rizve , Navid Kardan , Mubarak Shah

Computational capability often falls short when confronted with massive data, posing a common challenge in establishing a statistical model or statistical inference method dealing with big data. While subsampling techniques have been…

Methodology · Statistics 2024-10-31 Yixiao Ruan , Zan Li , Zhaohui Li , Dennis K. J. Lin , Qingpei Hu , Dan Yu

We propose a novel and computationally efficient approach for nonparametric conditional density estimation in high-dimensional settings that achieves dimension reduction without imposing restrictive distributional or functional form…

Econometrics · Economics 2025-10-14 Jianhua Mei , Fu Ouyang , Thomas T. Yang

Interval-censored multi state data is collected when the state of a subject is observed periodically. The analysis of such data using non-parametric multi-state models was not possible until recently, but is very desirable as it allows for…

Methodology · Statistics 2025-07-16 Daniel Gomon , Hein Putter

Semi-supervised multi-label learning (SSMLL) aims to address the challenge of limited labeled data in multi-label learning (MLL) by leveraging unlabeled data to improve the model's performance. While pseudo-labeling has become a dominant…

Machine Learning · Computer Science 2025-12-03 Bo Han , Zhuoming Li , Xiaoyu Wang , Yaxin Hou , Hui Liu , Junhui Hou , Yuheng Jia

Statistical power is often a concern for clustered RCTs due to variance inflation from design effects and the high cost of adding study clusters (such as hospitals, schools, or communities). While covariate pre-specification is the…

Methodology · Statistics 2020-05-07 Peter Z. Schochet

Estimating the causal effects of an intervention in the presence of confounding is a frequently occurring problem in applications such as medicine. The task is challenging since there may be multiple confounding factors, some of which may…

Methodology · Statistics 2018-11-28 Sonali Parbhoo , Mario Wieser , Volker Roth