中文
相关论文

相关论文: Adjusting for informative cluster size in pseudo-v…

200 篇论文

Collaboration between different data centers is often challenged by heterogeneity across sites. To account for the heterogeneity, the state-of-the-art method is to re-weight the covariate distributions in each site to match the distribution…

机器学习 · 统计学 2024-04-25 Tianyu Guo , Sai Praneeth Karimireddy , Michael I. Jordan

Propensity score methods are increasingly being used to reduce estimation bias of treatment effects for observational studies. Previous research has shown that propensity score methods consistently estimate the marginal hazard ratio for…

统计方法学 · 统计学 2019-11-19 Haodi Liang , Cecilia Cotton

Survival analysis is essential for clinical decision-making, as it allows practitioners to estimate time-to-event outcomes, stratify patient risk profiles, and guide treatment planning. Deep learning has revolutionized this field with…

机器学习 · 计算机科学 2026-02-03 Pinar Erbil , Alberto Archetti , Eugenio Lomurno , Matteo Matteucci

Extraordinary amounts of data are being produced in many branches of science. Proven statistical methods are no longer applicable with extraordinary large data sets due to computational limitations. A critical step in big data analysis is…

统计方法学 · 统计学 2019-06-27 HaiYing Wang , Min Yang , John Stufken

Unsupervised person re-identification (re-ID) has attracted increasing research interests because of its scalability and possibility for real-world applications. State-of-the-art unsupervised re-ID methods usually follow a clustering-based…

计算机视觉与模式识别 · 计算机科学 2022-07-08 Tianyang Liu , Yutian Lin , Bo Du

Sample size determination for cluster randomised trials (CRTs) is challenging as it requires robust estimation of the intra-cluster correlation coefficient (ICC). Typically, the sample size is chosen to provide a certain level of power to…

应用统计 · 统计学 2023-08-23 S. Faye Williamson , Svetlana V. Tishkovskaya , Kevin J. Wilson

We address the problem of validating the ouput of clustering algorithms. Given data $\mathcal{D}$ and a partition $\mathcal{C}$ of these data into $K$ clusters, when can we say that the clusters obtained are correct or meaningful for the…

机器学习 · 统计学 2023-02-02 Marina Meilă , Hanyu Zhang

Stepped wedge designs (SWDs) are increasingly used to evaluate longitudinal cluster-level interventions but pose substantial challenges for valid inference. Because crossover times are randomized, intervention effects are intrinsically…

统计方法学 · 统计学 2026-05-12 Fan Xia , K. C. Gary Chan , Emily Voldal , Avi Kenny , Patrick J. Heagerty , James P. Hughes

The association between multidimensional exposure patterns and outcomes is commonly investigated by first applying cluster analysis algorithms to derive patterns and then estimating the associations. However, errors in the underlying…

统计方法学 · 统计学 2023-05-31 Timm Intemann , Iris Pigeot

Survival regression is widely used to model time-to-events data, to explore how covariates may influence the occurrence of events. Modern datasets often encompass a vast number of covariates across many subjects, with only a subset of the…

统计方法学 · 统计学 2024-09-18 Abhishek Mandal , Abhisek Chakraborty

In many applications, data cluster. Failing to take the cluster structure into consideration generally leads to underestimated variances of point estimators and inflated type I errors in hypothesis tests. Many circumstance-dependent…

统计方法学 · 统计学 2025-07-21 Jiahua Chen , Pengfei Li , Yukun Liu , James V. Zidek

We propose autoregressive Bayesian semi-parametric models for waiting times between recurrent events. The aim is two-fold: inference on the effect of possibly time-varying covariates on the gap times and clustering of individuals based on…

应用统计 · 统计学 2016-07-28 Marta Tallarita , Maria De Iorio , Alessandra Guglielmi , James Malone-Lee

The reliance of text classifiers on spurious correlations can lead to poor generalization at deployment, raising concerns about their use in safety-critical domains such as healthcare. In this work, we propose to use counterfactual data…

机器学习 · 计算机科学 2024-01-10 Amir Feder , Yoav Wald , Claudia Shi , Suchi Saria , David Blei

Nonprobability samples have rapidly emerged to address time-sensitive priority topics in a variety of fields. While these data are timely, they are prone to selection bias. To mitigate selection bias, a large number of survey research…

统计方法学 · 统计学 2025-08-08 Kangrui Liu , Lingxiao Wang , Yan Li

Analyses of cluster randomized trials (CRTs) can be complicated by informative missing outcome data. Methods such as inverse probability weighted generalized estimating equations have been proposed to account for informative missingness by…

统计方法学 · 统计学 2023-04-13 Chia-Rui Chang , Rui Wang

Information-maximization clustering learns a probabilistic classifier in an unsupervised manner so that mutual information between feature vectors and cluster assignments is maximized. A notable advantage of this approach is that it only…

机器学习 · 统计学 2011-12-06 Masashi Sugiyama , Makoto Yamada , Manabu Kimura , Hirotaka Hachiya

Background: Subgroup analyses are frequently conducted in randomized clinical trials to assess evidence of heterogeneous treatment effect across patient subpopulations. Although randomization balances covariates within subgroups in…

统计方法学 · 统计学 2021-05-27 Siyun Yang , Fan Li , Laine E. Thomas , Fan Li

Fine-tuning large pretrained language models is a common approach for aligning them with human preferences, but noisy or off-target examples can dilute supervision. While small, well-chosen datasets often match the performance of much…

机器学习 · 计算机科学 2026-01-28 Ling Zhang , Xianliang Yang , Juwon Yu , Park Cheonyoung , Miran Lee , Lei Song , Jiang Bian

With continuous outcomes, the average causal effect is typically defined using a contrast of expected potential outcomes. However, in the presence of skewed outcome data, the expectation may no longer be meaningful. In practice the typical…

统计方法学 · 统计学 2023-02-06 Daisy A. Shepherd , Benjamin R. Baer , Margarita Moreno-Betancur

Clinical study populations often differ meaningfully from the broader populations to which results are intended to generalize. Weighting methods such as inverse probability of sampling weights (IPSW) reweight study participants to resemble…

统计方法学 · 统计学 2025-12-02 William Stewart , Carly L. Brantner , Elizabeth A. Stuart , Laine Thomas