中文
相关论文

相关论文: Variable Selection for Stratified Sampling Designs…

200 篇论文

This article introduces a subbagging (subsample aggregating) approach for variable selection in regression within the context of big data. The proposed subbagging approach not only ensures that variable selection is scalable given the…

统计方法学 · 统计学 2025-03-10 Xian Li , Xuan Liang , Tao Zou

Solving systems of Boolean equations is a fundamental task in symbolic computation and algebraic cryptanalysis, with wide-ranging applications in cryptography, coding theory, and formal verification. Among existing approaches, the Boolean…

密码学与安全 · 计算机科学 2026-04-21 Minzhong Luo , Yudong Sun , Yin Long

Fine stratification survey is useful in many applications as its point estimator is unbiased, but the variance estimator under the design cannot be easily obtained, particularly when the sample size per stratum is as small as one unit. One…

统计方法学 · 统计学 2026-03-05 Sepideh Mosaferi , Shonosuke Sugasawa

Stability and reproducibility are essential considerations in various applications of statistical methods. False Discovery Rate (FDR) control methods are able to control false signals in scientific discoveries. However, many FDR control…

统计方法学 · 统计学 2025-12-22 Jiajun Sun , Zhanrui Cai , Wei Zhong

Structured Latent Attribute Models (SLAMs) are a family of discrete latent variable models widely used in education, psychology, and epidemiology to model multivariate categorical data. A SLAM assumes that multiple discrete latent…

统计方法学 · 统计学 2021-07-12 Yuqi Gu , Gongjun Xu

The log-logistic regression model is one of the most commonly used accelerated failure time (AFT) models in survival analysis, for which statistical inference methods are mainly established under the frequentist framework. Recently,…

统计方法学 · 统计学 2023-10-11 Chengqian Xian , Camila P. E. de Souza , Wenqing He , Felipe F. Rodrigues , Renfang Tian

Bayesian inference for survival regression modeling offers numerous advantages, especially for decision-making and external data borrowing, but demands the specification of the baseline hazard function, which may be a challenging task. We…

We consider supervised learning problems where the features are embedded in a graph, such as gene expressions in a gene network. In this context, it is of much interest to automatically select a subgraph with few connected components; by…

机器学习 · 统计学 2013-09-20 Julien Mairal , Bin Yu

Selection bias arises when the probability that an observation enters a dataset depends on variables related to the quantities of interest, leading to systematic distortions in estimation and uncertainty quantification. For example, in…

We study the problem of efficiently estimating counts for queries involving complex filters, such as user-defined functions, or predicates involving self-joins and correlated subqueries. For such queries, traditional sampling techniques may…

数据库 · 计算机科学 2020-01-01 Brett Walenz , Stavros Sintos , Sudeepa Roy , Jun Yang

In this study, we introduce a novel methodological framework called Bayesian Penalized Empirical Likelihood (BPEL), designed to address the computational challenges inherent in empirical likelihood (EL) approaches. Our approach has two…

统计方法学 · 统计学 2025-03-04 Jinyuan Chang , Cheng Yong Tang , Yuanzheng Zhu

The identification of patient subgroups with comparable event-risk dynamics plays a key role in supporting informed decision-making in clinical research. In such settings, it is important to account for the inherent dependence that arises…

统计计算 · 统计学 2026-01-13 Alessandra Ragni , Lara Cavinato , Francesca Ieva

In traditional federated learning, a single global model cannot perform equally well for all clients. Therefore, the need to achieve the client-level fairness in federated system has been emphasized, which can be realized by modifying the…

机器学习 · 计算机科学 2025-10-09 Seok-Ju Hahn , Gi-Soo Kim , Junghye Lee

Clustered binary data with a large number of covariates have become increasingly common in many scientific disciplines. This paper develops an asymptotic theory for generalized estimating equations (GEE) analysis of clustered binary data…

统计理论 · 数学 2011-03-10 Lan Wang

In modern biomedical and econometric studies, longitudinal processes are often characterized by complex time-varying associations and abrupt regime shifts that are shared across correlated outcomes. Standard functional data analysis (FDA)…

统计方法学 · 统计学 2026-01-28 Baolin Chen , Mengfei Ran

The Expectation-Maximization (EM) algorithm is a commonly used method for finding the maximum likelihood estimates of the parameters in a mixture model via coordinate ascent. A serious pitfall with the algorithm is that in the case of…

统计计算 · 统计学 2018-08-31 Adrian O'Hagan , Arthur White

When drawing causal inferences about the effects of multiple treatments on clustered survival outcomes using observational data, we need to address implications of the multilevel data structure, multiple treatments, censoring and unmeasured…

统计方法学 · 统计学 2022-02-18 Liangyuan Hu , Jiayi Ji , Ronald D. Ennis , Joseph W. Hogan

When variable selection methods are applied to bootstrapped and multiply imputed datasets, the set of selected variables typically varies across iterations. Aggregating results via the union rule can lead to overly dense models. We propose…

统计方法学 · 统计学 2026-04-23 Johannes Bleher , Claudia Tarantola

Propensity score methods are widely used in observational studies for evaluating marginal treatment effects. The generalized propensity score (GPS) is an extension of the propensity score framework, historically developed in the case of…

应用统计 · 统计学 2020-07-07 Valérie Garès , Guillaume Chauvet , David Hajage

Causal inference typically assumes centralized access to individual-level data. Yet, in practice, data are often decentralized across multiple sites, making centralization infeasible due to privacy, logistical, or legal constraints. We…

统计方法学 · 统计学 2026-02-04 Rémi Khellaf , Aurélien Bellet , Julie Josse