English
Related papers

Related papers: Optimal two-phase sampling designs for generalized…

200 papers

This paper presents an adaptive sampling algorithm tailored for the optimization of parametrized dynamical systems using projection-based model order reduction. Unlike classical sampling strategies, this framework does not aim for a small…

Computational Engineering, Finance, and Science · Computer Science 2026-02-27 Marcel Warzecha , Sebastian Resch-Schopper , Gerhard Müller

Large-scale rare events data are commonly encountered in practice. To tackle the massive rare events data, we propose a novel distributed estimation method for logistic regression in a distributed system. For a distributed framework, we…

Methodology · Statistics 2023-04-06 Xuetong Li , Xuening Zhu , Hansheng Wang

One common approach for dose optimization is a two-stage design, which initially conducts dose escalation to identify the maximum tolerated dose (MTD), followed by a randomization stage where patients are assigned to two or more doses to…

Methodology · Statistics 2024-11-11 Yixuan Zhao , Rachael Liu , Jianchang Lin , Ying Yuan

Inverse probability weighting (IPW) methods are commonly used to analyze non-ignorable missing data under the assumption of a logistic model for the missingness probability. However, solving IPW equations numerically may involve…

Methodology · Statistics 2025-07-24 Pengfei Li , Jing Qin , Yukun Liu

Personalized medicine has received increasing attention among statisticians, computer scientists, and clinical practitioners. A major component of personalized medicine is the estimation of individualized treatment rules (ITRs). Recently,…

Methodology · Statistics 2015-08-14 Xin Zhou , Nicole Mayer-Hamblett , Umer Khan , Michael R. Kosorok

In big data analysis, a simple task such as linear regression can become very challenging as the variable dimension $p$ grows. As a result, variable screening is inevitable in many scientific studies. In recent years, randomized algorithms…

Methodology · Statistics 2019-02-13 Yu-Hsiang Cheng , Tzee-Ming Huang , Su-Yun Huang

There are now many options for doubly robust estimation; however, there is a concerning trend in the applied literature to believe that the combination of a propensity score and an adjusted outcome model automatically results in a doubly…

Inverse probability weighting (IPW) is widely used in many areas when data are subject to unrepresentativeness, missingness, or selection bias. An inevitable challenge with the use of IPW is that the IPW estimator can be remarkably unstable…

Methodology · Statistics 2021-11-29 Yukun Liu , Yan Fan

We develop sampling methods, which consist of Gaussian invariant versions of random walk Metropolis (RWM), Metropolis adjusted Langevin algorithm (MALA) and second order Hessian or Manifold MALA. Unlike standard RWM and MALA we show that…

Machine Learning · Statistics 2025-06-27 Michalis K. Titsias , Angelos Alexopoulos , Siran Liu , Petros Dellaportas

We study the problem of estimating a function of many parameters acquired by sensors that are distributed in space, e.g., the spatial gradient of a field. We restrict ourselves to a setting where the distributed sensors are probed with…

Quantum Physics · Physics 2018-10-03 T. J. Volkoff , Mohan Sarovar

Genetic risk prediction is an important component of individualized medicine, but prediction accuracies remain low for many complex diseases. A fundamental limitation is the sample sizes of the studies on which the prediction algorithms are…

Methodology · Statistics 2017-06-20 Sihai Dave Zhao

With the rapid development of new anti-cancer agents which are cytostatic, new endpoints are needed to better measure treatment efficacy in phase II trials. For this purpose, Von Hoff (1998) proposed the growth modulation index (GMI), i.e.…

Methodology · Statistics 2021-11-01 Li Chen , Mark Burkard , Jianrong Wu , Jill M. Kolesar , Chi Wang

We develop a model-based methodology for integrating gene-set information with an experimentally-derived gene list. The methodology uses a previously reported sampling model, but takes advantage of natural constraints in the…

Methodology · Statistics 2015-06-02 Zhishi Wang , Qiuling He , Bret Larget , Michael A. Newton

Homogenization is an important and crucial step to improve the usage of observational data for climate analysis. This work is motivated by the analysis of long series of GNSS Integrated Water Vapour (IWV) data which have not yet been used…

Methodology · Statistics 2020-05-12 Annarosa Quarello , Olivier Bock , Emilie Lebarbier

In a sequential multiple-assignment randomized trial (SMART), a sequence of treatments is given to a patient over multiple stages. In each stage, randomization may be done to allocate patients to different treatment groups. Even though…

Methodology · Statistics 2024-01-09 Rik Ghosh , Bibhas Chakraborty , Inbal Nahum-Shani , Megan E. Patrick , Palash Ghosh

Integrating non-probability samples into finite-population inference typically requires modeling unknown selection probabilities under a missing-at-random (MAR) assumption that is difficult to verify. We propose a design-based alternative…

Methodology · Statistics 2026-05-08 Andrius Čiginas , Ieva Burakauskaitė , Jae Kwang Kim

Semi-parametric methods are often used for the estimation of intervention effects on correlated outcomes in cluster-randomized trials (CRTs). When outcomes are missing at random (MAR), Inverse Probability Weighted (IPW) methods…

Methodology · Statistics 2016-01-27 Melanie Prague , Rui Wang , Alisa Stephens , Eric Tchetgen Tchetgen , Victor DeGruttola

Multi-arm trials are gaining interest in practice given the statistical and logistical advantages they can offer. The standard approach uses a fixed allocation ratio, but there is a call for making it adaptive and skewing the allocation of…

Methodology · Statistics 2026-01-16 Gianmarco Caruso , Pavel Mozgunov

Two-sample testing is a fundamental problem in statistics. Despite its long history, there has been renewed interest in this problem with the advent of high-dimensional and complex data. Specifically, in the machine learning literature,…

Methodology · Statistics 2019-11-19 Ilmun Kim , Ann B. Lee , Jing Lei

Two-phase outcome dependent sampling (ODS) is widely used in many fields, especially when certain covariates are expensive and/or difficult to measure. For two-phase ODS, the conditional maximum likelihood (CML) method is very attractive…

Methodology · Statistics 2022-12-21 Menglu Che , Peisong Han , Jerald F. Lawless