English

Retrospective score tests versus prospective score tests for genetic association with case-control data

Methodology 2025-07-29 v1

Abstract

Since the seminal work by Prentice and Pyke (1979), the prospective logistic likelihood has become the standard method of analysis for retrospectively collected case-control data, in particular for testing the association between a single genetic marker and a disease outcome in genetic case-control studies. When studying multiple genetic markers with relatively small effects, especially those with rare variants, various aggregated approaches based on the same prospective likelihood have been developed to integrate subtle association evidence among all considered markers. In this paper we show that using the score statistic derived from a prospective likelihood is not optimal in the analysis of retrospectively sampled genetic data. We develop the locally most powerful genetic aggregation test derived through the retrospective likelihood under a random effect model assumption. In contrast to the fact that the disease prevalence information cannot be used to improve the efficiency for the estimation of odds ratio parameters in logistic regression models, we show that it can be utilized to enhance the testing power in genetic association studies. Extensive simulations demonstrate the advantages of the proposed method over the existing ones. One real genome-wide association study is analyzed for illustration.

Keywords

Cite

@article{arxiv.2507.19893,
  title  = {Retrospective score tests versus prospective score tests for genetic association with case-control data},
  author = {Yukun Liu and Pengfei Li and Lei Song and Kai Yu and Jing Qin},
  journal= {arXiv preprint arXiv:2507.19893},
  year   = {2025}
}
R2 v1 2026-07-01T04:20:06.618Z