English
Related papers

Related papers: A Weighted U Statistic for Genetic Association Ana…

200 papers

We investigate the online detection of changepoints in the distribution of a sequence of observations using degenerate U-statistic-type processes. We study weighted versions of: an ordinary, CUSUM-type scheme, a Page-CUSUM-type scheme, and…

Statistics Theory · Mathematics 2025-10-28 Cooper Boniece , Lajos Horvath , Lorenzo Trapani

Genome-wide association studies (GWAS) suggests that a complex disease is typically affected by many genetic variants with small or moderate effects. Identification of these risk variants remains to be a very challenging problem.…

Methodology · Statistics 2014-01-21 Dongjun Chung , Can Yang , Cong Li , Joel Gelernter , Hongyu Zhao

We provide a view on high-dimensional statistical inference for genome-wide association studies (GWAS). It is in part a review but covers also new developments for meta analysis with multiple studies and novel software in terms of an…

Applications · Statistics 2020-02-17 Claude Renaux , Laura Buzdugan , Markus Kalisch , Peter Bühlmann

Accurate identification of breast lesion subtypes can facilitate personalized treatment and interventions. Ultrasound (US), as a safe and accessible imaging modality, is extensively employed in breast abnormality screening and diagnosis.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Shijing Chen , Xinrui Zhou , Yuhao Wang , Yuhao Huang , Ao Chang , Dong Ni , Ruobing Huang

The case-control design is often used to test associations between the case-control status and genetic variants. In addition to this primary phenotype a number of additional traits, known as secondary phenotypes, are routinely recorded and…

A large number of recent genome-wide association studies (GWASs) for complex phenotypes confirm the early conjecture for polygenicity, suggesting the presence of large number of variants with only tiny or moderate effects. However, due to…

Genomics · Quantitative Biology 2018-05-01 Mingwei Dai , Xiang Wan , Hao Peng , Yao Wang , Yue Liu , Jin Liu , Zongben Xu , Can Yang

Identifying disease-associated genes enables the development of precision medicine and the understanding of biological processes. Genome-wide association studies (GWAS), gene expression data, biological pathway analysis, and protein network…

Genomics · Quantitative Biology 2026-03-10 Muhammad Muneeb , David B. Ascher , YooChan Myung

In this paper, a robust weighted score for unbalanced data (ROWSU) is proposed for selecting the most discriminative feature for high dimensional gene expression binary classification with class-imbalance problem. The method addresses one…

Machine Learning · Statistics 2024-01-24 Zardad Khan , Amjad Ali , Saeed Aldahmani

Modern statistical analyses often involve testing large numbers of hypotheses. In many situations, these hypotheses may have an underlying tree structure that not only helps determine the order that tests should be conducted but also…

Methodology · Statistics 2019-03-19 Yunxiao Li , Yi-Juan Hu , Glen A. Satten

The q-Gaussians are a class of stable distributions which are present in many scientific fields, and that behave as heavy tailed distributions for an especific range of q values. The identification of these values, which are used in the…

Data Analysis, Statistics and Probability · Physics 2015-06-11 E. L de Santa Helena , C. M. Nascimento , G. J. L. Gerhardt

Genomic datasets generated with massively parallel sequencing methods have the potential to propel systematics in new and exciting directions, but selecting appropriate markers and methods is not straightforward. We applied two approaches…

Genomics · Quantitative Biology 2017-03-28 Michael G. Harvey , Brian Tilston Smith , Travis C. Glenn , Brant C. Faircloth , Robb T. Brumfield

Multi-trait genome-wide association studies (GWAS) use multi-variate statistical methods to identify associations between genetic variants and multiple correlated traits simultaneously, and have higher statistical power than independent…

Genomics · Quantitative Biology 2022-02-10 Muhammad Ammar Malik , Adriaan-Alexander Ludl , Tom Michoel

Federated learning enables collaborative training of deep learning models across institutions without sharing sensitive patient data. However, its performance is often limited by small datasets and non-independent, identically distributed…

Image and Video Processing · Electrical Eng. & Systems 2026-04-17 Hongyi Pan , Ziliang Hong , Gorkem Durak , Ziyue Xu , Ulas Bagci

High-dimensional data arise routinely in modern statistics, econometrics, finance, genomics, and machine learning. While a large body of existing methodology is developed under Gaussian or light-tailed assumptions, many real data sets…

Methodology · Statistics 2026-04-16 Long Feng

Clinical end-point traits are often characterized by quantitative or qualitative precursors and it has been argued that it may be statistically a more powerful strategy to analyze these precursor traits to decipher the genetic architecture…

Methodology · Statistics 2025-04-17 Soumya Sahu , Saurabh Ghosh

Small datasets are common in health research. However, the generalization performance of machine learning models is suboptimal when the training datasets are small. To address this, data augmentation is one solution. Augmentation increases…

Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with complex traits, and some variants are shown to be associated with multiple complex traits. Genetic covariance between two traits is defined…

Methodology · Statistics 2023-10-06 Jianqiao Wang , Sai Li , Hongzhe Li

Data selection seeks to identify a compact yet informative subset from large-scale training corpora, balancing sample quality against collection diversity. We formulate this problem as a Weighted Independent Set (WIS) on a similarity graph,…

Machine Learning · Computer Science 2026-05-21 Yuan Zhang , Lifeng Guo , Junwen Pan , Wenzhao Zheng , Wen Zhou , Kuan Cheng , Kurt Keutzer , Shanghang Zhang

We perform differential expression analysis of high-throughput sequencing count data under a Bayesian nonparametric framework, removing sophisticated ad-hoc pre-processing steps commonly required in existing algorithms. We propose to use…

Applications · Statistics 2017-05-04 Siamak Zamani Dadaneh , Xiaoning Qian , Mingyuan Zhou

Stepped wedge cluster randomized trials (SW-CRTs) have become increasingly popular and are used for a variety of interventions and outcomes, often chosen for their feasibility advantages. SW-CRTs must account for time trends in the outcome…

Methodology · Statistics 2024-07-16 Lee Kennedy-Shaffer , Victor De Gruttola , Marc Lipsitch