English
Related papers

Related papers: LMM-Lasso: A Lasso Multi-Marker Mixed Model for As…

200 papers

Constrained least squares regression is an essential tool for high-dimensional data analysis. Given a partition $\mathcal{G}$ of input variables, this paper considers a particular class of nonconvex constraint functions that encourage the…

Machine Learning · Statistics 2014-10-28 Fabian L. Wauthier , Peter Donnelly

We propose a generalization of the lasso that allows the model coefficients to vary as a function of a general set of modifying variables. These modifiers might be variables such as gender, age or time. The paradigm is quite general, with…

Methodology · Statistics 2018-01-11 Robert Tibshirani , Jerome Friedman

Imbalanced classification and spurious correlation are common challenges in data science and machine learning. Both issues are linked to data imbalance, with certain groups of data samples significantly underrepresented, which in turn would…

Machine Learning · Statistics 2026-02-10 Ryumei Nakada , Yichen Xu , Lexin Li , Linjun Zhang

The notion of multivariate total positivity has proved to be useful in finance and psychology but may be too restrictive in other applications. In this paper we propose a concept of local association, where highly connected components in a…

Methodology · Statistics 2022-02-10 Steffen Lauritzen , Piotr Zwiernik

We consider the problem of estimating multiple related but distinct graphical models on the basis of a high-dimensional data set with observations that belong to distinct classes. A motivating example occurs in the analysis of gene…

Methodology · Statistics 2012-07-12 Patrick Danaher , Pei Wang , Daniela M. Witten

Genome-Wide Association Studies (GWAS) help identify genetic variations in people with diseases such as Parkinson's disease (PD), which are less common in those without the disease. Thus, GWAS data can be used to identify genetic variations…

Genomics · Quantitative Biology 2023-04-07 Ali Amelia , Lourdes Pena-Castillo , Hamid Usefi

Blocking, a special case of rerandomization, is routinely implemented in the design stage of randomized experiments to balance the baseline covariates. This study proposes a regression adjustment method based on the least absolute shrinkage…

Methodology · Statistics 2024-11-15 Ke Zhu , Hanzhong Liu , Yuehan Yang

Collection of genotype data in case-control genetic association studies may often be incomplete for reasons related to genes themselves. This non-ignorable missingness structure, if not appropriately accounted for, can result in…

Methodology · Statistics 2024-07-12 Le Wang , Zhengbang Li , Ben Fitzpatrick , Clarice Weinberg , Jinbo Chen

Analysis of high-dimensional data is currently a popular field of research, thanks to many applications e.g. in genetics (DNA data in genomewide association studies), spectrometry or web analysis. At the same time, the type of problems that…

Methodology · Statistics 2018-05-25 Jozef Jakubik

Risk prediction models using genetic data have seen increasing traction in genomics. However, most of the polygenic risk models were developed using data from participants with similar (mostly European) ancestry. This can lead to biases in…

Machine Learning · Computer Science 2022-05-11 Prashnna K Gyawali , Yann Le Guen , Xiaoxia Liu , Hua Tang , James Zou , Zihuai He

Computer vision-based methods have valuable use cases in precision medicine, and recognizing facial phenotypes of genetic disorders is one of them. Many genetic disorders are known to affect faces' visual appearance and geometry. Automated…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Ömer Sümer , Fabio Hellmann , Alexander Hustinx , Tzung-Chien Hsieh , Elisabeth André , Peter Krawitz

Revealing relationships between genes and disease phenotypes is a critical problem in biomedical studies. This problem has been challenged by the heterogeneity of diseases. Patients of a perceived same disease may form multiple subgroups,…

Methodology · Statistics 2022-11-30 Yifan Sun , Ziye Luo , Xinyan Fan

We propose a resampling-based fast variable selection technique for detecting relevant single nucleotide polymorphisms (SNP) in a multi-marker mixed effect model. Due to computational complexity, current practice primarily involves testing…

Applications · Statistics 2025-04-30 Subhabrata Majumdar , Saonli Basu , Matt McGue , Snigdhansu Chatterjee

Symbolic regression (SR) aims to discover mathematical expressions from data, a task traditionally tackled using Genetic Programming (GP) through combinatorial search over symbolic structures. Latent Space Optimization (LSO) methods use…

Neural and Evolutionary Computing · Computer Science 2026-04-14 Benjamin Léger , Kazem Meidani , Christian Gagné

A class of multivariate mixed survival models for continuous and discrete time with a complex covariance structure is introduced in a context of quantitative genetic applications. The methods introduced can be used in many applications in…

Applications · Statistics 2014-05-06 Rafael Pimentel Maia , Per Madsen , Rodrigo Labouriau

In clinical trials, identification of prognostic and predictive biomarkers is essential to precision medicine. Prognostic biomarkers can be useful for the prevention of the occurrence of the disease, and predictive biomarkers can be used to…

Methodology · Statistics 2022-06-28 Wencan Zhu , Céline Lévy-Leduc , Nils Ternès

Social bias is shaped by the accumulation of social perceptions towards targets across various demographic identities. To fully understand such social bias in large language models (LLMs), it is essential to consider the composite of social…

Computation and Language · Computer Science 2024-06-07 Jisu Shin , Hoyun Song , Huije Lee , Soyeong Jeong , Jong C. Park

Whole and targeted sequencing of human genomes is a promising, increasingly feasible tool for discovering genetic contributions to risk of complex diseases. A key step is calling an individual's genotype from the multiple aligned short read…

Applications · Statistics 2012-06-29 Baiyu Zhou , Alice S. Whittemore

In many high dimensional classification or regression problems set in a biological context, the complete identification of the set of informative features is often as important as predictive accuracy, since this can provide mechanistic…

Machine Learning · Computer Science 2020-03-02 Yuxin Sun , Benny Chain , Samuel Kaski , John Shawe-Taylor

This thesis studies two problems in modern statistics. First, we study selective inference, or inference for hypothesis that are chosen after looking at the data. The motiving application is inference for regression coefficients selected by…

Machine Learning · Statistics 2015-07-02 Jason D. Lee
‹ Prev 1 8 9 10 Next ›