English
Related papers

Related papers: Double trouble: Predicting new variant counts acro…

200 papers

Genetic data are frequently categorical and have complex dependence structures that are not always well understood. For this reason, clustering and classification based on genetic data, while highly relevant, are challenging statistical…

Methodology · Statistics 2016-06-13 Gabriela Bettella Cybis , Marcio Valk , Silvia Regina Costa Lopes

Broadening eligibility criteria in cancer trials has been advocated to represent the true patient population more accurately. While the advantages are clear in terms of generalizability and recruitment, novel dose-finding designs are needed…

Applications · Statistics 2023-01-12 Rebecca B. Silva , Bin Cheng , Richard D. Carvajal , Shing M. Lee

In human microbiome studies, sequencing reads data are often summarized as counts of bacterial taxa at various taxonomic levels specified by a taxonomic tree. This paper considers the problem of analyzing two repeated measurements of…

Applications · Statistics 2017-02-17 Pixu Shi , Hongzhe Li

Phenotype variations define heterogeneity of biological and molecular systems, which play a crucial role in several mechanisms. Heterogeneity has been demonstrated in tumor cells. Here, samples from blood of patients affected from colon…

Biological Physics · Physics 2015-11-09 Giuseppina Simone

Cancer prognosis is often based on a set of omics covariates and a set of established clinical covariates such as age and tumor stage. Combining these two sets poses challenges. First, dimension difference: clinical covariates should be…

We consider the problem of model multiplicity in downstream decision-making, a setting where two predictive models of equivalent accuracy cannot agree on the best-response action for a downstream loss function. We show that even when the…

Machine Learning · Computer Science 2024-05-31 Ally Yalei Du , Dung Daniel Ngo , Zhiwei Steven Wu

The wisdom of crowds is the idea that the combination of independent estimates of the magnitude of some quantity yields a remarkably accurate prediction, which is always more accurate than the average individual estimate. In addition, it is…

Information Theory · Computer Science 2020-12-29 Davi A. Nobre , José F. Fontanari

Rapid technological advances have allowed for molecular profiling across multiple omics domains from a single sample for clinical decision making in many diseases, especially cancer. As tumor development and progression are dynamic…

Methodology · Statistics 2022-02-11 Dongyan Yan , Subharup Guha

Observational data about human behavior is often heterogeneous, i.e., generated by subgroups within the population under study that vary in size and behavior. Heterogeneity predisposes analysis to Simpson's paradox, whereby the trends…

Social and Information Networks · Computer Science 2022-12-16 Kristina Lerman

In this work, we investigate the population dynamics of tumor cells under therapeutic pressure. Although drug treatment initially induces a reduction in tumor burden, treatment failure frequently occurs over time due to the emergence of…

Probability · Mathematics 2025-10-02 Kevin Leder , Zicheng Wang , Xuanming Zhang

In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we…

Statistics Theory · Mathematics 2019-11-12 Dave Zachariah , Petre Stoica

The aim of this paper is to propose a novel estimation method of using genetic-predicted observations to estimate trans-ancestry genetic correlations, which describes how genetic architecture of complex traits varies among populations, in…

Methodology · Statistics 2022-03-24 Bingxin Zhao , Xiaochen Yang , Hongtu Zhu

Numerous challenges persist that delay clinical interpretation of human genetic variants, to name a few: (1) un- structured PubMed articles are the most abundant source of evidence, yet their variant annotations are difficult to query…

Information Retrieval · Computer Science 2016-02-10 Andrew J. McMurry

Genomic surveillance of infectious diseases allows monitoring circulating and emerging variants and quantifying their epidemic potential. However, due to the high costs associated with genomic sequencing, only a limited number of samples…

Large variability between cell lines brings a difficult optimization problem of drug selection for cancer therapy. Standard approaches use prediction of value for this purpose, corresponding e.g. to expected value of their distribution.…

Quantitative Methods · Quantitative Biology 2022-09-15 Jarek Duda

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding…

Machine Learning · Statistics 2016-10-04 Xin Gao , Raymond J. Carroll

A critical decision point when training predictors using multiple studies is whether studies should be combined or treated separately. We compare two multi-study prediction approaches in the presence of potential heterogeneity in…

Machine Learning · Statistics 2024-12-13 Zoe Guan , Giovanni Parmigiani , Prasad Patil

In cancer epidemiology using population-based data, regression models for the excess mortality hazard is a useful method to estimate cancer survival and to describe the association between prognosis factors and excess mortality. This method…

Methodology · Statistics 2019-04-19 Francisco J. Rubio , Bernard Rachet , Roch Giorgi , Camille Maringe , Aurelien Belot

Cancer remains one of the most challenging diseases to treat in the medical field. Machine learning has enabled in-depth analysis of rich multi-omics profiles and medical imaging for cancer diagnosis and prognosis. Despite these…

Machine Learning · Computer Science 2024-01-15 Lingchao Mao , Hairong Wang , Leland S. Hu , Nhan L Tran , Peter D Canoll , Kristin R Swanson , Jing Li

Genome-wide association analysis has generated much discussion about how to preserve power to detect signals despite the detrimental effect of multiple testing on power. We develop a weighted multiple testing procedure that facilitates the…

Statistics Theory · Mathematics 2007-06-13 Kathryn Roeder , Bernie Devlin , Larry Wasserman