English
Related papers

Related papers: A Goemans-Williamson type algorithm for identifyin…

200 papers

The use of machine learning algorithms in healthcare can amplify social injustices and health inequities. While the exacerbation of biases can occur and compound during the problem selection, data collection, and outcome definition, this…

Machine Learning · Computer Science 2024-02-26 Nabil Kahouadji

We consider joint estimation of multiple graphical models arising from heterogeneous and high-dimensional observations. Unlike most previous approaches which assume that the cluster structure is given in advance, an appealing feature of our…

Machine Learning · Statistics 2018-01-16 Botao Hao , Will Wei Sun , Yufeng Liu , Guang Cheng

Low-rank matrix approximations are often used to help scale standard machine learning algorithms to large-scale problems. Recently, matrix coherence has been used to characterize the ability to extract global information from a subset of…

Machine Learning · Statistics 2010-09-07 Mehryar Mohri , Ameet Talwalkar

Breast cancer molecular subtypes classification plays an import role to sort patients with divergent prognosis. The biomarkers used are Estrogen Receptor (ER), Progesterone Receptor (PR), HER2, and Ki67. Based on these biomarkers expression…

Machine Learning · Computer Science 2023-10-24 Matheus del-Valle , Emerson Soares Bernardes , Denise Maria Zezell

Cancer detection is one of the key research topics in the medical field. Accurate detection of different cancer types is valuable in providing better treatment facilities and risk minimization for patients. This paper deals with the…

Quantitative Methods · Quantitative Biology 2022-05-31 Yasamin Kowsari , Sanaz Nakhodchi , Davoud Gholamiangonabadi

Various approaches to gene selection for cancer classification based on microarray data can be found in the literature and they may be grouped into two categories: univariate methods and multivariate methods. Univariate methods look at each…

Quantitative Methods · Quantitative Biology 2015-06-18 Min Xu , Rudy Setiono

In modern drug development, the broader availability of high-dimensional observational data provides opportunities for scientist to explore subgroup heterogeneity, especially when randomized clinical trials are unavailable due to cost and…

Methodology · Statistics 2021-02-24 Xinzhou Guo , Linqing Wei , Chong Wu , Jingshen Wang

The discovery of disease subtypes is an essential step for developing precision medicine, and disease subtyping via omics data has become a popular approach. While promising, subtypes obtained from conventional approaches may not be…

Applications · Statistics 2023-09-28 Lingsong Meng , Zhiguang Huo

We develop a feature allocation model for inference on genetic tumor variation using next-generation sequencing data. Specifically, we record single nucleotide variants (SNVs) based on short reads mapped to human reference genome and…

Applications · Statistics 2015-09-15 Juhee Lee , Peter Müller , Kamalakar Gulukota , Yuan Ji

In cancer research, high-throughput profiling has been extensively conducted. In recent studies, the integrative analysis of data on multiple cancer patient groups/subgroups has been conducted. Such analysis has the potential to reveal the…

Methodology · Statistics 2022-12-01 Yifan Sun , Zhengyang Sun , Yu Jiang , Yang Li , Shuangge Ma

We introduce the Poisson Log-Normal Graphical Model for count data, and present a normality transformation for data arising from this distribution. The model and transformation are feasible for high-throughput microRNA (miRNA) sequencing…

Computation · Statistics 2017-08-16 David Sinclair , Giles Hooker

Heterogeneity in the cell population of cancer tissues poses many challenges in cancer diagnosis and treatment. Studying the heterogeneity in cell populations from gene expression measurement data in the context of cancer research is a…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-01-31 Anik Chaudhuri , Anwoy Mohanty , Manoranjan Satpathy

In a traditional Gaussian graphical model, data homogeneity is routinely assumed with no extra variables affecting the conditional independence. In modern genomic datasets, there is an abundance of auxiliary information, which often gets…

Methodology · Statistics 2023-08-16 Yabo Niu , Yang Ni , Debdeep Pati , Bani K. Mallick

Patients at a comprehensive cancer center who do not achieve cure or remission following standard treatments often become candidates for clinical trials. Patients who participate in a clinical trial may be suitable for other studies. A key…

Social and Information Networks · Computer Science 2024-11-07 Benjamin Smith , Tyler Pittman , Wei Xu

The accurate classification of lymphoma subtypes using hematoxylin and eosin (H&E)-stained tissue is complicated by the wide range of morphological features these cancers can exhibit. We present LymphoML - an interpretable machine learning…

In cancer research, the comparison of gene expression or DNA methylation networks inferred from healthy controls and patients can lead to the discovery of biological pathways associated to the disease. As a cancer progresses, its signalling…

Methodology · Statistics 2015-06-17 Da Ruan , Alastair Young , Giovanni Montana

Recent advances in cancer research largely rely on new developments in microscopic or molecular profiling techniques offering high level of detail with respect to either spatial or molecular features, but usually not both. Here, we present…

Randomization tests are a popular method for testing causal effects in clinical trials with finite-sample validity. In the presence of heterogeneous treatment effects, it is often of interest to select a subgroup that benefits from the…

Methodology · Statistics 2025-04-29 Zijun Gao

We consider comparisons of statistical learning algorithms using multiple data sets, via leave-one-in cross-study validation: each of the algorithms is trained on one data set; the resulting model is then validated on each remaining data…

Applications · Statistics 2015-06-02 Lorenzo Trippa , Levi Waldron , Curtis Huttenhower , Giovanni Parmigiani

Given a gene expression data array of a list of bladder cancer patients with their tumor states, it may be difficult to determine which genes can operate as disease markers when the array is large and possibly contains outliers and missing…