English
Related papers

Related papers: Inference on differences between classes using clu…

200 papers

This study proposes a method for imbalanced data classification based on deep probabilistic graphical models (DPGMs) to solve the problem that traditional methods have insufficient learning ability for minority class samples. To address the…

Machine Learning · Computer Science 2025-04-09 Yujia Lou , Jie Liu , Yuan Sheng , Jiawei Wang , Yiwei Zhang , Yaokun Ren

Cancer is a term that denotes a group of diseases caused by abnormal growth of cells that can spread in different parts of the body. According to the World Health Organization (WHO), cancer is the second major cause of death after…

Machine Learning · Computer Science 2023-01-31 Fadi Alharbi , Aleksandar Vakanski

In recent mutation studies, analyses based on protein domain positions are gaining popularity over gene-centric approaches since the latter have limitations in considering the functional context that the position of the mutation provides.…

Subgroup identification (clustering) is an important problem in biomedical research. Gene expression profiles are commonly utilized to define subgroups. Longitudinal gene expression profiles might provide additional information on disease…

Methodology · Statistics 2016-09-27 Jiehuan Sun , Jose D. Herazo-Maya , Naftali Kaminski , Hongyu Zhao , Joshua L. Warren

Generalized Class Discovery (GCD) aims to dynamically assign labels to unlabelled data partially based on knowledge learned from labelled data, where the unlabelled data may come from known or novel classes. The prevailing approach…

Machine Learning · Computer Science 2024-05-01 Ye Wang , Yaxiong Wang , Yujiao Wu , Bingchen Zhao , Xueming Qian

In their recent article, Madej et al. 1 proposed an original way to solve the recurrent issue of controlling for the false discovery rate (FDR) in peptide-spectrum-match (PSM) validation. Briefly, they proposed to derive a single precise…

Genomics · Quantitative Biology 2022-10-18 Lucas Etourneau , Thomas Burger

With the aim to propose a non parametric hypothesis test, this paper carries out a study on the Matching Error (ME), a comparison index of two partitions obtained from the same data set, using for example two clustering methods. This index…

Methodology · Statistics 2019-07-31 Mathias Bourel , Badih Ghattas , Meliza González

Researchers often have to deal with heterogeneous population with mixed regression relationships, increasingly so in the era of data explosion. In such problems, when there are many candidate predictors, it is not only of interest to…

Methodology · Statistics 2021-02-05 Yan Li , Chun Yu , Yize Zhao , Robert H. Aseltine , Weixin Yao , Kun Chen

Motivation. Understanding the pan-cancer mutational landscape offers critical insights into the molecular mechanisms underlying tumorigenesis. While patient-level machine learning techniques have been widely employed to identify tumor…

Machine Learning · Computer Science 2025-08-29 Yifan Dou , Adam Khadre , Ruben C Petreaca , Golrokh Mirzaei

The defect morphology is an essential aspect of the evolution of crystals' microstructure and its response to stress. Existing methods either only report defect concentration or characterize only some of the defect morphologies. The need…

Computational Physics · Physics 2021-04-21 Utkarsh Bhardwaj , Andrea E. Sand , Manoj Warrier

We introduce DiffKnock, a diffusion-based knockoff framework for high-dimensional feature selection with finite-sample false discovery rate (FDR) control. DiffKnock addresses two key limitations of existing knockoff methods: preserving…

Methodology · Statistics 2025-10-03 Heng Ge , Qing Lu

Motivation: Public and private repositories of experimental data are growing to sizes that require dedicated methods for finding relevant data. To improve on the state of the art of keyword searches from annotations, methods for…

Machine Learning · Statistics 2016-01-08 Paul Blomstedt , Ritabrata Dutta , Sohan Seth , Alvis Brazma , Samuel Kaski

This work was motivated by a twin study with the goal of assessing the genetic control of immune traits. We propose a mixture bivariate distribution to model twin data where the underlying order within a pair is unclear. Though estimation…

Methodology · Statistics 2025-07-21 Zonghui Hu , Pengfei Li , Dean Follmann , Jing Qin

Bayesian evidence ratios are widely used to quantify the statistical consistency between different experiments. However, since the evidence ratio is prior dependent, the precise translation between its value and the degree of…

Cosmology and Nongalactic Astrophysics · Physics 2021-11-17 V. Miranda , P. Rogozenski , E. Krause

Causal inference analyses often use existing observational data, which in many cases has some clustering of individuals. In this paper we discuss propensity score weighting methods in a multilevel setting where within clusters individuals…

Applications · Statistics 2020-12-24 Youjin Lee , Trang Q. Nguyen , Elizabeth A. Stuart

The prediction of phenotypic traits using high-density genomic data has many applications such as the selection of plants and animals of commercial interest; and it is expected to play an increasing role in medical diagnostics. Statistical…

Methodology · Statistics 2016-09-29 Marco Scutari , Ian Mackay , David Balding

Given only data generated by a standard confounding graph with unobserved confounder, the Average Treatment Effect (ATE) is not identifiable. To estimate the ATE, a practitioner must then either (a) collect deconfounded data;(b) run a…

Machine Learning · Statistics 2021-03-09 Kyra Gan , Andrew A. Li , Zachary C. Lipton , Sridhar Tayur

Understanding treatment effect heterogeneity is vital for scientific and policy research. However, identifying and evaluating heterogeneous treatment effects pose significant challenges due to the typically unknown subgroup structure.…

Methodology · Statistics 2024-11-05 Kwangho Kim , Jisu Kim , Larry A. Wasserman , Edward H. Kennedy

The high-throughput data generated by microarray experiments provides complete set of genes being expressed in a given cell or in an organism under particular conditions. The analysis of these enormous data has opened a new dimension for…

Computational Engineering, Finance, and Science · Computer Science 2012-11-12 Khalid Raza , Akhilesh Mishra

Detecting predictive biomarkers from multi-omics data is important for precision medicine, to improve diagnostics of complex diseases and for better treatments. This needs substantial experimental efforts that are made difficult by the…

Quantitative Methods · Quantitative Biology 2021-06-08 Betül Güvenç Paltun , Samuel Kaski , Hiroshi Mamitsuka