English
Related papers

Related papers: Differential protein expression and peak selection…

200 papers

Protein Language Models (PLMs) have emerged as performant and scalable tools for predicting the functional impact and clinical significance of protein-coding variants, but they still lag experimental accuracy. Here, we present a novel…

Differential analysis is a routine procedure in the statistical analysis toolbox across many applied fields, including quantitative proteomics, the main illustration of the present paper. The state-of-the-art limma approach uses a…

Methodology · Statistics 2025-12-12 Marie Chion , Arthur Leroy

Many genomic experiments, notably microarray experiments seeking to detect differential gene expression, involve calculating a large number of p-values. This leads to the multiple testing problem: when the number of null hypotheses is…

Quantitative Methods · Quantitative Biology 2007-05-23 David R. Bickel

The BayesBinMix package offers a Bayesian framework for clustering binary data with or without missing values by fitting mixtures of multivariate Bernoulli distributions with an unknown number of components. It allows the joint estimation…

Computation · Statistics 2017-07-03 Panagiotis Papastamoulis , Magnus Rattray

Data scarcity hinders deep learning for medical imaging. We propose a framework for breast cancer classification in thermograms that addresses this using a Diffusion Probabilistic Model (DPM) for data augmentation. Our DPM-based…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Sepehr Salem , M. Moein Esfahani , Jingyu Liu , Vince Calhoun

We propose a method for detecting differential gene expression that exploits the correlation between genes. Our proposal averages the univariate scores of each feature with the scores in correlation neighborhoods. In a number of real and…

Statistics Theory · Mathematics 2007-06-13 Robert Tibshirani , Larry Wasserman

Protein identification is one of the major task of Proteomics researchers. Protein identification could be resumed by searching the best match between an experimental mass spectrum and proteins from a database. Nevertheless this approach…

Biomolecules · Quantitative Biology 2008-12-18 Jean-Charles Boisson , Laetitia Jourdan , El-Ghazali Talbi

In recent years many sparse linear discriminant analysis methods have been proposed for high-dimensional classification and variable selection. However, most of these proposals focus on binary classification and they are not directly…

Methodology · Statistics 2015-04-23 Qing Mai , Yi Yang , Hui Zou

Intrinsically disordered regions of proteins play a crucial role in cell signaling and drug discovery. However, their high structural flexibility makes accurate residue-level prediction challenging. Existing methods often rely on…

Neural and Evolutionary Computing · Computer Science 2026-03-09 Shaokuan Wang , Pengshan Cui , Yining Qian , An-Yang Lu , Xianpeng Wang

Personalized medicine aims at identifying best treatments for a patient with given characteristics. It has been shown in the literature that these methods can lead to great improvements in medicine compared to traditional methods…

Machine Learning · Statistics 2018-11-26 Oleg Sysoev , Krzysztof Bartoszek , Eva-Charlotte Ekstrom , Katarina Ekholm Selling

The ability to characterize proteins at sequence-level resolution is vital to biological research. Currently, the leading method for protein sequencing is by liquid chromatography mass spectrometry (LC-MS) whereas proteins are reduced to…

Genomics · Quantitative Biology 2025-07-11 René L. Warren

The standard methods for detecting differential gene expression are mostly designed for analyzing a single gene expression experiment. When data from multiple related gene expression studies are available, separately analyzing each study is…

Methodology · Statistics 2013-11-07 Yingying Wei , Hongkai Ji

Linear shrinkage estimators of a covariance matrix --- defined by a weighted average of the sample covariance matrix and a pre-specified shrinkage target matrix --- are popular when analysing high-throughput molecular data. However, their…

Methodology · Statistics 2018-09-24 Harry Gray , Gwenaël G. R. Leday , Catalina A. Vallejos , Sylvia Richardson

In the context of machine learning, disparate impact refers to a form of systematic discrimination whereby the output distribution of a model depends on the value of a sensitive attribute (e.g., race or gender). In this paper, we propose an…

Information Theory · Computer Science 2018-05-14 Hao Wang , Berk Ustun , Flavio P. Calmon

Designing protein binders targeting specific sites, which requires to generate realistic and functional interaction patterns, is a fundamental challenge in drug discovery. Current structure-based generative models are limited in generating…

Machine Learning · Computer Science 2025-10-17 Zishen Zhang , Xiangzhe Kong , Wenbing Huang , Yang Liu

This paper introduces a kernel discrepancy-based framework for rerandomization to enhance the precision of causal inference in controlled experiments. We demonstrate that the kernel discrepancy is the key part of the variance upper bound…

Methodology · Statistics 2025-11-05 Yiou Li , Lulu Kang

Discovering patterns in data that best describe the differences between classes allows to hypothesize and reason about class-specific mechanisms. In molecular biology, for example, this bears promise of advancing the understanding of…

Machine Learning · Computer Science 2023-12-08 Nils Philipp Walter , Jonas Fischer , Jilles Vreeken

Sequence specific resonance assignment and secondary structure determination of proteins form the basis for variety of structural and functional proteomics studies by NMR. In this context, an efficient standalone method for rapid assignment…

Biological Physics · Physics 2013-09-05 Dinesh Kumar

DNA methylation datasets in cancer studies are comprised of measurements on a large number of genomic locations called cytosine-phosphate-guanine (CpG) sites with complex correlation structures. A fundamental goal of these studies is the…

Methodology · Statistics 2023-05-05 Chiyu Gu , Veerabhadran Baladandayuthapani , Subharup Guha

Data mining techniques have been used by researchers for analyzing protein sequences. In protein analysis, especially in protein sequence classification, selection of feature is most important. Popular protein sequence classification…

Databases · Computer Science 2012-11-22 Suprativ Saha , Rituparna Chaki