English
Related papers

Related papers: NBLDA: Negative Binomial Linear Discriminant Analy…

200 papers

Motivation: Modelling methods that find structure in data are necessary with the current large volumes of genomic data, and there have been various efforts to find subsets of genes exhibiting consistent patterns over subsets of treatments.…

Machine Learning · Computer Science 2016-09-15 Kerstin Bunte , Eemeli Leppäaho , Inka Saarinen , Samuel Kaski

Deep sequencing has become one of the most popular tools for transcriptome profiling in biomedical studies. While an abundance of computational methods exists for "normalizing" sequencing data to remove unwanted between-sample variations…

Genomics · Quantitative Biology 2022-01-14 Yannick Düren , Johannes Lederer , Li-Xuan Qin

Rapidly growing public gene expression databases contain a wealth of data for building an unprecedentedly detailed picture of human biology and disease. This data comes from many diverse measurement platforms that make integrating it all…

Genomics · Quantitative Biology 2014-10-16 Karolis Uziela , Antti Honkela

This paper proposes new linear regression models to deal with overdispersed binomial datasets. These new models, called tilted beta binomial regression models, are defined from the tilted beta binomial distribution, proposed assuming that…

Methodology · Statistics 2019-11-26 María Victoria Cifuentes-Amado , Edilberto Cepeda-Cuervo

Postulating that increasing linear energy transfer (LET) causes non-random clustering of lethal lesions to deviate from the Poisson distribution, we employ a non-Poisson approach as a more flexible alternative that accounts for…

Medical Physics · Physics 2020-09-22 M. Loan , M. Alameen , A. Bhat , M. Tantary

This paper presents Sparse Partitioning, a Bayesian method for identifying predictors that either individually or in combination with others affect a response variable. The method is designed for regression problems involving binary or…

Quantitative Methods · Quantitative Biology 2011-08-31 Doug Speed , Simon Tavaré

The evaluation of a match between the DNA profile of a stain found on a crime scene and that of a suspect (previously identified) involves the use of the unknown parameter $p=(p_1, p_2, ...)$, (the ordered vector which represents the…

Applications · Statistics 2015-09-21 Giulia Cereda

The paper proposes to employ deep convolutional neural networks (CNNs) to classify noncoding RNA (ncRNA) sequences. To this end, we first propose an efficient approach to convert the RNA sequences into images characterizing their…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Brian McClannahan , Krushi Patel , Usman Sajid , Cuncong Zhong , Guanghui Wang

This paper presents a new method to estimate systematic errors in the maximum-likelihood regression of count data. The method is applicable in particular to X-ray spectra in situations where the Poisson log-likelihood, or the Cash…

Instrumentation and Methods for Astrophysics · Physics 2023-05-03 M. Bonamente

Distance measures are part and parcel of many computer vision algorithms. The underlying assumption in all existing distance measures is that feature elements are independent and identically distributed. However, in real-world settings,…

Computer Vision and Pattern Recognition · Computer Science 2016-11-01 Muthukaruppan Swaminathan , Pankaj Kumar Yadav , Obdulio Piloto , Tobias Sjöblom , Ian Cheong

Genome annotation is an important issue in biology which has long been addressed with gene prediction methods and manual experiments requiring biological expertise. The expanding Next Generation Sequencing technologies and their enhanced…

Computation · Statistics 2013-07-02 Alice Cleynen , Michel Koskas , Emilie Lebarbier , Guillem Rigaill , Stephane Robin

Three-way data structures, characterized by three entities, the units, the variables and the occasions, are frequent in biological studies. In RNA sequencing, three-way data structures are obtained when high-throughput transcriptome…

Methodology · Statistics 2022-06-22 Anjali Silva , Steven J. Rothstein , Paul D. McNicholas , Xiaoke Qin , Sanjeena Subedi

Distribution regression refers to the supervised learning problem where labels are only available for groups of inputs instead of individual inputs. In this paper, we develop a rigorous mathematical framework for distribution regression…

Machine Learning · Computer Science 2021-09-30 Maud Lemercier , Cristopher Salvi , Theodoros Damoulas , Edwin V. Bonilla , Terry Lyons

Count data analysis is essential across diverse fields, from ecology and accident analysis to single-cell RNA sequencing (scRNA-seq) and metagenomics. While log transformations are computationally efficient, model-based approaches such as…

Methodology · Statistics 2024-11-14 Bastien Batardière , Julien Chiquet , Mahendra Mariadassou

Variable selection and classification are common objectives in the analysis of high-dimensional data. Most such methods make distributional assumptions that may not be compatible with the diverse families of distributions data can take. A…

Methodology · Statistics 2019-08-28 Weichang Yu , Lamiae Azizi , John T. Ormerod

Applying machine learning to biological sequences - DNA, RNA and protein - has enormous potential to advance human health, environmental sustainability, and fundamental biological understanding. However, many existing machine learning…

Machine Learning · Statistics 2023-04-11 Alan Nawzad Amin , Eli Nathan Weinstein , Debora Susan Marks

We propose a compressive classification framework for settings where the data dimensionality is significantly higher than the sample size. The proposed method, referred to as compressive regularized discriminant analysis (CRDA) is based on…

Machine Learning · Statistics 2020-11-13 Muhammad Naveed Tabassum , Esa Ollila

Count data are ubiquitous in ecology and the Poisson generalized linear model (GLM) is commonly used to model the association between counts and explanatory variables of interest. When fitting this model to the data, one typically proceeds…

Methodology · Statistics 2020-07-14 Harlan Campbell

High throughput technologies have become the practice of choice for comparative studies in biomedical applications. Limited number of sample points due to sequencing cost or access to organisms of interest necessitates the development of…

Methodology · Statistics 2018-07-17 Ariana Broumand , Siamak Zamani Dadaneh

High-throughput RNA-sequencing (RNA-seq) technologies are powerful tools for understanding cellular state. Often it is of interest to quantify and summarize changes in cell state that occur between experimental or biological conditions.…

Methodology · Statistics 2021-02-16 Andrew Jones , F. William Townes , Didong Li , Barbara E. Engelhardt
‹ Prev 1 3 4 5 6 7 10 Next ›