English
Related papers

Related papers: Conditional regression based on a multivariate zer…

200 papers

High dimensional and heterogeneous count data are collected in various applied fields. In this paper, we look closely at high-resolution sequencing data on the microbiome, which have enabled researchers to study the genomes of entire…

Methodology · Statistics 2024-01-12 Veronica Vinciotti , Pariya Behrouzi , Reza Mohammadi

Binomial data with unknown sizes often appear in biological and medical sciences. The previous methods either use the Poisson approximation or the quasi-likelihood approach. A full likelihood approach is proposed by treating unknown sizes…

Statistics Theory · Mathematics 2007-06-13 Wei Zhang

We present an approach for imputation of missing items in multivariate categorical data nested within households. The approach relies on a latent class model that (i) allows for household level and individual level variables, (ii) ensures…

Methodology · Statistics 2018-07-05 Olanrewaju Akande , Jerome Reiter , Andrés F. Barrientos

Using a sample from a population to estimate the proportion of the population with a certain category label is a broadly important problem. In the context of microbiome studies, this problem arises when researchers wish to use a sample from…

Methodology · Statistics 2019-02-08 Bryan D. Martin , Daniela Witten , Amy D. Willis

Structural transformation, the shift from agrarian economies to more diversified industrial and service-based systems, is a key driver of economic development. However, in low- and middle-income countries (LMICs), data scarcity and…

Applications · Statistics 2025-10-02 Ronald Katende

Bivariate ordered logistic models (BOLMs) are appealing to jointly model the marginal distribution of two ordered responses and their association, given a set of covariates. When the number of categories of the responses increases, the…

Statistics Theory · Mathematics 2014-07-08 Marco Enea , Gianfranco Lovison

Bayesian modelling for cost-effectiveness data has received much attention in both the health economics and the statistical literature in recent years. Cost-effectiveness data are characterised by a relatively complex structure of…

Statistics Theory · Mathematics 2013-07-22 Gianluca Baio

Building classification models that predict a binary class label on the basis of high dimensional multi-omics datasets poses several challenges, due to the typically widely differing characteristics of the data layers in terms of number of…

Methodology · Statistics 2020-08-04 Alessandra Cabassi , Denis Seyres , Mattia Frontini , Paul D. W. Kirk

Multivariate count data are commonly encountered through high-throughput sequencing technologies in bioinformatics, text mining, or in sports analytics. Although the Poisson distribution seems a natural fit to these count data, its…

Computation · Statistics 2020-04-16 Sanjeena Subedi , Ryan Browne

New technologies have enabled the investigation of biology and human health at an unprecedented scale and in multiple dimensions. These dimensions include a myriad of properties describing genome, epigenome, transcriptome, microbiome,…

Quantitative Methods · Quantitative Biology 2018-10-22 Marinka Zitnik , Francis Nguyen , Bo Wang , Jure Leskovec , Anna Goldenberg , Michael M. Hoffman

The construction of coherent prediction models holds great importance in medical research as such models enable health researchers to gain deeper insights into disease epidemiology and clinicians to identify patients at higher risk of…

Applications · Statistics 2024-01-17 Guanbo Wang , Sylvie Perreault , Robert W. Platt , Rui Wang , Marc Dorais , Mireille E. Schnitzer

Microbial identification is a central issue in microbiology, in particular in the fields of infectious diseases diagnosis and industrial quality control. The concept of species is tightly linked to the concept of biological and clinical…

Machine Learning · Statistics 2015-06-25 Kévin Vervier , Pierre Mahé , Jean-Baptiste Veyrieras , Jean-Philippe Vert

We propose a multiple imputation method to deal with incomplete categorical data. This method imputes the missing entries using the principal components method dedicated to categorical data: multiple correspondence analysis (MCA). The…

Methodology · Statistics 2015-06-01 Vincent Audigier , François Husson , Julie Josse

Logistic regression is a fundamental and widely used statistical method for modeling binary outcomes based on covariates. However, the presence of missing data, particularly in settings involving hybrid covariates (a mix of discrete and…

Methodology · Statistics 2025-06-05 Mohamed Cherifi , Xujia Zhu , Mohammed Nabil El Korso , Ammar Mesloub

Compositional data and multivariate count data with known totals are challenging to analyse due to the non-negativity and sum-to-one constraints on the sample space. It is often the case that many of the compositional components are highly…

Methodology · Statistics 2020-12-24 Janice L. Scealy , Andrew T. A. Wood

Compositional Data Analysis (CoDa) has gained popularity in recent years. This type of data consists of values from disjoint categories that sum up to a constant. Both Dirichlet regression and logistic-normal regression have become popular…

Methodology · Statistics 2024-06-25 Joaquín Martínez-Minaya , Haavard Rue

This paper studies the problem of statistical inference for genetic relatedness between binary traits based on individual-level genome-wide association data. Specifically, under the high-dimensional logistic regression models, we define…

Methodology · Statistics 2022-10-06 Rong Ma , Zijian Guo , T. Tony Cai , Hongzhe Li

We develop a new structured compartmental model for the coevolutionary dynamics between susceptible and infectious individuals in heterogeneous SI epidemiological systems. In this model, the susceptible compartment is structured by a…

Populations and Evolution · Quantitative Biology 2024-10-10 Tommaso Lorenzi , Elisa Paparelli , Andrea Tosin

We introduce a semi-parametric approach to ecological regression for disease mapping, based on modelling the regression M-quantiles of a Negative Binomial variable. The proposed method is robust to outliers in the model covariates,…

Methodology · Statistics 2014-08-14 Ray Chambers , Emanuela Dreassi , Nicola Salvati

While several Gaussian mixture models-based biclustering approaches currently exist in the literature for continuous data, approaches to handle discrete data have not been well researched. A multivariate Poisson-lognormal (MPLN) model-based…

Methodology · Statistics 2025-03-13 Caitlin Kral , Evan Chance , Ryan Browne , Sanjeena Subedi