中文
相关论文

相关论文: Conditional regression based on a multivariate zer…

200 篇论文

There is a keen interest in characterizing variation in the microbiome across cancer patients, given increasing evidence of its important role in determining treatment outcomes. Here our goal is to discover subgroups of patients with…

应用统计 · 统计学 2022-12-06 Yushu Shi , Liangliang Zhang , Kim-Anh Do , Robert Jenq , Christine Peterson

Owing to the advantages of increased accuracy and the potential to detect unseen patterns, provided by data mining techniques they have been widely incorporated for standard classification problems. They have often been used for high…

机器学习 · 计算机科学 2022-10-05 Anirudha Rayasam , Nagamma Patil

Many processes of scientific importance are characterized by time scales that extend far beyond the reach of standard simulation techniques. To circumvent this impediment a plethora of enhanced sampling methods has been developed. One…

Analyzing multivariate count data generated by high-throughput sequencing technology in microbiome research studies is challenging due to the high-dimensional and compositional structure of the data and overdispersion. In practice,…

应用统计 · 统计学 2023-11-03 Jingyan Fu , Matthew D. Koslovsky , Andreas M. Neophytou , Marina Vannucci

Advances in data collecting technologies in genomics have significantly increased the need for tools designed to study the genetic basis of many diseases. Effective statistical methods should excel in both prediction accuracy and biomarker…

统计方法学 · 统计学 2025-11-13 Anthony-Alexander Christidis , Stefan Van Aelst , Ruben Zamar

Logistic regression with unknown sizes has many important applications in biological and medical sciences. All models about this problem in the literature are parametric ones. A semiparametric regression model is proposed. This model…

统计理论 · 数学 2007-06-13 Wei Zhang

Calibrating mathematical models of biological processes is essential for achieving predictive accuracy and gaining mechanistic insight. However, this task remains challenging due to limited and noisy data, significant biological…

定量方法 · 定量生物学 2025-12-04 Piotr Gwiazda , Alexey Kazarnikov , Anna Marciniak-Czochra , Zuzanna Szymańska

We present a robust framework to perform linear regression with missing entries in the features. By considering an elliptical data distribution, and specifically a multivariate normal model, we are able to conditionally formulate a…

机器学习 · 计算机科学 2022-11-10 Alireza Aghasi , MohammadJavad Feizollahi , Saeed Ghadimi

We consider the problem of inferring the values of an arbitrary set of variables (e.g., risk of diseases) given other observed variables (e.g., symptoms and diagnosed diseases) and high-dimensional signals (e.g., MRI images or EEG). This is…

机器学习 · 统计学 2019-02-07 Hao Wang , Chengzhi Mao , Hao He , Mingmin Zhao , Tommi S. Jaakkola , Dina Katabi

We propose a comprehensive Bayesian joint modeling framework for zero-inflated longitudinal count data and time-to-event outcomes, explicitly incorporating a cure fraction to account for subjects who never experience the event. The…

统计方法学 · 统计学 2025-08-27 Taban Baghfalaki , Mojtaba Ganjali

We address the problem of survival regression modelling with multivariate responses and nonlinear covariate effects. Our model extends the proportional hazards model by introducing several weakly-parametric elements: the marginal baseline…

统计方法学 · 统计学 2025-10-16 Na Lei , Mark A. Wolters , Wenqing He

Data-driven discovery of governing equations from time-series data provides a powerful framework for understanding complex biological systems. Library-based approaches that use sparse regression over candidate functions have shown…

定量方法 · 定量生物学 2026-03-13 Yuxiang Feng , Niall M Mangan , Manu Jayadharan

Count-compositional data arise in many different fields, including high-throughput sequencing experiments, ecological surveys, and palaeoclimate studies, where a common, important goal is to understand how covariates relate to the observed…

统计方法学 · 统计学 2026-04-10 André F. B. Menezes , Andrew C. Parnell , Keefe Murphy

In binary classification, imbalance refers to situations in which one class is heavily under-represented. This issue is due to either a data collection process or because one class is indeed rare in a population. Imbalanced classification…

统计方法学 · 统计学 2022-01-07 Arezou Mojiri , Abbas Khalili , Ali Zeinal Hamadani

Observational longitudinal data on treatments and covariates are increasingly used to investigate treatment effects, but are often subject to time-dependent confounding. Marginal structural models (MSMs), estimated using inverse probability…

统计方法学 · 统计学 2020-02-11 Ruth H. Keogh , Shaun R. Seaman , Jon Michael Gran , Stijn Vansteelandt

In the last decade, the secondary use of large data from health systems, such as electronic health records, has demonstrated great promise in advancing biomedical discoveries and improving clinical decision making. However, there is an…

统计理论 · 数学 2021-03-25 Rui Duan , Yang Ning , Jiasheng Shi , Raymond J Carroll , Tianxi Cai , Yong Chen

The study of immune cellular composition has been of great scientific interest in immunology because of the generation of multiple large-scale data. From the statistical point of view, such immune cellular data should be treated as…

应用统计 · 统计学 2022-04-22 Jinkyung Yoo , Zequn Sun , Michael Greenacre , Qin Ma , Dongjun Chung , Young Min Kim

This paper considers the modeling of zero-inflated circular measurements concerning real case studies from medical sciences. Circular-circular regression models have been discussed in the statistical literature and illustrated with various…

统计方法学 · 统计学 2022-01-04 Jayant Jha , Prajamitra Bhuyan

Microbiome-based stratification of healthy individuals into compositional categories, referred to as "community types", holds promise for drastically improving personalized medicine. Despite this potential, the existence of community types…

定量方法 · 定量生物学 2016-04-27 Travis E. Gibson , Amir Bashan , Hong-Tai Cao , Scott T. Weiss , Yang-Yu Liu

Multivariate longitudinal data of mixed-type are increasingly collected in many science domains. However, algorithms to cluster this kind of data remain scarce, due to the challenge to simultaneously model the within- and between-time…

机器学习 · 统计学 2025-09-16 Francesco Amato , Julien Jacques